Constructs and sequences for enhanced gene expression
By using a combination of first and second promoters and intron sequences in the nucleic acid construct, the problems of poor transcription and expression in the prior art are solved, and significant enhancement of expression of the protein or peptide of interest and increase of transcripts are achieved, which is applicable to various cell systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PROTEONIC BIOTECHNOLOGY IP BV
- Filing Date
- 2014-12-24
- Publication Date
- 2026-07-31
AI Technical Summary
There is a lack of effective methods in the current technology to regulate the transcription of transcripts and the expression of proteins or peptides of interest in host cells, especially in cases of poor or high initial expression, making it difficult to achieve higher levels of transcription and expression.
By employing a nucleic acid construct containing first and second promoters and intron sequences flanking them, and by operatively linking these elements to the nucleotide sequence of interest, a nucleic acid construct is formed and transformed into cells to achieve enhanced transcription and expression.
It significantly increased the expression levels of proteins or peptides of interest, improved transcript production, and promoted the generation of clone lines and the proportion of high-production cell lines, especially in cases of initially poor or high expression, achieving higher levels of transcription and expression.
Smart Images

Figure CN106061511B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods for transcription and expression using nucleic acid constructs characterized by the presence of a promoter followed by an intron promoter. The invention further relates to said nucleic acid construct, expression vectors, and cells containing said construct, and their applications.
[0002] The present invention also relates to methods for using nucleotide sequences for transcription and optionally expression. The present invention further relates to said nucleotide sequences and constructs, expression vectors, and cells comprising said nucleotide sequences, and their applications. Background Technology
[0003] There is still a need in the art for alternative and preferred improved methods for regulating the transcription of transcripts and, optionally, the expression of proteins or peptides of interest in host cells. Summary of the Invention
[0004] This invention relates to a method for transcription and, optionally, for purifying the resulting transcript, comprising the following steps:
[0005] a) Provide a nucleic acid construct containing a first promoter, a second promoter, and a nucleotide sequence of interest, wherein the first and second promoters are operatively linked to the nucleotide sequence of interest, and wherein the second promoter is flanked by a first intron sequence upstream of the promoter and a second intron sequence downstream of the promoter; and,
[0006] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0007] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally,
[0008] d) Purify the resulting transcript.
[0009] The present invention further relates to a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0010] a) Provides a nucleic acid construct comprising a first promoter, a second promoter, and a nucleotide sequence encoding a protein or polypeptide of interest, wherein the first and second promoters are operatively linked to the nucleotide sequence encoding the protein or polypeptide of interest, and wherein the second promoter is flanked by a first intron sequence upstream of the promoter and a second intron sequence downstream of the promoter; and,
[0011] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0012] c) Inducing the transformed cells to express the protein or polypeptide of interest; and optionally, purifying the protein or polypeptide of interest.
[0013] Preferably, the first intron sequence includes at least a donor splice site and the second intron sequence includes at least an acceptor splice site. Furthermore, the nucleic acid construct of step a) of the method of the present invention includes the following nucleotide sequences shown herein in their relative positions in the 5' to 3' direction: (i) a first promoter; (ii) a first intron sequence including at least a donor splice site; (iii) a second promoter; (iv) a second intron sequence including at least an acceptor splice site; and (v) a nucleotide sequence encoding a protein or polypeptide of interest, wherein preferably the first promoter, the first intron sequence including at least a donor splice site, the second promoter, and the second intron sequence including at least an acceptor splice site are all operatively linked to the nucleotide sequence encoding the protein or polypeptide of interest.
[0014] Preferably, the first promoter has at least 50% identity over its entire length with nucleotides 1-969 of SEQ ID NO: 1 or nucleotides 1-614 of SEQ ID NO: 2. An overview of all SEQ ID NOs is given in Table 1. Preferably, the nucleotide sequence (comprising the first promoter and the first intron sequence (which contains at least a donor splice site)) has at least 50% identity over its entire length with SEQ ID NO: 1 or 2. Preferably, the second promoter has at least 50% sequence identity over its entire length with SEQ ID NO: 57 or SEQ ID NO: 58.
[0015] The present invention further relates to a nucleic acid construct comprising a first promoter and a second promoter, wherein the first and second promoters are configured to be operatively linked to an optional nucleotide sequence of interest, and wherein the second promoter is flanked by a first intron sequence upstream of the promoter and a second intron sequence downstream of the promoter. Preferably, the first intron sequence comprises at least a donor splice site and optionally the second intron sequence comprises at least a recipient splice site. Furthermore, preferably, the nucleic acid construct of the present invention comprises the following nucleotide sequences shown herein in the 5' to 3' orientation, in their relative positions: (i) a first promoter; (ii) a first intron sequence comprising at least a donor splice site; (iii) a second promoter; (iv) a second intron sequence comprising at least a recipient splice site; and optionally (v) a nucleotide sequence of interest, wherein preferably the first promoter, the first intron sequence comprising at least a donor splice site, the second promoter, and the second intron sequence comprising at least a recipient splice site are all configured to be operatively linked to the optional nucleotide sequence of interest.
[0016] Preferably, the first promoter has at least 50% identity with nucleotides 1-969 of SEQ ID NO: 1 or nucleotides 1-614 of SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence comprising the first promoter and the first intron sequence (which includes at least a donor splice site) has at least 50% identity with SEQ ID NO: 1 or 2 over its entire length. Preferably, the second promoter has at least 50% sequence identity with SEQ ID NO: 57 or SEQ ID NO: 58 over its entire length.
[0017] Preferably, the nucleic acid construct is an isolated construct. Preferably, the nucleic acid construct is a recombinant nucleic acid construct. Preferably, the optional nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, the protein or polypeptide of interest is a heterologous protein or polypeptide.
[0018] The present invention further relates to expression vectors comprising nucleic acid constructs or recombinant nucleic acid constructs as defined herein.
[0019] The present invention further relates to cells comprising nucleic acid constructs or recombinant nucleic acid constructs as defined herein, and / or expression vectors as defined herein.
[0020] The present invention also relates to the use of nucleic acid constructs or recombinant nucleic acid constructs as defined herein and / or expression vectors as defined herein and / or cells as defined herein for transcription of nucleotide sequences of interest.
[0021] The present invention further relates to the use of nucleic acid constructs or recombinant nucleic acid constructs as defined herein and / or expression vectors as defined herein and / or cells as defined herein for the expression of proteins or peptides of interest.
[0022] The present invention further relates to a method for transcription and optionally purification of the resulting transcripts, comprising the following steps:
[0023] a) Provides a nucleic acid construct of the present invention comprising an expression enhancing element, a heterologous promoter, and a nucleotide sequence of interest, wherein the expression enhancing element and the heterologous promoter are operatively linked to the nucleotide sequence of interest; and,
[0024] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0025] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally,
[0026] d) Purify the resulting transcript.
[0027] The present invention further relates to a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0028] a) Provide a nucleic acid construct comprising an expression-enhancing element, a heterologous promoter, and a nucleotide sequence encoding a protein or polypeptide of interest, wherein the expression-enhancing element and the heterologous promoter are operatively linked to the nucleotide sequence encoding the protein or polypeptide of interest; and,
[0029] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0030] c) Inducing the transformed cells to express a protein or polypeptide of interest; and optionally,
[0031] d) Purify the protein or polypeptide of interest.
[0032] Preferably, the nucleic acid construct of the method for transcription and / or expression, and optionally purification of transcripts and / or proteins or peptides of interest, further comprises additional expression regulatory elements operatively linked to the nucleotide sequence of interest and / or the nucleotide sequence encoding the protein or peptide of interest. Preferably, the additional expression regulatory element comprises an intron sequence. Preferably, the additional expression regulatory element comprises or is further an expression enhancing element. More preferably, the additional expression regulatory element further comprises a translation enhancing element.
[0033] The present invention further relates to nucleic acid molecules represented by nucleotide sequences containing the expression-enhancing elements of the present invention, i.e., nucleotide sequences having at least 50% identity with SEQ ID NO: 1 or 2 over their entire length. A summary of all SEQ ID NOs is given in Table 1. Preferably, the nucleic acid molecule is an isolated nucleic acid molecule. Preferably, the nucleic acid molecule or isolated nucleic acid molecule is represented by a nucleotide sequence, wherein the nucleotide sequence has at least 50% sequence identity with SEQ ID NO: 1 or 2 over its entire length. Preferably, the nucleic acid molecule or isolated nucleic acid molecule is represented by a nucleotide sequence, wherein the nucleotide sequence comprises a sequence derived from the Chinese hamster (Cricetulus griseus) gene for up to 8000 nucleotides of polyubiquitin (polyubiquitin protein). The present invention further relates to nucleic acid constructs comprising nucleic acid molecules of the present invention. Preferably, the nucleic acid construct is represented by a nucleotide sequence, wherein the nucleotide sequence further comprises a heterologous promoter, and preferably the expression-enhancing element and the heterologous promoter are configured to be operatively linked to an optional nucleotide sequence of interest. Preferably, the nucleic acid construct further comprises additional expression-regulating elements, wherein preferably the expression-enhancing element, the heterologous promoter, and the additional expression-regulating elements are configured to be operatively linked to the optional nucleotide sequence of interest. Preferably, the additional expression-regulating elements further comprise translation-enhancing elements. Preferably, the additional expression-regulating elements comprise intron sequences. Preferably, the optional nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, the protein or polypeptide of interest is a heterologous protein or polypeptide.
[0034] Preferably, the nucleic acid construct is a recombinant and / or isolated nucleic acid construct.
[0035] The present invention further relates to expression vectors comprising nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinant and / or isolated nucleic acid constructs as defined herein.
[0036] The present invention further relates to cells comprising nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinant nucleic acid constructs as defined herein and / or expression vectors as defined herein.
[0037] The present invention also relates to the use of nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinants as defined herein and / or isolated nucleic acid constructs and / or expression vectors and / or cells as defined herein for transcription of nucleotide sequences of interest.
[0038] The present invention further relates to the use of nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinants as defined herein and / or isolated nucleic acid constructs and / or expression vectors as defined herein and / or cells as defined herein for the expression of proteins or peptides of interest.
[0039] The present invention further relates to a method for transcription and optionally purification of the resulting transcripts, comprising the following steps:
[0040] a) Provide a nucleic acid construct comprising a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 over its entire length, and operably ligated to a nucleotide sequence of interest; and
[0041] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0042] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally,
[0043] d) Purify the resulting transcript.
[0044] The present invention further relates to a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0045] a) Provide a nucleic acid construct comprising a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 over its entire length, and being operatively ligated to a nucleotide sequence of interest and allowing cells to be contacted with the nucleic acid construct to obtain transformed cells; and
[0046] b) To induce the transformed cells to express the protein or polypeptide of interest; and optionally,
[0047] c) Purify the protein or polypeptide of interest.
[0048] Preferably, the nucleic acid construct of the method for transcription and / or expression, and optionally purification of transcripts and / or proteins or peptides of interest, further comprises an additional expression regulatory element operatively linked to the nucleotide sequence of interest and / or the nucleotide sequence encoding the protein or peptide of interest. Preferably, the additional expression regulatory element comprises an intron sequence. Preferably, the additional expression regulatory element further comprises a translation enhancement element.
[0049] The present invention further relates to nucleic acid molecules represented by nucleotide sequences, wherein the nucleotide sequences have at least 50% identity with SEQ ID NO: 88 over their entire length. Preferably, the nucleic acid molecule is an isolated nucleic acid molecule. The present invention further relates to nucleic acid constructs comprising the nucleic acid molecules of the present invention. Preferably, the nucleic acid construct is represented by a nucleotide sequence, which further comprises an optional nucleotide sequence of interest. Preferably, the nucleic acid construct further comprises additional expression regulatory elements, wherein preferably the expression enhancing elements are configured to be operatively linked to the optional nucleotide sequence of interest. Preferably, the additional expression regulatory elements further comprise translation enhancing elements. Preferably, the additional expression regulatory elements comprise intron sequences. Preferably, the optional nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, the protein or polypeptide of interest is a heterologous protein or polypeptide.
[0050] Preferably, the nucleic acid construct is a recombinant and / or isolated nucleic acid construct.
[0051] The present invention further relates to expression vectors comprising nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinant and / or isolated nucleic acid constructs as defined herein.
[0052] The present invention further relates to cells comprising nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinant nucleic acid constructs as defined herein and / or expression vectors as defined herein.
[0053] The present invention also relates to the use of nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinants as defined herein and / or isolated nucleic acid constructs and / or expression vectors and / or cells as defined herein for transcription of nucleotide sequences of interest.
[0054] The present invention further relates to the use of nucleic acid molecules or isolated nucleic acid molecules as defined herein and / or nucleic acid constructs or recombinant and / or isolated nucleic acid constructs as defined herein, and / or expression vectors as defined herein and / or cells as defined herein for the expression of proteins or peptides of interest.
[0055] Description of the present invention
[0056] The inventors have identified an expression construct for increasing the expression of a protein or peptide of interest. The expression construct of the present invention is characterized by two promoters operatively linked to the coding sequence of the protein or peptide of interest. The expression construct of the present invention typically comprises a first promoter, followed by a second promoter, the coding sequence of the protein or peptide of interest, and a polyadenylation sequence, wherein the second promoter is flanked by intron sequences. The promoter flanked by intron sequences is referred to herein as an intronic promoter. Additional expression regulatory sequences may be inserted upstream and downstream of the first and / or second promoters and / or downstream of the polyadenylation sequence. The inventors have surprisingly found that, compared to an expression construct comprising only one promoter operatively linked to the coding sequence, the expression construct of the present invention, comprising a promoter followed by an intronic promoter operatively linked to the coding sequence of the protein or peptide of interest, results in a significant increase in the expression of the protein or peptide of interest. The inventors have discovered that when the combination of the promoter and intron promoters of the present invention is used instead of a single promoter in expression constructs encoding these proteins, the expression of initially poorly expressed proteins increases to a significant level, as illustrated in the examples (more specifically in Example 1). In expression constructs for initially poorly expressed proteins, the combination of the promoter and intron promoters of the present invention promotes the generation of clonal lines and results in the generation of clonal lines with increased and associated expression levels, as illustrated in the examples (more specifically in Example 3). Furthermore, when the combination of the promoter and intron promoters of the present invention is used instead of a single promoter in expression constructs encoding these proteins, the expression of initially highly expressed proteins is further increased, as illustrated in the examples (more specifically in Example 5). In addition, as detailed below and illustrated in the appended examples, an increase in total mRNA and an increase in expression measured at protein levels have been observed. Furthermore, the percentage of high-production cell lines is significantly higher in stable transfection pools compared to pools with individual promoters operatively linked to coding sequences. Because the nucleotide sequences of this invention, comprising promoters and intron promoters (operatively linked to the nucleotide sequence of interest), result in increased transcription, this invention is not limited to the use of this sequence for protein and / or peptide expression and / or protein and / or peptide production, but extends to this combination of promoters and intron promoters for use in methods where higher levels of transcripts are required, such as methods for producing non-coding RNA transcripts, as further specified herein.Furthermore, as described in further detail herein, a further advantage of the present invention is that, in addition to increased transcriptional and / or expression levels of the protein or polypeptide of interest, the present invention also allows for the formation of different transcripts.
[0057] The inventors have identified expression-enhancing elements for increasing the expression of proteins or peptides of interest. This invention relates to such expression-enhancing elements. Compared to expressions utilizing similar expression constructs that differ only in that the expression-enhancing elements of this invention are absent, the application of the expression-enhancing elements of this invention in an expression construct (further comprising a heterologous promoter operatively linked to a sequence encoding a protein or peptide of interest) results in a significant increase in the expression of said protein or peptide of interest. The inventors have found that, after inserting the element into an expression construct encoding these proteins, the expression of initially poorly expressed proteins is increased to a significant level, as illustrated in the examples (more specifically in Example 1). In expression constructs for initially poorly expressed proteins, the insertion of the expression-enhancing element promotes the generation of clone lines and results in the generation of clone lines with relevant expression levels, as illustrated in the examples (more specifically in Example 3). Furthermore, after the insertion of the element into an expression construct encoding these proteins, the expression of initially highly expressed proteins is even further increased, as illustrated in the examples (more specifically in Example 5). Furthermore, as detailed below and illustrated by example in the appended examples, an increase in the total amount of mRNA and / or an increase in the expression measured at the protein level was observed. Because the expression-enhancing element of the present invention can lead to an increase in transcription, the present invention is not limited to the use of this element for protein and / or peptide expression and / or protein and / or peptide production, but extends to the use of this element in methods in which higher levels of transcripts are required, such as in methods for producing non-coding RNA transcripts, as further indicated herein.
[0058] The inventors have further identified nucleic acid molecules represented by nucleotide sequences, wherein the nucleotide sequences have at least 50% identity with SEQ ID NO: 88, for increasing the expression of proteins or peptides of interest. As illustrated in Example 11, the use of these nucleotide sequences is attractive.
[0059] First aspect
[0060] In a first aspect, the present invention provides a nucleic acid construct for increasing transcription and / or expression, comprising a first promoter and a second promoter configured to be operably linked to an optional nucleotide sequence of interest within the expression construct. "Optional" is understood herein to mean that it is not necessarily present in the expression construct. For example, such a nucleotide sequence of interest need not be present in a commercially available expression vector, but may be readily introduced by those skilled in the art prior to its use in the methods of the present invention.
[0061] Preferably, within this first aspect, the nucleotide construct of the present invention, comprising a first promoter and a second promoter, is capable of increasing transcription of the nucleotide sequence of interest under the control of the first and second promoters. Alternatively, or in conjunction with the increased transcription, the nucleotide construct is also preferably capable of increasing the expression of the protein or polypeptide of interest encoded by the nucleotide sequence of interest. Preferably, the transcriptional level is assessed in an expression system using the expression construct (which comprises the first and second promoters operatively linked to the nucleotide sequence of interest) and using a suitable assay, such as RT-qPCR. Preferably, within this first aspect, compared to transcription using a construct (the only difference being that the nucleotide sequence of interest is under the control of a separate promoter), when tested in systems exemplified in the embodiments included herein, the nucleotide construct of the present invention (which comprises the first and second promoters of the present invention) increases the transcription of the nucleotide sequence of interest by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which comprises a second promoter sequence operatively linked to the first promoter and the sequence to be tested). Transcription is preferably measured using RT-qPCR, and the transcription level is relative to the nucleotide sequence of interest, measured under the same conditions except that the first and second promoters are replaced by a separate promoter, preferably the CMV promoter represented by SEQ ID NO: 57, in the expression vector used.
[0062] Preferably, within this first aspect, in the expression system, expression levels are established using an expression construct comprising a first promoter and a second promoter operatively linked to a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and its expression is detected by suitable assays such as enzyme-linked immunosorbent assay (ELISA), Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Preferably, compared to the expression of the protein or polypeptide using the construct (the only difference being that the coding sequence of the protein or polypeptide of interest is under the control of a separate promoter, preferably when tested in a system as exemplified in the embodiments included herein), the first and second promoters of the present invention increase the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a first promoter and a second promoter sequence operatively linked to the nucleotide sequence of interest to be tested) Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, which is measured under the same conditions, except that the expression vector differs only in that, in the expression vector used, the first and second promoters are replaced by a separate promoter, preferably the CMV promoter represented by SEQ ID NO: 57.
[0063] Preferably, in the first aspect, the increase in protein or polypeptide expression is copy number independent, as established by a test suitable for determining copy number dependence performed by a technician (e.g., but not limited to the triple TaqMan assay further detailed in Example 4 of the invention).
[0064] Preferably, in the first aspect, the first promoter is located upstream of the second promoter or at the 5' site of the second promoter. Preferably, as defined herein, the second promoter lacks a sequence element that serves as a transcription terminator. Transcription terminators, well known to those skilled in the art, are sequences that cause premature termination of transcription, such as, but not limited to, stable hairpin structures, repeating sequences such as long terminal repeats (LTRs) or Alu repeats, polyadenylated motifs, and transposable elements.
[0065] In the context of the first aspect of the invention, a promoter is a promoter capable of initiating transcription in a selected host cell. As used herein, promoters include tissue-specific promoters, tissue-preferred promoters, cell-type-specific promoters, inducible promoters, and constitutive promoters as defined herein in the definition section. Promoters that may be included within the first or second promoter as defined herein are promoters that can be used for transcription of nucleotide sequences of interest and / or expression of proteins or peptides of interest (preferably in mammalian cells), and include, but are not limited to, human or mouse cytomegalovirus (CMV) promoters, simian virus (SV40) promoters, human or mouse ubiquitin C (UBC) promoters, human or mouse or rat elongation factor α (EF1-a) promoters, mouse or hamster β-actin promoters, or hamster rpS21 promoters. The Tet-Off and Tet-On effector elements upstream of minimal promoters, such as CMV promoters, are examples of inducible mammalian promoters. Examples of suitable yeast and fungal promoters include the Leu2 promoter, galactose (Gal1 or Gal7) promoter, alcohol dehydrogenase I (ADH1) promoter, glucosyl amylase (Gla) promoter, triose phosphate isomerase (TPI) promoter, translation elongation factor EF-Iα (TEF2) promoter, glyceraldehyde-3-phosphate dehydrogenase (gpdA) promoter, alcohol oxidase (AOX1) promoter, or glutamate dehydrogenase (gdhA) promoter. An example of a strong, ubiquitous promoter used for expression in plants is the cauliflower mosaic virus (CaMV) 35S promoter.
[0066] In embodiments within the first aspect, the first and second promoters are similar promoters. Preferably, the first promoter has at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to the second promoter.
[0067] In another embodiment of the first aspect, the first promoter and the second promoter are different promoters. Preferably, the first promoter has less than 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% sequence identity relative to the second promoter.
[0068] In a preferred embodiment of the first aspect, the first promoter sequence comprises or consists of the UBC promoter or the CCT8 promoter, and the second promoter comprises or consists of the CMV promoter, or otherwise. Preferably, the first promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-969 of SEQ ID NO: 1 or nucleotides 1-614 of SEQ ID NO: 2. Preferably, the second promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57.
[0069] In the first aspect, preferred are the nucleotide sequences of the present invention, wherein the first promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotide 1-969 of SEQ ID NO: 1, and the second promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57.
[0070] In the first aspect, the nucleotide sequence of the present invention is also preferred, wherein the first promoter comprises or is composed of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-614 of SEQ ID NO: 2, and the second promoter comprises or is composed of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57.
[0071] In the first aspect, a preferred nucleotide sequence of the invention is also provided, wherein the first promoter comprises or is composed of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 58 or preferably SEQ ID NO: 57, and wherein the second promoter comprises or is composed of the following sequence, wherein the sequence has nucleotides 1-969 of SEQ ID NO 1 or has nucleotides of SEQ ID NO 1. The nucleotides 1-614 of 2 have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
[0072] Preferably, within the nucleotide sequence of the first aspect of the present invention, the first promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57, and the second promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-969 of SEQ ID NO: 1. Also preferred are the nucleotide sequences of the present invention, wherein the first promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57, and the second promoter comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-614 of SEQ ID NO: 2.
[0073] It should be understood that, within the first aspect, the first and / or second promoters are not solely composed of promoter enhancer sequences selected from the group consisting of SEQ ID NO: 52-54. Preferably, the first and second promoters are not solely composed of promoter enhancer sequences selected from the group consisting of SEQ ID NO: 52-54. Preferably, the nucleotide sequences of the present invention do not contain or are not composed of SEQ ID NO: 55 or 56.
[0074] In a preferred embodiment of the first aspect, the second promoter is flanked by a first intron sequence at the 5' site or upstream of the second promoter and a second intron sequence at the 3' site or downstream of the second promoter. "Flanked" herein is understood to mean located between the specified sequences, which may optionally be separated by 1-50, 1-60, 1-70, 1-80, 1-90, 1-100, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-800, 1-900, 1-1000, 1-5,000, or 1-100,000 nucleotides, which are understood to include the 5'-UTR. An intron sequence is understood to be at least a portion of the nucleotide sequence of an intron. Preferably, the first intron sequence contains at least a donor splice site or splice site GT at the 5' site or upstream of the second promoter. The donor splice site is understood herein as a splice site, as defined herein, which, when bound to the recipient splice site, results in the formation of an intron (as defined in the definition section). Preferably, if, by RNA splicing, using a suitable assay for detecting intron splicing, such as, but not limited to, reverse transcriptase polymerase chain reaction (RT-PCR), followed by RT-PCR size or sequence analysis, at least 2%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the initial RNA loses this sequence, then the nucleotide sequence is an intron. Preferred donor splicing sites of the present invention are MWG-[cut]-GTRAGK or MAR-[cut]-GTRAGK (in the case of mammalian cells), AG-[cut]-GTAWK (in the case of plant cells), [cut]-GTAWGTT (in the case of yeast cells), and RG-[cut]-GTRAG (in the case of insect cells). “[cut]” will be understood herein as a specific cleavage site where splicing will occur. Intron splicing can be functionally evaluated using assays as described in detail under “Introns” in the definition section. Most preferably, the donor splicing site contained within the first intron sequence of the present invention is CTG-[cut]-GTGAGG or AAA-[cut]-GTGAGG. Preferably, the first intron sequence consists of the donor splicing site or splicing site GT. Preferably, as defined herein, the first intron sequence contains a single donor splicing site. Preferably, as defined below, the first intron sequence does not contain a receptor splicing site.
[0075] Preferably, in the first aspect, the second intron sequence contains at least one receptor splice site at the 3' position or downstream of the promoter, which is understood herein as a splice site AG, preferably preceded by a pyrimidine-rich sequence or a polypyrimidine fragment nucleotide sequence, and optionally separated from the splice site AG by 1-50 nucleotides, and optionally further includes a branch site containing the sequence YTNAY at the 5' position of the polypyrimidine fragment nucleotide sequence, wherein the branch site may have the nucleotide sequence CYGAC. A receptor splice site is understood herein as a splice site that, when bound to a donor splice site contained within the construct, results in the formation of an intron as defined in the defining portion. Preferably, the receptor splice site or splice site AG has the sequence [Y-enriched]-NYAG-[cleavage]. Preferably, the receptor splice site or splice site AG has the sequence [Y-enrichment]-NYAG-[cleave]-R (in the case of mammalian cells), [Y-enrichment]-DYAG-[cleave]-R or [Y-enrichment]-DYAG-[cleave]-RW (in the case of plant cells), [Y-enrichment]-AYAG-[cleave] (in the case of yeast cells), and [Y-enrichment]-NYAG-[cleave] (in the case of insect cells). “[Y-enrichment]” will be understood herein as a polypyrimidine segment, preferably defined as a continuous sequence of at least 10 nucleotides containing at least 6, 7, 8, 9, or preferably 10 pyrimidine nucleotides. Preferably, the receptor splice site or splice site GT has the sequence YAG-[cleave]-R. Preferably, the second intron sequence contains a single receptor splice site. In embodiments, the second intron sequence does not contain a donor splice site (as defined herein). In alternative embodiments, the second intron sequence includes both a donor splicing site and a recipient splicing site as defined herein. Most preferably, the second intron sequence is an intron as defined in the definition section. Preferably, the intron sequences of the second promoter and flanking second promoters are configured to form an intron promoter (see reference). Figure 1For those skilled in the art, an intron promoter is referred to as a promoter located within an intron sequence. Preferably, the intron promoter is an intron as defined in the definition section. Preferably, the boundary of the intron promoter of the present invention is formed by a donor splicing site of the intron sequence at the 5' site or upstream of the second promoter of the present invention and a acceptor splicing site of the intron sequence at the 3' site or downstream of the second promoter of the present invention. The intron promoter of the present invention may have a length equivalent to or similar to a naturally occurring intron, preferably equivalent to or similar to a naturally occurring intron in a host cell or organism as defined herein. Preferably, as defined herein, the length of the intron promoter is at most 12,000 nucleotides. Preferably, the first intron sequence at the 5' site or upstream of the second promoter is located at the 3' site or downstream of the first promoter. Preferably, the first promoter and the second promoter, the intron sequences flanking the second promoter, and the nucleotide sequence encoding the protein or polypeptide of interest are configured such that the first promoter is upstream of the second promoter, wherein the intron sequences flanking the second promoter form an intron promoter, and wherein the first promoter and the second promoter are both configured to be upstream and operatively linked to the nucleotide sequence encoding the protein or polypeptide of interest. Figure 1 Intron promoters may contain additional expression-enhancing elements, but preferably, within the first and second intron sequences as defined herein, in addition to donor and acceptor splicing sites as defined herein, the intron promoter does not contain further splicing sites. Preferably, one or more expression-enhancing sequences are contained within the first and / or second promoter. Without wishing to be bound by any theory, transcription starting from either of the two promoters can result in different transcripts (pre-mRNA), which, after splicing, result in different mRNAs, such as those in... Figure 1 As shown in the figure. To support this theory, the inventors have discovered that different transcripts can be formed using the constructs of this invention (for this purpose, see reference). Figure 1 (See Example 8 and Figure 10). Furthermore, a four-nucleotide mutation in the intron promoter (which prevents proper intron splicing) severely weakens the increased activity (for this purpose, refer to...). Figure 9(And Example 7), which also supports the theory that both promoters are active in the construct. Therefore, a further advantage of the invention is that, in addition to increasing transcription of the nucleotide sequence of interest and / or increasing the expression level of the protein or polypeptide of interest, the invention also forms distinct transcripts. “Distinct transcripts” herein refers to transcripts that are structurally different, i.e., have different nucleotide sequences. Therefore, a further advantage of the invention is the ability to direct or redirect the splicing of the nucleotide sequence of interest. Depending on the location of the intron splicing site, the transcripts can have different UTR sequences and / or different coding sequences. It is also possible to form only one type of transcript, for example, when the 5'-UTR sequences of the first and second intron sequences are identical. The formation of distinct transcripts can be assessed using any suitable method known to those skilled in the art, such as, but not limited to, rapid amplification of cDNA terminal polymerase chain reaction (RACE-PCR).
[0076] Preferably, in the first aspect, the first intron sequence has at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with nucleotides 970-1449 of SEQ ID NO: 1 at the 5' site or upstream of the second promoter, or has at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with nucleotides 667-1228 of SEQ ID NO: 2, preferably containing at least a donor splice site or a splice site GT.
[0077] Preferably, in the first aspect, the intron sequence downstream of the second promoter or at the 3' site comprises the following nucleotide sequence, which has the same nucleotides as SEQ ID NO: 14 nucleotides 171-277, SEQ ID NO: 19 nucleotides 171-274, SEQ ID NO: 20 nucleotides 133-210, SEQ ID NO: 21 nucleotides 134-211, SEQ ID NO: 22 nucleotides 134-226, SEQ ID NO: 23 nucleotides 134-226, SEQ ID NO: 24 nucleotides 133-225, SEQ ID NO: 25 nucleotides 134-226, SEQ ID NO: 26 nucleotides 146-257 or SEQ ID NO: 14. The nucleotides 147-223 of NO:27 have at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and preferably contain at least a receptor splice site AG (preceded by a TC-rich nucleotide sequence), optionally separated from the splice site AG by 1-50 nucleotides and a branching site, which contains the sequence YTNAY or CYGAC at the 5' position of the TC-rich nucleotide sequence.
[0078] In a preferred embodiment of the first aspect, the first promoter is side-mounted with the first intron sequence at its 3' site. In an embodiment, the first promoter and the first intron sequence are not aligned in nature, but are aligned in the construct of the invention through recombination. In another embodiment, the sequence (which includes the first promoter with the first intron sequence side-mounted at its 3' site) is a naturally occurring sequence. In a preferred embodiment, the nucleotide sequence comprising the first promoter and the first intron sequence according to the invention is a sequence derived from the UBC ubiquitin gene. Preferably, the sequence is derived from the mammalian UBC ubiquitin gene. More preferably, the nucleotide sequence comprising the first promoter and the first intron sequence according to the invention is a Chinese hamster homologous gene derived from the human UBC ubiquitin gene, the gene being represented as a hamster species gene for polyubiquitin, or CRUPUQ (GenBank D63782). In a preferred embodiment, the nucleotide sequence derived from CRUPUQ comprises the first promoter and first intron sequence of the present invention, and is a contiguous sequence of at least 500, 600, 700, 800, 900, 1000 or 1117 nucleotides in length, preferably at least 1449 nucleotides in length of SEQ ID NO: 1. Preferably, the nucleotide sequence having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1 has a length of at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides. Most preferably, the length of the sequence is 1449 nucleotides. Preferably, the length of the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides, and it has at least 65% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the length of the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides, and it has at least 70% identity with SEQ ID NO: 1 throughout its entire length.Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 75% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 80% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 85% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 90% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 95% identity with SEQ ID NO: 1 over its entire length. Also preferred is a sequence of at most 8000 nucleotides having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1.
[0079] In a further preferred embodiment of the first aspect, the nucleotide sequence comprising the first promoter and first intron sequence according to the invention is a sequence derived from the CCT8 gene. Preferably, the sequence is derived from the mammalian CCT8 gene. More preferably, the nucleotide sequence comprising the first promoter and first intron sequence according to the invention is derived from the human or Homo sapiens CCT8 gene. In a preferred embodiment, the nucleotide sequence derived from the CCT8 gene comprising the first promoter and first intron sequence according to the invention is a contiguous sequence of at least 500, 600, 700, 791, or 1223 nucleotides in length of SEQ ID NO: 2, preferably at least 1228 nucleotides in length. Preferably, the nucleotide sequence having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 2 is at most 8000 nucleotides in length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length. Most preferably, the sequence is 1228 nucleotides in length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 65% identity with SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 70% identity with SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 75% identity with SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 80% identity with SEQ ID NO: 2 over its entire length.Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 85% identity with SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 90% identity with SEQ ID NO: 2 over its entire length. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 95% identity with SEQ ID NO: 2 over its entire length. Also preferred are sequences of up to 8000 nucleotides having at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 2 over its entire length.
[0080] Preferably, within the first aspect, the nucleotide sequence comprising a first promoter and a first intron sequence as defined herein has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity over its entire length with any of SEQ ID NOs: 1, 2, and 59-61. Preferably, the nucleotide sequence comprising a first promoter and a first intron sequence as defined herein comprises or consists of any of the following sequences selected from the group consisting of SEQ ID NOs: 1, 2, and 59-61. Most preferably, the nucleotide sequence comprising a first promoter and a first intron sequence as defined herein comprises or consists of any of the following sequences selected from the group consisting of SEQ ID NOs: 1, 2, and 59.
[0081] Preferably, the nucleotide construct of the first aspect further comprises one or more additional expression regulatory sequences, wherein preferably the first promoter, the intron sequence, and optionally the additional expression regulatory sequences as defined herein are all configured to be operatively linked to an optional nucleotide sequence of interest. “Additional expression regulatory sequence” will be understood herein as a sequence or element other than the first and / or second promoter and / or the first and / or second intron sequences as defined herein, and may be additional expression-enhancing sequences and / or different expression-enhancing sequences. Additional expression regulatory sequences included in this invention may be, but are not limited to, transcriptional and / or translational regulation of genes, including but not limited to 5′-UTR, 3′-UTR, enhancers, promoters, introns, polyadenylation signals, and chromatin control elements such as S / MAR (scaffold / matrix attachment region), ubiquitous chromatin opening element, cytosine phosphate diester guanine islands, and STAR (stabilizing and anti-repressor elements) and any derivatives thereof. Other optional regulatory sequences that may be present in the nucleic acid constructs of the present invention include, but are not limited to, coding nucleotide sequences of homologous and / or heterologous nucleotide sequences, including iron response elements (IREs), translation cis-regulatory elements (TLREs), or uORFs in the 5' UTR and poly(U) segments in the 3' UTR. Such one or more additional expression regulators, preferably enhancing elements, may be located anywhere in the construct, preferably directly aligned or contained within the first and / or second promoters.
[0082] A further preferred regulatory sequence within the first aspect comprises or is composed of a translation-enhancing element. Preferably, the translation-enhancing element increases the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% compared to the expression of the protein or polypeptide using the construct (where the construct differs only in that it does not contain the translation-enhancing element), preferably when tested in the system exemplified in the embodiments included herein. More specifically, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is preferably measured in mammalian cell systems, most preferably in CHO cells, using a pcDNA3.1 expression vector (which contains the translation-enhancing element to be tested and a CMV promoter represented by SEQ ID NO: 57 (operably linked to the nucleotide sequence of interest)). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the translation-enhancing element to be tested.
[0083] Preferably, in the first aspect, the translation enhancement element comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity over its entire length with any one of SEQ ID NO: 3-51, or the translation enhancement element comprises or consists of the following nucleotide sequences:
[0084] i) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence (containing at least 8 consecutive C or T nucleotides), at least 3 A-rich nucleotide sequences (containing at least 5 consecutive A nucleotides), or a GT-rich nucleotide sequence (containing at least 10 nucleotides, of which at least 80% are G or T nucleotides).
[0085] ii) A TC-rich nucleotide sequence (containing at least 8 consecutive C or T nucleotides), at least 3 A-rich nucleotide sequences (containing at least 5 consecutive A nucleotides), and a GT-rich nucleotide sequence (containing at least 10 nucleotides, of which at least 80% are G or T nucleotides), wherein the first nucleotide sequence does not contain a GAA repeating nucleotide sequence; or
[0086] iii) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence (containing at least 8 consecutive C or T nucleotides), at least 3 A-rich nucleotide sequences (containing at least 5 consecutive A nucleotides), and a GT-rich nucleotide sequence (containing at least 10 nucleotides, of which at least 80% are G or T nucleotides), wherein the GAA repeating nucleotide sequence is located at the 3' position of any one or more of the TC-rich nucleotide sequence, the A-rich nucleotide sequence, and / or the GT-rich nucleotide sequence.
[0087] GAA repeat nucleotide sequences are defined herein as containing at least three GAA repeat sequences. GAA repeat nucleotide sequences may contain incomplete GAA repeat sequences. A GAA repeat nucleotide sequence has at least 50% sequence identity, at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, or 100% sequence identity with nucleotides 14-50 of SEQ ID NO: 3. An incomplete GAA repeat sequence may contain the nucleotide sequence (GAA)3ATAA(GAA)8.
[0088] TC-rich nucleotide sequences are defined herein as having at least 70%, 80%, 90%, or 100% sequence identity with nucleotides 54-68 of SEQ ID NO: 3.
[0089] A-rich nucleotide sequences are defined herein as having at least 70%, 80%, 90%, or 100% sequence identity with any one of nucleotides 77-87, 93-105, 111-121, 126-132, or 152-169 of SEQ ID NO: 3.
[0090] GT-rich nucleotide sequences are defined herein as having at least 70%, 80%, 90%, or 100% sequence identity with nucleotides 133-148 of SEQ ID NO: 3.
[0091] Preferably, within the first aspect, the translation-enhancing sequence comprises or consists of sequences having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 19. Preferably, the translation-enhancing sequence is located downstream of the second promoter sequence of the invention or at the 3' site and upstream of an optional nucleotide sequence encoding the protein or polypeptide of interest or at the 5' site. Preferably, the translation-enhancing sequence is located downstream of the second promoter sequence of the invention or at the 3' site and upstream of a second intron sequence (as defined herein) or at the 5' site.
[0092] Most preferably, in the first aspect, the nucleic acid construct of the first aspect of the present invention (which comprises a first promoter, a first intron sequence, a second promoter, and a second intron sequence) has at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 73, 74, 75, or 76.
[0093] Preferably, within the first aspect, the nucleic acid construct of the present invention further comprises a nucleotide sequence of interest (operably linked to and / or under the control of the first and second promoters) and optional additional expression regulatory sequences as defined herein. It should be understood that the first and second promoters and the optional additional expression regulatory sequences are all configured to be operably linked to the same single nucleotide sequence of interest. In a preferred embodiment, the nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. The protein or polypeptide of interest may be a homologous protein or polypeptide, but in a preferred embodiment of the invention, the protein or polypeptide of interest is a heterologous protein or polypeptide. The nucleotide sequence encoding a heterologous protein or polypeptide may be wholly or partially derived from any source known in the art, including bacterial or viral genomes or episomes, eukaryotic nuclear or plasmid DNA, cDNA, or chemically synthesized DNA. The nucleotide sequence encoding the protein or polypeptide of interest may constitute a continuous coding region, or it may include one or more introns bound by appropriate splicing sites, and it may further consist of fragments derived from different sources (naturally occurring or synthetic). The nucleotide sequence encoding the protein or polypeptide of interest according to the method of the present invention is preferably a full-length nucleotide sequence, but may also be the functionally active portion or other portion of the full-length nucleotide sequence. When the nucleotide sequence encoding the protein or polypeptide of interest is expressed at a specific location in a cell or tissue, the nucleotide sequence encoding the protein or polypeptide of interest may also contain a signal sequence that guides the protein or polypeptide of interest. Furthermore, the nucleotide sequence encoding the protein or polypeptide of interest may also contain sequences that facilitate protein purification and detection, by means of, for example, Western blotting and ELISA (e.g., c-myc or polyhistidine sequences).
[0094] In the context of this invention, the proteins or peptides of interest may have industrial or pharmaceutical applications. Examples of proteins or peptides with industrial applications include enzymes, such as lipases (e.g., for use in the detergent industry), proteases (particularly for use in the detergent industry, brewing, etc.), cell wall degrading enzymes (e.g., cellulase, pectinase, β-1,3 / 4- and β-1,6-glucanase, rhamnogalacturonase, mannanase, xylanase, amylopectinase, galactanase, esterase, etc., for use in fruit processing, brewing, etc., or for use in animal feed), phytases, phospholipases, glycosidases (e.g., amylase, β-glucosidase, arabinofuranase, rhamnosidase, apigeninase, etc.), and milk enzymes (e.g., rennet). Mammal and preferably human proteins or peptides and / or enzymes with therapeutic, cosmetic, or diagnostic applications include, but are not limited to, insulin, serum albumin (HSA), lactoferrin, hemoglobin α and β, tissue plasminogen activator (tPA), cytopoietin (EPO), tumor necrosis factor (TNF), BMP (bone morphogenetic protein), growth factors (G-CSF, GM-CSF, M-CSF, PDGF, EGF, etc.), and peptide hormones (e.g., calcitonin, somatostatin, growth hormone, follicle-stimulating hormone). The vaccine includes, but is not limited to, FSH, interleukin (IL-x), interferon (IFN-γ), phosphatases, antibodies, and antibody-like proteins, such as, but not limited to, multispecific antibody-like DART (dual-affinity repositioning) and trisomy proteins, as well as antibody fragments such as Fc, Fab, Fab2, Fv, and scFv. It also includes, for example, bacterial and viral antigens used as vaccines, including, for example, heat-labile toxin B subunits, cholera toxin B subunits, envelope surface proteins of Hepatitis B virus, capsid proteins of norovirus, glycoprotein B of human cytomegalovirus, glycoprotein S, interferon, and transmissible gastroenteritis coronavirus receptors, etc. Further includes genes encoding mutants or analogs of said proteins.
[0095] In the context of this invention, in alternative embodiments, the nucleotide sequence of interest is not a protein or polypeptide coding sequence, but may be a functional nucleotide sequence, such as, but not limited to, a sequence encoding non-coding RNA, wherein non-coding RNA is understood to be RNA that does not encode a protein or polypeptide. Preferably, the non-coding RNA is a reference sequence or regulatory molecule that can regulate gene expression or regulate the activity or localization of a protein or polypeptide. For example, the non-coding RNA may be an antisense RNA or miRNA molecule. Because the first and second promoters of this invention are considered to operate at the transcriptional level, i.e., the increase in expression resulting from sequences of this invention containing the first and second promoters (as shown herein) is considered to be an increase in transcription, the constructs of this invention can also be used to generate increased levels of transcripts, as well as transcripts with different sequences. The transcriptional level can be quantified using conventional transcription quantification methods known to those skilled in the art, such as, but not limited to, Northern blotting and RT-qPCR.
[0096] Second aspect
[0097] In a second aspect, the present invention provides an expression vector comprising a nucleic acid construct according to the first aspect of the invention. The nucleic acid construct according to the invention is preferably a vector, particularly a plasmid, granulosome, or bacteriophage, or a nucleotide sequence (linear or circular) derived from single-stranded or double-stranded DNA or RNA of any origin, wherein many nucleotide sequences have been ligated or recombined into a unique construct capable of introducing any of the nucleotide sequences (sense or antisense orientations) of the invention into cells. The choice of vector depends on the subsequent recombination method and the host cell used. The vector can be a self-replicating vector or can be replicated together with the chromosome into which it has been integrated. Preferably, the vector contains a selection marker. Useful markers depend on the chosen host cell and are well known to those skilled in the art, and are selected from, but not limited to, selection markers as defined in the third aspect of the invention. A preferred expression vector is the pcDNA3.1 expression vector. Preferred selection markers are neomycin resistance genes, zeocin resistance genes, and blasticidin resistance genes.
[0098] Third aspect
[0099] In a third aspect, as defined herein, the present invention provides cells comprising nucleic acid constructs according to the first aspect of the invention and / or expression vectors according to the second aspect of the invention.
[0100] In the context of this invention, cells can be mammalian cells, including human cells, plant cells, animal cells, insect cells, fungal cells, yeast cells, or bacterial cells. Recombinant host cells (such as mammalian cells, including human, plant, animal, insect, fungal, or bacterial cells, containing one or more copies of the nucleic acid construct according to the invention) are another subject of this invention. A host cell is a cell that contains a nucleic acid construct, such as a vector, and supports the replication and / or expression of the nucleic acid construct. Examples of bacteria are Gram-positive bacteria, such as several species of the genera *Bacillus*, *Streptomyces*, and *Staphylococcus*, or Gram-negative bacteria, such as several species of the genera *Escherichia* and *Pseudomonas*. Fungal cells include yeast cells. Expression in yeast can be achieved by utilizing yeast strains such as *Pichia pastoris*, *Saccharomyces cerevisiae*, and *Hansenula polymorpha*. Other fungal cells of interest include filamentous fungal cells, such as *Aspergillus niger*, *Trichoderma reesei*, etc. In addition, insect cells, such as cells or cell lines from Drosophila melanogaster, Fall Armyworm, and Spodoptera litura, including but not limited to S2, Sf9, Sf21, and High Five cells, can be used as host cells. Alternatively, suitable expression systems may be baculovirus systems or expression systems utilizing mammalian cells, such as CHO, COS, CPK (pig kidney), MDCK, BHK, Sp2 / 0, NSO, and Vero cells. Suitable human cells or human cell lines are astrocytes, adipocytes, chondrocytes, endothelial cells, epithelial cells, fibroblasts, hair, keratinocytes, melanocytes, osteoblasts, skeletal muscle cells, smooth muscle cells, stem cells, synovial cells, or cell lines. Examples of suitable human cell lines also include HEK 293 (human embryonic kidney), HeLa, Per.C6, CAP (a cell line derived from human primary amniotic fluid cells), and melanoma cells. In this embodiment, the human cells are not embryonic stem cells.
[0101] Therefore, another aspect of the invention relates to genetically modified host cells, preferably genetically modified host cells by the methods of the invention, wherein the host cell comprises a nucleic acid construct as defined herein. The host cell is a genetically modified cell. The term "host cell" can be replaced by a modified cell, a transformed cell, or a recombinant cell, or a modified host cell, a transformed host cell, or a recombinant host cell. Suitable bacteria for transformation methods in plants include *Agrobacterium tumefaciens* and *Agrobacterium rhizogenes*.
[0102] As autonomously replicating elements, nucleic acid constructs are preferably stably maintained, or more preferably, integrated into the genome of a host cell, typically at random locations within the host cell's genome, such as through non-homologous recombination. Stably transformed host cells are generated using known methods. The term stable transformation refers to methods that expose cells to the transfer and binding of exogenous DNA to their genome. These methods include, but are not limited to, the transfer of purified DNA via cationic lipid reagents and polyvinylimide (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation of protoplasts, and microinjection or the use of silica fibers to facilitate DNA permeation and transfer into host cells.
[0103] Alternatively, the protein or peptide of interest can be expressed in a host cell, such as a mammalian cell, which depends on transient expression from the vector.
[0104] The nucleic acid constructs according to the invention preferably also include marker genes, which can provide selection or screening capabilities in treated host cells. Selectivity markers are generally preferred for host transformation events but are not applicable to all host cells. The nucleic acid constructs disclosed herein may also include nucleotide sequences encoding marker products. Marker products can be used to determine whether the construct or a portion thereof has been delivered to cells and expressed after delivery. Examples of marker genes include, but are not limited to, the *E. coli* lacZ gene, which encodes β-galactosidase, and genes encoding green fluorescent protein.
[0105] In the context of this invention, examples of suitable selective markers for mammalian cells include, but are not limited to, dihydrofolate reductase (DHFR), glutathione synthase (GS), thymidine kinase, neomycin, neomycin analogue G418, hygromycin, isoprothiolane, bleomycin, and puromycin.
[0106] Other suitable selective markers include, but are not limited to, antibiotic resistance genes, metabolic genes, auxotrophic genes, or herbicide-resistant genes, which, when inserted into cultured host cells, confer the ability of these cells to withstand exposure to antibiotic resistance genes. Metabolic or auxotrophic marker genes enable transformed cells to synthesize essential components (typically amino acids), allowing cells to grow in media lacking these components. Another type of marker gene is one that can be screened by histochemical or biochemical analysis, even if the gene cannot be selected. A suitable marker gene found for such host cell transformation experiments is the luciferase gene. Luciferase catalyzes the oxidation of luciferin, resulting in the production of oxidized luciferin and light. Therefore, the use of the luciferase gene provides a convenient assay for detecting the expression of introduced DNA in host cells through histochemical analysis of the cells. In examples of transformation processes, the nucleotide sequence to be expressed in host cells can be tandemly coupled to the luciferase gene. The tandem construct can be transformed into host cells, and the expression of luciferase in the resulting host cells can be analyzed. The advantage of this marker is the non-destructive nature of the substrate and subsequent assay applications.
[0107] When such selectivity markers are successfully transferred into host cells, the transformed host cells can survive under selective pressure. Two distinct categories of selectivity systems exist. The first category is based on cell metabolism and uses mutant cell lines that lack the ability to grow independently of supplemental medium. Two non-restrictive examples are CHO DHFR cells and mouse LTK cells. These cells lack the ability to grow without the addition of nutrients such as thymidine or hypoxanthine. They cannot survive unless the missing nucleotides are provided in supplemental medium because these cells lack certain genes necessary for the complete nucleotide synthesis pathway. An alternative to supplemental medium is to introduce the complete DHFR or TK gene into cells lacking the corresponding gene, thereby altering their growth requirements. Individual cells not transformed with the DHFR or TK gene will not survive in non-supplemental medium.
[0108] The second category is dominant selection, which refers to selection schemes used in any cell type and that do not require the use of mutant cell lines. These schemes typically use drugs to inhibit the growth of host cells. Those cells with novel genes will express proteins that deliver resistance and survive after undergoing selection. Examples of such dominant selection use the drugs neomycin (Southem P. and Berg, P., J. Molec. Appl. Genet. 1: 327 (1982)), mycophenolic acid (Mulligan, RC and Berg, P. Science 209: 1422 (1980)), or hygromycin (Sugden, B. et al., Mol. Cell. Biol. 5: 410-413 (1985)). The above three examples employ bacterial genes under eukaryotic control to deliver resistance to the appropriate drugs G418 or neomycin (geneticin), xgpt (mycophenolic acid), or hygromycin, respectively. Other drugs include the neomycin analogue G418 and puromycin. Other useful markers are dependent on the selected host cell and are well known to those skilled in the art.
[0109] When transformed host cells are obtained by means of the method according to the invention (see below), host tissues can be regenerated from said transformed cells in a suitable culture medium, which may optionally contain antibiotics or biocides known in the art for the selection of transformed cells.
[0110] Preferably, selection is employed, using selection marker genes present on nucleic acid constructs as defined herein to determine the obtained transformed host tissue.
[0111] Fourth aspect
[0112] In a fourth aspect, the present invention provides a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0113] a) Provides a nucleic acid construct according to a first aspect of the invention, comprising a nucleotide sequence encoding a protein or polypeptide of interest; and
[0114] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0115] c) Inducing the transformed cells to express the protein or polypeptide of interest; and optionally...
[0116] d) Purify the protein or polypeptide of interest.
[0117] The method of this invention can be in vitro or ex vivo. It can be applied to cell culture, organism culture, or tissue culture. Alternatively, expression adjacent to host cells can be performed in a cell-free translation system using RNA derived from the nucleic acid constructs of this invention to generate proteins or peptides of interest. The method of this invention can be carried out on cultured cells.
[0118] According to step b), the technician is able to transform the cells. Transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate DNA permeation and transfer into host cells.
[0119] In step c), the transformed cells are made to express a protein or polypeptide of interest, and optionally, the protein or polypeptide is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to the expression of the protein or polypeptide of interest. Techniques for expressing or overexpressing a protein or polypeptide of interest are well understood by those skilled in the art. The method of the invention also includes methods in which the transformed cells are not required to undergo specific conditions that lead to the expression of the protein or polypeptide of interest, but where the protein or polypeptide of interest is expressed automatically (e.g., constitutively).
[0120] In the context of this invention, the purification step depends on the expressed protein or peptide and the host cell used, but may include the isolation of the protein or peptide. When applied to proteins / peptides, the term "isolation" indicates that the protein or peptide was found under conditions different from its native environment. In a preferred form, the isolated protein or peptide is substantially free of other proteins, particularly other homologous proteins. It is preferred to provide a protein or peptide in a form with a purity greater than 40%, more preferably greater than 60%. Even more preferably, it is preferred to provide a highly purified form of the protein or peptide, i.e., greater than 80% purity, more preferably greater than 95% purity, and even more preferably greater than 99% purity, as determined by SDS-PAGE. If desired, the nucleotide sequence encoding the protein or peptide of interest can be ligated to a heteronucleotide sequence to encode a fusion protein or peptide, thereby facilitating, for example, protein purification and protein detection in Western blotting and ELISA. Suitable hetero sequences include, but are not limited to, nucleotide sequences encoding proteins, such as, for example, glutathione S-transferase, maltose-binding protein, metal-binding polyhistidine, green fluorescent protein, luciferase, and β-galactosidase. Proteins or peptides can also be coupled to non-peptide carriers, tags, or labels, which facilitate the tracking of proteins or peptides in vivo and in vitro, and enable the identification and quantification of protein or peptide bindings to substrates. Such tags, labels, or carriers are well known in the art and include, but are not limited to, biotinylate, radiolabels, and fluorescent labels.
[0121] Preferably, the method of this fourth aspect of the invention increases the expression of the protein or polypeptide of interest. Preferably, in the expression system, an expression construct according to the first aspect of the invention is used to establish the expression level, wherein the expression construct comprises first and second promoters according to the first aspect of the invention, operably linked to a nucleotide sequence encoding the protein or polypeptide of interest. Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by a suitable assay, such as ELISA, Western blotting, or (depending on the identity of the protein or polypeptide of interest) any suitable protein identification and / or quantification assay known to those skilled in the art. Preferably, compared to methods that (distinguishing only from step a) use a construct that does not contain the first and second promoters and is operatively linked to a single promoter, the method of the present invention increases the expression of a protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%, preferably when tested in the systems exemplified in the examples included herein. More specifically, in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (containing the first and second promoter sequences to be tested and operatively linked to the nucleotide sequence of interest) Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that, in the expression vector used, the first and second promoters are replaced by separate promoters, preferably the CMV promoter represented by SEQ ID NO: 57.
[0122] Fifth aspect
[0123] In a fifth aspect, the present invention provides a method for expressing a protein or polypeptide of interest in an organism, comprising the following steps:
[0124] a) Provides a nucleic acid construct according to a first aspect of the invention, comprising a nucleotide sequence encoding a protein or polypeptide of interest; and
[0125] b) Contacting the nucleic acid construct with target cells and / or target tissues of an organism to obtain transformed target cells and / or transformed target tissues, thereby causing the transformed cells to express proteins or peptides of interest; and optionally...
[0126] c) to develop the transformed target cells and / or target tissues into transformed organisms; and, optionally...
[0127] d) To induce the transformed organism to express a protein or polypeptide of interest, for example, by subjecting the transformed organism to conditions that lead to the expression of the protein or polypeptide of interest, and optionally by recovering the protein or polypeptide.
[0128] In the context of this invention, the target cell can be an embryonic target cell, such as an embryonic stem cell, for example, derived from non-human mammals, such as cattle, pigs, etc. Preferably, the target cell is not a human embryonic stem cell. In the case of multicellular fungi, such a target cell can be a fungal cell that can be proliferated into the multicellular fungus. When transformed plant tissues or plant cells (e.g., leaves, stem segments, roots, but also protoplasts or plant cells in suspension culture) are obtained by means of the method according to the invention, the whole plant can be regenerated from the transformed tissues or cells in a suitable culture medium, wherein the culture medium may optionally contain antibiotics or biocides known in the art for selecting transformed cells. The method of the invention can, preferably in mammals, and most preferably in humans, be applied to nucleic acid-based inoculation and / or gene therapy. Covered within the invention are therapeutic methods including the method of this aspect, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. The invention also relates to constructs of the first aspect of the invention for therapeutic purposes, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. Furthermore, the present invention relates to the application of the construct of the first aspect of the invention to the manufacture of pharmaceuticals, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide.
[0129] Furthermore, a portion of this invention pertains to non-human transformed organisms. These organisms are transformed using nucleotide sequences, recombinant nucleic acid constructs, or vectors according to the invention, and are capable of producing polypeptides of interest. This includes non-human transgenic organisms such as transgenic non-human mammals, transgenic plants (including propagation, harvesting, and tissue material of said transgenic plants, including, but not limited to, leaves, roots, stems, and flowers), multicellular fungi, and the like.
[0130] Preferably, the method of this fifth aspect of the invention results in the increased expression of a protein or polypeptide of interest in the organism or at least in one tissue, organelle, or cellular organ of the organism. Preferably, in the expression system, an expression construct according to the first aspect of the invention is used to establish the expression level, wherein the expression construct comprises first and second promoters (operably linked to nucleotide sequences encoding the protein or polypeptide of interest) according to the first aspect of the invention. Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by a suitable assay such as ELISA, Western blotting, or, depending on the identity of the protein or polypeptide of interest, any suitable protein identification and / or quantitative assay known to those skilled in the art. Preferably, compared to the following methods (which differ only in that the construct used in step a) does not contain the first and second promoters and is operatively linked to a single promoter), the method of the present invention, when tested in the systems exemplified in the embodiments included herein, increases the expression of a protein or polypeptide in the organism or at least in one tissue, organelle, or organ of the organism by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a first promoter and a second promoter sequence to be tested operatively linked to the nucleotide sequence of interest) Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the nucleotide sequence of interest, measured under the same conditions except that, in the expression vector used, the first and second promoters are replaced by a separate promoter, preferably the CMV promoter represented by SEQ ID NO: 57. Preferably, the increase in protein or peptide expression is copy number independent, as established by a person skilled in the art through tests suitable for determining copy number dependence, such as, but not limited to, the triple TaqMan test further described in Example 4 of the invention.
[0131] Sixth aspect
[0132] In a sixth aspect, the present invention provides a method for transcription and, optionally, purification of the resulting transcript, comprising the following steps:
[0133] a) Provides a nucleic acid construct according to a first aspect of the invention, comprising a nucleotide sequence of interest; and
[0134] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0135] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally...
[0136] d) Purify the resulting transcript.
[0137] In a preferred embodiment of the method according to the invention, a nucleic acid construct as defined above is used. The method of the invention can be an in vitro or ex vivo method. The method of the invention can be applied to cell culture, organism culture, or tissue culture. Preferably in mammals, preferably in humans, the method of the invention can be applied to nucleic acid-based immunotherapy and / or gene therapy. Covered within the invention are methods for treatment, which include or constitute the methods of this aspect, wherein the nucleotide sequence of interest encodes a therapeutic transcript. The invention also relates to constructs of the first aspect of the invention for use solely in treatment, wherein the nucleotide sequence of interest encodes a therapeutic transcript. Furthermore, the invention relates to constructs of the first aspect of the invention used in the manufacture of pharmaceutical agents, wherein the nucleotide sequence of interest encodes a therapeutic transcript.
[0138] According to step b), the technician is able to transform the cells. The transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate the permeation and transfer of DNA into the host cells.
[0139] In step c), the transformed cells are made to produce a transcript of the nucleotide sequence of interest, and optionally the produced transcript is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to transcription of the nucleotide sequence of interest. Techniques for transcribing the nucleotide sequence of interest are well understood by those skilled in the art. The method of the invention also includes methods in which the transformed cells are not required to undergo specific conditions that lead to transcription of the nucleotide sequence of interest, but where the nucleotide sequence of interest is transcribed automatically (e.g., constitutively).
[0140] The purification steps depend on the resulting transcript. The term "isolation" indicates that the transcript was found under conditions different from its native environment. In a preferred form, the isolated transcript is substantially free of other cellular components, particularly other homologous cellular components, such as homologous proteins. Transcripts in a pure form greater than 40% are preferably provided, more preferably greater than 60%. Even more preferably, highly purified transcripts are provided, i.e., greater than 80% purity, more preferably greater than 95% purity, and even more preferably greater than 99% purity, as determined by Northern blotting.
[0141] Preferably, the method of this aspect of the invention increases transcription of the nucleotide sequence of interest. Preferably, in an expression system, transcription levels are established using an expression construct according to the first aspect of the invention (which comprises first and second promoters according to the first aspect of the invention operably linked to the nucleotide sequence of interest). Preferably, transcription of the nucleotide sequence of interest is detected by a suitable assay, such as RT-qPCR. Preferably, compared to the method described below (which differs only in that the construct used in step a) does not contain the first and second promoters and is operably linked to separate promoters), the method of the invention increases transcription by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% when tested in a system as exemplified in the embodiments included herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the first and second promoter sequences to be tested and is operatively linked to the nucleotide sequence of interest) . Transcription is preferably measured using RT-qPCR and the transcription level is relative to the nucleotide sequence of interest, measured under the same conditions, except that in the expression vector used, the first and second promoters are replaced by separate promoters, preferably by the CMV promoter represented by SEQ ID NO: 57.
[0142] Seventh aspect
[0143] In a seventh aspect, the present invention provides a method for splicing or redirecting nucleotide sequences of interest and optionally purifying the resulting transcripts, comprising the following steps:
[0144] a) Provides a nucleic acid construct according to a first aspect of the invention, comprising a nucleotide sequence of interest; and
[0145] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0146] c) Causing the transformed cells to produce a transcript containing a nucleotide sequence of interest, thereby resulting in the production of the transcript; and optionally...
[0147] d) Purify the resulting transcript.
[0148] Preferably, in this aspect, the nucleic acid construct used in step a) comprises a nucleotide sequence upstream of or at the 5' site of the second intron sequence of the invention, which differs from the nucleotide sequence upstream of or at the 5' site of the first intron sequence of the invention. Preferably, the nucleotide sequence between the first promoter and the 5' site of the first intron sequence differs from the nucleotide sequence between the second promoter and the 5' site of the second intron sequence by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% in nucleotide sequence. Preferably, the method of this aspect of the invention (where such nucleic acid constructs are used) results in the production of two different transcripts differing in the nucleotide sequence at the 5' site of the transcripts. If the nucleotide sequence of interest is located downstream of the second intron sequence or at the 3' site, the resulting transcript will differ in the sequence upstream of the nucleotide sequence of interest, as can be detected by any suitable test known to those skilled in the art, such as, but not limited to, 5' RACE-PCR.
[0149] In a preferred embodiment of this aspect, the nucleotide sequence of interest is a sequence encoding a protein or polypeptide of interest. Using the constructs of the present invention, the methods of this aspect can be used to generate two proteins or polypeptides with different N-termini. Preferably, the methods of this aspect are used to generate a first protein or polypeptide comprising a first N-terminus and a second protein or polypeptide comprising a second N-terminus, wherein preferably, the first nucleotide sequence encoding the first N-terminus is located directly upstream of the first intron sequence or at the 5' site, and the second nucleotide sequence encoding the second N-terminus is located directly upstream of the second intron sequence or at the 5' site. Preferably, the first nucleotide sequence encoding the first N-terminus is located downstream of the first promoter or at the 3' site and upstream of the first intron sequence or at the 5' site. Preferably, the second nucleotide sequence encoding the second N-terminus is located downstream of the second promoter or at the 3' site and upstream of the second intron sequence or at the 5' site. Preferably, the nucleic acid construct further comprises a nucleotide sequence encoding a C-terminus, which is located downstream of the second intron sequence or at the 3' site. In this embodiment, it is desirable that the second intron sequence is an intron as defined in the definition section. The differences between the N-termini may be limited to signal sequences and result in the same expression of a protein or peptide, although the localization of the protein or peptide may differ. If performed in an expression system as previously defined herein, the method of this embodiment preferably results in the production of two proteins or peptides of interest, wherein the first protein or peptide will contain a first N-terminus linked to the C-terminus and the second protein or peptide will contain a second N-terminus linked to the C-terminus, detectable by any suitable assay known to those skilled in the art, such as, but not limited to, ELISA, Western blotting, or, depending on the identity of the protein or peptide of interest, any suitable protein identification and / or quantification assay known to those skilled in the art. Preferably, the assay used to detect two different or distinct proteins or peptides produced is particularly suitable for distinguishing the different proteins or peptides produced, for example, using a detection antibody that specifically binds to the first or second N-terminus of the produced protein or peptide.
[0150] Eighth aspect
[0151] In an eighth aspect, the present invention provides the application of nucleic acid constructs according to the first aspect of the invention and / or the application of expression vectors according to the second aspect of the invention and / or the application of cells according to the third aspect of the invention for the expression of proteins or peptides of interest.
[0152] Ninth aspect
[0153] In a ninth aspect, the present invention provides nucleic acid constructs according to the first aspect of the invention and / or expression vectors according to the second aspect of the invention and / or cells according to the third aspect of the invention for use as pharmaceutical agents. The invention also relates to treatment methods comprising administering nucleic acid constructs according to the first aspect of the invention and / or expression vectors according to the second aspect of the invention and / or cells according to the third aspect of the invention, wherein preferably, the administration is administered to mammals, more preferably to humans. Preferably, the treatment is nucleic acid-based immunotherapy and / or gene therapy, preferably in mammals, most preferably in humans. Furthermore, the present invention relates to the application of nucleic acid constructs according to the third aspect of the invention and / or the application of expression vectors according to the second aspect of the invention and / or the application of cells according to the third aspect of the invention for the preparation of pharmaceutical agents. Preferably, the pharmaceutical agent is for nucleic acid-based immunotherapy and / or gene therapy, preferably in mammals, most preferably in humans.
[0154] Tenth aspect
[0155] In a tenth aspect, the present invention provides nucleic acid molecules represented by nucleotide sequences (which contain or constitute expression-enhancing elements) for increasing transcription of the nucleotide sequence of interest and / or expression of the protein or polypeptide of interest. Preferably, the expression-enhancing element of the present invention is capable of increasing transcription of the nucleotide sequence of interest and / or expression of the protein or polypeptide of interest. Preferably, in this aspect, the expression-enhancing element of the present invention capable of increasing transcription of the nucleotide sequence of interest and / or expression of the protein or polypeptide of interest is located upstream of the promoter of the nucleotide sequence of interest or at the 5' site.
[0156] Preferably, in this respect, in the expression system, an expression construct (which contains the expression-enhancing element operatively linked to the nucleotide sequence of interest) is used, and transcriptional levels are established using a suitable assay, such as RT-qPCR. Preferably, compared to transcriptional levels using a construct (which differs only in that it does not contain the expression-enhancing element), preferably as illustrated by examples included herein, the expression-enhancing element of the present invention increases transcription by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, and most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains an expression-enhancing element to be tested operably linked to the nucleotide sequence of interest and a CMV promoter represented by SEQ ID NO: 57). Preferably, the transcription level of the nucleotide sequence of interest is measured under the same conditions, except that the expression vector used does not contain the expression-enhancing element to be tested, and transcription is measured using RT-qPCR and transcription levels.
[0157] Preferably, in this respect, the expression level is established in the expression system using an expression construct (which contains the expression-enhancing element operatively linked to a nucleotide sequence encoding a protein or polypeptide of interest). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and its expression is detected by suitable assays, such as enzyme-linked immunosorbent assay (ELISA), Western blotting, or, depending on the identity of the protein or polypeptide of interest, any suitable protein identification and / or quantification assay known to those skilled in the art. Preferably, compared to the expression of the protein or polypeptide using the construct (which differs only in that it does not contain the expression-enhancing element), the expression-enhancing element of the present invention preferably increases the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% when tested in a system as exemplified in the embodiments included herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the expression-enhancing element to be tested operatively linked to the nucleotide sequence of interest and a CMV promoter represented by SEQ ID NO: 57). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the expression-enhancing element to be tested.
[0158] Preferably, in this respect, the nucleic acid molecule is an isolated nucleic acid molecule as defined herein. In a preferred embodiment, the expression-enhancing element is a sequence derived from the UBC ubiquitin gene. Preferably, the expression-enhancing element is derived from the mammalian UBC ubiquitin gene. More preferably, the expression-enhancing element is a Chinese hamster homologous gene derived from the human UBC ubiquitin gene, the gene being represented as a hamster gene for polyubiquitin, or CRUPUQ (GenBank D63782).
[0159] In a preferred embodiment, the expression-enhancing element derived from CRUPUQ comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-969 of SEQ ID NO: 1. Preferably, the sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-969 of SEQ ID NO: 1 is a promoter defined in the definition section. Preferably, the promoter is capable of inducing transcription of a nucleotide sequence of interest and / or expression of a protein or polypeptide of interest encoded by the nucleotide sequence in a host cell as defined below. In a further preferred embodiment, the expression-enhancing element derived from CRUPUQ comprises or is composed of an intron sequence. An intron sequence is understood to be at least a portion of the nucleotide sequence of an intron. Preferably, the intron sequence comprises at least a donor splice site or a splice site GT. A donor splice site is understood herein to be a splice site that, when bound to a receptor splice site as defined herein, results in the formation of an intron as defined in the defining section. Preferably, if, by RNA splicing and using a suitable assay for detecting intron splicing, such as, but not limited to, reverse transcriptase polymerase chain reaction (RT-PCR), followed by RT-PCR size or sequence analysis, at least 2%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the initial RNA loses this sequence, then the nucleotide sequence is an intron. Preferred donor splicing sites of the present invention are MWG-[cleave]-GTRAGK (in the case of mammalian cells), AG-[cleave]-GTAWK (in the case of plant cells), [cleave]-GTAWGTT (in the case of yeast cells), and RG-[cleave]-GTRAG (in the case of insect cells). “[cleave]” is understood herein to refer to the specific cleavage site where splicing will occur. Intron splicing can be functionally assessed using assays as described in detail in the definition section under “Introns”. Most preferably, the donor splicing site included in the expression enhancement element of the present invention is CTG-[cut]-GTGAGG. Preferably, the intron sequence covered within the expression enhancement element consists of the donor splicing site or the splicing site GT. Preferably, the intron sequence covered within the expression enhancement element does not contain the receptor splicing site as defined below.Preferably, the expression-enhancing element comprising the intron sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 970-1449 of SEQ ID NO: 1. In a preferred embodiment, the expression-enhancing element comprises or consists of a promoter and intron sequences as defined herein. Preferably, the expression enhancement element of the present invention has at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1 over its entire length. Preferably, the expression-enhancing element of the present invention, having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1 throughout its entire length, comprises promoter and intron sequences as defined herein. Preferably, the expression-enhancing element for increasing expression is a contiguous sequence of at least 500, 600, 700, 800, 900, 1000, 1100, or 1117 nucleotides in length, preferably at least 1449 nucleotides in length as defined in SEQ ID NO: 1. Preferably, the expression-enhancing element having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1 is at most 8000 nucleotides in length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length. Most preferably, the sequence is 1449 nucleotides in length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 65% identity with SEQ ID NO: 1 over its entire length.Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length, and has at least 70% identity with SEQ ID NO: 1 over its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length, and has at least 75% identity with SEQ ID NO: 1 over its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 80% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 85% identity with SEQ ID NO: 1 throughout its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 90% identity with SEQ ID NO: 1 throughout its length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1449 nucleotides in length and has at least 95% identity with SEQ ID NO: 1 throughout its length. Preferably, the expression-enhancing element comprises or consists of the sequence represented by SEQ ID NO: 1.
[0160] In a further preferred embodiment of this aspect, the expression-enhancing element of the present invention is a sequence derived from the CCT8 gene. Preferably, the element is derived from the mammalian CCT8 gene. More preferably, the expression-enhancing element is derived from the human or Homo sapiens CCT8 gene.
[0161] In a preferred embodiment within the described aspect, the expression-enhancing element derived from the Homo sapiens CCT8 gene comprises or consists of the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-614 of SEQ ID NO: 2. Preferably, the sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 1-614 of SEQ ID NO: 2 is a promoter defined in the definition section. In a further preferred embodiment, the expression-enhancing element derived from the Homo sapiens CCT8 gene comprises or is composed of an intron sequence. Preferably, the intron sequence is an intron sequence previously defined herein that contains at least a donor splice site or splice site GT as previously defined herein. Preferably, the donor splice site has the sequence MAR-[cleave]-GTRAGK, most preferably AAA-[cleave]-GTGAGG. Preferably, the intron sequence encompassed within the expression-enhancing element is composed of the donor splice site or splice site GT. Preferably, the expression-enhancing element comprising the intron sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides 667-1228 of SEQ ID NO: 2.
[0162] In a preferred embodiment within the stated aspects, the expression-enhancing element comprises or is composed of promoter and intron sequences as defined herein. Preferably, the expression-enhancing element of the present invention has at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 2 along its entire length. Preferably, the expression-enhancing element of the present invention, having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 2 throughout its entire length, comprises a promoter and intron sequences as defined herein. Preferably, the expression-enhancing element for increasing expression is a contiguous sequence of at least 500, 600, 700, 791, or 1223 nucleotides in length of SEQ ID NO: 2, preferably at least 1228 nucleotides in length. Preferably, the expression-enhancing element having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 2 is at most 8000 nucleotides in length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length. Most preferably, the sequence is 1228 nucleotides in length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 65% identity with SEQ ID NO: 2 over its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 70% identity with SEQ ID NO: 2 over its entire length.Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 75% identity with SEQ ID NO: 2 over its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 80% identity with SEQ ID NO: 2 over its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 85% identity with SEQ ID NO: 2 throughout its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 90% identity with SEQ ID NO: 2 throughout its entire length. Preferably, the expression-enhancing element is at most 8000, 7000, 6000, 5000, 4000, 3000, 2000, 1900, 1800, 1700, 1600, 1500, or 1228 nucleotides in length and has at least 95% identity with SEQ ID NO: 2 over its entire length. Preferably, the expression-enhancing element comprises or consists of the sequence represented by SEQ ID NO: 2.
[0163] Further preferred are nucleotide sequences comprising expression-enhancing elements as defined herein, having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity throughout their length with any of SEQ ID NO: 59-61. Preferably, the expression-enhancing element comprises or consists of any sequence selected from the group consisting of SEQ ID NO: 59-61. Most preferably, the expression-enhancing element comprises or consists of the sequence represented by SEQ ID NO: 59.
[0164] Eleventh aspect
[0165] In an eleventh aspect, the present invention provides a nucleic acid construct comprising a nucleic acid molecule according to a tenth aspect of the invention. The nucleic acid construct of the present invention comprises an expression-enhancing element according to a tenth aspect of the invention. Preferably, the nucleic acid construct is a recombinant and / or isolated construct as defined herein. Preferably, the nucleic acid construct further comprises a heteropromoter, wherein, as defined below, the expression-enhancing element and the heteropromoter are preferably configured to be operatively linked to an optional nucleotide sequence of interest. “Heteropromoter” will be understood herein as a promoter that is not naturally operatively linked to the expression-enhancing element of the present invention, i.e., the adjacent sequences comprising the expression-enhancing element and the heteropromoter are not naturally occurring neighboring sequences but can be synthesized as recombinant sequences.
[0166] Preferably, in this respect, the heterologous promoter is located within the nucleic acid construct of the invention, downstream of the expression-enhancing element of the invention or at the 3' site. Preferably, the heterologous promoter is located within the construct of the invention, upstream of the nucleotide sequence encoding the protein or polypeptide of interest of the invention or at the 5' site. In embodiments of the invention, a heterologous promoter is a promoter capable of initiating transcription in a selected host cell. Heterologous promoters as used herein include tissue-specific promoters, tissue-preferred promoters, cell-type-specific promoters, inducible promoters, and constitutive promoters as defined herein. Preferably, in mammalian cells, heterologous promoters and / or regulatory sequences that can be used in the expression of the polypeptide according to the invention include, but are not limited to, human or mouse cytomegalovirus (CMV) promoters, simian virus (SV40) promoters, human or mouse ubiquitin C promoters, human or mouse or rat elongation factor α (EF1-a) promoters, mouse or hamster β-actin promoters, or hamster rpS21 promoters. Examples of inducible mammalian promoters formed by the Tet-Off and Tet-On elements upstream of minimal promoters, such as the CMV promoter. Suitable examples of yeast and fungal promoters include the Leu2 promoter, the galactose (Gal1 or Gal7) promoter, the alcohol dehydrogenase I (ADH1) promoter, the glucoamylase (Gla) promoter, the triose phosphate isomerase (TPI) promoter, the translation elongation factor EF-Iα (TEF2) promoter, the glyceraldehyde-3-phosphate dehydrogenase (gpdA) promoter, the alcohol oxidase (AOX1) promoter, or the glutamate dehydrogenase (gdhA) promoter. An example of a ubiquitous promoter used for expression in plants is the cauliflower mosaic virus (CaMV) 35S promoter. Preferably, the nucleic acid construct of the present invention comprises a heterologous promoter represented by the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 58. More preferably, the nucleic acid construct of this aspect of the present invention comprises a heterologous promoter represented by the following sequence, wherein the sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 57.
[0167] In a further preferred embodiment within this aspect, the nucleic acid construct of the present invention further comprises one or more additional expression regulatory elements, wherein preferably the expression enhancing element, the heterologous promoter, and the one or more additional expression regulatory elements are configured to be operatively linked to optional nucleotide sequences of interest as defined below. “Additional expression regulatory element” will be understood herein as an element other than the expression enhancing element and / or promoter as defined above, and in its broadest sense, it can be an additional expression enhancing element or different expression enhancing elements or expression regulatory elements. The additional expression regulatory elements included in the present invention can participate in the transcriptional and / or translational regulation of genes, including, but not limited to, 5′-UTR, 3′-UTR, enhancers, promoters, intron sequences, polyadenylation signals, and chromatin control elements such as scaffold / matrix attachment regions, ubiquitous chromatin opening elements, cytosine phosphate-guanine pairs, and stabilizing and anti-repressor elements and any derivatives thereof. Other optional regulatory elements that may be present in the nucleic acid constructs of the present invention include, but are not limited to, coding nucleotide sequences of homologous and / or heterologous nucleotide sequences, including iron response elements (IREs), translation cis-regulatory elements (TLREs), or uORFs in the 5'UTR and poly(U) segments in the 3'UTR.
[0168] Preferably, in this respect, the additional expression regulatory element comprises or is composed of an intron sequence as defined herein. Preferably, the intron sequence encompassed within the additional expression regulatory element comprises at least a receptor splice site, which is understood herein to comprise a splice site AG, preferably preceded by a polypyrimidine fragment nucleotide sequence and optionally separated from the splice site AG by 1-50 nucleotides, and optionally further comprises a branching site comprising the sequence YTNAY at the 5' position of the polypyrimidine fragment nucleotide sequence, wherein the branching site may have the nucleotide sequence CYGAC. The receptor splice site is understood herein to be a splice site that, when bound to a donor splice site encompassed within the construct, results in the formation of an intron as defined in the defining portion. Preferably, the receptor splice site or splice site AG has the sequence [Y-rich]-NYAG-[cleavage]. Preferably, the receptor splice site or splice site AG has the sequence [Y-rich]-NYAG-[cleavage]-R (in the case of mammalian cells, where the host cell is), [Y-rich]-DYAG-[cleavage]-R or [Y-rich]-DYAG-[cleavage]-RW (in the case of plant cells, where the host cell is), [Y-rich]-AYAG-[cleavage] (in the case of yeast cells, where the host cell is), and [Y-rich]-NYAG-[cleavage] (in the case of insect cells, where the host cell is). “[Y-rich]” is understood herein to mean a polypyrimidine segment, preferably defined as at least 10 nucleotides comprising a continuous sequence of at least 6, 7, 8, 9, or preferably 10 pyrimidine nucleotides. Preferably, the receptor splice site or splice site GT has the sequence YAG-[cleavage]-R. Preferably, the intron sequence encompassed within the additional expression regulatory element consists of the receptor splice site or splice site AG. The intron sequence preferably comprises or consists of the following nucleotide sequences, wherein the nucleotide sequences have at least 80% of the following nucleotide sequences: nucleotides 171-277 of SEQ ID NO: 14, nucleotides 171-274 of SEQ ID NO: 19, nucleotides 133-210 of SEQ ID NO: 20, nucleotides 134-211 of SEQ ID NO: 21, nucleotides 134-226 of SEQ ID NO: 22, nucleotides 134-226 of SEQ ID NO: 23, nucleotides 133-225 of SEQ ID NO: 24, nucleotides 134-226 of SEQ ID NO: 25, nucleotides 146-257 of SEQ ID NO: 26, or nucleotides 147-223 of SEQ ID NO: 27, or nucleotides 970-1449 of SEQ ID NO: 1, or nucleotides 667-1228 of SEQ ID NO: 2.90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity. Preferably, the intron sequence included within the additional expression regulation element further includes donor splicing sites as defined herein. Even more preferably, the intron sequence encompassed within the additional expression regulation element is an intron as defined in the definition section. Most preferably, the intron sequence encompassed within the additional expression regulatory element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequence has at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with nucleotides 171-277 of SEQ ID NO: 14, 171-274 of SEQ ID NO: 19, 133-210 of SEQ ID NO: 20, 134-211 of SEQ ID NO: 21, 134-226 of SEQ ID NO: 22, 134-226 of SEQ ID NO: 23, 133-225 of SEQ ID NO: 24, 134-226 of SEQ ID NO: 25, 146-257 of SEQ ID NO: 26, or 147-223 of SEQ ID NO: 27.
[0169] In this regard, expression regulatory elements, which are translation enhancement elements, are also preferred. Preferably, compared with the expression of the protein or polypeptide using the construct (the only difference being that it does not contain the translation enhancement element), the translation enhancement element preferably increases the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% when tested in a system as illustrated in the examples included herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the translation enhancement element to be tested and the CMV promoter represented by SEQ ID NO: 57 and is operatively linked to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the translation-enhancing element to be tested.
[0170] Preferably, in this respect, the translation enhancement element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 3-51 throughout its length. Preferably, the translation enhancement element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequence has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 19 throughout its length. More preferably, the translation enhancement element comprises or consists of the following nucleotide sequence, wherein the nucleotide sequence has at least 90% identity with SEQ ID NO: 3-51 over its entire length. In this respect, a translation enhancement element comprising or consisting of the following nucleotide sequence is also preferred, wherein the nucleotide sequence comprises:
[0171] i) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence containing at least 8 consecutive C or T nucleotides, at least 3 A-rich nucleotide sequences containing at least 5 consecutive A nucleotides, or a GT-rich nucleotide sequence containing at least 10 nucleotides, wherein at least 80% are G or T nucleotides;
[0172] ii) A TC-rich nucleotide sequence comprising at least 8 consecutive C or T nucleotides, at least 3 A-rich nucleotide sequences comprising at least 5 consecutive A nucleotides, and a GT-rich nucleotide sequence comprising at least 10 nucleotides, wherein at least 80% are G or T nucleotides, and the expression-enhancing element does not contain a GAA repeating nucleotide sequence; or
[0173] iii) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence containing at least 8 consecutive C or T nucleotides, at least 3 A-rich nucleotide sequences containing at least 5 consecutive A nucleotides, or a GT-rich nucleotide sequence containing at least 10 nucleotides, wherein at least 80% are G or T nucleotides, wherein the GAA repeating nucleotide sequence is located at the 3' position of any one or more of the TC-rich nucleotide sequence, the A-rich nucleotide sequence, and / or the GT-rich nucleotide sequence.
[0174] In a first aspect of the invention, GAA repeating nucleotide sequences, TC-rich nucleotide sequences, A-rich nucleotide sequences, and GT-rich nucleotide sequences have been defined herein. These definitions also apply herein.
[0175] Preferably, within said aspect, the additional expression modulation element comprises translation enhancement elements and intron sequences as defined herein.
[0176] Preferably, within the aforementioned aspects, the additional expression regulatory element is located within the nucleic acid construct of the present invention and downstream of the heterologous promoter or at the 3' site. Preferably, the additional expression regulatory element is located within the nucleic acid construct of the present invention and upstream of the nucleic acid sequence encoding the protein or polypeptide of interest or at the 5' site. Furthermore, preferably, the nucleic acid construct of the present invention comprises the following nucleotide sequences shown herein in their relative positions, in the 5' to 3' direction: (i) an expression enhancing element; (ii) a heterologous promoter; optionally (iii) an additional expression regulatory element; and optionally (iv) a nucleotide sequence of interest, wherein preferably the expression enhancing element, the heterologous promoter, and the additional expression regulatory element are all configured to be operatively linked to the optional nucleotide sequence of interest as defined below. It should be understood that the expression enhancing element, the heterologous promoter, and optionally the additional expression regulatory element of the nucleic acid construct of the present invention are all configured to be operatively linked to the same individual nucleotide sequence of interest.
[0177] When the expression-enhancing element of the present invention is combined with additional expression-regulating elements as defined herein in an expression construct for expressing a protein or peptide of interest, the inventors have discovered unexpected synergistic effects. In stable transfection pools with both the expression-enhancing element and the additional expression-regulating element, protein yields are significantly higher than those expected based on the individual effects of adding either element. Preferably, the increase in protein or peptide expression is copy number independent, as established by a person skilled in the art through tests suitable for determining copy number dependence, such as, but not limited to, the triple TaqMan test, as further detailed in Example 4 of the present invention. Preferably, the nucleic acid construct is a recombinant and / or isolated construct as defined herein. Preferably, the nucleic acid construct further comprises a nucleotide sequence of interest operably linked to and / or, under the control of the expression-enhancing element, the heterologous promoter, and optionally the additional expression-regulating elements as defined herein. The presence of the nucleotide sequence of interest is optional. “Optional” is understood herein to mean that it is not necessarily present in the expression construct. For example, such a nucleotide sequence of interest need not be present in a commercial expression vector, but may be readily introduced by a person skilled in the art prior to its use in the methods of the present invention. It should be understood that the expression-enhancing element, the heterologous promoter, and optionally the additional expression-regulating element are all configured to be operatively linked to the same individual nucleotide sequence of interest. Preferably, the nucleic acid construct of the tenth aspect of the invention has at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO 73, 74, 75, or 76.
[0178] In a preferred embodiment of this aspect, the nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. The protein or polypeptide of interest may be a homologous protein or polypeptide, but in a preferred embodiment of the invention, the protein or polypeptide of interest is a heterologous protein or polypeptide. The nucleotide sequence encoding a heterologous protein or polypeptide may be wholly or partially derived from any source known in the art, including bacterial or viral genomes or episomes, eukaryotic or plasmid DNA, cDNA, or chemically synthesized DNA. The nucleotide sequence encoding the protein or polypeptide of interest may constitute a continuous coding region or it may include one or more introns bound by appropriate splicing sites. It may further consist of segments derived from different sources, either naturally occurring or synthetic. According to the method of the invention, the nucleotide sequence encoding the protein or polypeptide of interest is preferably a full-length nucleotide sequence, but may also be the functionally active portion or other portion of the full-length nucleotide sequence. The nucleotide sequence encoding the protein or polypeptide of interest may also include a signal sequence that directs when the protein or polypeptide of interest is expressed at a specific location in a cell or tissue. In addition, the nucleotide sequence encoding the protein or polypeptide of interest may also contain a sequence that facilitates protein purification and detection through methods such as Western blotting and ELISA (e.g., c-myc or polyhistidine sequences).
[0179] The proteins or polypeptides of interest in this respect have been defined earlier in the first aspect of the invention.
[0180] In an alternative embodiment, the nucleotide sequence of interest is not a coding sequence for a protein or polypeptide, but may be a functional nucleotide sequence. This alternative embodiment of this aspect has already been defined earlier in the first aspect of the invention.
[0181] Twelfth aspect
[0182] In a twelfth aspect, the present invention provides an expression vector comprising a nucleic acid construct according to an eleventh aspect of the invention. The expression vector of the present invention is preferably a plasmid, granulosome, or bacteriophage, or a nucleotide sequence of linear or circular single-stranded or double-stranded DNA or RNA derived from any source, wherein many nucleotide sequences have been ligated or recombined into a unique construct capable of introducing any of the sense or antisense oriented nucleotide sequences of the present invention into cells. The choice of vector depends on the subsequent recombination method and the host cell used. The vector may be a self-replicating vector or may be replicated together with the chromosome into which it has been integrated. Preferably, the vector contains a selection marker. Useful markers are dependent on the chosen host cell and are well known to those skilled in the art and selected from, but not limited to, selection markers defined in the third aspect of the invention. A preferred expression vector is the pcDNA3.1 expression vector. Preferred selection markers are neomycin resistance genes, bleomycin resistance genes, and blasicidin resistance genes.
[0183] Thirteenth aspect
[0184] In a thirteenth aspect, the present invention provides cells as defined herein, comprising nucleic acid molecules according to the tenth aspect of the invention and / or nucleic acid constructs according to the eleventh aspect of the invention and / or expression vectors according to the twelfth aspect of the invention. The cell type in the context of this aspect is the same as the cell type defined in the context of the third aspect.
[0185] Therefore, another aspect of the invention relates to host cells preferably genetically modified by the methods of the invention, wherein the host cells contain nucleic acid constructs as defined above in the thirteenth aspect. Suitable bacteria for transformation methods in plants include *Agrobacterium tumefaciens* and *Agrobacterium rhizogenes*.
[0186] In the context of this thirteenth aspect, the nucleic acid construct refers to the nucleic acid construct of the third aspect: it is preferably stably maintained as an autonomously replicating element, or more preferably, the nucleic acid construct is integrated into the genome of the host cell, in which case the construct is typically integrated, for example, by non-homologous recombination, at a random location in the host cell's genome. Stably transformed host cells are produced by known methods. The terminology, the definition of stable transformation, and the methods covered for stable transformation have already been provided under the third aspect.
[0187] Alternatively, the protein or peptide of interest can be expressed in a host cell, such as a mammalian cell, depending on transient expression from the vector.
[0188] The nucleic acid constructs in this regard preferably also contain marker genes, which can provide selection or screening capabilities in the treated host cells.
[0189] All definitions relating to selective markers and types of selective markers have been provided in the third aspect, including examples of luciferase genes used as selective markers, examples of first-class markers based on cell metabolism, and examples of dominant selection (the use of mutant cell lines lacking the ability to grow independently of supplemental culture medium). These also apply to the thirteenth aspect of the invention.
[0190] When transformed host cells are obtained by means of the method according to the invention (see below), in a suitable culture medium, which may optionally contain selected antibiotics or biocides known in the art for transforming cells, host tissues can be regenerated from said transformed cells.
[0191] Preferably, the obtained transforming host tissue is determined by using the selection of selectable marker genes present on the nucleic acid construct as defined herein.
[0192] Fourteenth aspect
[0193] In a fourteenth aspect, the present invention provides a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0194] a) Providing a nucleic acid construct according to the eleventh aspect of the invention, comprising a nucleotide sequence encoding a protein or polypeptide of interest; and
[0195] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0196] c) Inducing the transformed cells to express the protein or polypeptide of interest; and optionally...
[0197] d) Purify the protein or polypeptide of interest.
[0198] In a preferred embodiment of the method according to the inventors, a nucleic acid construct as defined above in the eleventh aspect of the invention is used. The method of the invention can be in vitro or ex vivo. The method of the invention can be applied to cell culture, organism culture, or tissue culture. Alternatively, expression immediately adjacent to host cells can be used in a cell-free translation system to generate proteins or peptides of interest using RNA derived from the nucleic acid constructs of the invention. The method of the invention can be used in cultured cells.
[0199] According to step b), the technician is able to transform the cells. Transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate the permeation and transfer of DNA into host cells.
[0200] In step c), the transformed cells are made to express a protein or peptide of interest, and optionally, the protein or peptide is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to the expression of the protein or peptide of interest. Techniques for expressing or overexpressing a protein or peptide of interest are well understood by those skilled in the art. The method of the invention also includes methods in which the transformed cells are not required to be subjected to specific conditions that lead to the expression of the protein or peptide of interest, but where the protein or peptide of interest is expressed automatically (e.g., constitutively).
[0201] The purification steps and the definitions involved in these steps, as well as the definition of the isolated protein or peptide, are identical to those defined in the methods of the fourth aspect and have been previously defined herein. If desired, as defined in the methods of the fourth aspect, a nucleotide sequence encoding the protein or peptide of interest may be ligated to a heteronucleotide sequence to encode a fusion protein or peptide, thereby facilitating protein purification and detection (e.g., with Western blotting and in ELISA). Suitable heteroseeds include, but are not limited to, nucleotide sequences encoding proteins, such as, for example, glutathione S-transferase, maltose-binding proteins, metal-binding polyhistidines, green fluorescent protein, luciferase, and β-galactosidase. Proteins or peptides may also be coupled to non-peptide carriers, tags, or labels that facilitate the tracking of proteins or peptides in vivo and in vitro and enable the identification and quantification of protein or peptide binding to substrates. Such tags, labels, or carriers are well known in the art and include, but are not limited to, biotin, radiolabels, and fluorescent labels.
[0202] Preferably, the method of this fourteenth aspect of the invention results in increased expression of the protein or polypeptide of interest. Preferably, in the expression system, the expression level is established using an expression construct according to the eleventh aspect of the invention (which comprises an expression-enhancing element and a heterologous promoter operably linked to a nucleotide sequence encoding the protein or polypeptide of interest according to the eleventh aspect of the invention). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by suitable assays, such as ELISA, Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Preferably, compared to methods that differ only in that the construct used in step a does not contain the expression-enhancing element, the method of the present invention, when tested in systems exemplified in the embodiments included herein, increases the expression of a protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the expression-enhancing element to be tested operatively linked to the nucleotide sequence of interest and a CMV promoter represented by SEQ ID NO: 57). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the expression-enhancing element to be tested. Preferably, the increase in protein or peptide expression is copy number independent, as established by a person skilled in the art using tests suitable for determining copy number dependence, such as, but not limited to, the triple TaqMan test further described in Example 4 of the invention.
[0203] Fifteenth aspect
[0204] In a fifteenth aspect, the present invention provides a method for expressing a protein or polypeptide of interest in an organism, comprising the following steps:
[0205] a) Provide a nucleic acid construct according to aspect eleven, comprising a nucleotide sequence encoding a protein or polypeptide of interest according to the present invention; and
[0206] b) Contacting the target cells and / or target tissues of an organism with the nucleic acid construct to obtain transformed target cells and / or transformed target tissues, thereby enabling the transformed cells to express proteins or peptides of interest; and optionally...
[0207] c) Develop the transformed target cells into a transformed organism; and optionally
[0208] d) To induce the transformed organism to express a protein or polypeptide of interest, for example, by subjecting the transformed organism to conditions that lead to the expression of the protein or polypeptide of interest; and optionally, to recover the protein or polypeptide.
[0209] The target cells can be embryonic target cells, such as embryonic stem cells, derived from non-human mammals, such as cattle, pigs, etc. Preferably, the target cells are not human embryonic stem cells. In the case of multicellular fungi, such target cells can be fungal cells that can be proliferated into the multicellular fungi. When transforming plant tissues or plant cells (e.g., leaves, stem segments, roots, but also protoplasts or plant cells in suspension culture) are obtained using the method according to the invention, the whole plant can be regenerated from the transformed tissues or cells in a suitable culture medium, which may optionally contain antibiotics or biocides known in the art for selecting transformed cells. This method of the invention can be applied to nucleic acid-based inoculation and / or gene therapy, preferably in mammals, and most preferably in humans. Covered within the invention are therapeutic methods incorporating this aspect, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. The invention also relates to a construct for treatment according to the eleventh aspect of the invention, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. Furthermore, the present invention relates to the use of the constructs of the eleventh aspect of the invention for the manufacture of pharmaceutical preparations, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide.
[0210] Furthermore, embodiments of the present invention involve non-human transformed organisms. These organisms are transformed using the nucleotide sequences, recombinant nucleic acid constructs, or vectors according to the present invention, and are capable of producing polypeptides of interest. This includes non-human transgenic organisms, such as transgenic non-human mammals, transgenic plants (including propagation, harvesting, and tissue materials of said transgenic plants, including but not limited to leaves, roots, stems, and flowers), multicellular fungi, etc.
[0211] Preferably, the method of this aspect of the invention increases the expression of a protein or polypeptide of interest in the organism or at least in one tissue, organelle, or organ of the organism. Preferably, in the expression system, the expression level is established using an expression construct according to the eleventh aspect of the invention (which comprises an expression-enhancing element and a heterologous promoter operably linked to a nucleotide sequence encoding a protein or polypeptide of interest according to the eleventh aspect of the invention). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by suitable assays, such as ELISA, Western blotting, or, depending on the identity of the protein or polypeptide of interest, any suitable protein identification and / or quantitative assay known to those skilled in the art. Preferably, compared to methods that differ only in that the construct used in step a does not contain the expression-enhancing element, this method of the invention, when tested in the organism or at least in one tissue or organelle of the organism, as illustrated in the embodiments included herein, increases the expression of a protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the expression-enhancing element to be tested and the CMV promoter represented by SEQ ID NO: 57, and which is operatively linked to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the expression-enhancing element to be tested.
[0212] Preferably, the increase in protein or polypeptide expression is copy number independent, as established by a technician using tests suitable for determining copy number dependence, such as, but not limited to, the triple TaqMan test further detailed in Example 4 of the invention.
[0213] Sixteenth aspect
[0214] In a sixteenth aspect, the present invention provides a method for transcription and optionally purification of the obtained transcript, comprising the following steps:
[0215] a) Provide a nucleic acid construct according to aspect eleven, comprising the nucleotide sequence of interest of the present invention; and
[0216] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0217] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally...
[0218] d) Purify the resulting transcript.
[0219] In a preferred embodiment of this aspect of the invention, a nucleic acid construct as defined above in the eleventh aspect is used. The method of the invention can be an in vitro or ex vivo method. The method of the invention can be applied to cell culture, organism culture, or tissue culture. The method of the invention can be applied to nucleic acid-based inoculation and / or gene therapy, preferably in mammals, particularly humans. Covered within the invention are methods for treatment, which include or constitute the methods of this aspect, wherein the nucleotide sequence of interest encodes a therapeutic transcript. The invention also relates to a construct for treatment according to the eleventh aspect of the invention, wherein the nucleotide sequence of interest encodes a therapeutic transcript. Furthermore, the invention relates to the use of the construct of the eleventh aspect of the invention for the manufacture of pharmaceutical agents, wherein the nucleotide sequence of interest encodes a therapeutic transcript.
[0220] According to step b), the technician is able to transform the cells. Transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate the permeation and transfer of DNA into host cells.
[0221] In step c), the transformed cells are made to produce a transcript of the nucleotide sequence of interest, and optionally the obtained transcript is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to transcription of the nucleotide sequence of interest. Techniques for transcribing the nucleotide sequence of interest are well understood by those skilled in the art. The method of the invention also includes methods in which the transformed cells are not required to undergo specific conditions that lead to transcription of the nucleotide sequence of interest, but where the nucleotide sequence of interest is transcribed automatically (e.g., constitutively).
[0222] The purification steps depend on the resulting transcript. The term "isolation" indicates that the transcript was found under conditions different from its native environment. In a preferred form, the isolated transcript is substantially free of other cellular components, particularly other homologous cellular components, such as homologous proteins. It is preferable to provide transcripts in a pure form greater than 40%, more preferably greater than 60%. Even more preferably, it is preferable to provide transcripts in a highly purified form, i.e., a purity greater than 80%, more preferably greater than 95%, and even more preferably greater than 99%, as determined by Northern blotting.
[0223] Preferably, the method of this aspect of the invention increases transcription of the nucleotide sequence of interest. Preferably, in the expression system, transcriptional levels are established using an expression construct according to the second aspect of the invention (which comprises an expression-enhancing element operatively linked to the nucleotide sequence of interest according to the second aspect of the invention and a heterologous promoter). Preferably, transcription of the nucleotide sequence of interest is detected by a suitable assay, such as RT-qPCR. Preferably, compared to methods that differ only in that the construct used in step a does not contain the expression-enhancing element, the method of the invention, preferably when tested in the systems exemplified in the embodiments included herein, increases transcription by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, and most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains an expression-enhancing element to be tested operably linked to the nucleotide sequence of interest and a CMV promoter represented by SEQ ID NO: 57). Transcription is preferably measured using RT-qPCR, and the transcription level is relative to the transcription level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector used does not contain the expression-enhancing element to be tested.
[0224] Seventeenth aspect
[0225] In a seventeenth aspect, the present invention provides the application of nucleic acid molecules according to the tenth aspect of the invention and / or the application of nucleic acid constructs according to the eleventh aspect of the invention and / or the application of expression vectors according to the twelfth aspect of the invention and / or the application of cells according to the thirteenth aspect of the invention for transcription of nucleotide sequences of interest and / or expression of proteins or polypeptides of interest.
[0226] Eighteenth aspect
[0227] In an eighteenth aspect, the present invention provides nucleic acid molecules according to the tenth aspect of the invention and / or nucleic acid constructs according to the eleventh aspect of the invention and / or expression vectors according to the twelfth aspect of the invention and / or cells according to the thirteenth aspect of the invention for use as pharmaceutical agents. The invention also relates to treatment methods comprising administering nucleic acid molecules according to the tenth aspect of the invention and / or nucleic acid constructs according to the eleventh aspect of the invention and / or expression vectors according to the twelfth aspect of the invention and / or cells according to the thirteenth aspect of the invention, wherein preferably the administration is administered to mammals, more preferably to humans. Preferably, the treatment is based on nucleic acid inoculation and / or gene therapy, preferably in mammals, most preferably in humans. Furthermore, the present invention relates to the application of nucleic acid molecules according to the tenth aspect of the invention and / or nucleic acid constructs according to the eleventh aspect of the invention and / or expression vectors according to the twelfth aspect of the invention and / or cells according to the thirteenth aspect of the invention for the preparation of pharmaceutical agents. Preferably, the pharmaceutical agent is for nucleic acid-based inoculation and / or gene therapy, preferably in mammals, most preferably in humans.
[0228] Nineteenth aspect
[0229] In a nineteenth aspect, the present invention provides a nucleic acid molecule represented by a nucleotide sequence having at least 50% identity with SEQ ID NO: 88, for increasing transcription of the nucleotide sequence of interest and / or expression of the protein or polypeptide of interest. In the context of aspects nineteen through twenty-seven, the percentage of identity is preferably assessed relative to the full length of SEQ ID NO: 88. However, assessment of the percentage of identity relative to a portion of SEQ ID NO: 88 as defined in the section entitled "Defined" is not excluded. The nucleotide sequence of the present invention is capable of increasing transcription of the nucleotide sequence of interest and / or expression of the protein or polypeptide of interest. The nucleic acid molecule represented by the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 may be referred to as a transcriptional regulatory sequence.
[0230] Preferably, in this respect, in the expression system, the transcription level is established using an expression construct (which contains the nucleotide sequence operably linked to the nucleotide sequence of interest) and a suitable assay such as RT-qPCR. Preferably, compared to the transcription level using a construct (wherein the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 has been replaced by a substitute sequence), the nucleotide sequence of the present invention increases transcription by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%, preferably as illustrated in Example 11 included herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the nucleotide sequence of the present invention operably linked to the nucleotide sequence of interest). Transcription is preferably measured using RT-qPCR, and the transcription level is relative to the transcription level of the nucleotide sequence of interest, which is measured under the same conditions except that, in the expression vector used, the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 has been replaced by a substitute sequence, preferably as illustrated in Example 11 included herein.
[0231] Preferably, in this respect, in the expression system, the expression level is established using an expression construct (which contains the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and is operatively linked to a nucleotide sequence encoding a protein or polypeptide of interest). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by suitable assays, such as enzyme-linked immunosorbent assay (ELISA), Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Preferably, compared to the expression of the protein or polypeptide using the construct (the only difference being that the nucleotide sequence has been replaced by a substitution sequence), preferably when tested in a system exemplified in Example 11 included herein, the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 increases the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and is operatively linked to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the isolated nucleic acid molecule as defined herein. In a preferred embodiment, the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 is a sequence derived from the UBC ubiquitin gene. Preferably, the nucleotide sequence is derived from the mammalian UBC ubiquitin gene. More preferably, the nucleotide sequence is a Chinese hamster homologous gene derived from the human UBC ubiquitin gene, represented as the Hamster gene for polyubiquitin, or CRUPUQ (GenBank D63782).
[0232] In a preferred embodiment, the nucleotide sequence comprises or consists of a sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 88 throughout its entire length. Preferably, the sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 88 throughout its entire length includes a promoter as defined in the definition section. Preferably, in a host cell as defined below, the promoter is capable of inducing transcription of the nucleotide sequence of interest and / or expression of a protein or polypeptide of interest encoded by the nucleotide sequence. Preferably, the nucleotide sequence used to increase expression is a contiguous sequence of at least 1450, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, or 2617 nucleotides in length, preferably at least 2617 nucleotides in length as specified in SEQ ID NO: 88. Preferably, the nucleotide sequence having at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 88 is at most 8000 nucleotides in length. Preferably, the length of the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, or 2617 nucleotides. Most preferably, the length of the sequence is 2617 nucleotides. Preferably, the nucleotide sequence is at most 8000, 7000, 6000, 5000, 4000, 3000, or 2617 nucleotides in length, and has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with SEQ ID NO: 88 throughout its entire length.
[0233] Twentieth aspect
[0234] In a twentieth aspect, the present invention provides a nucleic acid construct comprising a nucleic acid molecule according to the nineteenth aspect of the invention. The nucleic acid construct of the present invention comprises a nucleotide sequence according to the nineteenth aspect of the invention. Preferably, the nucleic acid construct is a recombinant and / or isolated construct as defined herein. Preferably, the nucleic acid construct further comprises an optional nucleotide sequence of interest as defined below, wherein the nucleotide sequence of the present invention is operatively linked to the optional nucleic acid sequence of interest.
[0235] In a further preferred embodiment within this aspect, the nucleic acid construct of the present invention further comprises one or more additional expression regulatory elements, wherein preferably the nucleotide sequence and the one or more additional expression regulatory elements are configured to be operatively linked to optional nucleotide sequences of interest as defined below. “Additional expression regulatory elements” will be understood herein as elements other than the nucleotide sequences as defined above, which may be additional expression regulatory elements or different expression regulatory elements, or additional expression enhancing elements or different expression enhancing elements. The additional expression regulatory elements included in the present invention can participate in the transcriptional and / or translational regulation of genes, including, but not limited to, 5′-UTR, 3′-UTR, enhancers, promoters, introns, polyadenylation signals, and chromatin control elements such as scaffold / matrix attachment regions, panchromatin opening elements, cytosine phosphate-guanine pairs, as well as stabilizing and anti-repressor elements and any derivatives thereof. Other expression regulatory elements that may be present in the nucleic acid constructs of this invention, as included in this invention, include, but are not limited to, iron-reactive elements (IREs), translation cis-regulatory elements (TLREs), or uORFs in the 5' UTR and poly(U) segments in the 3' UTR.
[0236] In this regard, expression regulatory elements, which are translation enhancement elements, are also preferred. Preferably, compared to the expression of the protein or polypeptide using the construct (which differs only in that it does not contain the translation enhancement element), the translation enhancement element preferably increases the expression of the protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% when tested in the systems illustrated in the embodiments herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains the translation enhancement element to be tested and a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operatively ligated to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the expression vector does not contain the translation-enhancing element to be tested.
[0237] Preferably, in this respect, the translation enhancement element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequences have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 3-51 throughout their entire length. Preferably, the translation enhancement element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequences have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 19 throughout their entire length. More preferably, the translation enhancement element comprises or consists of the following nucleotide sequences, wherein the nucleotide sequences have at least 90% identity with SEQ ID NO: 3-51 over their entire length. Also preferred in this respect is a translation enhancement element comprising or consisting of the following nucleotide sequences, wherein the nucleotide sequences include:
[0238] i) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence containing at least 8 consecutive C or T nucleotides, at least 3 A-rich nucleotide sequences containing at least 5 consecutive A nucleotides, or a GT-rich nucleotide sequence containing at least 10 nucleotides, wherein at least 80% are G or T nucleotides;
[0239] ii) A TC-rich nucleotide sequence comprising at least 8 consecutive C or T nucleotides, at least 3 A-rich nucleotide sequences comprising at least 5 consecutive A nucleotides, and a GT-rich nucleotide sequence comprising at least 10 nucleotides, wherein at least 80% are G or T nucleotides, and the expression-enhancing element does not contain a GAA repeating nucleotide sequence; or
[0240] iii) A GAA repeating nucleotide sequence, a TC-rich nucleotide sequence containing at least 8 consecutive C or T nucleotides, a 3'-rich nucleotide sequence containing at least 5 consecutive A nucleotides, a GT-rich nucleotide sequence containing at least 10 nucleotides, wherein at least 80% are G or T nucleotides, wherein the GAA repeating nucleotide sequence is located at the 3' of any one or more of the TC-rich nucleotide sequence, the A-rich nucleotide sequence, and / or the GT-rich nucleotide sequence.
[0241] In a first aspect of the invention, GAA repeating nucleotide sequences, TC-rich nucleotide sequences, A-rich nucleotide sequences, and GT-rich nucleotide sequences have been defined herein. These definitions also apply herein.
[0242] Preferably, within the aforementioned aspects, the additional expression regulatory element is located within the nucleic acid construct of the present invention and has at least 50% identity with SEQ ID NO: 88. Preferably, the additional expression regulatory element is located within the nucleic acid construct of the present invention and is upstream of the nucleic acid sequence encoding the protein or polypeptide of interest or at the 5' site. Furthermore, preferably, the nucleic acid construct of the present invention comprises the following nucleotide sequences shown herein in their relative positions, in the 5' to 3' direction:
[0243] Optionally (i) the expression of the preferred enhancement element,
[0244] (ii) Having a nucleotide sequence with at least 50% identity to SEQ ID NO: 88
[0245] Optionally (iii) additional expression regulation elements, and
[0246] Optionally (iv) the nucleotide sequence of interest,
[0247] Preferably, the expression-enhancing element having at least 50% identity with the nucleotide sequence of SEQ ID NO: 88 and the additional expression-regulating element are configured to be operatively linked to the optional nucleotide sequence of interest as defined below. It should be understood that the expression-enhancing element of the nucleic acid construct of the present invention, the nucleotide sequence having at least 50% identity with SEQ ID NO: 88, and optionally the additional expression-regulating element are all configured to be operatively linked to the same individual nucleotide sequence of interest.
[0248] The presence of the nucleotide sequence of interest is optional. "Optional" will be understood herein to mean that it is not necessarily present in the expression construct. For example, such a nucleotide sequence of interest does not need to be present in a commercial expression vector, but can be readily introduced by those skilled in the art prior to its use in the methods of the present invention. It should be understood that the expression-enhancing element, the nucleotide sequence having at least 50% identity with SEQ ID NO: 88, and optionally the additional expression-regulating element are all configured to be operatively linked to the same individual nucleotide sequence of interest.
[0249] In a preferred embodiment within this aspect, the nucleotide sequence of interest is a nucleotide sequence encoding a protein or polypeptide of interest. The protein or polypeptide of interest may be a homologous protein or polypeptide, but in a preferred embodiment of the invention, the protein or polypeptide of interest is a heterologous protein or polypeptide. The nucleotide sequence encoding a heterologous protein or polypeptide may be wholly or partially derived from any source known in the art, including bacterial or viral genomes or episomes, eukaryotic or plasmid DNA, cDNA, or chemically synthesized DNA. The nucleotide sequence encoding the protein or polypeptide of interest may constitute a continuous coding region or it may include one or more introns bound by appropriate splicing sites. It may further consist of segments derived from different sources, either naturally occurring or synthetic. The nucleotide sequence encoding the protein or polypeptide of interest according to the method of the invention is preferably a full-length nucleotide sequence, but may also be the functionally active portion or other portion of the full-length nucleotide sequence. The nucleotide sequence encoding the protein or polypeptide of interest may also include a signal sequence that directs when the protein or polypeptide of interest is expressed at a specific location in a cell or tissue. In addition, the nucleotide sequence encoding the protein or polypeptide of interest may also contain sequences that facilitate protein purification and detection, such as Western blotting and ELISA (e.g., c-myc or polyhistidine sequences).
[0250] In a first aspect of the invention, proteins or polypeptides of interest in this aspect have been defined earlier herein.
[0251] In an alternative embodiment, the nucleotide sequence of interest is not a sequence encoding a protein or polypeptide, but may be a functional nucleotide sequence. This alternative embodiment of this aspect has been defined earlier in the first aspect of the invention herein.
[0252] Twenty-first aspect
[0253] In a twenty-first aspect, the present invention provides an expression vector comprising a nucleic acid construct according to the twentieth aspect of the invention. The expression vector of the present invention is preferably a nucleotide sequence derived from plasmids, granules, or bacteriophages of any origin, or linear or circular single-stranded or double-stranded DNA or RNA, wherein many nucleotide sequences have been ligated or recombined into a unique construct capable of introducing any of the sense or antisense nucleotide sequences of the present invention into cells. The choice of vector depends on the subsequent recombination method and the host cell used. The vector can be a self-replicating vector or can be replicated together with the chromosome into which it has been integrated. Preferably, the vector contains a selection marker. Useful markers are dependent on the chosen host cell and are well known to those skilled in the art and selected from, but not limited to, selection markers defined in the third aspect of the invention. A preferred expression vector is the pcDNA3.1 expression vector. Preferred selection markers are neomycin resistance genes, bleomycin resistance genes, and blast fungicide resistance genes.
[0254] Twenty-second aspect
[0255] In a twenty-second aspect, as defined herein, the invention provides a cell comprising a nucleic acid molecule according to the nineteenth aspect of the invention and / or a nucleic acid construct according to the twentieth aspect of the invention and / or an expression vector according to the twenty-first aspect of the invention. The cell type in the context of this aspect is the same as the cell type defined in the context of the third aspect.
[0256] Therefore, another aspect of the invention relates to host cells preferably genetically modified by the method of the invention, wherein the host cells contain nucleic acid constructs as defined in the twentieth aspect above. Suitable bacteria for transformation methods in plants include *Agrobacterium tumefaciens* and *Agrobacterium rhizogenes*.
[0257] In the context of this twentieth aspect, the nucleic acid construct is the same as that in the third aspect: it is preferably stably maintained as an autonomously replicating element, or more preferably, the nucleic acid construct is integrated into the genome of the host cell, in which case the construct is typically integrated, for example, by non-homologous recombination, at a random location in the host cell's genome. Stably transformed host cells are produced by known methods. The definition of the term stable transformation and the methods covered for stable transformation have already been provided under the third aspect.
[0258] Alternatively, the protein or peptide of interest can be expressed in a host cell, such as a mammalian cell, depending on transient expression from a vector.
[0259] The nucleic acid constructs in this regard preferably also contain marker genes, which can provide selection or screening capabilities in the treated host cells.
[0260] In the third aspect, all definitions relating to selective markers and types of selective markers have been provided, including examples of luciferase genes used as selective markers, examples of first-class markers based on cell metabolism, and examples of dominant selection (the application of mutant cell lines lacking the ability to grow independently of supplemental culture medium). These also apply to the thirteenth aspect of the invention.
[0261] When transformed host cells are obtained by means of the method according to the invention (see below), in a suitable culture medium, which may optionally contain selected antibiotics or biocides known in the art for transforming cells, host tissues can be regenerated from said transformed cells.
[0262] Preferably, the obtained transformation host tissue is determined by selecting and utilizing selection marker genes (when present on nucleic acid constructs as defined herein).
[0263] Twenty-third aspect
[0264] In a twenty-third aspect, the present invention provides a method for expressing and optionally purifying a protein or polypeptide of interest, comprising the following steps:
[0265] a) Provides a nucleic acid construct according to the twentieth aspect of the invention, comprising a nucleotide sequence encoding a protein or polypeptide of interest; and
[0266] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0267] c) Inducing the transformed cells to express the protein or polypeptide of interest; and optionally...
[0268] d) Purify the protein or polypeptide of interest.
[0269] In a preferred embodiment of the method according to the invention, a nucleic acid construct as defined above in the twentieth aspect of the invention is used. The method of the invention can be in vitro or ex vivo. The method of the invention can be applied to cell culture, organism culture, or tissue culture. Alternatively, expression of the protein or polypeptide of interest immediately adjacent to the host cell can be generated in a cell-free translation system using RNA derived from the nucleic acid construct of the invention. The method of the invention can be used in cultured cells.
[0270] According to step b), the technician is able to transform the cells. Transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate the permeation and transfer of DNA into host cells.
[0271] In step c), the transformed cells are made to express a protein or peptide of interest, and optionally, the protein or peptide is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to the expression of the protein or peptide of interest. Those skilled in the art will appreciate the techniques used to express or overexpress a protein or peptide of interest. The method of the invention also includes methods in which the transformed cells are not required to undergo specific conditions that lead to the expression of the protein or peptide of interest, but where the protein or peptide of interest is expressed automatically (e.g., constitutively).
[0272] The purification steps and the definitions of those steps, as well as the definition of the isolated protein or peptide, are the same as those defined in the fourth aspect and earlier herein. If desired, as defined in the fourth aspect, a nucleotide sequence encoding the protein or peptide of interest may be coupled to a heteronucleotide sequence to encode a fusion protein or peptide, thereby facilitating protein purification and detection, for example, in Western blotting and ELISA. Suitable hetero sequences include, but are not limited to, nucleotide sequences encoding proteins, such as, for example, glutathione S-transferase, maltose-binding proteins, metal-binding polyhistidines, green fluorescent protein, luciferase, and β-galactosidase. Proteins or peptides may also be coupled to non-peptide carriers, tags, or labels that facilitate the tracking of proteins or peptides in vivo and in vitro, and enable the identification and quantification of protein or peptide binding to substrates. Such tags, labels, or carriers are well known in the art and include, but are not limited to, biotin, radiolabels, and fluorescent labels.
[0273] Preferably, the method of this twenty-third aspect of the invention increases the expression of the protein or polypeptide of interest. Preferably, in the expression system, the expression level is established using an expression construct according to the twenty-first aspect of the invention (the expression construct comprising a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operably linked to a nucleotide sequence encoding the protein or polypeptide of interest). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and its expression is detected by suitable assays, such as ELISA, Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Preferably, compared to a method (which differs only in step a) in which a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 has been replaced by an alternative sequence, preferably one of those alternative sequences described in Example 11), more preferably when tested in a system exemplified in Example 11 included herein, the method of the present invention increases the expression of a protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operatively linked to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the nucleotide sequence in the expression vector having at least 50% identity with SEQ ID NO: 88 has been replaced by an alternative sequence, preferably one of those alternative sequences described in Example 11.
[0274] Twenty-fourth aspect
[0275] In a twenty-fourth aspect, the present invention provides a method for expressing a protein or polypeptide of interest in an organism, comprising the following steps:
[0276] a) Provide a nucleic acid construct according to aspect 20, comprising a nucleotide sequence encoding a protein or polypeptide of interest according to the present invention; and
[0277] b) Contacting the target cells and / or target tissues of an organism with the nucleic acid construct to obtain transformed target cells and / or transformed target tissues, thereby enabling the transformed cells to express proteins or peptides of interest; and optionally...
[0278] c) Develop the transformed target cells into a transformed organism; and optionally
[0279] d) To induce the transformed organism to express a protein or polypeptide of interest, for example, by subjecting the transformed organism to conditions that lead to the expression of the protein or polypeptide of interest, and optionally by recovering the protein or polypeptide.
[0280] The target cells can be embryonic target cells, such as embryonic stem cells, derived from non-human mammals, such as cattle, pigs, etc. Preferably, the target cells are not human embryonic stem cells. In the case of multicellular fungi, such target cells can be fungal cells that can be proliferated into the multicellular fungi. When transformed plant tissues or plant cells (e.g., leaves, stem segments, roots, and also protoplasts or plant cells in suspension culture) are obtained by means of this method according to the invention, the whole plant can be regenerated from the transformed tissues or cells in a suitable culture medium (which may optionally contain antibiotics or biocides of choice known in the art for transforming cells). This method of the invention can be applied to nucleic acid-based inoculation and / or gene therapy, preferably in mammals, and most preferably in humans. Covered within the invention are therapeutic methods including this aspect, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. The invention also relates to a construct for treatment of the twentieth aspect of the invention, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide. Furthermore, the present invention relates to a construct of the twentieth aspect of the invention used in the manufacture of a pharmaceutical agent, wherein the protein or polypeptide of interest is a therapeutic and / or immunogenic protein or polypeptide.
[0281] Furthermore, embodiments of the present invention involve non-human transformed organisms. These organisms are transformed using the nucleotide sequences, recombinant nucleic acid constructs, or vectors according to the present invention, and are capable of producing polypeptides of interest. This includes non-human transgenic organisms, such as transgenic non-human mammals, transgenic plants (including propagation, harvesting, and tissue materials of said transgenic plants, including but not limited to leaves, roots, stems, and flowers), multicellular fungi, etc.
[0282] Preferably, the method of this aspect of the invention increases the expression of a protein or polypeptide of interest in the organism or at least in one tissue, organelle, or cell of the organism. Preferably, in the expression system, the expression level is established using an expression construct according to the twenty-first aspect of the invention (the expression construct comprising a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operably linked to a nucleotide sequence encoding a protein or polypeptide of interest). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by suitable assays such as ELISA, Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Preferably, compared to a method that differs only in that, in step a, a construct is used in which a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 has been replaced by a substitution sequence, preferably one of those substitution sequences described in Example 11), this method of the invention, when tested in the system exemplified in Example 11 included herein, increases the expression of a protein or polypeptide by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% in the organism or at least in one tissue or organelle of the organism. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the expression of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operatively linked to the nucleotide sequence of interest). Expression is preferably measured by measuring the conversion of any suitable alkaline phosphatase substrate, and the expression level is relative to the expression level of the nucleotide sequence of interest, measured under the same conditions except that the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 in the expression vector has been replaced by an alternative sequence, preferably one of those alternative sequences described in Example 11.
[0283] Twenty-fifth aspect
[0284] In a twenty-fifth aspect, the present invention provides a method for transcription and optionally purification of the obtained transcript, comprising the following steps:
[0285] a) Provide a nucleic acid construct according to aspect 20, comprising the nucleotide sequence of interest of the present invention; and
[0286] b) Contacting cells with the nucleic acid construct to obtain transformed cells; and
[0287] c) Inducing the transformed cells to produce a transcript containing the nucleotide sequence of interest; and optionally...
[0288] d) Purify the resulting transcript.
[0289] In a preferred embodiment of this method according to the invention, a nucleic acid construct as defined above in the twentieth aspect is used. The method of the invention can be an in vitro or ex vivo method. The method of the invention can be applied to cell culture, organism culture, or tissue culture. The method of the invention can be applied to nucleic acid-based inoculation and / or gene therapy, preferably in mammals, particularly humans. Covered within the invention are methods for treatment comprising or composed of this aspect, wherein the nucleotide sequence of interest encodes a therapeutic transcript. The invention also relates to constructs of the twentieth aspect of the invention for use in treatment, wherein the nucleotide sequence of interest encodes a therapeutic transcript. Furthermore, the invention relates to constructs of the twentieth aspect of the invention for use in the manufacture of pharmaceutical preparations, wherein the nucleotide sequence of interest encodes a therapeutic transcript.
[0290] According to step b), the technician is able to transform the cells. Transformation methods used in step b) include, but are not limited to, transferring purified DNA via cationic lipid reagents and polyethyleneimine (PEI), calcium phosphate coprecipitation, microparticle bombardment, electroporation and microinjection of protoplasts, or using silica fibers to facilitate the permeation and transfer of DNA into host cells.
[0291] In step c), the transformed cells are made to produce a transcript of the nucleotide sequence of interest, and optionally the obtained transcript is subsequently recovered. For example, the transformed cells may be subjected to conditions that lead to transcription of the nucleotide sequence of interest. Techniques for transcribing the nucleotide sequence of interest are well understood by those skilled in the art. The method of the invention also includes methods in which the transformed cells are not required to undergo specific conditions that lead to transcription of the nucleotide sequence of interest, but where the nucleotide sequence of interest is transcribed automatically (e.g., constitutively).
[0292] The purification steps depend on the obtained transcript. The term "isolation" indicates that the transcript was found under conditions different from its native environment. In a preferred form, the isolated transcript is substantially free of other cellular components, particularly other homologous cellular components, such as homologous proteins. Transcripts in a pure form greater than 40% are preferably provided, more preferably greater than 60%. Even more preferably, highly purified transcripts are provided, i.e., with a purity greater than 80%, more preferably greater than 95%, and even more preferably greater than 99%, as determined by Northern blotting.
[0293] Preferably, the method of this aspect of the invention increases the transcription of the nucleotide sequence of interest. Preferably, in the expression system, transcription levels are established using an expression construct according to a second aspect of the invention, which comprises a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operatively linked to the nucleotide sequence of interest. Preferably, transcription of the nucleotide sequence of interest is detected by a suitable assay, such as RT-qPCR. Preferably, compared to methods that differ only in the use of constructs in step a, wherein the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 has been replaced by a substituted sequence, preferably one of those substituted sequences as in the technique of Example 11, the method of the present invention preferably increases transcription by at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% when tested in systems exemplified in Example 11, which is included herein. More specifically, preferably in mammalian cell systems, most preferably in CHO cells, the transcription of the nucleotide sequence of interest encoding secreted alkaline phosphatase (SeAP) is measured using a pcDNA3.1 expression vector (which contains a nucleotide sequence having at least 50% identity with SEQ ID NO: 88 and operatively linked to the nucleotide sequence of interest). Transcription is preferably measured using RT-qPCR, and the transcription level is relative to the transcription level of the nucleotide sequence of interest, which is measured under the same conditions except that the nucleotide sequence having at least 50% identity with SEQ ID NO: 88 in the expression vector used has been replaced by an alternative sequence, preferably one of those alternative sequences described in Example 11.
[0294] Twenty-sixth aspect
[0295] In a twenty-sixth aspect, the present invention provides the application of nucleic acid molecules according to the nineteenth aspect of the invention and / or the application of nucleic acid constructs according to the twentieth aspect of the invention and / or the application of expression vectors according to the twenty-first aspect of the invention and / or the application of cells according to the twenty-second aspect of the invention for transcription of nucleotide sequences of interest and / or expression of proteins or polypeptides of interest.
[0296] Twenty-seventh aspect
[0297] In a twenty-seventh aspect, the present invention provides nucleic acid molecules according to the nineteenth aspect of the invention and / or nucleic acid constructs according to the twentieth aspect of the invention and / or expression vectors according to the twenty-first aspect of the invention and / or cells according to the twenty-second aspect of the invention for use as pharmaceutical agents. The invention also relates to treatment methods comprising administering nucleic acid molecules according to the nineteenth aspect of the invention and / or nucleic acid constructs according to the twentieth aspect of the invention and / or expression vectors according to the twenty-first aspect of the invention and / or cells according to the twenty-second aspect of the invention, wherein preferably the administration is administered to mammals, more preferably to humans. Preferably, the treatment is based on nucleic acid inoculation and / or gene therapy, preferably in mammals, most preferably in humans. Furthermore, the present invention relates to the application of nucleic acid molecules according to the nineteenth aspect of the invention and / or the application of nucleic acid constructs according to the twentieth aspect of the invention and / or the application of expression vectors according to the twenty-first aspect of the invention and / or the application of cells according to the twenty-second aspect of the invention for the preparation of pharmaceutical agents. Preferably, the pharmaceutical agent is for nucleic acid-based inoculation and / or gene therapy, preferably in mammals, most preferably in humans.
[0298] definition
[0299] As used herein, the term "nucleic acid" refers to a naturally occurring or synthetic oligonucleotide or polynucleotide, whether DNA or RNA or a DNA-RNA hybrid, single-stranded or double-stranded, sense or antisense, capable of hybridizing to a complementary nucleic acid via Watson-Crick base pairing. The nucleic acids of the present invention are preferably modified by including at least 1, 2, 3, 4, 5, 10, 20, 30, or 50 nucleotide mutations (compared to their naturally occurring counterparts). Preferably, the nucleic acids of the present invention do not occur in nature. The nucleic acids of the present invention may also include nucleotide analogs (e.g., BrdU) and nonphosphodiester nucleotide bonds (e.g., peptide nucleic acid (PNA) or thiodiester bonds). In particular, nucleic acids may include, but are not limited to, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, ssRNA, dsRNA, non-coding RNA, hnRNA, pre-mRNA, mature mRNA, or any combination thereof. As used herein, the terms "nucleic acid sequence" and "nucleotide sequence" are interchangeable and have their usual meanings in the art. The terms above refer to DNA or RNA molecules in single-stranded or double-stranded form. "Separated nucleic acid sequence" refers to a nucleic acid sequence that is no longer in the natural environment from which it was isolated. Nucleic acid molecules are represented by nucleotide sequences. Furthermore, elements (e.g., but not limited to expression-enhancing elements and transcriptional regulatory elements) are represented by nucleotide sequences.
[0300] A “recombinant construct” (or chimeric construct) refers to any nucleic acid sequence or molecule, particularly a nucleic acid sequence, molecule, or gene, that is not normally found in species in nature, containing one or more portions of the nucleic acid sequence that are not associated with each other in nature. For example, a recombinant construct may contain a promoter that is not associated in nature with some or all of the transcribed or other regulatory regions contained within the recombinant construct. The term “recombinant construct” is understood to include expression constructs in which a promoter or expression regulatory sequence is operatively linked to one or more sense sequences (e.g., coding sequences) or to antisense (inverse complement of the sense strand) or inverted repeat sequences (sense and antisense, thereby forming double-stranded RNA after transcription), or to any other sequence encoding a functional RNA molecule.
[0301] A “nucleic acid construct” is defined as a polynucleotide isolated from a naturally occurring gene or modified to contain a polynucleotide fragment that is bound or juxtaposed in a manner not normally found in nature. Optionally, the polynucleotide present in the nucleic acid construct is operatively linked to one or more control sequences that direct the production or transcription of a nucleotide sequence of interest and / or the expression of a peptide or polypeptide of interest in a cell or individual.
[0302] In this document, “vector” or “plasmid” is understood to refer to an artificial (usually circular) nucleic acid molecule derived from the use of recombinant DNA technology and used to deliver exogenous DNA into a host cell. Vectors typically contain further genetic elements to facilitate their use in molecular cloning, such as, for example, selectivity markers, multiple cloning sites, etc. (see below). Nucleic acid constructs can also be part of a recombinant viral vector for protein expression: in plants or plant cells (e.g., vectors derived from cauliflower mosaic virus, CaMV, or tobacco mosaic virus, TMV) or in mammalian organisms or mammalian cell systems (e.g., vectors derived from Moloney mouse leukemia virus (MMLV; a retrovirus), lentivirus, adeno-associated virus (AAV), or adenovirus (AdV)).
[0303] "Transformed cell" is a term that refers to a new individual cell (or organism) resulting from the introduction of at least one nucleic acid molecule, particularly a chimeric or recombinant construct encoding a desired protein or nucleic acid sequence (which, after transcription, produces antisense RNA for silencing a target gene / gene family). The host cell can be a plant cell, bacterial cell (e.g., Agrobacterium strains), fungal cell (including yeast cells), animal cell (including insects, mammals), etc. Transformed cells can contain nucleic acid constructs as extrachromosomal (free) replicating molecules, as non-replicating molecules, or recombinant constructs integrated into the nuclear or organelle DNA of the host cell. As used herein, the term "organism" encompasses all organisms, including those with more than one cell type, i.e., multicellular organisms, and includes multicellular fungi.
[0304] "Transformation" and "transformed" refer to the transfer of nucleic acid sequences, typically nucleic acid sequences containing a recombinant construct or gene of interest (GOI), into the nuclear genome of a cell to produce "transgenic" cells or organisms containing the transgene. The introduced nucleic acid sequence is usually, but not always, integrated into the host genome. When the introduced nucleic acid sequence is not integrated into the host genome, it is said to be "transfection," "transiently transfected," and "transfected." For the purposes of this patent specification, the terms "transformation," "transiently transfected," and "transfection" are used interchangeably and refer to the stable or transient presence of the nucleic acid sequence entering the cell or organism. When the cell is a bacterial cell, the above terms generally refer to an extrachromosomal self-replicating vector with selective antibiotic resistance.
[0305] In the case of amino acid sequences or nucleic acid sequences, "sequence identity" or "identity" is defined herein as the relationship between two or more amino acid (peptide, polypeptide, or protein) sequences or two or more nucleic acid (nucleotide, polynucleotide) sequences, as determined by sequence comparison. In the art, "identity" also means the degree of sequence correlation between amino acid or nucleotide sequences (as the case may be), as determined by matching between such sequence chains. Within the scope of this invention, sequence identity with a specific sequence represented by a particular SEQ ID NO preferably refers to sequence identity over the entire length of the specific polypeptide or polynucleotide sequence indicated by the particular SEQ ID NO. However, sequence identity with a specific sequence represented by a particular SEQ ID NO may also mean assessing sequence identity relative to a portion of the SEQ ID NO. A portion may mean at least 50%, 60%, 70%, 80%, 90%, or 95% of the length of the SEQ ID NO. The sequence information provided herein should not be interpreted so narrowly as requiring the inclusion of misidentified bases. Those skilled in the art can identify such misidentified bases and know how to correct such errors.
[0306] Any nucleotide sequence capable of hybridizing to the nucleotide sequence of the present invention is defined as a portion of the cis-acting element of the present invention. Strict hybridization conditions are defined herein as conditions at about 65°C, in a solution containing about 1M salt, preferably 6x SSC, or any other solution with equivalent ionic strength, such that at least 25, preferably 50, 75, or 100, and most preferably 150 or more nucleotides of nucleic acid sequence are hybridized, and washed at 65°C in a solution containing about 0.1M salt, preferably 0.2x SSC, or any other solution with equivalent ionic strength. Preferably, hybridization is performed overnight, i.e., at least 10 hours, and preferably with washing for at least 1 hour after at least two changes of the washing solution. These conditions generally result in sequence-specific hybridization with about 90% or greater sequence identity or at least 90% sequence identity. Mild hybridization conditions are defined herein as conditions that hybridize a nucleic acid sequence of at least 50, preferably 150 or more nucleotides in a solution containing about 1 M salt, preferably 6x SSC, or any other solution with equivalent ionic strength, at about 45°C or 45°C, followed by washing at room temperature in a solution containing about 1 M salt, preferably 6x SSC, or any other solution with equivalent ionic strength. Preferably, hybridization is performed overnight, i.e., at least 10 hours, and preferably, washing is performed for at least 1 hour, with the washing solution changed at least twice. These conditions will generally enable specific hybridization of sequences with up to 50% sequence identity. Those skilled in the art will be able to modify these hybridization conditions to specifically identify sequences with 50% to 90% identity.
[0307] "Identity" can be readily calculated by known methods, including but not limited to those described in: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heine, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991; and Carillo, H. and Lipman, D., SIAM J. Applied Math., 48:1073 (1988).
[0308] Preferred methods for determining identity are designed to provide the maximum match between the tested sequences. Methods for determining identity and similarity are compiled in publicly available computer programs. Preferred computer program methods for determining identity and similarity between two sequences include, for example, the GCG package (Devereux, J., et al., Nucleic Acids Research 12(1):387(1984)), BestFit, BLASTP, BLASTN, and FASTA (Altschul, S., et al., J.Mol.Biol.215:403-410(1990)). The BLAST X program is publicly available from NCBI and other sources (BLAST Manual, Altschul, S., et al., NCBI NLM NIH Bethesda, MD20894; Altschul, S., et al., J.Mol.Biol.215:403-410(1990). The well-known Smith-Waterman algorithm can also be used to determine identity.
[0309] Preferred parameters for peptide sequence comparison include the following algorithms: Needleman / Wunsch algorithm (Needleman and Wunsch), J. Mol. Biol. 48: 443-453 (1970); comparison matrix; BLOSSUM62, from Hentikov and Hentikov, Proc. Natl. Acad. Sci. USA. 89: 10915-10919 (1992); gap penalty: 12; and gap length penalty: 4. Programs available for these parameters are publicly available as the "Ogap" program from the Genetics Computer Group, Madison, WI. The above parameters are the default parameters for amino acid comparison (along with no penalty for terminal vacancies).
[0310] Preferred parameters for nucleic acid comparison include the following algorithm: Needleman / Wunsch algorithm, J. Mol. Biol. 48: 443-453 (1970); comparison matrix: match = +10, mismatch = 0; gap penalty: 50; gap length penalty: 3. This is available as a gap program from the Genetics Computer Group in Madison, Wisconsin. The above are the default parameters for nucleic acid comparison.
[0311] For nucleic acid comparison, the preferred procedure and parameters for assessing identity are calculated using the EMBOSS Needle nucleotide alignment algorithm with the following parameters: the complete DNA matrix, which has the following vacancy penalties: open = 10; extended = 0.5, as performed in Example 9.
[0312] In cases where the source is a specific naturally occurring gene or sequence, the term "derived from" is defined herein as a gene or sequence that has been chemically synthesized and / or isolated and / or purified from a naturally occurring gene or sequence. Techniques for the chemical synthesis, isolation, and / or purification of nucleic acid molecules are well known in the art. Generally, a derived sequence is a partial or a portion of a naturally occurring gene or sequence. Optionally, the derived sequence contains nucleic acid substitutions or mutations, preferably resulting in a sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical in length to a portion of a naturally occurring gene, or a sequence or partial sequence.
[0313] As used herein, "polypeptide" refers to any peptide, oligopeptide, polypeptide, gene product, expression product, or protein. Polypeptides consist of a sequence of amino acids. The term "polypeptide" encompasses naturally occurring or synthetic molecules. Polypeptides are represented by their amino acid sequence. Polynucleotides are represented by their nucleotide sequence. Polypeptides are represented by their amino acid sequence.
[0314] When used to describe the relationship between a given nucleic acid or polypeptide molecule and a given host organism or host cell, the term "homologous" should be understood to mean that, in nature, the nucleic acid or polypeptide molecule is produced by a host cell or organism of the same species, preferably the same variant or strain. If homologous to the host cell, the nucleic acid sequence of interest, preferably encoding a polypeptide, will typically be operatively linked to another promoter sequence, or, if applicable, another secretion signaling sequence and / or termination sequence (different from its natural environment).
[0315] When used to describe the correlation between two nucleic acid sequences, the term "homologous" means that a single-stranded nucleic acid sequence can hybridize with a complementary single-stranded nucleic acid sequence. The degree of hybridization can depend on many factors, including the degree of identity between sequences as discussed later, and hybridization conditions such as temperature and salt concentration. Preferably, the region of identity is greater than 5 bp, and more preferably, the region of identity is greater than 10 bp.
[0316] When used to describe the relationship between a given (recombinant) nucleic acid or polypeptide molecule and a given host organism or host cell, the term "heterogeneous" should be understood to mean a nucleic acid or polypeptide molecule from a foreign cell that is not naturally present as part of the organism, cell, genome, or DNA or RN sequence in which it is present, or whose location in the cell or genome or DNA or RN sequence differs from its location in nature. For the cell in which they are introduced, the heterologous nucleic acid or protein is not endogenous but has been acquired from another cell or synthesized or recombined.
[0317] When used to indicate the correlation between two nucleic acid sequences, the term "heterologous sequence" or "heterologous nucleic acid" refers to a heterologous sequence or nucleic acid that is not a naturally found, operable, neighboring sequence that can be linked to another sequence. As used herein, the term "heterologous" may mean "recombinant." A "recombinant" refers to a genetic entity that differs from genetic entities commonly found in nature. When applied to nucleotide sequences or nucleic acid molecules, this means that the nucleotide sequence or nucleic acid molecule is the product of various combinations of cloning, restriction, and / or ligation steps and other methods (which result in constructs that differ from sequences or molecules found in nature).
[0318] "Operably linked" is defined herein as a configuration in which a control or regulatory sequence is appropriately placed at a position relative to a nucleotide sequence of interest that preferably encodes a polypeptide of interest, to control or regulate the sequence, directing or influencing the transcription and / or production or expression of the nucleotide sequence of interest (preferably encoding the peptide or polypeptide of the invention in a cell and / or in a host). For example, if a promoter is capable of initiating or regulating transcription or expression of a coding sequence, then the promoter is operably linked to the coding sequence, in which case the coding sequence should be understood as being "under the control of the promoter." When one or more nucleotide sequences and / or elements contained in a construct are defined herein as "configured to be operably linked to an optional nucleotide sequence of interest," the nucleotide sequences and / or elements are understood to be operably linked to the nucleotide sequence of interest after being configured in the construct such that the nucleotide sequence of interest is present in the construct.
[0319] A "promoter" is a nucleic acid sequence located upstream or at the 5′ end of the translation start codon within the reading frame (or protein-coding region) of a gene, and is involved in the recognition and binding of RNA polymerase II and other proteins (trans-acting transcription factors) to initiate transcription. The term promoter refers to a nucleic acid fragment used to control the transcription of one or more genes, located upstream of the transcription start site of the gene, and structurally defined by binding sites for DNA-dependent RNA polymerase, transcription start sites, and any other DNA sequences, including, but not limited to, transcription factor binding sites, repressor and activator protein binding sites, and any other sequences of nucleotides known to those skilled in the art, to directly or indirectly regulate the amount of transcription by the promoter. A promoter does not include a transcription start site (TSS) but is located at the nucleotide-1 end of the transcription site, and does not include such nucleotide sequences that become untranslated regions such as the 5′-UTR in the transcribed mRNA. The promoters of the present invention can be tissue-specific promoters, tissue-preferred promoters, cell-type-specific promoters, inducible promoters, and constitutive promoters. Tissue-specific promoters are promoters that initiate transcription only in certain tissues and refer to DNA sequences that provide recognition signals to RNA polymerases and / or other factors required to accurately initiate transcription and / or control the expression of coding sequences within certain tissues or in certain cells of the aforementioned tissues. Tissue-specific expression can occur only in individual tissues or in combinations of tissues. Tissue-preferred promoters are promoters that preferentially initiate transcription in certain tissues. Cell-type-specific promoters are promoters that primarily drive expression in certain cell types. Inducible promoters are promoters that, in response to an inducer, are capable of activating the transcription of one or more DNA sequences or genes. DNA sequences or genes are not transcribed in the absence of an inducer. Activation of inducible promoters is established by applying an inducer. Constitutive promoters are promoters that are active under many environmental conditions and in many different tissue types. Preferably, in an expression system, an expression construct is used that contains the promoter operatively linked to the nucleotide sequence of interest, and the ability to initiate transcription is established using a suitable assay, such as RT-qPCR or Northern blotting. Compared to transcription using a construct (the only difference being that it does not contain the promoter), if transcripts can be detected or if an increase in transcript levels of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000% is observed, then the promoter is capable of initiating transcription.In a further preferred embodiment, in the expression system, the ability to initiate expression is established using an expression construct (which includes the promoter operatively linked to a nucleotide sequence encoding a protein or polypeptide of interest). Preferably, the protein or polypeptide of interest is a secreted protein or polypeptide, and the expression of the protein or polypeptide of interest is detected by suitable assays, such as ELISA, Western blotting, or any suitable protein identification and / or quantification assay known to those skilled in the art, depending on the identity of the protein or polypeptide of interest. Expression using the construct (which differs only in that it does not contain the promoter) is considered to be initiated if the protein or polypeptide of interest can be detected, or if an increase in expression level of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500%, or 2000%. As the first and second promoters of the present invention, induced or constitutive promoters or combinations thereof may be used in the present invention.
[0320] An "intron" is a nucleotide sequence within the initial RNA transcript that is removed by RNA splicing or intron splicing, resulting in the final mature RNA product. The presence of intron splicing can be assessed using any suitable method known to those skilled in the art, such as, but not limited to, reverse transcriptase polymerase chain reaction (RT-PCR), followed by size or sequence analysis of the RT-PCR product. Preferably, the nucleotide sequence is an intron if, by RNA splicing and using a suitable test for detecting intron splicing (as indicated above), at least 2%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the initial RNA loses this sequence. Preferably, the intron comprises a splice site GT at the 5' end of the nucleotide sequence and a splice site AG at the 3' end of the nucleotide sequence, the splice site AG being preceded by a pyrimidine-rich nucleotide sequence or a polypyrimidine segment, and optionally separated from the splice site AG by 1-50 nucleotides. The intron may further contain a branching site, which is the sequence YTNAY contained at the 5' side of the polypyrimidine segment. This branching site may have the nucleotide sequence CYGAC. An "intron sequence" is understood to be the nucleotide sequence of at least a portion of the intron.
[0321] "Expression" will be understood to include any step involved in the production of peptides or polypeptides, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.
[0322] Optionally, a promoter represented by a nucleotide sequence present in the nucleic acid construct may be operatively linked to another nucleotide sequence encoding a peptide or polypeptide as defined herein.
[0323] The expression vector can be any vector that can be readily subjected to recombinant DNA methods and can result in the expression of the nucleotide sequence of the coding polypeptide of the present invention in cells and / or in the host.
[0324] As used herein, "5′-UTR" is a sequence that begins with nucleotide 1 of the mRNA and ends with nucleotide -1 of the start codon. It is possible that the regulatory portion of the promoter is contained within a nucleotide sequence that becomes the 5′-UTR; however, in this case, the 5′-UTR is still not part of the promoter as defined herein.
[0325] The term "control sequence" is defined herein as encompassing all components that are essential or advantageous for the expression of a polynucleotide or polypeptide. Each control sequence may be native or exogenous for a nucleic acid sequence having or encoding a polynucleotide or polypeptide. Such control sequences include, but are not limited to, leader sequences, optimal translation initiation sequences (as described in Kozak, 1991, J. Biol. Chem. 266: 19867-19870), polyadenylation sequences, propeptide sequences, pre-propeptide sequences, promoters, signal sequences, and transcription terminators. Control sequences include at least a promoter and transcription and translation termination signals.
[0326] Control sequences can provide linkers to introduce specific restriction sites, which facilitate the connection of control sequences to the coding regions of nucleic acid sequences encoding polypeptides.
[0327] The control sequence can be a suitable promoter sequence (a nucleic acid sequence recognized by the host cell for the expression of the nucleic acid sequence). The promoter sequence contains a transcription control sequence that mediates the expression of the polypeptide. The promoter can be any nucleic acid sequence that exhibits transcriptional activity in cells including mutant, truncated, and heterozygous promoters, and can acquire genes that encode extracellular or intracellular polypeptides (homological or heterologous to the aforementioned cells).
[0328] The control sequence can also be a suitable transcription terminator sequence, a sequence recognized by the host cell to terminate transcription. The terminator sequence is operatively linked to the 3′ end of the nucleic acid sequence of interest, preferably encoding a polypeptide of interest. Any terminator that is functional in the cell can be used in this invention.
[0329] The control sequence can also be a suitable leader sequence (the untranslated region of mRNA, which is important for translation through the host cell). The leader sequence is operatively linked to the 5′ end of the nucleic acid sequence of interest, preferably encoding a polypeptide of interest. Any leader sequence that is functional in the cell can be used in this invention.
[0330] The control sequence can also be a polyadenylated sequence (a sequence operatively linked to the 3′ end of a nucleic acid sequence). Furthermore, when transcribed, this sequence is recognized by the host cell as a signal to add adenine residues to the transcribed mRNA. Any polyadenylated sequence that is functional in the cell can be used in this invention.
[0331] In this document and its claims, the verb "comprising" and its inflections are used in their non-limiting sense to indicate inclusion of items following the word, but not excluding items not specifically mentioned. Furthermore, the verb "consisting of" may be replaced by "consisting substantially of," meaning that a nucleic acid construct, vector, or cell product or composition, or nucleic acid molecule, peptide, or polypeptide, as defined herein, may contain additional components different from the specifically identified components; said additional components do not alter the distinctive features of the invention. Moreover, the use of the indefinite articles "a" or "an" to refer to an element does not exclude the possibility of more than one element, unless the context explicitly requires that only one element exists. Therefore, the indefinite articles "a" or "an" generally mean "at least one."
[0332] All patents and literature references cited in this specification are incorporated herein by reference in their entirety.
[0333] Table 1: Sequence Identification
[0334]
[0335] Attached Figure Description
[0336] Figure 1 A schematic diagram of the intron promoter construct and different transcripts. The construct contains two promoters. Transcription from promoter 1 produces a primary transcript containing the intron, which has the complete sequence of promoter 2 and adjacent 5' and 3' splicing sites. Following intron splicing, the primary transcript results in the production of mRNA of the intron (transcription 1) without encoding a "gene". Transcription from promoter 2 also results in the production of mRNA (transcription 2) encoding the same "gene".
[0337] Figure 2. Schematic diagram of the EEE1(A) and EEE2(B) elements, illustrating some characteristics of the UBC and CCT8 genes relative to their promoter activity in the genomic context. Characteristics include predicted transcription start site (TSS), 5'-UTR, exon, and intron information.
[0338] Figure 3 A schematic diagram of an expression vector for an Ig light chain (igLC) with an EEE1 sequence integrated upstream of the CMV promoter.
[0339] Figure 4. Comparison of HuMab1 produced by CHO-S pools stably transfected with the reference or EEE1 construct. Expression vectors without (A) and with (B) additional expression regulatory elements. Bars indicate the average exhaustion titer of the four pools derived from two independent transfections. Pools were grown in 125 ml shake flasks in 30 ml CD FortiCHO selective medium.
[0340] Figure 5 Comparison of HuMab1 produced by CHO-S pools stably transfected with the reference or EEE2 construct. Bars represent the average depletion titer from four pools derived from two independent transfections. Pools were grown in 30 ml CD FortiCHO selective medium in 125 ml shake flasks.
[0341] Figure 6 Analysis was performed on HuMab1 cells produced by stable transfection of the top-10CHO-S clone cell line with the EEE1-TEE construct containing EEE1. Cells were grown in 125 ml shake flasks and in 30 ml CD FortiCHO selective medium.
[0342] Figure 7. Comparison of HuMab2 production from CHO-S pools stably transfected with EEE1 in a reference vector (A, left panel) and a vector with additional regulatory elements (A, right panel). Bars represent the average depletion titer of the four pools derived from two independent transfections. Pools were grown in 30 ml CD FortiCHO selective medium in 125 ml shake flasks. (B) Analysis of HuMab2 production by top-12 reference and EEE1-TEE CHO-S clone cell lines. Cells were grown in 30 ml CD FortiCHO selective medium in 125 ml shake flasks.
[0343] Figure 8Comparison of SeAP activity in depleted medium of CHO-S pools stably transfected with the Reference or EEE1 construct. Expression vectors without (left panel) and with (right panel) additional expression regulatory elements. Bars represent the average activity of the four pools derived from two independent transfections, measured using the SEAP reporter gene assay kit (Abcam). Pools were grown in 125 ml shake flasks in 30 ml CD FortiCHO selective medium.
[0344] Figure 9 Comparison of SeAP activity in depleted medium of CHO-S pools stably transfected with constructs containing different forms of EEE1 elements. The bars represent the average activity of four pools derived from two independent transfections, measured using the SEAP reporter gene assay kit (Abcam). Pools were grown in 125 ml shake flasks in 30 ml CD FortiCHO selection medium.
[0345] Figure 10. 5'-RACE amplification of the 5' end of the heavy chain transcript from a CHO-S clone stably transfected with the EEE1-TEE construct expressing HuMab1 (A). Two bands were detected on an agarose gel, corresponding to transcripts generated via the CMV promoter (transcript 1) and the UBC promoter (transcript 2). The size difference between transcript 1 and transcript 2 is illustrated in the schematic diagram of the intron promoter construct and the different transcripts (B). The construct contains two promoters, UBC and CMV. The CMV promoter is ligated to the TEE sequence, which also contains a short intron. Transcription via the CMV promoter after intron splicing results in the production of mRNA with the TEE as the 5'-UTR (transcript 1). The UBC promoter is ligated to a partial UBC 5'-UTR region, which contains the 5' splice donor site and precedes the CMV promoter. Transcription via the UBC promoter results in the production of mRNA with the UBC 5'-UTR sequence (transcript 2). Large introns spliced from the primary transcript move from the 5'-splicing donor sequence in the UBC sequence to the 3'-splicing acceptor site in the TEE and contain the complete CMV sequence.
[0346] Figure 11 The effect of EEE in Pichia pastoris expressing recombinant human interleukin-8. Each bar represents the average expression of 10 independent clones.
[0347] The invention will be explained in more detail in the following embodiments section, with reference to the accompanying drawings. The embodiments are for illustrative purposes only and are not intended to limit the invention in any way. Example
[0348] The expression-enhancing element represented by SEQ ID NO: 1 is based on the Chinese hamster (U. spp.) ubiquitin-C (UBC) gene. It contains the predicted promoter sequence and a portion of the 5'-untranslated region (Fig. 2A). The expression-enhancing element represented by SEQ ID NO: 2 is based on the human CCT8 gene (containing the TCP1 chaperone protein, subunit 8). It contains the predicted promoter sequence, a portion of the 5'-untranslated region, a short sequence encoding 27 amino acids, and a portion of the first intron (Fig. 2B).
[0349] Example 1
[0350] The expression plasmid was constructed based on the pcDNA3.1 expression vector (SEQ ID NO: 71). The vector was modified by removing f1-ori. The coding sequences of the IgG1 (HuMab1) heavy chain (represented by the sequence, which is at least 96% identical to SEQ ID NO: 68) and light chain (represented by the sequence, which is at least 99% identical to SEQ ID NO: 67) genes were inserted into the above vector (the light chain coding sequence was inserted into the vector represented by SEQ ID NO: 65 and the heavy chain coding sequence was inserted into the vector represented by SEQ ID NO: 66), thereby generating a reference construct. To generate the EEE1 (expression enhancement element 1) vector, the EEE1 sequence (SEQ ID NO: 1) was inserted upstream of the CMV promoter (SEQ ID NO: 57). Figure 3 EEE2 (SEQ ID NO: 2) was introduced in a similar manner, resulting in the EEE2 vector. A vector with additional expression regulatory elements was generated by replacing pcDNA3.15'-UTR with SEQ ID NO: 19 (transcriptional enhancement element; TEE).
[0351] Following the manufacturer's instructions, CHO-S cells (Life Technologies) were maintained. Repeated transfections were performed using 3E7 cells, 50 μg linear DNA, and FreeStyle MAX reagent (Life Technologies). The transfected pools were divided into two portions and then selected in CD FortiCHO cultures supplemented with 8 mM glutamine and 800 μg / ml G418. The selected pools were seeded in 30 ml of the same medium at a density of 3E5 cells / ml in 125 ml shake flasks. The HuMab1 depletion titer was determined by ELISA (Figure 4). For precise determination, the depletion titer of the reference pool was too low. This indicated poor antibody expression. The EEE1 pool produced approximately 1 μg / ml (Figure 4A). Similar effects were observed in vectors with additional expression regulatory elements. In these vectors, the introduction of EEE1 increased production from approximately 0.2 μg / ml to 9–12 μg / ml in stable pools from three independent transfection experiments (Figure 4B). These data indicate that poorly expressed antibodies can be expressed at significantly higher levels by introducing the EEE1 element.
[0352] Example 2
[0353] The effects of introducing the EEE2 element were investigated in vectors with additional expression regulatory elements (see Example 1, Figure 4B). CHO-S cells were transfected with a reference or with an EEE2 construct (as previously described). The antibody depletion titer of the stably transfected EEE2 pool was equal to that of the reference pool ( Figure 5 It is more than 20 times higher than that.
[0354] Example 3
[0355] Previously generated EEE1-TEE cells and a reference cell line were seeded in six 96-well plates in CD FortiCHO selective medium at a density of 0.5 cells / well. The reference cells showed impaired growth and thus no HuMab1 production compared to the EEE1-TEE clones. The EEE1-TEE clones showed normal growth and HuMab1 production (see below). One hundred clones of the EEE1-TEE line were evaluated for HuMab1 production in microtiter plates. The ten clones with the highest specific productivity were expanded to 125 ml shake flasks. The clones were seeded in 30 ml of CD FortiCHO selective medium at a density of 3E5 cells / ml, and the HuMab1 depletion titer was determined by ELISA. Figure 6 Cloning yielded up to 0.25 μg / ml HuMab1. These data indicate that EEE1 can promote the generation of clone lines and result in clone lines with relevant expression levels.
[0356] Example 4
[0357] The copy number of antibodies expressing the EEE containing the clone was determined. The PrimerExpress program (Life Technologies) was used to design Taqman primers and probes specific for the heavy and light chains of HuMab1 and β-2 microglobulins. Primers were bound in a triple Taqman assay to measure gene copies in gDNA samples from the EEE1-TEE HuMab1 clone and pool. Gene copy number was compared with HuMab1 titer (Table 2). Clonal cell lines producing similar HuMab1 titers had different numbers of light and heavy chain gene copies (clones 1 and 2). Furthermore, clones producing very different HuMab1 titers had similar gene copy numbers (clones 3 and 4). In the pool, relatively high numbers of light and heavy chain genes matched relatively low expression levels. These data (Table 2) indicate that there is no correlation between the copy number of EEE-containing genes and HuMab1 expression levels.
[0358] Table 2: IgG1 titer and gene copy number
[0359]
[0360] Example 5
[0361] The HuMab1 heavy and light chain genes of the previous embodiments were replaced with the heavy and light chain genes (SEQ ID NO: 69 and 70) encoding a biosimilar antibody (HuMab2, derived from DrugBank accession number DB00072). The construct was used to generate the previously described CHO-S pool. The depletion titer was determined using ELISA. Data (6.3 μg / ml, without enhancement elements) indicated that this antibody was produced at a higher level than the antibody from the previous embodiments. The introduction of no additional expression regulatory elements into EEE resulted in a 3.7-fold increase (Fig. 7A, left panel), while in the modified vector, the increase was 7-fold (Fig. 7A, right panel). The data also indicate a synergistic effect between EEE and additional expression regulatory elements, as the introduction of independent additional expression regulatory elements resulted in a 40% increase. Clonal lines were isolated from the reference and EEE1-TEE pools (as previously described). The best EEE1-TEE clone produced a 3-fold higher HuMab2 titer compared to the best reference clone (Fig. 7B). These data demonstrate that the EEE1 element can be successfully applied to enhance the expression of recombinant proteins from stable cell lines.
[0362] Example 6
[0363] The HuMab1 light chain gene from the construct of Example 1 was replaced with the gene encoding secreted alkaline phosphatase (SeAP, SEQ ID NO: 72). The construct was used to generate the CHO-S pool as previously described. SeAP activity was measured in depleted medium using the SEAP reporter gene assay kit (Abcam). The EEE1 pool showed twice the activity compared to the reference pool. Figure 8 Compared to the reference pool, the increase in the EEE1-TEE pool was almost fourfold. These data indicate that EEE1 enhances the expression of single subunit non-antibody proteins in transfected cell lines.
[0364] Example 7
[0365] All SeAP constructs used in Example 6 contained a CMV promoter. Two TEE vector variants were prepared containing a human EF-1α promoter instead of CMV (SEQ ID NO: 79). The constructs were used to generate the CHO-S pool (as previously described). SeAP activity was measured in depleted medium using the SEAP reporter gene assay kit (Abcam). Compared to a reference EF-1α promoter pool without EEE1, the EEE1-TEE pool with EF-1α as an intron promoter produced 2.8-fold higher SeAP activity. These data indicate that EEE1 enhances protein expression in intron promoter constructs when the intron promoter is not a CMV promoter, such as the EF-1α promoter.
[0366] Example 8
[0367] The EEE1 element of the EEE1 SeAP expression vector is replaced by the following variants: 1. EEE1-80 represented by SEQ ID NO: 60 with a 290 bp truncation from the 5' end; 2. EEE1-60 represented by SEQ ID NO: 61 with a 580 bp truncation from the 5' end; 3. EEE1-50 represented by SEQ ID NO: 62 with a 725 bp truncation from the 5' end; 4. EEE1-Xt represented by SEQ ID NO: 59 with an 800 bp extension from the genomic C. griseus UBC sequence at the 5' end; 5. EEE1-SL (SEQ ID NO: 63) with splice donor and recipient sites for all major predicted mutations. In the supernatant of cells transfected with the EEE1 element, SeAP activity was set at 100%, decreasing to 39% activity in the absence of the EEE1 element. Figure 9The 5' truncation of the EEE1-80 and EEE1-60 constructs gradually reduced activity, but still showed enhanced activity compared to the construct without EEE. The EEE1-50 element reduced SeAP activity to 40%, similar to the construct without EEE. The EEE1-Xt construct showed an almost 40% increase in activity compared to the EEE1 construct. The data suggest that sequences with greater than 50% identity relative to EEE1 can serve as expression-enhancing elements. The almost 40% increase in activity produced by the EEE1-Xt construct compared to the EEE1 construct indicates that additional enhancer sequences are located in regions upstream of the genomic sequence from which EEE1 is derived. The 4nt mutation in the EEE1-SL construct severely impaired the activity of the EEE1 element, preventing proper intron splicing and resulting in a significant reduction in SeAP expression compared to the EEE1 construct.
[0368] Example 9
[0369] The EEE1 element in the EEE1 SeAP expression vector is replaced by nine variants of the EEE1 element, which can be grouped based on two different types of mutations. The first type of EEE1 variant (EEE1-A) involves variations within either the EEE1 or EEE1-Xt element, with more than 30% of the mutated nucleotides, each located in one of three separate regions, each consisting of at least 244 bp. The second type of EEE1 variant (EEE1-B) also has the same size (1,449 bp) as the EEE1 element with at least 96% sequence identity and contains mutations that target different functional sequences within the EEE1 sequence. The different mutations are listed in Table 3.
[0370] Table 3: Modifications of EEE1
[0371]
[0372] 1) Identity was calculated using the EMBOSS needle nucleotide alignment algorithm with the following parameters: DNA full matrix with the following gap penalties: open = 10; extended = 0.5.
[0373] 2) % Identity with EEE1-Xt calculated as EEE1-A1
[0374] 3) In Example 8, it is referred to as EEE1-SL
[0375] SeAP activity was measured in the supernatant of cells transfected with different variants. The activity of cells with the EEE1 element was used as a reference (100%). In this experiment, the activity was 24% without the EEE1 element. SeAP activity in cells transfected with the EEE1-A2 and EEE1-A3 constructs decreased to 75% and 48%, respectively, relative to EEE1 (Table 4). This is higher than the 24% activity observed with the construct without EEE1 in this experiment. The EEE1-A1 construct reduced SeAP activity to 30% relative to its based EEE1-Xt construct, which is still higher than the construct without EEE1, which produced only 18% SeAP activity relative to the EEE1-Xt construct. The data suggest that EEE1 variants with as little as 72% overall identity and 50% local identity relative to the genomic UBC sequence can serve as expression-enhancing elements.
[0376] Compared to the EEE1 construct, cells transfected with the EEE1-B1 to B6 constructs showed a reduction in SeAP activity of up to 42% (Table 4). The data indicate that mutations in regions that predict the intron promoter activity of EEE1 elements can significantly limit the ability of EEE1 elements to enhance expression. For example, mutations involving 4nt intron-splicing resulted in a 38% reduction in SeAP titer (EEE1-B6). Mutations in CpG from different groups also led to reduced SeAP titers (B1, B2, B4).
[0377] Table 4: SeAP activity of EEE1 variants
[0378]
[0379] 1) Numerical values represent the average activity of four pools derived from two independent transfections, as measured using the SEAP reporter gene assay kit (Abcam). Pools were grown in 125 ml shake flasks in 30 ml CD FortiCHO selection medium.
[0380] Example 10
[0381] The EEE1-TEE CHO-S clone from Example 3 was grown and harvested during the logarithmic growth phase. Total RNA was isolated from the cells using the AllPrep DNA / RNA Mini Kit (Qiagen). cDNA was synthesized using the Epicentre Exact Start Eukaryotic mRNA 5' and 3' RACE Kit. The first-strand cDNA was amplified using 5' RACE primers from a kit containing heavy chain-specific primers (SEQ ID NO: 64) and ZymoTaq DNA polymerase. PCR products were analyzed on a 1.2% agarose gel, showing two separate bands (Fig. 10A), which were isolated and inserted into the PCR4-TOPO vector (Life Technologies), respectively. Sequencing analysis showed that the upper band seen on the agarose gel corresponded to transcripts initiated from the CMV promoter. The lower band corresponded to transcripts initiated from the UBC promoter. Both products had correctly spliced predicted intron sequences. The size difference corresponded to different lengths of the 5'-UTR, as shown in Fig. 10B. Data indicate that both promoters contribute to transcription.
[0382] Example 11
[0383] CHO-S pools stably transfected with constructs having three different individual promoters were compared by measuring SeAP activity in the supernatant. Pools were grown in 125 ml shake flasks in 30 ml CD FortiCHO selection medium, and SeAP activity was then measured in four pools / constructs derived from two independent transfections using the SEAP reporter gene assay kit (Abcam). Constructs contained the CMV promoter (Example 6), the EF-1α promoter (Example 7), or the UBC promoter (Example 11). The UBC promoter produced 2.7-fold higher SeAP activity compared to the CMV promoter construct. The UBC promoter produced 6.0-fold higher SeAP activity compared to the EF-1α promoter. The data indicate that expression is higher with the UBC promoter alone compared to either the CMV promoter or the EF-1α promoter alone.
[0384] Example 12: Methanol-induced secretion of IL-8 in Pichia pastoris GS115 integrative transformant
[0385] A plasmid for stable transformation of Pichia pastoris with a human interleukin-8 (hIL-8) expression construct was generated in plasmid pPIC9K (Life Technologies). Insertion of the hIL-8 gene into pPIC9K resulted in plasmid pPNic384 (SEQ ID NO: 77), which contains the hIL-8 gene under the control of the AOX1 promoter. Insertion of the EEE1 sequence upstream of the AOX1 promoter as the AatII-AleI fragment in pPNic384 (SEQ ID NO: 78) yielded plasmid pPNic602.
[0386] The expression vector was linearized by digestion with SaiI and transformed into Pichia pastoris strain GS115 (as recommended) via electroporation (invitrogen, 2008). Transformants were placed on RDB agar plates (regenerated glucose medium, a histidine-deficient medium). Large colonies were observed after incubation at 30°C for 48 hours. A control transformation without DNA was performed, resulting in no colonies. Ten colonies / constructs were randomly picked from the transformation plate and grown to saturation in 800 μl of BMG (buffered basal medium containing 1% glycerol) in 2 ml deep-well plates. The plates were maintained in an incubator (infors-HT Microton) at 30°C and 1000 rpm for 18 hours. The culture had an absorbance density of 5–10 absorbance units at 600 nm. Cells were harvested and the medium was replaced in 2 ml deep-well plates with 800 μl of BMM (buffered basal medium containing 0.5% methanol). Cells were grown in a shaking incubator, and 0.5% methanol (final concentration) was added to the culture every 24 hours to maintain induction. After 72 hours of methanol induction, the culture supernatant was collected and tested using the ALISA hIL8 kit (Perkin Elmer) to measure the yield of secreted hIL8. Data showed ( Figure 11 The study found a significant difference in IL8 yield between the reference and EEE1 transformants, suggesting that an EEE1 sequence upstream of the promoter improves hIL8 yield compared to expression plasmids without the EEE1 sequence. sequence list <110> R1 B3 Holdings <120> Constructs and sequences for enhanced gene expression <130> P6045820PCT <150> EP 13199873 <151> 2013-12-31 <150> EP 13199875 <151> 2013-12-31 <160>88 <170>PatentIn version 3.3 <210>1 <211>1449 <212>DNA <213>Artificial <220> <223>EEE1 <400>1 tttcaggcaa ccagagctac atagtgagat cctgtctcaa caaaaataaa ataatctaag 60 gcttcaaagg gttcaatctc ttaggtagct aaatatgaac aaaatttggg aaatgtgacc 120 ttttccttag tgacagtcag atagaacctt ctcgagtgca aggacaccaa gtgcaaacag 180 gctcaagaac agcctggaaa ggtctagtgc tatggggctt caggtcgaat gccaactgtt 240 ttcaagaact gtgtggattt ttctgcctgt aacgaattca gattcatttt tcaaaactcg 300 gggagagttt tcccccttta taattttttt tttaaattta ttaaactttg tttcgttccc 360 cttgttttga gaattgcaga gtcatccacc ctgtcacagt gccagggagc tcagggatgg 420 gcccaggggc ctggcggggc tgaaggggct ggggaagcga gggctccaaa gggaccccag 480 tgtggcagga gccaaagccc taggtcccta gaacgcagag gccaccggga ccccccagac 540 ggggtaagcg ggtgggtgtc tggggcgcga agccgcactg cgcatgcgcc gaggtccgct 600 ccggccgcgc tgatccaagc cgggttctcg cgccgacctg gtcgtgattg acaagtcaca 660 cacgctgatc cctccgcggg gccgcacagg gtcacagcct ttcccctccc cacaaagccc 720 cctactctct gggcaccaca cacgaacatt ccttgagcgt gaccttgttg gctctagtca 780 ggcgcctccg gtgcagagac tggaacggcc ttgggaagta gtccctaacc gcatttccgc 840 ggagggatcg tcgggagggc gtggcttctg aggattatat aaggcgactc cgggcgggtc 900 ttagctagtt ccgtcggaga cccgagttca gtcgccgctt ctctgtgagg actgctgccg 960 ccgccgctgg tgaggagaag ccgccgcgct tggcgtagct gagagacggg gagggggcgc 1020 ggacacgagg ggcagcccgc ggcctggacg ttctgtttcc gtggcccgcg aggaaggcga 1080 ctgtcctgag gcggaggacc cagcggcaag atggcggcca agtggaagcc tgaggggata 1140 ggcgagcggc cctgaggcgc tcgacggggt tgggggggaa gcaggcccgc gaggcagctg 1200 cagccgggaa cgtgcggcca accccttatt ttttttgacg ggttgcgggc cgtaggtgcc 1260 tccgaagtga gagccgtggg cgtttgactg tcgggagagg tcggtcggat tttcatccgt 130 tgctaaagac ggaagtgcga ctgagacggg aagggggggg agtcggttgg tggcggttga 1380 acctggacta aggcgcacat gacgtcgcgg tttctatggg ctcataatgg gtggtgagga 1440 catttccct 1449 <210>2 <211>1228 <212>DNA <213>Artificial <220> <223>EEE2 <400>2 gtaaagcaga tcacacagaa tatggcacac ttgagcactt gatgtgtact acattactct 60 tagtgacgac tttaattatc gtgcgcattc ccagcgcttc ctatggtgcc caacacagag 120 cggacgccta gagacaattt tgggggatgg ggcagatgct ctgcctcggg aaaaaaaaag 180 cacacctgcc ctgacgttgg tggctgggtc tggaagatac gtggaaatta agctaaggat 240 gtgtggcttc cagatcaaaa accgcaaaaa tctaacgccg tgactactga ctacggtcag 300 agagcacaga ctggagcaac ctctcacggc ctgggctgtc tgcgcgtgcg tgagccagaa 360 acccgagggg ctccctgggc ccgccctatc gatcgacccg atcggggatc gtcagcttgg 420 ttctggccac agaggttgct cttctcgcga tgcttcagac ctggcggcag ggaagggtg 480 ggctaattgg agagccagga agagcgtgag gcggccccac gctgctttcc cagaaggctg 540 tgcgtgctcc tcgcttctc cgcggtcttc cgagcggtcg cgtgaactgc ttccagcagg 600 ctggccatgg cgcttcacgt tcccaaggct ccgggctttg cccagatgct caaggaggga 660 gcgaaagtaa gggctgaagg aaaggaatga ggtgggagcg tcagcatagg gctgcgggg 720 cggcggcgaa gtagggggc ctactaacgg gctgagcgtg ctgccctggc tcagcggccg 780 ggggaagaga agattccaga aagggaggtg attttggaag ggctcggcca ccggagcctg 840 cgggcacttc tcttcttccg cgaccgggag aaggccgagg gatcggcggc acgatcgaca 900 ttgtcacct tgaaggtgga cggatgtgaa gccgcgcgtg cgttttgcct ccatccgtaa 960 atggggctaa ggccccgtcac ccttaaaagga ggttgtgagg gtgaaattga atacgtaga 1020 tgaaattgtc ttgagaactg cgacgtcgat tatcacatag ctcgcgagtt gtaggatggg 1080 gaaacgag aactagccga tccagagaag agagtgggaa aaaggccgg gtcttggttg 1140 cttgcttccc agtgagaaac atacggcttt cagcttagtt gacagaagcc atgcgttgta 1200 gccaaatgag ttccggtccc aacttatg 1228 <210>3 <211>173 <212>DNA <213>Artificial <220> <223>TEE <400>3 caagctctag caggaagaag aaataagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aataaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaataa aaaaaacaac gtc 173 <210>4 <211>175 <212>DNA <213>Artificial <220> <223>TEE <400>4 caagctctag caggaagaag aaataagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aataaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaataa aaaaaacaac gtccc 175 <210>5 <211>189 <212>DNA <213>Artificial <220> <223>TEE <400>5 agatcactag aagcttcaag ctctagcagg aagaagaaat aagaagaaga agaagaagaa 60 gaagaagcgt ctcctcttct tcttgtgaga gtaaaaaata aaactcccaa aaaaaagaaa 120 atcatcaaaa aaacaaattt caaaaagagt ttttgtgttt ggggattaaa gaataaaaaa 180 aacaacgcc 189 <210>6 <211>189 <212>DNA <213>Artificial <220> <223>TEE <400>6 agatcactag aagcttcaag ctctagcagg aagaagaaag aagaagaaga agaagaagaa 60 gaagaagcgt ctcctcttct tcttgtgaga gtaaaaaaga aaactcccaa aaaaaagaaa 120 atcatcaaaa aaacaaattt caaaaagagt ttttgtgttt ggggattaaa gaagaaaaaa 180 aacaacgcc 189 <210>7 <211>191 <212>DNA <213>Artificial <220> <223>TEE <400>7 agatcactag aagcttcaag ctctagcagg aagaagaaat aagaagaaga agaagaagaa 60 gaagaagcgt ctcctcttct tcttgtgaga gtaaaaaata aaactcccaa aaaaaagaaa 120 atcatcaaaa aaacaaattt caaaaagagt ttttgtgttt ggggattaaa gaataaaaaa 180 aacaacaggc c 191 <210>8 <211>284 <212>DNA <213>Artificial <220> <223>TEE <400>8 ctttttcgca acgggtttgc cgccagaaca caggtgtcgt gaggaattag cttggtacta 60 atacgactca ctatagggag acccaagctg gctaggtaag cttggtaccc aagctctagc 120 aggaagaaga aataagaaga agaagaagaa gaagaagaag cgtctcctct tcttcttgtg 180 agagtaaaaa ataaaactcc caaaaaaaag aaaatcatca aaaaaacaaa tttcaaaaag 240 agtttttgtg tttggggatt aaagaataaa aaaaacaacg tccc 284 <210>9 <211>230 <212>DNA <213>Artificial <220> <223>TEE <400>9 aacccactgc ttactggctt atcgaaatta atacgactca ctatagggag acccaagctc 60 tagcaggaag aagaaataag aagaagaaga agaagaagaa gaagcgtctc ctcttcttct 120 tgtgagagta aaaaataaaa ctcccaaaaa aaagaaaatc atcaaaaaaa caaatttcaa 180 aaagagtttt tgtgtttggg gattaaagaa taaaaaaaac aacctccacc 230 <210>10 <211>177 <212>DNA <213>Artificial <220> <223>TEE <400>10 caagctctag cagcaacaac aaataacaac aacaacaaca acaacaacaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aataaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaataa aaaaaacaac ctccacc 177 <210>11 <211>191 <212>DNA <213>Artificial <220> <223>TEE <400>11 agatcactag aagcttcaag ctctagcagg aagaagaaat aagaagaaga agaagaataa 60 gaagaagcgt ctcgtcttct tcttgtgaga gtaaaaaata aaactcccaa aaaaaataaa 120 atcatcaaaa aaagaaattt caaaaagagt ttttgtgttt ggggattaaa gaataaaaaa 180 aacaacaggc c 191 <210>12 <211>189 <212>DNA <213>Artificial <220> <223>TEE <400>12 agatcactag aagcttcaag ctctagcagg aagaagaaat aataagaaga agaagaataa 60 gaagaagcgt ctcctcttct tcttgtgaga gtaaaaaata aaactcccaa aaaaaataaa 120 atcatcaaaa aaataaattt caaaaagagt ttttgtgttt ggggattaaa gaataaaaaa 180 aacaacgcc 189 <210>13 <211>173 <212>DNA <213>Artificial <220> <223>TEE <400>13 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaakaa aaaaaacaac gtc 173 <210>14 <211>274 <212>DNA <213>Artificial <220> <223>UN2 <400>14ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaakaa aaaaaacaac aggtgagtaa 180 gcgcagttgt cgtctcttgc ggtgccgttg ctggttctca caccttttag gtctgttctc 240 gtcttccgtt ctgactctct ctttttcgtt gcag 274 <210>15 <211>123 <212>DNA <213>Artificial <220> <223>UN1dGAA <400>15 ggcgtctcct cttcttcttg tgagagtaaa aaataaaact cccaaaaaaa akaaaatcat 60 caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka aaaaaacaac 120 gtc 123 <210>16 <211>257 <212>DNA <213>Artificial <220> <223>UN2dGAA <400>16 aaagtatcaa caaaaaagct tcgtctcctc ttcttcttgt gagagtaaaa aakaaaactc 60 ccaaaaaaaa kaaaatcatc aaaaaaacaa atttcaaaaa gagtttttgt gtttgtaagt 120 caggactcta gctttctact gtagtatcct ctaaaggact gctgttctgt gcaccccctt 180 cctttgttta tcatagcgca cgacaagagt actaactaat taacttaggg ggattaaaga 240 akaaaaaaaa caacaaa 257 <210>17 <211>173 <212>DNA <213>Artificial <220> <223>R3 <400>17 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa akaagaagaa gaagaagaag aagaagaagc ctc 173 <210>18 <211>384 <212>DNA <213>Artificial <220> <223>fUN1 <400>18 caagctctag caggaagaag aaataagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aataaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaataa aaaaaacaac gtctggacaa 180 accacaacta gaatgcagtg aaaaaaaatgc tttatttgtg aaatttgtga tgctattgct 240 ttattgtaa ccattataag ctgcaataaa caagttaaca acaacaattg cattcatttt 300 atgtttcagg ttcaggggga ggtgtgggag gttttttaaa gcaagtaaaa cctctacaaa 360 tgtggtaaaa tcgataagga tccg 384 <210> 19 <211> 277 <212> DNA <213> Artificial <220> <223> UN2‑2 <400> 19 caagctctag caggaagaag aaakaagaag aagaagaagaa agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaagaa aaaaaacaac aggtgagtaa 180 gcgcagttgt cgtctcttgc ggtgccgttg ctggttctca caccttttag gtctgttctc 240 gtttccgtt ctgactctct ctttttcgtt gcaggcc 277 <210> 20 <211> 252 <212> DNA <213> Artificial <220> <223> UN2‑3 <400> 20 aagctctagc aggaagaaga aakaagaaga agaagaagaa gaagaagaag cgtctcctct 60 tcttcttgtg agagtaaaaa akaaaactcc caaaaaaaak aaaatcatca aaaaaacaaa 120 tttcaaaaag agtaggtaag attatctctt cccaaaattg attacttttt tattgaacaa 180 ttattaacca atcatggctt aacgaaaaac aggttttgtg tttggggatt aaagaakaaa 240 aaaaacaaaa ca 252 <210>21 <211>254 <212>DNA <213>Artificial <220> <223>UN2-4 <400>21 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtaggtaa gattatctct tcccaaaatt gattactttt attattgaac 180 aattactaac atttcatggc ttaacgaaaa acaggttttg tgtttgggga ttaaagaaka 240 aaaaaaacaa aaca 254 <210>22 <211>266 <212>DNA <213>Artificial <220> [[ID=�4]]<223>UN2-5 <400>22 It should be noted that in the original text, there seems to be a small error in ID=44 where it is written as "�" instead of "4". This has been corrected in the translation.caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtgaggta agattatcga tatttaaatt atttatttct tcttttccat 180 atttcaaaaa gagtgaggta agattatcga tatttaaatt atttatttct tcttttccat 180 ttttttggct aacattttcc atggttttat gatatcatgc aggtacgttt tgtgtttggg 240 ttttttggct aacattttcc atggttttat gatatcatgc aggtacgttt tgtgtttggg 240 gattaaagaa kaaaaaaaac aaaaca 266 gattaaagaa kaaaaaaaac aaaaca 266 <210>23<210>23 <211>266<211>266 <212>DNA<212>DNA <213>人工<213>Artificial <220><220> <223>UN2‑6<223>UN2-6 <000 / S1131><400>23<400>23 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtgaggta agattatcga tatttaaatt atttatttct tcttttccat 180 atttcaaaaa gagtgaggta agattatcga tatttaaatt atttatttct tcttttccat 180 ttttttggct aacattttcc taggttttat tatatctagc aggtacgttt tgtgtttggg 240 ttttttggct aacattttcc taggttttat tatatctagc aggtacgttt tgtgtttggg 240 [ gattaaagaa kaaaaaaaac aaaaca 2( gattaaagaa kaaaaaaaac aaaaca 266 <210>24 <210>24 <211>265 <211>265 <212>DNA <212>DNA <213>人工 <213>Artificial <220> <220> <223>UN2‑7 <223>UN2-7 <400>24 aagctctagc aggaagaaga aakaagaaga agaagaagaa gaagaagaag cgtctcctct 60 tcttcttgtg agagtaaaaa akaaaactcc caaaaaaaak aaaatcatca aaaaaacaaa 120 tttcaaaaag agtgaggtaa gattatcgat atttaaatta tttatttctt cttttccatt 180 tttttggcta acattttcct aggttttatt atatctagca ggtacgtttt gtgtttgggg 240 attaaagaak aaaaaaaaca aaaca 265 <210>25 <211>266 <212>DNA <213>Artificial <220> <223>UN2-8 <400>25 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtgaggta agattatcga tatttaaatt atttatttct tcttttccat 180 ttttttggct aacattttcc taggttttat tatatctagc aggtacgttt tgtgtttggg 240 gattaaagaa kaaaaaaaac aaaacc 266 [[ID=3,7]]<,210>26 <211>287 <212>DNA <213>Artificial <220> <223>UN2-9 <400>26 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttgtaagt caggactcta gctttctact gtagtatcct 180 ctaaaggact gctgttctgt gcaccccctt cctttgttta tcatagcgca cgacaagagt 240 actaactaat taacttaggg ggattaaaga akaaaaaaaa caacaaa 287 <210>27 <211>251 <212>DNA <213>Artificial <220> <223>UN2-10 <400>27 caagctctag caggaagaag aaakaagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttgggtaa gtaattgcct tactcggaaa ataatcaatc 180 atcatactaa cgcaagaggc gctgatattg cggttataca gggattaaag aakaaaaaaa 240 acaacgtcac c 251 <210>28 <211>143 <212>DNA <213>Artificial <220> <223>UN1dGAA-2 <400>28 aaagtatcaa caaaaaagct tcgtctcctc ttcttcttgt gagagtaaaa aakaaaactc 60 ccaaaaaaaa kaaaatcatc aaaaaaacaa atttcaaaaa gagtttttgt gtttggggat 120 taaagaakaa aaaaaacaac aaa 143 <210>29 <211>123 <212>DNA <213>Artificial <220> <223>UN1dGAA-3 <400>29 tcgtctcctc ttcttcttgt gagagtaaaa aakaaaactc ccaaaaaaaa kaaaatcatc 60 aaaaaaacaa atttcaaaaa gagtttttgt gtttggggat taaagaakaa aaaaaacaac 120 gtc 123 <210>30 <211>134 <212>DNA <213>Artificial <220> <223>UN1dGAA-4 <400>30 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa cgcc 134 <210>31 <211> 140 <212> DNA <213> artificial <220> <223> UN1dGAA‑5 <400> 31 ggcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa 60 aaakaaaatc atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa 120 kaaaaaaaac aacctccacc 140 <210> 32 <211> 130 <212> DNA <213> artificial <220> <223> UN1dGAA‑6 <400> 32 caagcgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa aaakaaaatc 60 atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa kaaaaaaaac 120 aacctccacc 130 <210> 33 <211> 178 <212> DNA <213> artificial <220> <223> UN2dGAA‑2 <400> 33 cgtctcctct tcttcttgtg agagtaaaa akaaaactcc caaaaaaaak aaaatcatca 60 aaaaaacaaa tttcaaaaag agttttgtg tttggggatt aaagaakaaa aaaaacaacc 120 tcgtgcgtgt tgccgattcg cgtacgaata cgccttgtgc tgacacttct gtagcacc 178 <210>34 <211>238 <212>DNA <213>Artificial <220> <223>UN2dGAA-3 <400>34 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa caggtgagta agcgcagttg tcgtctcttg cggtgccgtt gctggttctc 180 acacctttta ggtctgttct cgtcttccgt tctgactctc tctttttcgt tgcaggcc 238 <210>35 <211>229 <212>DNA <213>Artificial <220> <223>UN2dGAA-4 <400>35 atcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa 60 aaakaaaatc atcaaaaaaa caaatttcaa aaagagtgag gtaagattat cgatatttaa 120 attatttatt tcttcttttc catttttttg gctaacattt tcctaggttt tattatatct 180 agcaggtacg ttttgtgttt ggggattaaa gaakaaaaaa aacaaaaca 229 <210>36 <211>188 <212>DNA <213>Artificial <220> <223>UN1 Rearrangement <400>36 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa akaagaagaa gaagaagaag aagaagaagg gcggccgccc 180 ccttcacc 188 <210>37 <211>173 <212>DNA <213>Artificial <220> <223>UN1 Rearrangement - 2 <400>37 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa akaagaagaa gaagaagaag aagaagaagc gcc 173 <210>38 <211>154 <212>DNA <213>Artificial <220> <223>UN1 Rearrangement - 3 <400>38 gaagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 gaagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa gaagaagaag cgcc 154 aaaaaaacaa ggaagaagaa gaagaagaag cgcc 154 <210>39<210>39 <211>179<211>179 <212>DNA<212>DNA <213>人工<213>Artificial <220><220> <223>UN1重排‑4<223>UN1 Rearrangement-4 <400>39<400>39 ggcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa 60 ggcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa 60 aaakaaaatc atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa 120 aaakaaaatc atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa 120 kaaaaaaaac aaggaagaag aaakaagaag aagaagaaga agaagaagaa gcctccacc 179 kaaaaaaaac aaggaagaag aaakaagaag aagaagaaga agaagaagaa gcctccacc 179 <210>40<210>40 <211>160 <211>160 <212>DNA <212>DNA <213>人工 <213>Artificial <220> <220> <223>UN1重排‑5 <223>UN1 Rearrangement-5 <400>40 <400>40 ggcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaataaaa ctcccaaaaa 60 ggcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaataaaa ctcccaaaaa 60 aaagaaaatc atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa 120 aaagaaaatc atcaaaaaaa caaatttcaa aaagagtttt tgtgtttggg gattaaagaa 120 taaaaaaaac aaggaagaag aagaagaaga agcctccacc 160 taaaaaaaac aaggaagaag aagaagaaga agcctccacc 160 <210>41 <210>41 <211>177 <212>DNA <213>Artificial <220> <223>UN1 Rearrangement-6 <400>41 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa akaagaagaa gaagaagaag aagaagaagc ctccacc 177 <210>42 <211>277 <212>DNA <213>Artificial <220> <223>UN2 Rearrangement-1 <400>42 caagctctag cacgtctcct cttcttcttg tgagagtaaa aaakaaaact cccaaaaaaa 60 akaaaatcat caaaaaaaca aatttcaaaa agagtttttg tgtttgggga ttaaagaaka 120 aaaaaaacaa ggaagaagaa akaagaagaa gaagaagaag aagaagaagc aggtgagtaa 180 gcgcagttgt cgtctcttgc ggtgccgttg ctggttctca caccttttag gtctgttctc 240 gtcttccgtt ctgactctct ctttttcgtt gcaggcc 277 ;<210>43 <211>267 <212>DNA <213>Artificial <220> <223>UN2 Rearrangement - 2 <400>43 atcaagctct agcacgtctc ctcttcttct tgtgagagta aaaaakaaaa ctcccaaaaa 60 aaakaaaatc atcaaaaaaa caaatttcaa aaagagtgag gtaagattat cgatatttaa 120 attatttatt tcttcttttc catttttttg gctaacattt tcctaggttt tattatatct 180 agcaggtacg ttttgtgttt ggggattaaa gaakaaaaaa aacaaggaag aagaaakaag 240 aagaagaaga agaagaagaa gaaaaca 267 <210>44 <211>264 <212>DNA <213>Artificial <220> <223>CAA1 <400>44 caagctctac caccaagaac aaacaacaac aacatatata aaacaacaac caccatctcc 60 tcttcttctt gtcaactcca aaatcaaact cccaaaaaaa agcaaatcat caaaagtgag 120 gtaagattat cgatatttaa attatttatt tcttcttttc catttttttg gctaacattt 180 tcctaggttt tattatatct agcaggtacg aaatttcaaa caacaacaac aaacaacaaa 240 caacattaac atcatatcaa aacc 264 <210>45 <211>188 <212>DNA <213>Artificial <220> <223>CAA2 <400>45 caacctctac caccaacaac aaacaacaac aacaacaaca acaacaacaa ccctctccac 60 atctccctct cagagtaaaa aacaaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gacttcttct cattccttat taaagaacaa aaaaaacaag gcggccgccc 180 ccttcacc 188 <210>46 <211>207 <212>DNA <213>Artificial <220> <223>CAA3 <400>46 gtatttttac aacaattacc aacaacaaca aacaacaaac aacattacaa ttactattta 60 caattacaag cgtctcctct tcttcttgtg agagtaaaaa ataaaactcc caaaaaaaag 120 aaaatcatca aaaaaacaaa tttcaaaaag agtttttgtg tttggggatt aaagaataaa 180 aaaaacaagg cggccgcccc cttcacc 207 <210>47 <211>188 <212>DNA <213>Artificial <220> <223>CAA4 <400>47 caagctctac caccaagaac aaacaacaac aacatatata aaacaacaac caccatctcc 60 tcttcttctt gtcaactcca aaatcaaact cccaaaaaaa agcaaatcat caaaaccaca 120 aatttcaaac aacaacaaca aacaacaaac aacattaaca tcatatcaag gcggccgccc 180 ccttcacc 188 <210>48 <211>175 <212>DNA <213>Artificial <220> <223>CAA5 <400>48 caacctctac caccaacaac aaacaacaac aacaacaaca acaacaacaa ccctctccac 60 atctccctct cagagtaaaa aacaaaattg acaaaaaaaa gattttataa taaaaacaaa 120 tttcaaaaag aattcaactc attcaatatt acaacaagaa caaaggaggt cacat 175 <210>49 <211>177 <212>DNA <213>Artificial <220> <223>CAA6 <400>49 caagctctac caccaagaac aaacaacaac aacatatata aaacaacaac caccatctcc 60 tcttcttctt gtcaactcca aaatcaaact cccaaaaaaa agcaaatcat caaaaccaca 120 aatttcaaac aacaacaaca aacaacaaac aacattaaca tcatatcaac ctccacc 177 <210>50 <211>175 <212>DNA <213>artificial <220> <223>TATA1 <400>50 caagctctag caggaagaag aaataagaag aagaagaaga agaagaagaa gcgtctcctc 60 ttcttcttga cagagtaaaa aataactttt ataataaaga aaatcatcaa aaaaacaaat 120<00,01426>ttcaaaaaga gtttttgtgt ttggggatta aagaataaaa aaaaggaggt cacat 175 <210>51 <211>177 <212>DNA<00014,30><213>artificial <220> <223>TATA2 <400>51 caagctctag caggaagaag aaataagaag aagtatataa aagaagaaga agcgtctcct 60 cttcttcttg tgaagtaaaa aataaaactc ccaaaaaaaa gaaaatcatc aaaaaaacaa 120 atttcaaaaa gagtttttgt gtttggggat taaagaataa aaaaaacaac ctccacc 177 <210>52 <211>315 <212>DNA <213>artificial<00014,41><220> <223>fragment <400>52 ggagttccgc gttacataac ttacggtaaa tggcccgcct ggctgaccgc ccaacgaccc 60 ccgcccattg acgtcaataa tgacgtatgt tcccatagta acgccaatag ggactttcca 120 ttgacgtcaa tgggtggagt atttacggta aactgcccac ttggcagtac atcaagtgta 180 tcatatgcca agtacgcccc ctattgacgt caatgacggt aaatggcccg cctggcatta 240 tgcccagtac atgaccttat gggactttcc tacttggcag tacatctacg tattagtcat 300 cgctattacc atggt 315 <210>53 <211>303 <212>DNA <213>Artificial <220> <223>Fragment <400>53 ggcctccgcg ccgggttttg gcgcctcccg cgggcgcccc cctcctcacg gcgagcgctg 60 ccacgtcaga cgaagggcgc agcgagcgtc ctgatccttc cgcccggacg ctcaggacag 120 cggcccgctg ctcataagac tcggccttag aaccccagta tcagcagaag gacattttag 180 gacgggactt gggtgactct agggcactgg ttttctttcc agagagcgga acaggcgagg 240 aaaagtagtc ccttctcggc gattctgcgg agggatctcc gtggggcggt gaacgccgat 300 gat 303 <210>54 <211>305 <212>DNA <213>Artificial <220> <223>Fragment <400>54 cgttacataa cttacggtaa atggcccgcc tggctgaccg cccaacgacc cccgcccatt 60 gacgtcaata atgacgtatg ttcccatagt aacgccaata gggactttcc attgacgtca 120 atgggtggag tatttacggt aaactgccca cttggcagta catcaagtgt atcatatgcc 180 aagtacgccc cctattgacg tcaatgacgg taaatggccc gcctggcatt atgcccagta 240 catgacctta tgggactttc ctacttggca gtacatctac gtattagtca tcgctattac 300 catga 305 <210>55 <211>1428 <212>DNA <213>Artificial <220> <223>Fragment <400>55 cgttacataa cttacggtaa atggcccgcc tggctgaccg cccaacgacc cccgcccatt 60 gacgtcaata atgacgtatg ttcccatagt aacgccaata gggactttcc attgacgtca 120 atgggtggag tatttacggt aaactgccca cttggcagta catcaagtgt atcatatgcc 180 aagtacgccc cctattgacg tcaatgacgg taaatggccc gcctggcatt atgcccagta 240 catgacctta tgggactttc ctacttggca gtacatctac gtattagtca tcgctattac 300 catgaattgg tttgatctga ttataaccta ggtcgaggaa ggtttcttca actcaaattc 360 atccgcctga taattttctt atattttcct aaagaaggaa gagaagcgca tagaggagaa 420 gggaaataat tttttaggag cctttcttac ggctatgagg aatttggggc tcagttgaaa 480 agcctaaact gcctctcggg aggttgggcg cggcgaacta ctttcagcgg cgcacggaga 540 cggcgtctac gtgaggggtg ataagtgacg caacactcgt tgcataaatt tgcctccgcc 600 agcccggagc atttaggggc ggttggcttt gttgggtgag cttgtttgtg tccctgtggg 660 tggacgtggt tggtgattgg caggatcctg gtatccgcta acaggtactg gcccgcagcc 720 gtaacgacct tgggggggtg tgagaggggg gaatgggtga ggtcaaggtg gaggcttctt 780 ggggttgggt gggccgctga ggggagggcg tgggggaggg gagggcgagg tgacgcggcg 840 ctgggccttt ccgggacagt gggccttgtt gacctgaggg gggcgagggc ggttggcgcg 900 cgcgggttga cggaaactaa cggacgccta accgatcggc gattctgtcg agtttacttc 960 gcggggaagg cggaaaagag gtagtttgtg tggtttctgg aagcctttac tttggaatct 1020 cagtgtgaga aaggtgcccc ttcttgtgtt tcaatgggat ttttatttcg cgagtcttgt 1080 gggtttggtt ttgttttcag tttgcctaac accgtgctta ggtttgaggc agattggagt 1140 tcggtcgggg gagtttgaat atccggaaca gttagtgggg aaagctgtgg acgattggta 1200 agagagcgct ctggattttc cgctgttgac gttgaaacct tgaatgacga atttcgtatt 1260 aagtgactta gccttgtaaa attgagggga ggcttgcgga atattaacgt atttaaggca 1320 ttttgaagga atagttgcta attttgaaga atattaggtg taaaagcaag aaatacaatg 1380 atcctgaggt gacacgctta tgttttactt ttaaactagg tcagcatg 1428 <210>56 <211>1055 <212>DNA <213>Fragment <400>56 ggagttccgc gttacataac ttacggtaaa tggcccgcct ggctgaccgc ccaacgaccc 60 ccgcccattg acgtcaataa tgacgtatgt tcccatagta acgccaatag ggactttcca 120 ttgacgtcaa tgggtggagt atttacggta aactgcccac ttggcagtac atcaagtgta 180 tcatatgcca agtacgcccc ctattgacgt caatgacggt aaatggcccg cctggcatta 240 tgcccagtac atgaccttat gggactttcc tacttggcag tacatctacg tattagtcat 300 cgctattacc atggtcgagg tgagccccac gttctgcttc actctcccca tctcccccc 360 ctccccaccc ccaattttgt attatttat tttttaatta ttttgtgcag cgatgggggc 420 gggggggggg ggggggcgcg gccaggcggg gcggggcggg gcgaggggcg gggcggggcg 480 aggcggagag gtgcggcggc agccaatcag agcggcgcgc tccgaaagtt tcctttatg 540 gcgaggcggc ggcggcggcg gccctataaa aagcgaagcg cgcggcgggc gggagtcgct 600 gcgcgctgcc ttcgccccgt gccccgctcc gccgccgcct cgcgccgccc gccccggctc 660 tgactgaccg cgttactaaaa acaggtaagt ccggcctccg cgccgggttt tggcgcctcc 720 cgcgggcgcc cccctcctca cggcgagcgc tgccacgtca gacgaagggc gcagcgagcg 780 tcctgatcct tccgcccgga cgctcaggac agcggcccgc tgctcataag actcggcctt 840 agaaccccag tatcagcaga aggacatttt aggacgggac ttgggtgact ctagggcact 900 ggttttcttt ccagagagcg gaacaggcga ggaaaagtag tcccttctcg gcgattctgc 960 ggagggatct ccgtggggcg gtgaacgccg atgatgcctc tactaaccat gttcatgttt 1020 tctttttttt tctacaggtc ctgggtgacg aacag 1055 <210>57 <211>753 <212>DNA <213>Artificial <220> <223>CMV promoter <400>57 tcaatattgg ccattagcca tattattcat tggttatata gcataaatca atattggcta 60 ttggccattg catacgttgt atctatatca taatatgtac atttatattg gctcatgtcc 120 aatatgaccg ccatgttggc attgattatt gactagttat taatagtaat caattacggg 180 gtcattagtt catagcccat atatggagtt ccgcgttaca taacttacgg taaatggccc 240 gcctggctga ccgcccaacg acccccgccc attgacgtca ataatgacgt atgttcccat 300 agtaacgcca atagggactt tccattgacg tcaatgggtg gagtatttac ggtaaactgc 360 ccacttggca gtacatcaag tgtatcatat gccaagtccg ccccctattg acgtcaatga 420 cggtaaatgg cccgcctggc attatgccca gtacatgacc ttacgggact ttcctacttg 480 gcagtacatc tacgtattag tcatcgctat taccatagtg atgcggtttt ggcagtacac 540 caatgggcgt ggatagcggt ttgactcacg gggatttcca agtctccacc ccattgacgt 600 caatgggagt ttgttttggc accaaaatca acgggacttt ccaaaatgtc gtaataaccc 660 cgccccgttg acgcaaatgg gcggtaggcg tgtacggtgg gaggtctata taagcagagc 720 tcgtttagtg aaccgtcaga tcactagaag ctt 753 <210>58 <211>69 <212>DNA <213>Artificial <220> <223>Minimal cMV promoter <400>58 taggcgtgta cggtgggagg tctatataag cagagctcgt ttagtgaacc gtcagatcac 60 tagaagctt 69 <210>59 <211>2248 <212>DNA <213>Artificial [[ID=...]]<220> <223>EEE1-Xt <400>59 tggtgaccct gtctcaaaaa accctcaaaa agtgttggga ttagtggcat gcaccaccat 60 tcccaccaaa ggtttatttt tataatatg tgtgtgagtg tgtatcacta tgagtatatg 120 tcaatatgtg tcaatgtccc cagggacatt taaagagccc ctgaagctgg agtcataggc 180 cattatgaac tgcctgacat ggctaatggg aattgaactc agattttctg gaagttatac 240 ctgctcttac tgctgagcca tgtctctgaa gaccccaggg attttttttt ttttttgaga 300 caggtatttt ctgtatagcc ctggctgtcc tgaaagcact ctctatatgt agaccaggct 360 tgcctggagc ttggatatgc acctgcttct gcctcaggaa tggtgggatt gaaggtgtgc 420 accaccacat ccgctaacat gcacaattct taatgggttt atatcttatt taatgaatga 480 aaggtttggg ggatggatgt agcttaatgg aaaatgactg aagatttcaa ttaaaaatct 540 ggggcttagc tgcgcggtg gtggtgcctg cctttagtcc cagtactggg gaggcagagg 600 aaggaggatc tctgtgagtt cgaggccagc tggtctataa cgtgagttcc aggacagcca 660 gagatacaca gacaaaccct gtctcaccaa aaaaacaa caacaacaac aaaaatctg 720 ggacgtaggc ttggtgtggt ggcacacatt ttgattccag cacttggaag gaagaggcct 780 gcatggtcta catagcttgt ttcaggcaac cagagctaca tagtgagatc ctgtctcaac 840 aaaaataaaa taatctaagg cttcaaaggg ttcaatctct taggtagcta aatatgaaca 900 aaatttggga aatgtgacct tttccttagt gacagtcaga tagaaccttc tcgagtgcaa 960 ggacaccaag tgcaaacagg ctcaagaaca gcctggaaag gtctagtgct atggggcttc 1020 aggtcgaatg ccaactgttt tcaagaactg tgtggatttt tctgcctgta acgaattcag 1080 attcattttt caaaactcgg ggagagtttt ccccctttat aatttttttt ttaaatttat 1140 taaactttgt ttcgttcccc ttgttttgag aattgcagag tcatccaccc tgtcacagtg 1200 ccagggagct cagggatggg cccaggggcc tggcggggct gaaggggctg gggaagcgag 1260 ggctccaaag ggaccccagt gtggcaggag ccaaagccct aggtccctag aacgcagagg 1320 ccaccgggac cccccagacg gggtaagcgg gtgggtgtct ggggcgcgaa gccgcactgc 1380 gcatgcgccg aggtccgctc cggccgcgct gatccaagcc gggttctcgc gccgacctgg 1440 tcgtgattga caagtcacac acgctgatcc ctccgcgggg ccgcacaggg tcacagcctt 1500 tcccctcccc acaaagcccc ctactctctg ggcaccacac acgaacattc cttgagcgtg 1560 accttgttgg ctctagtcag gcgcctccgg tgcagagact ggaacggcct tgggaagtag 1620 tccctaaccg catttccgcg gagggatcgt cgggagggcg tggcttctga ggattatata 1680 aggcgactcc gggcgggtct tagctagttc cgtcggagac ccgagttcag tcgccgcttc 1740 tctgtgagga ctgctgccgc cgccgctggt gaggagaagc cgccgcgctt ggcgtagctg 1800 agagacgggg agggggcgcg gacacgaggg gcagcccgcg gcctggacgt tctgtttccg 1860 tggcccgcga ggaaggcgac tgtcctgagg cggaggaccc agcggcaaga tggcggccaa 1920 gtggaagcct gaggggatag gcgagcggcc ctgaggcgct cgacggggtt gggggggaag 1980 caggcccgcg aggcagctgc agccgggaac gtgcggccaa ccccttattt tttttgacgg 2040 gttgcgggcc gtaggtgcct ccgaagtgag agccgtgggc gtttgactgt cgggagaggt 2100 cggtcggatt ttcatccgtt gctaaagacg gaagtgcgac tgagacggga agggggggga 2160 gtcggttggt ggcggttgaa cctggactaa ggcgcacatg acgtcgcggt ttctatgggc 2220 tcataatggg tggtgaggac atttccct 2248 <210>60 <211>1159 <212>DNA <213>Artificial <220> <223>EEE1-80 <400>60 tcaaaactcg gggagagttt tcccccttta taattttttt tttaaattta ttaaactttg 60 tttcgttccc cttgttttga gaattgcaga gtcatccacc ctgtcacagt gccagggagc 120 tcagggatgg gcccaggggc ctggcggggc tgaaggggct ggggaagcga gggctccaaa 180 gggaccccag tgtggcagga gccaaagccc taggtcccta gaacgcagag gccaccggga 240 ccccccagac ggggtaagcg ggtgggtgtc tggggcgcga agccgcactg cgcatgcgcc 300 gaggtccgct ccggccgcgc tgatccaagc cgggttctcg cgccgacctg gtcgtgattg 360 acaagtcaca cacgctgatc cctccgcggg gccgcacagg gtcacagcct ttcccctccc 420 cacaaagccc cctactctct gggcaccaca cacgaacatt ccttgagcgt gaccttgttg 480 gctctagtca ggcgcctccg gtgcagagac tggaacggcc ttgggaagta gtccctaacc 540 gcatttccgc ggagggatcg tcgggagggc gtggcttctg aggattatat aaggcgactc 600 cgggcgggtc ttagctagtt ccgtcggaga cccgagttca gtcgccgctt ctctgtgagg 660 cgggcgggtc ttagctagtt ccgtcggaga cccgagttca gtcgccgctt ctctgtgagg 660 actgctgccg ccgccgctgg tgaggagaag ccgccgcgct tggcgtagct gagagacggg 720 actgctgccg ccgccgctgg tgaggagaag ccgccgcgct tggcgtagct gagagacggg 720 gagggggcgc ggacacgagg ggcagcccgc ggcctggacg ttctgtttcc gtggcccgcg 780 gagggggcgc ggacacgagg ggcagcccgc ggcctggacg ttctgtttcc gtggcccgcg 780 aggaaggcga ctgtcctgag gcggaggacc cagcggcaag atggcggcca agtggaagcc 840 aggaaggcga ctgtcctgag gcggaggacc cagcggcaag atggcggcca agtggaagcc 840 tgaggggata ggcgagcggc cctgaggcgc tcgacggggt tgggggggaa gcaggcccgc 900 tgaggggata ggcgagcggc cctgaggcgc tcgacggggt tgggggggaa gcaggcccgc 900 gaggcagctg cagccgggaa cgtgcggcca accccttatt ttttttgacg ggttgcgggc 960 gaggcagctg cagccgggaa cgtgcggcca accccttatt ttttttgacg ggttgcgggc 960 cgtaggtgcc tccgaagtga gagccgtggg cgtttgactg tcgggagagg tcggtcggat 1020 cgtaggtgcc tccgaagtga gagccgtggg cgtttgactg tcgggagagg tcggtcggat 1020 tttcatccgt tgctaaagac ggaagtgcga ctgagacggg aagggggggg agtcggttgg 1080 tttcatccgt tgctaaagac ggaagtgcga ctgagacggg aagggggggg agtcggttgg 1080 tggcggttga acctggacta aggcgcacat gacgtcgcgg tttctatggg ctcataatgg 1140 tggcggttga acctggacta aggcgcacat gacgtcgcgg tttctatggg ctcataatgg 1140 gtggtgagga catttccct 1159 gtggtgagga catttccct 1159 <210>61 <211>869 <212>DNA <213>Artificial <220> <223>EEE1-60 <400>61 cgcatgcgcc gaggtccgct ccggccgcgc tgatccaagc cgggttctcg cgccgacctg 60 gtcgtgattg acaagtcaca cacgctgatc cctccgcggg gccgcacagg gtcacagcct 120 ttcccctccc cacaaagccc cctactctct gggcaccaca cacgaacatt ccttgagcgt 180 gaccttgttg gctctagtca ggcgcctccg gtgcagagac tggaacggcc ttgggaagta 240 gtccctaacc gcatttccgc ggagggatcg tcgggagggc gtggcttctg aggattatat 300 aaggcgactc cgggcgggtc ttagctagtt ccgtcggaga cccgagttca gtcgccgctt 360 ctctgtgagg actgctgccg ccgccgctgg tgaggagaag ccgccgcgct tggcgtagct 420 gagagacggg gagggggcgc ggacacgagg ggcagcccgc ggcctggacg ttctgtttcc 480 gtggcccgcg aggaaggcga ctgtcctgag gcggaggacc cagcggcaag atggcggcca 540 agtggaagcc tgaggggata ggcgagcggc cctgaggcgc tcgacggggt tgggggggaa 600 gcaggcccgc gaggcagctg cagccgggaa cgtgcggcca accccttatt ttttttgacg 660 ggttgcgggc cgtaggtgcc tccgaagtga gagccgtggg cgtttgactg tcgggagagg 720 tcggtcggat tttcatccgt tgctaaagac ggaagtgcga ctgagacggg aagggggggg 780 agtcggttgg tggcggttga acctggacta aggcgcacat gacgtcgcgg tttctatggg 840 ctcataatgg gtggtgagga catttccct 869 <210>62 <211>724 <212>DNA <213>Artificial <220> <223>EEE1-50 <400>62 tctctgggca ccacacacga acattccttg agcgtgacct tgttggctct agtcaggcgc 60 ctccggtgca gagactggaa cggccttggg aagtagtccc taaccgcatt tccgcggagg 120 gatcgtcggg agggcgtggc ttctgaggat tatataaggc gactccgggc gggtcttagc 180 tagttccgtc ggagacccga gttcagtcgc cgcttctctg tgaggactgc tgccgccgcc 240 gctggtgagg agaagccgcc gcgcttggcg tagctgagag acggggaggg ggcgcggaca 300 cgaggggcag cccgcggcct ggacgttctg tttccgtggc ccgcgaggaa ggcgactgtc 360 ctgaggcgga ggacccagcg gcaagatggc ggccaagtgg aagcctgagg ggataggcga 420 gcggccctga ggcgctcgac ggggttgggg gggaagcagg cccgcgaggc agctgcagcc 480 gggaacgtgc ggccaacccc ttattttttt tgacgggttg cgggccgtag gtgcctccga 540 agtgagagcc gtgggcgttt gactgtcggg agaggtcggt cggattttca tccgttgcta 600 aagacggaag tgcgactgag acgggaaggg gggggagtcg gttggtggcg gttgaacctg 660 gactaaggcg cacatgacgt cgcggtttct atgggctcat aatgggtggt gaggacattt 720 ccct 724 <210>63 <211>1449 <212>DNA <213>Artificial <220> <223>EEE1-SL <400>63 tttcaggcaa ccagagctac atagtgagat cctgtctcaa caaaaataaa ataatctaag 60 gcttcaaagg gttcaatctc ttaggtagct aaatatgaac aaaatttggg aaatgtgacc 120 ttttccttag tgacagtcag atagaacctt ctcgagtgca aggacaccaa gtgcaaacag 180 gctcaagaac agcctggaaa ggtctagtgc tatggggctt caggtcgaat gccaactgtt 240 ttcaagaact gtgtggattt ttctgcctgt aacgaattca gattcatttt tcaaaactcg 300 gggagagttt tcccccttta taattttttt tttaaattta ttaaactttg tttcgttccc 360 cttgttttga gaattgcaga gtcatccacc ctgtcacagt gccagggagc tcagggatgg 420 gcccaggggc ctggcggggc tgaaggggct ggggaagcga gggctccaaa gggaccccag 480 tgtggcagga gccaaagccc taggtcccta gaacgcagag gccaccggga ccccccagac 540 ggggaaagcg gttgggtgtc tggggcgcga agccgcactg cgcatgcgcc gaggtccgct 600 ccggccgcgc tgatccaagc cgggttctcg cgccgacctg gtcgtgattg acaagtcaca 660 cacgctgatc cctccgcggg gccgcacagg gtcacagcct ttcccctccc cacaaagccc 720 cctactctct gggcaccaca cacgaacatt ccttgagcgt gaccttgttg gctctagtca 780 ggcgcctccg gtgcagagac tggaacggcc ttgggaagta gtccctaacc gcatttccgc 840 ggagggatcg tcgggagggc gtggcttctg aggattatat aaggcgactc cgggcgggtc 900 ttagctagtt ccgtcggaga cccgagttca gtcgccgctt ctctgtgagg actgctgccg 960 ccgccgctgc tgaggagaag ccgccgcgct tggcgtagct gagagacggg gagggggcgc 1020 ggacacgagg ggcagcccgc ggcctggacg ttctgtttcc gtggcccgcg aggaaggcga 1080 ctgtcctgag gcggaggacc cagcggcaag atggcggcca agtggaagcc tgaggggata 1140 ggcgagcggc cctgaggcgc tcgacggggt tgggggggaa gcaggcccgc gaggcagctg 1200 cagccgggaa cgtgcggcca accccttatt ttttttgacg ggttgcgggc cgtaggtgcc 1260 tccgaattga gagccgtggg cgtttgactg tcgggagagg tcggtcggat tttcatccgt 1320 tgctaaagac ggaagtgcga ctgagacggg aagggggggg agtcggttgg tggcggttga 1380 acctggacta aggcgcacat gacgtcgcgg tttctatggg ctcataatgg gtggtgagga 1440 catttccct 1449 <210>64 <211>33 <212>DNA <213>Artificial <220> <223>Hc RAcE primer[[ID={25]]<00]01711><400>64 gctggtgccc aggtccttag cgcaatagta cac 33 <210>65 <211>4128 <212>DNA <213>Artificial <220> <223>Light chain vector sequence without coding sequence <400>65 tgcaggcggc cgctttcagg caaccagagc tacatagtga gatcctgtct caacaaaaat 60 aaaataatct aaggcttcaa agggttcaat ctcttaggta gctaaatatg aacaaaattt 120 gggaaatgtg accttttcct tagtgacagt cagatagaac cttctcgagt gcaaggacac 180 caagtgcaaa caggctcaag aacagcctgg aaaggtctag tgctatgggg cttcaggtcg 240 aatgccaact gttttcaaga actgtgtgga tttttctgcc tgtaacgaat tcagattcat 300 ttttcaaaac tcggggagag ttttccccct ttataatttt ttttttaaat ttattaaact 360 ttgtttcgtt ccccttgttt tgagaattgc agagtcatcc accctgtcac agtgccaggg 420 agctcaggga tgggcccagg ggcctggcgg ggctgaaggg gctggggaag cgagggctcc 480 aaagggaccc cagtgtggca ggagccaaag ccctaggtcc ctagaacgca gaggccaccg 540 ggacccccca gacggggtaa gcggggtgggt gtctggggcg cgaagccgca ctgcgcatgc 600 gccgaggtcc gctccggccg cgctgatcca agccgggttc tcgcgccgac ctggtcgtga 660 ttgacaagtc acacacgctg atccctccgc ggggccgcac agggtcacag cctttcccct 720 ccccacaaag ccccctactc tctgggcacc acacacgaac attccttgag cgtgaccttg 780 ttggctctag tcaggcgcct ccggtgcaga gactggaacg gccttgggaa gtagtcccta 840 accgcatttc cgcggaggga tcgtcgggag ggcgtggctt ctgaggatta tataaggcga 900 ctccgggcgg gtcttagcta gttccgtcgg agacccgagt tcagtcgccg cttctctgtg 960 aggactgctg ccgccgccgc tggtgaggag aagccgccgc gcttggcgta gctgagagac 1020 ggggaggggg cgcggacacg aggggcagcc cgcggcctgg acgttctgtt tccgtggccc 1080 gcgaggaagg cgactgtcct gaggcggagg acccagcggc aagatggcgg ccaagtggaa 1140 gcctgagggg ataggcgagc ggccctgagg cgctcgacgg ggttgggggg gaagcaggcc 1200 cgcgaggcag ctgcagccgg gaacgtgcgg ccaacccctt attttttttg acgggttgcg 1260 ggccgtaggt gcctccgaag tgagagccgt gggcgtttga ctgtcgggag aggtcggtcg 1320 gattttcatc cgttgctaaa gacggaagtg cgactgagac gggaaggggg gggagtcggt 1380 tggtggcggt tgaacctgga ctaaggcgca catgacgtcg cggtttctat gggctcataa 1440 tgggtggtga ggacatttcc ctgtttaaac ttaaacaagt ttgtacaaaa aagcaggcta 1500 gatcttcaat attggccatt agccatatta ttcattggtt atatagcata aatcaatatt 1560 ggctattggc cattgcatac gttgtatcta tatcataata tgtacattta tattggctca 1620 tgtccaatat gaccgccatg ttggcattga ttatgacta gttattaata gtaatcaatt 1680 acggggtcat tagttcatag cccatatatg gagttccgcg ttacataact tacggtaaat 1740 ggcccgcctg gctgaccgcc caacgacccc cgcccattga cgtcaataat gacgtatgtt 1800 cccatagtaa cgccaatagg gactttccat tgacgtcaat gggtggagta tttacggtaa 1860 actgcccact tggcagtaca tcaagtgtat catatgccaa gtccgccccc tattgacgtc 1920 aatgacggta aatggcccgc ctggcattat gcccagtaca tgaccttacg ggactttcct 1980 acttggcagt acatctacgt attagtcatc gctattacca tagtgatgcg gttttggcag 2040 tacaccaatg ggcgtggata gcggtttgac tcacggggat ttccaagtct ccaccccatt 2100 gacgtcaatg ggagtttgtt ttggcaccaa aatcaacggg actttccaaa atgtcgtaat 2160 aaccccgccc cgttgacgca aatgggcggt aggcgtgtac ggtgggaggt ctatataagc 2220 agagctcgtt tagtgaaccg tcagatcact agaagcttaa tacgactcac tatagggaga 2280 cccaagctgg ctagcgttta aacgggccct ctagtaacgg ccgccagtgt gctggaattc 2340 ggcttaactc tagaccatgg ggcgcgccgg ttcagcctcg actgtgcctt ctagttgcca 2400 gccatctgtt gtttgcccct cccccgtgcc ttccttgacc ctggaaggtg ccactcccac 2460 tgtcctttcc tataaaatg aggaaattgc atcgcattgt ctgagtaggt gtcattctat 2520 tctggggggt ggggtggggc aggacagcaa gggggaggat tgggaagaca atagcaggca 2580 tgctggggat gcggtgggct ctatggcttc tgaggcggaa agaaccagct ggatccatcc 2640 gttagatatc tgtggaatgt gtgtcagtta gggtgtggaa agtccccagg ctccccagca 2700 ggcagaagta tgcaaagcat gcatctcaat tagtcagcaa ccaggtgtgg aaagtcccca 2760 ggctccccag caggcagaag tatgcaaagc atgcatctca attagtcagc aaccatagtc 2820 ccgcccctaa ctccgcccat cccgcccta actccgccca gttccgccca ttctccggcc 2880 catgcctgac taattttttt tatttatgca gaggccgagg ccgcctctgc ctctgagcta 2940 ttccagaagt agtgaggagg cttttttgga ggcctaggct tttgcaaaaa gctcccggga 3000 gcttgtatat ccattttcgg atctgatcaa gagacaggat gaggatcgtt tcacatgatt 3060 gaacaagatg gattgcacgc aggttctccg gccgcttggg tggagaggct attcggctat 3120 gactgggcac aacagacaat cggctgctct gatgccgccg tgttccggct gtcagcgcag 3180 gggcgcccgg ttctttttgt caagaccgac ctgtccggtg ccctgaatga actgcaggac 3240 gaggcagcgc ggctatcgtg gctggccacg acgggcgttc cttgcgcagc tgtgctcgac 3300 gttgtcactg aagcgggaag ggactggctg ctattgggcg aagtgccggg gcaggatctc 3360 ctgtcatctc accttgctcc tgccgagaaa gtatccatca tggctgatgc aatgcggcgg 3420 ctgcatacgc ttgatccggc tacctgccca ttcgaccacc aagcgaaaca tcgcatcgag 3480 cgagcacgta ctcggatgga agccggtctt gtcgatcagg atgatctgga cgaagagcat 3540 caggggctcg cgccagccga actgttcgcc aggctcaagg cgcgcatgcc cgacggcgag 3600 gatctcgtcg tgacacatgg cgatgcctgc ttgccgaata tcatggtgga aaatggccgc 3660 ttttctggat tcatcgactg tggccggctg ggtgtggcgg accgctatca ggacatagcg 3720 ttggctaccc gtgatattgc tgaagagctt ggcggcgaat gggctgaccg cttcctcgtg 3780 ctttacggta tcgccgctcc cgattcgcag cgcatcgcct tctatcgcct tcttgacgag 3840 ttcttctagg taccacgaga tttcgattcc accgccgcct tctatgaaag gttgggcttc 3900 ggaatcgttt tccgggacgc cggctggatg atcctccagc gcggggatct catgctggag 3960 ttcttcgccc accccaactt gtttattgca gcttataatg gttacaaata aagcaatagc 4020 atcacaaatt tcacaaataa agcatttttt tcactgcatt ctagttgtgg tttgtccaaa 4080[[ID=gggaaatgtg accttttcct tagtgacagt cagatagaac cttctcgagt gcaaggacac 180 caagtgcaaa caggctcaag aacagcctgg aaaggtctag tgctatgggg cttcaggtcg 240 aatgccaact gttttcaaga actgtgtgga tttttctgcc tgtaacgaat tcagattcat 300 ttttcaaaac tcggggagag ttttccccct ttataatttt ttttttaaat ttattaaact 360 ttgtttcgtt ccccttgttt tgagaattgc agagtcatcc accctgtcac agtgccaggg 420 agctcaggga tgggcccagg ggcctggcgg ggctgaaggg gctggggaag cgagggctcc 480 aaagggaccc cagtgtggca ggagccaaag ccctaggtcc ctagaacgca gaggccaccg 540 ggacccccca gacggggtaa gcggggtgggt gtctggggcg cgaagccgca ctgcgcatgc 600 gccgaggtcc gctccggccg cgctgatcca agccgggttc tcgcgccgac ctggtcgtga 660 ttgacaagtc acacacgctg atccctccgc ggggccgcac agggtcacag cctttcccct 720 ccccacaaag ccccctactc tctgggcacc acacacgaac attccttgag cgtgaccttg 780 ttggctctag tcaggcgcct ccggtgcaga gactggaacg gccttgggaa gtagtcccta 840 accgcatttc cgcggaggga tcgtcgggag ggcgtggctt ctgaggatta tataaggcga 900 ctccgggcgg gtcttagcta gttccgtcgg agacccgagt tcagtcgccg cttctctgtg 960 aggactgctg ccgccgccgc tggtgaggag aagccgccgc gcttggcgta gctgagagac 1020 ggggaggggg cgcggacacg aggggcagcc cgcggcctgg acgttctgtt tccgtggccc 1080 gcgaggaagg cgactgtcct gaggcggagg acccagcggc aagatggcgg ccaagtggaa 1140 gcctgagggg ataggcgagc ggccctgagg cgctcgacgg ggttgggggg gaagcaggcc 1200 cgcgaggcag ctgcagccgg gaacgtgcgg ccaacccctt attttttttg acgggttgcg 1260 ggccgtaggt gcctccgaag tgagagccgt gggcgtttga ctgtcgggag aggtcggtcg 1320 gattttcatc cgttgctaaa gacggaagtg cgactgagac gggaaggggg gggagtcggt 1380 tggtggcggt tgaacctgga ctaaggcgca catgacgtcg cggtttctat gggctcataa 1440 tgggtggtga ggacatttcc ctgtttaaac ttaaacaagt ttgtacaaaa aagcaggcta 1500 gatcttcaat attggccatt agccatatta ttcattggtt atatagcata aatcaatatt 1560 ggctattggc cattgcatac gttgtatcta tatcataata tgtacattta tattggctca 1620 tgtccaatat gaccgccatg ttggcattga ttatgacta gttattaata gtaatcaatt 1680 acggggtcat tagttcatag cccatatatg gagttccgcg ttacataact tacggtaaat 1740 ggcccgcctg gctgaccgcc caacgacccc cgcccattga cgtcaataat gacgtatgtt 1800 cccatagtaa cgccaatagg gactttccat tgacgtcaat gggtggagta tttacggtaa 1860 actgcccact tggcagtaca tcaagtgtat catatgccaa gtccgccccc tattgacgtc 1920 aatgacggta aatggcccgc ctggcattat gcccagtaca tgaccttacg ggactttcct 1980 acttggcagt acatctacgt attagtcatc gctattacca tagtgatgcg gttttggcag 2040 tacaccaatg ggcgtggata gcggtttgac tcacggggat ttccaagtct ccaccccatt 2100 gacgtcaatg ggagtttgtt ttggcaccaa aatcaacggg actttccaaa atgtcgtaat 2160 aaccccgccc cgttgacgca aatgggcggt aggcgtgtac ggtgggaggt ctatataagc 2220 agagctcgtt tagtgaaccg tcagatcact agaagcttaa tacgactcac tatagggaga 2280 cccaagctgg ctagcgttta aacgggccct ctagtaacgg ccgccagtgt gctggaattc 2340 ggcttaactc tagaccatgg ggcgcgccgg ttcagcctcg actgtgcctt ctagttgcca 2400 gccatctgtt gtttgcccct cccccgtgcc ttccttgacc ctggaaggtg ccactcccac 2460 tgtcctttcc taataaaatg aggaaattgc atcgcattgt ctgagtaggt gtcattctat 2520 tctggggggt ggggtggggc aggacagcaa gggggaggat tgggaagaca atagcaggca 2580 tgctggggat gcggtgggct ctatggcttc tgaggcggaa agaaccagct ggatccatcc 2640 gttagat 2647 <210>67 <211>235 <212>PRT <213>Artificial <220> <223>HuMab1 protein light chain <400>67 Met Gly Trp Ser Cys Ile Ile Leu Phe Leu Val Ala Thr Ala Thr Gly 1 5 10 15 Val His Ser Ala Gln Asp Ile Gln Met Thr Gln Ser Pro Ser Ser Val 20 25 30 Ser Ala Ser Val Gly Asp Arg Val Thr Ile Thr Cys Arg Ala Ser Gln 35 40 45 Gly Ile Ser Ser Trp Leu Ala Trp Tyr Gln Gln Lys Pro Gly Lys Ala 50 55 60 Pro Lys Leu Leu Ile Tyr Ala Ala Ser Ser Leu Gln Ser Gly Val Pro 65 70 75 80 Ser Arg Phe Ser Gly Ser Gly Ser Gly Thr Asp Phe Thr Leu Thr Ile 85 90 95 Ser Ser Leu Gln Pro Glu Asp Phe Ala Thr Tyr Tyr Cys Gln Gln Ala 100 105 110 Asn Asn Phe Pro Leu Thr Phe Gly Gly Gly Thr Lys Val Glu Ile Lys 115 120 125 Arg Thr Val Ala Ala Pro Ser Val Phe Ile Phe Pro Pro Ser Asp Glu 130 135 140 Gln Leu Lys Ser Gly Thr Ala Ser Val Val Cys Leu Leu Ash Ash Phe 145 150 155 160 Tyr Pro Arg Glu Ala Lys Val Gln Trp Lys Val Asp Ash Ala Leu Gln 165 170 175 Ser Gly Asn Ser Gln Glu Ser Val Thr Glu Gln Asp Ser Lys Asp Ser 180 185 190 Thr Tyr Ser Leu Ser Ser Thr Leu Thr Leu Ser Lys Ala Asp Tyr Glu 195 200 205 Lys His Lys Val Tyr Ala Cys Glu Val Thr His Gln Gly Leu Ser Ser 210 215 220 Pro Val Thr Lys Ser Phe Asn Arg Gly Glu Cys 225 230 235 <210>68 <211>474 <212>PRT <213>Artificial <220> <223>Heavy chain of HuMab1 protein <400>68 Met Gly Trp Ser Cys Ile Ile Leu Phe Leu Val Ala Thr Ala Thr Gly 1 5 10 15 Val His Ser Glu Val Gln Leu Leu Glu Ser Gly Gly Gly Leu Val Gln 20 25 30 Pro Gly Gly Ser Leu Arg Leu Ser Cys Ala Ala Ser Gly Phe Thr Phe 35 40 45 Ser Asn Tyr Ala Met Ser Trp Val Arg Gln Ala Pro Gly Lys Gly Leu 50 55 60 Glu Trp Val Ser Ala Ile Ser Ala Ser Gly His Ser Thr Tyr Leu Ala 65 70 75 80 Asp Ser Val Lys Gly Arg Phc Thr Ile Ser Arg Asp Asn Ser Lys Asn 85 90 95 Thr Leu Tyr Leu Gln Met Asn Ser Leu Arg Ala Glu Asp Thr Ala Val 100 105 110 Tyr Tyr Cys Ala Lys Asp Arg Glu Val Thr Met Ile Val Val Leu Asn 115 120 125 Gly Gly Phe Asp Tyr Trp Gly Gln Gly Thr Arg Val Thr Val Ser Ser 130 135 140 Ala Ser Thr Lys Gly Pro Ser Val Phe Pro Leu Ala Pro Ser Ser Lys 145 150 155 160 Ser Thr Ser Gly Gly Thr Ala Ala Leu Gly Cys Leu Val Lys Asp Tyr 165 170 175 Phe Pro Glu Pro Val Thr Val Ser Trp Asn Ser Gly Ala Leu Thr Ser 180 185 190 Gly Val His Thr Phe Pro Ala Val Leu Gln Ser Ser Gly Leu Tyr Ser 195 200 205 Leu Ser Ser Val Val Thr Val Pro Ser Ser Ser Leu Gly Thr Gln Thr 210 215 220 Tyr Ile Cys Asn Val Asn His Lys Pro Ser Asn Thr Lys Val Asp Lys 225 230 235 240 Arg Val Glu Pro Lys Ser Cys Asp Lys Thr His Thr Cys Pro Pro Cys 245 250 255 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 260 265 270 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 275 280 285 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 290 295 300 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 305 310 315 320 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 325 330 335 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 340 345 350 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 355 360 365 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Glu Glu 370 375 380 Met Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 385 390 395 400 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 405 410 415 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 420 425 430 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 435 440 445 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 450 455 460 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys 465 470 <210> 69 <211> 233 <212> PRT <213> artificial <220> <223> HuMab2 protein <400> 69 Met Gly Trp Ser Cys Ile Ile Leu Phe Leu Val Ala Thr Ala Thr Gly 1 5 10 15 Val His Ser Asp Ile Gln Met Thr Gln Ser Pro Ser Ser Leu Ser Ala 20 25 30 Ser Val Gly Asp Arg Val Thr Ile Thr Cys Arg Ala Ser Gln Asp Val 35 40 45 Asn Thr Ala Val Ala Trp Tyr Gln Gln Lys Pro Gly Lys Ala Pro Lys 50 55 60 Leu Leu Ile Tyr Ser Ala Ser Phe Leu Tyr Ser Gly Val Pro Ser Arg 65 70 75 80 Phe Ser Gly Ser Arg Ser Gly Thr Asp Phe Thr Leu Thr Ile Ser Ser 85 90 95 Leu Gln Pro Glu Asp Phe Ala Thr Tyr Tyr Cys Gln Gln His Tyr Thr 100 105 110 Thr Pro Pro Thr Phe Gly Gln Gly Thr Lys Val Glu Ile Lys Arg Thr 115 120 125 Val Ala Ala Pro Ser Val Phe Ile Phe Pro Pro Ser Asp Glu Gln Leu 130 135 140 Lys Ser Gly Thr Ala Ser Val Val Cys Leu Leu Asn Asn Phe Tyr Pro 145 150 155 160 Arg Glu Ala Lys Val Gln Trp Lys Val Asp Asn Ala Leu Gln Ser Gly 165 170 175 Asn Ser Gln Glu Ser Val Thr Glu Gln Asp Ser Lys Asp Ser Thr Tyr 180 185 190 Ser Leu Ser Ser Thr Leu Thr Leu Ser Lys Ala Asp Tyr Glu Lys His 195 200 205 Lys Val Tyr Ala Cys Glu Val Thr His Gln Gly Leu Ser Ser Pro Val 210 215 220 Thr Lys Ser Phe Asn Arg Gly Glu Cys 225 230 <210>70 <211>470 <212>PRT <213>Artificial <220> <223>HuMab2 heavy chain <400>70 Met Gly Trp Ser Cys Ile Ile Leu Phe Leu Val Ala Thr Ala Thr Gly 1 5 10 15 Val His Ser Glu Val Gln Leu Val Glu Ser Gly Gly Gly Leu Val Gln 20 25 30 Pro Gly Gly Ser Leu Arg Leu Ser Cys Ala Ala Ser Gly Phe Asn Ile 35 40 45 Lys Asp Thr Tyr Ile His Trp Val Arg Gln Ala Pro Gly Lys Gly Leu 50 55 60 Glu Trp Val Ala Arg Ile Tyr Pro Thr Asn Gly Tyr Thr Arg Tyr Ala 65 70 75 80 Asp Ser Val Lys Gly Arg Phe Thr Ile Ser Ala Asp Thr Ser Lys Asn 85 90 95 Thr Ala Tyr Leu Gln Met Ash Ser Leu Arg Ala Glu Asp Thr Ala Val 100 105 110 Tyr Tyr Cys Ser Arg Trp Gly Gly Asp Gly Phe Tyr Ala Met Asp Tyr 115 120 125 Trp Gly Gln Gly Thr Leu Val Thr Val Ser Ser Ala Ser Thr Lys Gly 130 135 140 Pro Ser Val Phe Pro Leu Ala Pro Ser Ser Lys Ser Thr Ser Gly Gly 145 150 155 160 Thr Ala Ala Leu Gly Cys Leu Val Lys Asp Tyr Phe Pro Glu Pro Val 165 170 175 Thr Val Ser Trp Asn Ser Gly Ala Leu Thr Ser Gly Val His Thr Phe 180 185 190 Pro Ala Val Leu Gln Ser Ser Gly Leu Tyr Ser Leu Ser Ser Val Val 195 200 205 Thr Val Pro Ser Ser Ser Leu Gly Thr Gln Thr Tyr Ile Cys Asn Val 210 215 220 Asn His Lys Pro Ser Asn Thr Lys Val Asp Lys Lys Val Glu Pro Pro 225 230 235 240 Lys Ser Cys Asp Lys Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu 245 250 255 Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp 260 265 270 Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val Val Val Asp 275 280 285 Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly 290 295 300 Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn 305 310 315 320 Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu His Gln Asp Trp 325 330 335 Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro 340 345 350 Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu 355 360 365 Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn 370 375 380 Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile 385 390 395 400 Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr 405 410 415 Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys 420 425 430 Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys 435 440 445 Ser Val Met His Glu Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu 450 455 460 Ser Leu Ser Pro Gly Lys 465 470 <210>71 <211>5428 <212>DNA <213>Artificial <220> <223>pcDNA3.1(+) cloning vector <400>71 gacggatcgg gagatctccc gatcccctat ggtgcactct cagtacaatc tgctctgatg 60 ccgcatagtt aagccagtat ctgctccctg cttgtgtgtt ggaggtcgct gagtagtgcg 120 cgagcaaaat ttaagctaca acaaggcaag gcttgaccga caattgcatg aagaatctgc 180 ttagggttag gcgttttgcg ctgcttcgcg atgtacggg...
Claims
1. A method for transcription, comprising the following steps: a) Providing a nucleic acid construct comprising, in the 5' to 3' direction: a first promoter, a first intron sequence containing a donor splicing site, a second promoter, a second intron sequence containing a recipient splicing site, and a separate nucleotide sequence of interest, wherein the first promoter and the second promoter are constitutive promoters and operatively linked to the same separate nucleotide sequence of interest, and wherein the second promoter is an intron promoter flanked by a first intron sequence located 1-1000 nucleotides upstream of the second promoter and a second intron sequence located 1-1000 nucleotides downstream of the second promoter; wherein the sequence comprising the first promoter and the first intron sequence is as shown in SEQ ID NO:1, and wherein the second intron sequence is as shown in SEQ ID NO:19, wherein the second promoter is as shown in SEQ ID NO:57 or as shown in SEQ ID NO:58 along its entire length; and b) Contacting cells with the nucleic acid construct to obtain transformed cells; and c) Cause the transformed cells to produce a transcript of the nucleotide sequence of interest.
2. A method for transcription and purification of the resulting transcript, comprising the following steps: a) Providing a nucleic acid construct comprising, in the 5' to 3' direction: a first promoter, a first intron sequence containing a donor splicing site, a second promoter, a second intron sequence containing a recipient splicing site, and a separate nucleotide sequence of interest, wherein the first promoter and the second promoter are constitutive promoters and operatively linked to the same separate nucleotide sequence of interest, and wherein the second promoter is an intron promoter flanked by a first intron sequence located 1-1000 nucleotides upstream of the second promoter and a second intron sequence located 1-1000 nucleotides downstream of the second promoter; wherein the sequence comprising the first promoter and the first intron sequence is as shown in SEQ ID NO:1, and wherein the second intron sequence is as shown in SEQ ID NO:19, wherein the second promoter is as shown in SEQ ID NO:57, or wherein the second promoter is as shown in SEQ ID NO:58; and b) Contacting cells with the nucleic acid construct to obtain transformed cells; and c) Causing the transformed cells to produce a transcript of the nucleotide sequence of interest; and d) Purify the resulting transcript.
3. A method for expressing a protein or polypeptide of interest, comprising steps a) and b) of claim 1, wherein the nucleotide sequence of interest encodes the protein or polypeptide of interest; and additionally, c) Cause the transformed cells to express the protein or polypeptide of interest.
4. A method for expressing and purifying a protein or polypeptide of interest, comprising steps a) and b) of claim 1, wherein the nucleotide sequence of interest encodes the protein or polypeptide of interest; and additionally, c) Inducing the transformed cells to express the protein or polypeptide of interest; and d) Purify the protein or polypeptide of interest.
5. A nucleic acid construct comprising, in the 5' to 3' direction: a first promoter, a first intron sequence containing a donor splicing site, a second promoter, and a second intron sequence containing a recipient splicing site, wherein the first and second promoters are constitutive promoters, and wherein the second promoter is an intron promoter flanked by a first intron sequence located 1-1000 nucleotides upstream of the second promoter and a second intron sequence located 1-1000 nucleotides downstream of the second promoter. The sequence comprising the first promoter and the first intron sequence is shown in SEQ ID NO:1, and wherein the second intron sequence is shown in SEQ ID NO:19, wherein the second promoter is shown in SEQ ID NO:57 or 58.
6. A nucleic acid construct comprising, in the 5' to 3' direction: a first promoter, a first intron sequence including a donor splicing site, a second promoter, a second intron sequence including a recipient splicing site, and a nucleotide sequence of interest, wherein the first and second promoters are constitutive promoters and configured to be operatively ligated to the same individual nucleotide sequence of interest, and wherein the second promoter is an intron promoter flanked by a first intron sequence located 1-1000 nucleotides upstream of the second promoter and a second intron sequence located 1-1000 nucleotides downstream of the second promoter, wherein... The sequence comprising the first promoter and the first intron sequence is shown in SEQ ID NO:1, and wherein the second intron sequence is shown in SEQ ID NO:19, wherein the second promoter is shown in SEQ ID NO:57 or 58.
7. The nucleic acid construct of claim 6, wherein the nucleic acid construct further comprises an additional expression regulatory sequence, wherein the additional expression regulatory sequence, the first promoter, and the second promoter are configured to be operatively linked to the nucleotide sequence of interest.
8. The nucleic acid construct of claim 7, wherein the additional expression regulatory sequence comprises an intron.
9. The nucleic acid construct according to claim 6, wherein the nucleotide sequence of interest encodes a protein or polypeptide of interest.
10. The nucleic acid construct according to claim 9, wherein the protein or polypeptide of interest is a heterologous protein or polypeptide.
11. An expression vector comprising a nucleic acid construct according to any one of claims 5-10.
12. A cell comprising a nucleic acid construct according to any one of claims 5-10.
13. Use of the nucleic acid construct according to any one of claims 5-10 for transcription and / or expression of a nucleotide sequence of interest.
14. A cell comprising the expression vector according to claim 11.
15. Use of the cell of claim 14 for transcription and / or expression of a nucleotide sequence of interest.
16. Use of the expression vector of claim 11 for transcription and / or expression of a nucleotide sequence of interest.