Methods of Eukaryotic Gene Expression

JP2024523829A5Pending Publication Date: 2025-06-27CAMBRIDGE ENTERPRISE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023575545
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-21
Filing Date
2022-06-20
Publication Date
2025-06-27

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method of adapting or modifying a complementary DNA (cDNA) sequence for expression in a eukaryotic cell. A nucleic acid molecule is provided that comprises a cDNA sequence that includes two or more splicing consensus motifs that divide the cDNA sequence into exon regions of 50 to 1200 nucleotides. A heterologous intron is then inserted into the splicing consensus motif of the cDNA sequence, each heterologous intron including a 3' region that has a GC content equal to or lower than the GC content of the 5' region of the exon region immediately downstream. This produces a nucleic acid molecule that includes a modified cDNA sequence for expression in a eukaryotic cell. Methods, recombinant nucleic acids that include the cDNA sequence, expression vectors that include the recombinant nucleic acid, and eukaryotic cells that include the recombinant nucleic acid or expression vector are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the manipulation of transgene cDNA sequences to increase expression in eukaryotic cells. [Background technology]

[0002] Knowledge of the cis-acting elements required for gene expression has been accumulated over decades, beginning with an initial understanding of bacteriophage and bacterial systems, and extending these to eukaryotic viruses and finally to eukaryotic genomes. The knowledge has been gradually enhanced and refined by transferring ectopic transcription units from one genome to another. Initially cultured mammalian cell lines were used for this purpose, but since the early 1980s transgenic animals have provided a convenient assay system to explore the regulatory aspects of transgene expression. Transgenic mice have been used to define cis-acting regulatory elements in terms of their ability to direct appropriate levels of expression in the correct tissues and times. Most conveniently, this investigation has used transgenes obtained from other species (LacZ, GFP, etc.) that label the cells in which expression occurs (Chalfie et al. 1994, Schmidt et al. 1998). Through such methods, promoters and enhancers responsive to endogenous regulatory circuits have been determined for many genes. In addition, other important elements such as locus control regions, often found substantially away from genes, have been recognized that allow copy number-dependent gene expression for transgenes integrated at ectopically located sites. At the nucleotide level, enhancement has been achieved by optimizing translation, e.g., the "Kozak" consensus start site (Kozak 1984) is almost universally used.

[0003] Mammalian genes are typically large, their coding sequences distributed across tens to hundreds of kilobases of genomic DNA, and the regulatory elements required to maximize transgene expression can often be located substantially away from the transcription unit. Thus, transgenes designed to express such sequences are typically reduced to their minimal size by removing sequences whose contribution to gene expression is uncertain or not understood, such as introns and 5' and 3' untranslated sequences, even though they are characteristic of virtually all mammalian genes. Such "trimming" of transgenes has the advantage that the transgene can be squeezed into viral vector systems such as adeno-associated virus, which have packaging size limitations.

[0004] When used as naked DNA, smaller size may result in more efficient transfection, either as a result of more cells taking up the DNA and / or more copies being inserted into the host cell genome. Larger transgene copy numbers are often considered advantageous, as this may, in principle, result in higher levels of gene expression. In fact, methods for selecting cells with increased copy numbers of transfected DNA are often used when gene expression levels have commercial interest. Examples of this include the use of genes such as DHFR and GS, which can be used to select clones with amplified copies of the transgene located directly upstream of the selection cassette (Urlaub et al. 1980, Cockett et al. 1990).

[0005] Other methods to improve gene expression include the use of regulatory sequences that are better suited to the target cells, i.e., using a promoter from the Chinese hamster genome to drive expression in CHO cells. Removal of prokaryotic sequences is also considered advantageous to prevent loss of transgene expression (Haruyama et al. 2009). Similarly, coding sequences can be "optimized" to introduce a balance of codons that is more similar to that of the cell line / organism species of interest, rather than that used by the origin species (Gustafsson et al. 2004). By removing rare codons, in principle, translation rates can be improved, but this can have other less desirable characteristics, as folding of complex molecules may be more rate-limiting than translation itself.

[0006] Despite these and other innovations, the process of isolating high-yielding cell lines with stable expression over many generations is tedious, slow and expensive, typically requiring screening of thousands of clones to find one with the right characteristics. Such cell lines are always derived empirically, but in some cases, the characteristics of the integration site are also investigated - for example, transgene integrations in so-called "methylation canyons" may be less susceptible to silencing than integrations in more methylated regions.

[0007] The importance of intron sequences in the context of eukaryotic gene expression was recognized over 40 years ago (Hamer et al. 1979), and since then, a variety of related processes including early gene transcription, rate of transcription, polyadenylation, nuclear export, RNA editing, translation efficiency, and mRNA decay have been shown to be influenced by introns (Le Hir et al. 2003, Shaul 2017). This understanding has also led to the current common practice of including 5'UTR introns in standard transgene expression systems.

[0008] Although there are many examples of intron-mediated expression enhancement, our understanding in this field is still incomplete and various contradictory results have been reported. For example, in some cases, different introns placed identically within a single gene had opposite effects on protein expression (Bourdon et al. 2001), and sometimes the same intron placed within different positions of a cDNA sequence also had opposite results (Buchman et al. 1988, Bourdon et al. 2001). There are examples of introns negatively affecting gene expression directly or indirectly (Gromak 2012, Jin et al. 2017), and the magnitude of intron-dependent positive effects also varies greatly, from almost nothing to more than a 400-fold increase in mRNA levels (Buchman et al. 1988, Bourdon et al. 2001). In an attempt to understand the underlying contradictions, a recent publication concluded that introns only improve expression of AT-rich cDNA sequences and do not benefit GC-rich sequences (Mordstein et al. 2020).

[0009] Although most endogenous genes in higher eukaryotes contain many introns (Piovesan et al. 2019), the expression benefits of adding multiple introns to transgenes remain controversial and have not been implemented in common practice. Several reports have described expression enhancement using constructs with multiple endogenous introns, also known as minigenes (Virts et al. 2001). The use of two heterologous introns in mammalian cells (Lacy-Hulbert et al. 2001) and multiple introns in plants (Marillonnet et al. 2010, Grutzner et al. 2021) has been reported to improve mRNA and protein expression, but the basis of the effect described by Lacy-Hulbert et al. is not understood and appears to be specific to the reported case and cannot be applied to similar situations. Various reports have detailed that the addition of more introns did not provide any additional benefit to expression levels (Crane et al. 2019). US9708636B2 (Enenkel2017) reports the insertion of one or more artificial introns to enhance gene expression and advises using preferably only one intron to reduce the risk of alternative splicing. Furthermore, their examples of intronization are limited to the location of endogenous introns within the cDNA.

[0010] It is remarkable that 20 years after the human genome was decoded and most gene structures defined at the nucleotide level, the underlying rules that allow transcripts to be correctly spliced ​​are not understood. Although intron-exon boundaries are highly or completely conserved in species as distant as humans and mice (approximately 100 million years of evolution), it remains impossible to predict with certainty where in the genomic sequence an intron will be located without access to the mRNA sequence and aligning it to the genome. Thus, it is not possible to de novo design gene structures that reliably and reproducibly produce designed spliced ​​products in an experimental setting. The fact that intron / exon junctions are so highly conserved across species teaches that there is a very strong evolutionary selection to maintain the status quo. Moreover, this conservation poses a very serious obstacle to deciphering the rules that allow cells to decide what to splice and what to keep in the transcript.

[0011] Recent reports indicate that the definition of endogenous introns and exons in humans and other vertebrates is not uniform and different splicing factors are used in different genomic contexts (Amit et al. 2012, Lemaire et al. 2019), highlighting the fact that introns are not uniform in genomes and may not function well in different genomic contexts such as transgenes. Amit et al. observed that genes in low GC% genomic regions tend to have large AT-rich introns with clear GC% gradients at intron-exon junctions (Amit et al. 2012). Wang et al. (2014) arXiv:1404.2487 [q-bio.GN] reported that grouping exons by the GC content of the adjacent introns indicates that the average exon size is positively correlated with GC content. Summary of the Invention

[0012] The present inventors have developed methods to modify transgenes to increase their expression in eukaryotic cells through the incorporation of multiple heterologous introns to generate exonic regions of defined length with a defined gradient of GC content across intron / exon boundaries. These methods can be useful for in vitro and in vivo expression of proteins, such as recombinant protein production, gene therapy, and nucleic acid or virus-based vaccination. These methods can also be useful for generating transgenic animals, for example, or reprogramming or engineering cells, such as T cells and other immune cells, for example, through recombinant expression of chimeric antigen receptors or other antigen receptors, in in vitro and in vivo transfection systems.

[0013] A first aspect of the invention provides a method of adapting or modifying a complementary DNA (cDNA) sequence for expression in a eukaryotic cell, comprising the steps of: providing a nucleic acid molecule comprising a cDNA sequence, providing a cDNA sequence comprising two or more splicing consensus motifs dividing the cDNA sequence into exon regions of 50 to 1200 nucleotides; inserting heterologous introns into splicing consensus motifs of a cDNA sequence, each heterologous intron comprising a 3' region having a GC content equal to or lower than the GC content of the 5' region of the immediately downstream exonic region; thereby producing a nucleic acid molecule comprising the modified cDNA sequence for expression in a eukaryotic cell.

[0014] A second aspect of the invention is a recombinant nucleic acid comprising a cDNA sequence for expression in a eukaryotic cell, The cDNA sequence contains two or more heterologous introns and three or more exon regions of 50 to 1200 nucleotides; A recombinant nucleic acid is provided in which each heterologous intron comprises a 3' region having a GC content that is equal to or lower than the GC content of the 5' region of the immediately downstream exon region.

[0015] A third aspect of the invention provides an expression vector comprising a recombinant nucleic acid of the second aspect.

[0016] A fourth aspect of the invention provides a eukaryotic cell comprising a recombinant nucleic acid of the second aspect or an expression vector of the third aspect.

[0017] Other aspects and embodiments of the invention are described in more detail below. [Brief description of the drawings]

[0018]

Figure 1

Figure 2

Figure 3

Figure 4-1

Figure 4-2

Figure 5-1

Figure 5-2

Figure 6-1

Figure 6-2

[0019] The methods described herein relate to the modification of transgenes for expression in eukaryotic cells. The transgenes may comprise cDNA sequences. Heterologous introns are inserted into the splicing consensus motifs of the cDNA sequences, so that the cDNA sequences are divided into exon regions of defined length. All or part of each heterologous intron nucleic acid has a sequence with a GC content that is equal to or lower than the GC content of all or part of the exon region immediately downstream. In some embodiments, a gradient of GC content can be generated across the intron / exon boundaries of the modified cDNA sequences.

[0020] The modified cDNA sequences produced as described herein may exhibit increased expression in eukaryotic cells relative to unmodified cDNA sequences. In some embodiments, the amount of potential splicing that occurs when the modified cDNA sequence is expressed in a eukaryotic cell may be less than the amount that occurs when the unmodified cDNA sequence is expressed. This reduction in potential splicing may result in increased production of correctly spliced ​​transcripts and increased expression in eukaryotic cells. For example, the modified cDNA sequence may exhibit at least a 10%, at least a 20%, at least a 30%, at least a 40%, at least a 50%, at least a 100%, at least a 200%, or at least a 500% increase in expression relative to the unmodified cDNA sequence.

[0021] Expression of the cDNA sequences can be determined by any suitable technique, either at the mRNA or protein expression level.

[0022] In some embodiments, expression of a cDNA sequence may be determined by measuring the level or amount of mRNA transcribed from the cDNA, for example, the steady-state transcript number of full-length cytoplasmic mRNA transcribed from the cDNA may be compared to a standard or set of standards.

[0023] Cytoplasmic full-length mRNA can be captured by standard techniques such as RNA sequencing, either with no amplification, with low amplification, or with controls for amplification bias. In some embodiments, Shashimi plots can be used to visualize read density across exons and splicing artifacts.

[0024] In other embodiments, expression of a cDNA sequence may be determined by measuring the level or amount of protein produced from the cDNA sequence. For example, the level or amount of secreted protein may be determined as molecules per cell per day compared to a standard or set of standards. The level or amount of protein may be determined using routine techniques such as ELISA or surface plasmon resonance (SPR), Western blot, mass spectrometry, size exclusion chromatography (SEC), and comparison to a standard curve. In some embodiments, biological activity may be assessed in comparison to a standard. For example, factor VIII may be quantified in a thrombin generation assay [TGA], and viral proteins such as viral spike proteins may be quantified in a pseudotyped virus assay. The level or amount of protein retained on the surface of the cells may be determined by any suitable technique, such as antibody staining and a shift in the average intensity of the transfected cell population. Improved expression may also be indicated by higher transfection efficiency, when more cells reach the threshold at which the transgene product is detectable in the assay.

[0025] The cDNA sequences described herein are the nucleotide sequences of the exons of a gene.cDNA may correspond to the sequence of mRNA expressed as DNA bases.cDNA may be produced by any suitable technique and is not limited to sequences produced by reverse transcription of mRNA.

[0026] The cDNA sequence can be expressed to produce a gene product such as a protein or a non-coding RNA molecule, e.g., shRNA or long non-coding RNA (IncRNA). The cDNA sequence of a non-coding RNA can consist of a non-coding nucleotide sequence that is transcribed but not translated in eukaryotic cells.

[0027] In some preferred embodiments, the cDNA sequence may include a coding sequence that codes for the amino acid sequence of a protein. The cDNA sequence may be transcribed and translated in eukaryotic cells after expression of the cDNA to produce the encoded protein. The cDNA sequence may further include one or more non-coding sequences that are transcribed but not translated in eukaryotic cells. The non-coding sequences may include 5' and 3' untranslated regions (UTRs) and poly A tails. In some embodiments, the cDNA sequence may lack endogenous introns from a gene. For example, an unmodified cDNA sequence may consist of a contiguous nucleotide sequence of an exon of a gene. In other embodiments, the cDNA sequence may further include one or more endogenous introns from a gene. Suitable endogenous introns exhibit the GC content and spacing of heterologous introns described herein. For example, in addition to two or more heterologous introns, the modified cDNA sequence described herein may further include one or more endogenous introns.

[0028] The coding sequence of a cDNA sequence may code for a gene product such as a protein. The cDNA sequence may code for any protein for which increased expression or overexpression is desired. Suitable gene products include therapeutic proteins such as clotting factors, enzymes, toxins, hormones, antibody molecules, cytokines, receptors such as PD-1, e.g., T-cell receptors, and chimeric antigen receptors. In other cases, suitable gene products include proteins with non-therapeutic applications, such as industrially relevant proteins, e.g., proteins involved in the manufacture of chemicals, flavors, and foods. Modifications of the cDNA sequences described herein may be useful in maximizing yields in the manufacture of therapeutic or non-therapeutic proteins, or in increasing the expression of therapeutic or non-therapeutic proteins in vivo. Other suitable gene products include antigenic proteins, such as viral, bacterial, and parasitic protein antigens, and tumor antigens. Viral protein antigens may include coronavirus proteins, such as the coronavirus spike (S) protein (e.g., SARS-CoV-2 S protein). Tumor antigens may include tumor-specific and tumor-associated antigens. Other suitable gene products include research proteins, e.g., gene editing proteins such as Cas9, and fluorescent proteins such as GFP.

[0029] The cDNA sequence may be any suitable length to code for the gene product of interest. For example, a suitable cDNA sequence may be 200 nucleotides or more, 240 nucleotides or more, 300 nucleotides or more, 400 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 1500 nucleotides or more, or 2000 nucleotides or more. In some embodiments, longer cDNA sequences, such as 1000 nucleotides or more, may be preferred for intronization as described herein.

[0030] The cDNA sequence suitable for modification as described herein may be derived from any source. For example, the cDNA sequence may be an artificial sequence, an archaeal sequence, a viral sequence, a bacterial sequence, or a eukaryotic sequence, such as a mammalian sequence, such as a protozoan or mesozoan sequence. In some embodiments, the cDNA sequence suitable for modification as described herein may be derived from a source that is not exposed to a cell nucleus, such as a bacterial cDNA sequence or a cytoplasmic viral cDNA sequence.

[0031] Suitable cDNA sequences can be codon-optimized for expression in host eukaryotic cells.For example, the codons in the cDNA sequence of the cDNA can be modified to reflect the codon usage bias of the host eukaryotic cell.Technologies for codon optimization are readily available in the art.

[0032] The cDNA sequences described herein may be operably linked to suitable regulatory elements to form a transgene.

[0033] The cDNA sequence is modified by the incorporation of a heterologous intron as described herein. The incorporation of a heterologous intron as described herein may be referred to as "intronization." The intronized cDNA sequence may be transcribed in a eukaryotic cell to produce a pre-mRNA molecule that contains a heterologous intron. The intron is then removed from the pre-mRNA during splicing in the eukaryotic cell to generate an mRNA molecule that contains the cDNA sequence for translation along with a 5' CAP, 5' and 3' untranslated regions (UTRs), and a polyA tail.

[0034] Heterologous nucleic acid is a nucleic acid that is foreign to a particular gene or other biological system and does not naturally occur in that system. Heterologous nucleic acid, such as a heterologous intron, can be introduced into a gene or other biological system using artificial means, e.g., recombinant techniques. For example, a heterologous intron is inserted into the cDNA sequence of a gene at a position that does not naturally occur there.

[0035] Heterologous introns can be artificial or naturally occurring. For example, heterologous introns can naturally occur in a gene different from the cDNA sequence. The different gene can be the same or different species as the cDNA sequence, for example, the different gene can be a corresponding gene in a species different from the cDNA sequence. In some embodiments, heterologous introns can naturally occur in the same gene in the same species as the cDNA sequence, but can be inserted at a different position within the cDNA sequence. For example, the order of introns in the modified cDNA sequence can be changed relative to the gene in which the intron and cDNA sequence naturally occur.

[0036] The cDNA sequences modified as described herein can be expressed in eukaryotic cells. Suitable eukaryotic cells include higher eukaryotic cells, such as higher plant cells, insect cells and organelle cells, such as mammalian cells.

[0037] Suitable eukaryotic cells include isolated cell lines used for the production of recombinant proteins, e.g., mammalian cells such as Chinese hamster ovary (CHO) cells, baby hamster kidney cells (BHK), mouse myeloma cells (NS / O), and human embryonic kidney (HEK) cells.

[0038] Other suitable eukaryotic cells include in vivo host cells, e.g., cells in a human or non-human individual. Expression of the cDNA sequences modified as described herein in in vivo host cells can be useful, for example, for gene therapy, immunotherapy such as vaccination, and for the production of transgenic non-human animals.

[0039] Other suitable eukaryotic cells include ex vivo host cells, e.g., cells obtained from a human or non-human individual. Expression of cDNA sequences modified as described herein in ex vivo host cells can be useful, for example, in the production of cells for cell therapy, such as hematopoietic stem cells and immune cells, such as T cells and NK cells.

[0040] Suitable eukaryotic cells include isolated cell lines used for the industrial production of recombinant proteins, for example yeast cells such as S. cerevisiae cells or Pichia pastoris cells, as well as insect cells such as Trichoplusia ni cells.

[0041] The cDNA sequence of the transgene is modified as described herein to more closely correspond to the structure of the endogenous gene in eukaryotic cells. Without being bound by theory, mimicking the endogenous gene structure may reduce the amount of potential splicing that occurs during expression of the cDNA sequence in eukaryotic systems and increase the amount of gene product produced. The modified cDNA sequence may be of any suitable length for cloning and delivery into eukaryotic cells.

[0042] Heterologous introns divide a cDNA sequence into exon regions, with each heterologous intron having an upstream (5') and downstream (3') exon region. Splicing of the heterologous intron during expression in a eukaryotic cell removes the intron and reconnects the exon regions to generate an mRNA molecule that contains the exon regions in a contiguous sequence.

[0043] The number of heterologous introns inserted into a cDNA sequence depends on the size of the cDNA sequence and the number of introns required to divide it into exon regions of 50 to 1200 nucleotides. For example, a cDNA sequence can be modified to contain 2, 3, 4, 5, 6, 7, 8, 9, 10 or more heterologous introns.

[0044] A cDNA sequence suitable for modification as described herein may contain 2, 3, 4, 5, 6, 7, 8, 9, 10 or more splicing consensus motifs. A splicing consensus motif is a site at which a heterologous intron is inserted into a cDNA sequence. A heterologous intron may be inserted into a splicing consensus motif within the cDNA sequence or a UTR of the cDNA sequence.

[0045] A splicing consensus motif is a nucleotide sequence within a cDNA sequence that includes a donor splice site exonic element occurring at the 5' end of an intron (5' exonic element) and an acceptor splice site exonic element occurring at the 3' end of an intron (3' exonic element). A heterologous intron may be inserted into the splicing consensus motif between the 5' and 3' exonic elements to generate an intronized cDNA sequence that includes a heterologous intron with a donor splice site at the 5' end and an acceptor splice site at the 3' end. The splicing consensus motif may be frame independent and may be present in any reading frame of the cDNA sequence. Suitable splicing consensus motifs are known in the art and may include the nucleotide sequence (C / A / G)AG↑G(T / N)(T / N), preferably CAG↑GTT (the insertion site of the heterologous intron between the 5' and 3' exonic elements is indicated). Other suitable splicing consensus motifs include ATG↑AAT, CAG↑GTT, GAG↑ATT, CAG↑GCC, CAG↑GAT, GAA↑GCG, GTT↑CAA, CAT↑ATG, and CAG↑GAT. Splicing consensus motifs can be readily identified in cDNA sequences using standard techniques.

[0046] The splicing consensus motifs may divide the cDNA sequence into exon regions of 50 to 1200 nucleotides in length, more preferably 80 to 380 nucleotides in length. In some preferred embodiments, the exon regions in the modified cDNA sequence may be 50 to 250 or 100 to 150 nucleotides in length.

[0047] The exon region may be an artificial exon generated in a cDNA sequence by inserting a heterologous intron into a consensus splicing motif. The cDNA sequence is divided by a heterologous intron into exon regions that together code a gene product. In some embodiments, the cDNA sequence may include one or more endogenous introns that define one or more of the exon regions of the modified cDNA sequence.

[0048] In some embodiments, suitable splicing consensus motifs for dividing a cDNA sequence into exon regions may exist in the cDNA sequence or may exist in advance. The method described herein may include identifying splicing consensus motifs in a cDNA sequence. Sequence analysis tools for identifying splicing consensus motifs are readily available in the art.

[0049] In other embodiments, the cDNA sequence may lack one or more of the splicing consensus motifs required to divide the cDNA sequence into exon regions. The splicing consensus motifs may be generated in the cDNA sequence by the introduction of one or more mutations to modify the existing cDNA sequence. Preferably, the one or more mutations generate one or more splicing consensus motifs without modifying the sequence of the encoded protein. In some embodiments, the one or more mutations may also optimize the codons in the cDNA sequence for expression in eukaryotic cells. In other embodiments, the one or more mutations may modify the sequence of the encoded protein, for example, to increase or modify its activity.

[0050] The heterologous intron can be inserted between the 5' and 3' exon elements of the splicing consensus motif of the cDNA sequence.

[0051] A suitable heterologous intron may be 30-400 nucleotides in length, preferably 60-120 nucleotides in length, or 80-100 nucleotides in length. The optimal intron length may depend on the eukaryotic host cell and may be optimized for expression in any particular eukaryotic host cell.

[0052] A heterologous intron is 5' splicing donor sequence, 3' splicing acceptor sequence, Polypyrimidine tract (PPT), Branch point sequences, and It may include a 3' region that has a GC content equal to or lower than the 5' region of the exon region immediately downstream of the splicing consensus motif into which the intron is inserted.

[0053] The heterologous intron may comprise a splicing donor sequence and a splicing acceptor sequence at the 5' and 3' ends of the intron, respectively. The splicing donor sequence defines the 5' end of the intron, and the splicing acceptor sequence defines the 3' end of the intron. A suitable splicing donor sequence may, for example, comprise a GT dinucleotide. A suitable splicing donor sequence may, for example, comprise an AG dinucleotide. The splicing donor and splicing acceptor sequences of the heterologous intron may be optimized for the eukaryotic cell in which the cDNA sequence is expressed.

[0054] The heterologous intron may further comprise a polypyrimidine tract (PPT). The polypyrimidine tract may be located upstream of the 3' end of the heterologous intron, for example, 5-40 nucleotides upstream of the 3' end. The polypyrimidine tract may comprise a sequence of 15-20 nucleotides rich in pyrimidines (C and U). Suitable PPTs include 5'-UUUUUUUCCCUUUUUUUCC-3' and variants thereof. Other suitable PPTs are known in the art (see, for example, Wagner et al 2001 Mol Cell Biol 21(10):3281-3288, WO2017 / 171654A1).

[0055] The heterologous intron may further comprise a branchpoint sequence. The branchpoint sequence may be located upstream of the 3' end of the intron nucleic acid, for example, 20-50 nucleotides upstream of the 3' end. The branchpoint sequence may comprise the sequence YURAC or YNURAC, where R=purine, Y=pyrimidine, and N=any nucleotide. Suitable branchpoint sequences include 5'-UACUAACA-3' and are known in the art (see, for example, Gao et al Nucl Acid Res 2008 36(7)2257-2267, US20060094675).

[0056] GC content is the proportion of guanine or cytosine nucleotides in a nucleic acid sequence (i.e., (G+C) / total nucleotides), generally expressed as a percentage (GC%). In some embodiments, the insertion of a heterologous intron as described herein may generate a GC content gradient between the heterologous intron and the immediately downstream exonic region (i.e., the exonic region immediately adjacent to the 3' end of the heterologous intron). For example, a heterologous intron inserted into a splicing consensus motif may create a GC content gradient between the 3' region of the heterologous intron and the 5' region of the following exonic region. The heterologous intron may include a 3' region that has a lower GC content than the 5' region of the immediately downstream exonic region. In other embodiments, the heterologous intron may include a 3' region that has a GC content that is the same as the 5' region of the immediately downstream exonic region. A GC content gradient may not be generated between the heterologous intron and the immediately downstream exonic region by the insertion of a heterologous intron as described herein.

[0057] GC content can be measured starting from a junction in the 3' to 5' direction for introns and in the 5' to 3' direction for exons. Suitable tools for measuring GC content are readily available in the art.

[0058] The 3' region of a heterologous intron inserted into a cDNA sequence may have a GC content equal to or at least 1%, at least 2%, at least 4%, at least 6%, at least 8%, at least 10%, at least 15%, or at least 20% lower than the 5' region of the immediately downstream exon region. In some embodiments, the 3' region of a heterologous intron inserted into a cDNA sequence may have a GC content that is 0%-46%, 2%-40%, or 5%-35% lower than the 5' region of the immediately downstream exon region.

[0059] The size of the 3' region of the intron and the 5' region of the downstream exon region (i.e., the window in which the GC content is determined) can be 30 nucleotides or more, 40 nucleotides or more, 50 nucleotides or more, 60 nucleotides or more, 70 nucleotides or more, 80 nucleotides or more, 90 nucleotides or more, or 100 nucleotides or more. In some embodiments, the GC content can be determined throughout the intron and downstream exon region (i.e., the 3' region of the intron and the 5' region of the downstream exon region can consist of the entire intron and exon region, respectively). The GC content of the 3' region of the heterologous intron can be equal to or lower than the 5' region of the immediately downstream exon region described herein for any size of 3' and 5' regions.

[0060] In some preferred embodiments, the 3' region of the heterologous intron and the 5' region of the downstream exon region consist of 30 nucleotides. For example, the 30 nucleotides at the 3' end of the heterologous intron may have a GC content that is equal to or lower than the 30 nucleotides at the 5' end of the downstream exon region, preferably up to 30%, 40%, 45%, 50%, or 60% lower.

[0061] The sequence of the heterologous intron depends on the position in the cDNA sequence into which it is inserted. The GC content of the 5' region of the exon region downstream of the splicing consensus motif can be determined. Thus, the intron sequence for insertion into the splicing consensus motif can be designed to include a 3' region that has a GC content equal to or lower than that of the 5' region of the exon region downstream of the splicing consensus motif, as described herein.

[0062] In some embodiments, the nucleotide sequence of a heterologous intron can be found in a naturally occurring intron, for example, an intron from a different gene or a different location in the same gene.

[0063] In other embodiments, the nucleotide sequence of the heterologous intron may be artificial, i.e., not found in naturally occurring introns. The artificial intron sequence may be designed using any convenient technique. For example, splice donor and splice acceptor sites may be located at the 5' and 3' ends of the nascent intron sequence. A branch point may be introduced in the middle of the nascent sequence. A random combination of T's and C's may be added to the nascent sequence to generate a pyrimidine tract of about 20 nucleotides. A random sequence of 50 or more nucleotides may be added between the pyrimidine tract and the branch point. Additional nucleotides may be added between the splice donor site and the branch point of the nascent sequence. The additional nucleotides may be a random sequence with an A / T content adjusted to generate a GC% content equal to or lower than the 5' region of the exon region downstream of the splicing consensus motif into which the intron is inserted. A suitable artificial intron may be 80-85 nucleotides in length. Suitable intron sequences for use herein are highlighted (in lower case) in SEQ ID NOs: 1-30.

[0064] Suitable heterologous introns for insertion into the splicing consensus motifs can be produced using standard synthetic or recombinant techniques. The methods described herein can include providing a heterologous intron for insertion into two or more splicing consensus motifs in a cDNA sequence.

[0065] In addition to the insertion of a heterologous intron, one or more additional mutations may be introduced into the cDNA sequence, for example, to remove potential splice sites. Potential splice sites may be identified by computational prediction tools that are readily available in the art (see, for example, Alternative Splice Site Predictor (Wang M. and Marin A. (2006)) Gene 366:219-227). Potential splice sites are preferably removed without altering the sequence of the gene product.

[0066] In some embodiments, the methods described herein can include providing a nucleic acid comprising a cDNA sequence and inserting a heterologous intron into the cDNA sequence of the nucleic acid described herein to generate a nucleic acid comprising a modified cDNA sequence. The heterologous intron can be synthesized and inserted using standard techniques.

[0067] In other embodiments, a cDNA sequence modified to contain a heterologous intron can be designed and a nucleic acid containing the modified cDNA sequence synthesized or assembled. For example, methods of adapting a cDNA sequence for expression in eukaryotic cells include: (i) providing a cDNA sequence, the cDNA sequence comprising two or more splicing consensus motifs that divide the cDNA sequence into exon regions of 50 to 1200 nucleotides; (ii) generating heterologous introns for insertion into each splicing consensus motif of a cDNA sequence, each of said introns comprising a 3' region having a GC content equal to or lower than the GC content of the 5' region of the immediately downstream exonic region; (iii) generating a modified cDNA sequence that includes the generated intron inserted into a splicing consensus motif; and (iv) synthesizing a nucleic acid molecule containing the modified cDNA sequence.

[0068] Steps 1-3 can be computer-implemented, for example, using standard sequence analysis software tools.

[0069] Examples of cDNA sequences modified as described herein are shown in SEQ ID NO:6, SEQ ID NO:7, SEQ ID NOs:9-15, SEQ ID NOs:17-21, SEQ ID NO:25, SEQ ID NO:27, and SEQ ID NOs:28-30.

[0070] Also provided are cDNA sequences, nucleic acids, and transgenes modified as described herein. The recombinant nucleic acids described herein can include cDNA sequences for expression in eukaryotic cells, The cDNA sequence contains two or more heterologous introns and three or more exon regions of 50 to 1200 base pairs; Each such heterologous intron comprises a 3' region that has a GC content equal to or lower than the 5' region of the immediately downstream exon region.

[0071] The cDNA sequence of the recombinant nucleic acid can be produced by the methods described herein. In some embodiments, the recombinant nucleic acid or transgene comprising the modified cDNA sequence described herein can be directly inserted into the genome of a eukaryotic cell. For example, the modified cDNA sequence can be knocked into an endogenous locus. Suitable techniques for random or targeted insertion into a genome are well known in the art, and include, for example, CRISPR, Lox / Cre, or transposon-based techniques.

[0072] In other embodiments, the recombinant nucleic acid or transgene comprising the modified cDNA sequence described herein may be cloned and / or incorporated into a nucleic acid construct or vector, such as an expression vector. For example, the cDNA sequence may be operably linked to one or more control elements or regulatory sequences capable of directing the expression of the cDNA sequence. Suitable control elements or regulatory sequences for driving the expression of heterologous nucleic acid cDNA sequences in eukaryotic cells, preferably mammalian cells, are well known in the art and include constitutive promoters, e.g., viral promoters such as CMV or SV40, and tissue-specific promoters, e.g., promoters such as the human thyroxine-binding globulin (TBG) promoter, or system-specific promoters, such as hypoxia-responsive promoters.

[0073] Further provided are constructs in the form of vectors (e.g., expression vectors), transcription or expression cassettes, or other delivery systems, such as plasmids, viral vectors, e.g., phage or phagemid vectors, that contain the adapted or intronized cDNA sequences described herein. For example, the modified or intronized cDNA sequences can be contained in an expression vector. Suitable expression vectors can be selected or constructed as appropriate, containing appropriate regulatory sequences, including promoter sequences, terminator fragments, polyadenylation sequences, enhancer sequences, marker genes, and other sequences. The vectors can also include sequences such as origins of replication, promoter regions, and selectable markers that allow their selection, expression, and replication in bacterial hosts, such as E. coli.

[0074] Preferred vectors will be tropic for the cell type in which expression is required and will contain appropriate control and regulatory elements to enhance specific expression in that cell type.

[0075] The vector may be a plasmid, a virus, such as a phage, or a phagemid, as appropriate. For example, cosmids, BACs, or YACs may be used to accommodate long modified cDNA sequences. For further details, see, for example, Molecular Cloning: a Laboratory Manual: 3rd edition, Russell et al., 2001, Cold Spring Harbor Laboratory Press. Many known techniques and protocols for the manipulation of nucleic acids, for example in the preparation of nucleic acid constructs, mutagenesis, sequencing, DNA introduction into cells, and gene expression, are described in detail in Current Protocols in Molecular Biology, Ausubel et al. eds. John Wiley & Sons, 1992. In some preferred embodiments, the expression vector may be a viral vector, such as a lentivirus or adeno-associated virus (AAV) vector.

[0076] Recombinant nucleic acids, transgenes, or expression vectors can be introduced into eukaryotic cells. Introduction can use any available technique. Suitable techniques can depend on the vector and cell type and can include calcium phosphate transfection, DEAE-dextran, electroporation, liposome-mediated transfection, and transduction using retroviruses or other viruses, e.g., vaccines.

[0077] Nucleic acids can be introduced into host eukaryotic cells using viral or plasmid-based systems. Plasmid systems can be maintained episomally or integrated into the host cell or into artificial chromosomes. Integration can be by either random or targeted integration of one or more copies at single or multiple loci.

[0078] The introduction can be followed by causing or allowing expression of the modified cDNA sequence, for example, by culturing host cells under conditions for expression of the gene.

[0079] Also provided are recombinant eukaryotic cells, e.g., recombinant mammalian cells, that contain a recombinant nucleic acid or vector having a modified cDNA sequence described herein. The cDNA sequence can be expressed in the cell to produce a gene product.

[0080] Systems for cloning and expressing nucleic acids in a variety of different eukaryotic host cells are well known. Suitable host cells include mammalian, insect, and yeast systems. Mammalian cell lines available in the art for the expression of heterologous proteins include Chinese hamster ovary (CHO) cells, baby hamster kidney cells (BHK), mouse myeloma cells (NS / O), and human embryonic kidney (HEK) cells, as well as many others.

[0081] 1. A method for expressing a cDNA sequence in a eukaryotic cell, comprising: modifying the cDNA sequence by a method described herein to produce a modified cDNA sequence and incorporating the modified cDNA sequence into an expression vector; introducing the expression vector into a eukaryotic cell; and causing or allowing expression from the modified cDNA sequence to produce a gene product.

[0082] The cDNA sequence may code for a gene product. After production by expression of the nucleic acid containing the modified cDNA sequence, the gene product may be isolated and / or purified using any suitable technique and then used as appropriate. For example, the production method may further include formulating the product into a composition that includes at least one additional component, such as a pharma- ceutically acceptable excipient.

[0083] Other aspects and embodiments of the invention provide those aspects and embodiments described above in which the term "comprising" is replaced by the term "consisting of," as well as those aspects and embodiments described above in which the term "comprising" is replaced by the term "consisting essentially of."

[0084] As used herein, the term "downstream" refers to the 5' to 3' direction of the nucleic acids described herein, and the term "upstream" as used herein refers to the 3' to 5' direction of the nucleic acids described herein.

[0085] Reference to a nucleotide sequence set forth herein includes DNA molecules having the particular sequence and includes RNA molecules having the particular sequence in which U is replaced by T, unless the context requires otherwise.

[0086] It should be understood that the present application discloses all combinations of any of the above aspects and embodiments with each other unless the context requires otherwise.Similarly, the present application discloses all combinations of preferred and / or optional features alone or with any of the other aspects unless the context requires otherwise.

[0087] Modifications of the above embodiments, further embodiments and modifications thereof will be apparent to those of skill in the art upon reading this disclosure and, therefore, are within the scope of the present invention.

[0088] All documents and sequence database entries mentioned herein are incorporated by reference in their entirety for all purposes.

[0089] As used herein, "and / or" should be interpreted as a specific disclosure of each of the two specified features or components, with or without the other. For example, "A and / or B" should be interpreted as a specific disclosure of (i) A, (ii) B, and (iii) each of A and B, as if each were individually set forth herein.

[0090] Experimental Materials and Methods Transgene constructs Table 1 provides an overview of all transgene constructs with associated 5' and 3' elements and the plasmid backbone used. The complete DNA sequences of these constructs are shown below. The wild type (wt) SARS-CoV-2 S protein CDS sequence refers to the S protein cDNA sequence from the Wuhan-Hu-1 isolate (Genbank: MN908947.3), while "18F" refers to the removal of the last 18 amino acids of the S protein C-terminus (ER retention sequence) and the addition of the FLAG tag. The DNA sequence of the codon-optimized (co) SARS-CoV-2 S protein was obtained from the National Institute for Biological Standards and Control website (nibsc.org, CFAR#100976). mCherry CDS refers to the “synthetic construct monomeric red fluorescent protein gene” (Genbank:AY678264.1), in which the stop codon has been changed from “TAA” to “TGA”, while human ACE2 CDS refers to the “Homo sapiens angiotensin-converting enzyme 2, mRNA transcript variant 2” (Genbank:NM_021804.3).

[0091] All constructs were assembled (Gibson Assembly Master Mix, NEB) by Gibson cloning of relevant PCR products and / or custom-designed gene blocks (gBlocks Gene Fragments, Integrated DNA Technologies) according to the manufacturer's protocols.

[0092] GC% calculations overall and at intron-exon junctions The GC% of the constructs was calculated using a sliding window of 30 base pairs (bp) across the sequence and reset at each element (intron or exon) to highlight their GC% differences. For a given 30bp window, the frequency of G and C nucleotides was measured (equal to the total number of G and C nucleotides in the sequence divided by the sequence length). The window then slides 1bp, moving in the 5' to 3' direction. The last measurement of the element is calculated when the sliding window hits the start of the next element. The window then jumps 30bp to start measuring the GC% at the start of the next element. This gap in the GC% measurement is visualized as a dashed line in Figure 2A.

[0093] The GC% at the intron-exon junctions was calculated for all intron-exon pairs (intron followed by exon). GC% was measured as above for various sequence lengths starting from the junction in the 3' to 5' direction for introns and 5' to 3' direction for exons, shown for the 50 bp segment in Figure 3A. All calculations were performed using in-house python scripts and plotted in R.

[0094] cell line 293FT cells were obtained from the laboratory of Dr. Kosuke Yusa. 293FT.Cas9 cell lines were generated via lentiviral integration of the EF1a-Cas9-T2A-BlastR construct at a low MOI to achieve single copy integration. To generate cell lines permissive for spike-pseudotyped lentivirus infection, 293FT.Cas9 cells were engineered to stably express the SARS-CoV-2 receptors ACE2 and TMPRSS2. PiggyBac transposition was used to integrate the EF1a-ACE2-T2A-TMPRSS2 construct, followed by single cell cloning. This resulted in 293FT.Cas9.ACE2 / TMPRSS2 clonal cell lines. Clones C10 and D10 were used in this study. 293T cells were obtained from the laboratory of Dr. Ravindra Gupta and were primarily used for spike-pseudotyped lentivirus production. The JM8 mouse embryonic stem cell line was derived from a B57BL / 6N blastocyst (Pettitt et al. 2009). MC38 cells were purchased from Kerafast (catalogue 2388609). All cell lines have been tested negative for mycoplasma contamination.

[0095] Cell culture conditions All lines were maintained at 5% CO2 and 37°C unless otherwise stated. 293FT, 293T, and MC38 cell lines were routinely cultured in M10 medium (DMEM, 10% FBS, and 2mM L-glutamine). Cas9-expressing cell lines were maintained in M10 supplemented with 10 μg / mL blasticidin. JM8 cells were maintained in M15 medium (DMEM, 15% FBS, 100 μM b-mercaptoethanol, and 2mM L-glutamine) on a layer of irradiated feeder fibroblasts (SNL76 / 7).

[0096] Cell transfection Cell transfections were performed using Lipofectamine LTX reagent (Invitrogen) according to the manufacturer's instructions. For 6-well format transfections, Lipofectamine::DNA complexes were formulated using 750 ng DNA, 5 μL Plus reagent, and 10 μL Lipofectamine LTX. These were then used to transfect 1.5 million cells per reaction. For analysis of transgene expression, cells were typically harvested with trypsin 48 hours after transfection. Samples were either kept as frozen cell pellets for cDNA analysis or used directly for flow cytometry assays. For MC38 cells, transfections were performed using an Amaxa Nucleofector (Lonza) according to the manufacturer's instructions, using program H-022 and Nucleofector kit V. For transfections, 0.41 pmol of DNA was transfected into 1 million cells per reaction.

[0097] cDNA analysis RNA was extracted from frozen cell pellets using the RNeasy Mini Kit (Qiagen) according to the manufacturer's recommendations and treated with ezDNase (ThermoFisher) before applying oligo(dT)-guided first strand cDNA synthesis using Superscript VI reverse transcriptase (ThermoFisher). RT-PCR was performed using GoTaq Green Master Mix (Promega) following the recommended protocol. The PCR primers used to capture the full length of the investigated transgenes are listed in Table 2.

[0098] PCR products were both visualized on agarose gels and 'TA cloned using the TA Cloning Kit (ThermoFisher) with pCR2.1 vector and OneShot TOP10 Chemically Competent E.coli according to the kit's instructions. After overnight growth on LB plates containing 100 μg / ml ampicillin at 37°C, single colonies were picked in 20 μl of PBS and the respective vector inserts were PCR amplified with M13F (GTAAAACGACGGCCAGT) and M13R (CAGGAAACAGCTATGAC) primers using GoTaq Green Master Mix. These PCR products were purified using AmPure XP magnetic beads (Beckman Coulter) according to the manufacturer's recommendations and submitted for Sanger sequencing (supplied by Source BioScience Inc) using the M13F / M13R primers mentioned above. On average, 24 clones per construct were evaluated by PCR and a further 8 clones were selected for Sanger sequencing. All reads were mapped back to the original construct DNA sequence using SnapGene software to assess individual mRNA splicing events.

[0099] Flow cytometry assay Cells were harvested 48 hours after transfection using trypsin dissociation. For analysis of mCherry expression, cells were directly assessed by flow cytometry. If surface staining was required, upon harvesting, cells were washed twice with staining buffer (see Table 3). They were then incubated with the appropriate dilution of primary antibody (in staining buffer) for 30 minutes at the indicated temperature. Cells were washed twice and incubated with secondary antibody (1:500) for 30 minutes on ice (for unconjugated primary antibodies). After another set of two washes, cells were analyzed by flow cytometry using Cytoflex (BD Biosciences). Data analysis was performed using FlowJo software (BD Biosciences). S protein expression data are plotted as % of positively stained cells (Figure 2). mCherry is shown as % of mCherry expressing cells (Figure 4) and median population mCherry intensity normalized to the intronless construct to highlight the changes in population intensity (Figure 6). The same visualization is used for the S protein and ACE2 constructs in FIG.

[0100] Pseudotyped lentivirus production Pseudotyped lentiviruses were produced by transfection of 293T cells using lipofectamine LTX according to the manufacturer's instructions. All S protein constructs were tested using three independent virus productions. Briefly, one day before transfection, one million 293T cells were seeded in gelatinized 6-well plates. For transfection, 1 μg of lentiviral transfer vector (pCSGW-GFP) was mixed with 0.72 μg of gag-pol expression plasmid p8.9 and 68.33 fmol of S protein expression construct in 500 μL of optiMEM medium, followed by addition of 2 μL of PLUS reagent and incubation at room temperature for 5 minutes. Then, 6 μL of Lipofectamine LTX reagent was added to the mixture and incubated for 10 minutes. The medium was aspirated from the plate, and the Lipofectamine:DNA complex was added dropwise and topped off with 1.5 mL of M10. Production was carried out at 32°C with 5% CO2. The medium was changed to 2.5 mL fresh M10 the following morning and the supernatant was harvested after 56 hours. The virus-containing supernatant was centrifuged at 500g for 5 min to remove cell debris and was either used directly for infection of permissive cell lines or aliquoted and frozen at -80°C.

[0101] Transduction of permissive cell lines Transduction was performed in 96-well plates in duplicate for each independent virus sample. For pseudotyped lentivirus titration, a dilution series was prepared ranging from 100% virus-containing supernatant to a 1:500 dilution in a total volume of 200 μL_M10 medium. 293FT.Cas9.ACE2 / TMPRSS2 clonal cell lines were harvested by trypsinization and resuspended at a density of 70.000 cells per 30 μL. Then, 30 μL were seeded per well, mixed and incubated at 37 °C. Viral infection efficiency was measured after 48-72 h and assessed by the percentage of GFP-positive cells in flow cytometry. Data were analyzed using FlowJo software (BD Biosciences). Pseudotyped lentivirus infection assay data are presented either as % cells infected with the full dose of pseudotyped virus (Figure 5) or at a 1:500 dilution normalized to the infection rate of the intron-less construct (Figure 4).

[0102] result The addition of multiple introns to the SARS-CoV-2 spike protein results in alternatively spliced ​​mRNA products. The wild-type (wt) SARS-CoV-2 spike (S) protein coding sequence (CDS) has proven difficult to express as a transgene (Chen Ling 2020), as has its relative SARS-CoV spike protein (Callendret et al. 2007). To improve its expression, we generated two constructs with introns added to the wt S CDS. Although various intron insertion sites exist in the endogenous gene (and here in the functional transgene), there is a slight preference in the human canonical splicing consensus motif for the sequence "(C / A)AG" before the intron and "G(T / N)(T / N)" after the intron (Sibley et al. 2016). For optimal placement of these introns, we looked for the presence of the 'CAGGTT' nucleotide sequence in the wt S protein or the opportunity to achieve that sequence using codon optimization. In the wt S protein, the amino acid sequence "SGW" at positions 256-258 is encoded by TCA-G|-intron-|GT-TGG', containing the desired sequence (underlined) and thus providing an intron insertion site at the position indicated at G257. A codon-optimized insertion site opportunity is available in the amino acid sequence "DRL" at positions 1184-1186, where the original nucleotide sequence: "GAC-CGC-CTC" can be codon-optimized to the optimal intron insertion site: "GAC-AG|-intron-|G-TTG" at amino acid R1185. The first generated construct (SEQ ID NO:1 (P91), FIG. 1A) had EF1-α intron A (sequence from the EF1-α promoter) inserted between R1185 and the 5'UTR β-globin intron. The second construct (SEQ ID NO:2 (P92), FIG. 1B) had a hybrid chicken β-actin / Minute Virus (sequence from the CBh promoter, (Gray et al. 2011)) with a mouse intron inserted at G257.

[0103] S protein expression was measured 48 hours after transfection into Hek293 cells, but was not detectable on the surface of the cells (data not shown). To investigate why, RT-PCR was performed, which detected several strongly preferred alternatively spliced ​​cDNA products. Sanger sequencing identified these products as consisting of the correct outer exons of the S CDS construct, but the inner coding sequence was not incorporated into the mature transcript. It was removed via alternative splicing, using either canonical or cryptic splicing sites within the introduced intron (Figure 1A-B). These introns are frequently and very successfully spliced ​​within the 5'UTR context, but are not automatically recognized as independent intronic units outside that context and will interact with adjacent introns.

[0104] Since the introns used above do not occur together in the same gene in the natural environment, we then tested consecutive introns from endogenous human genes. The gene PRR36 (Genbank: NM_001190467) was identified as a potentially good intron donor due to its short intron but similar length CDS in relation to the S protein. We generated a vector that inserted all PRR36 introns into the S CDS, maintaining their endogenous 5' to 3' order and their nucleotide sequence context (3 bp before and after the intron, if possible). To some extent, the exon length was consistent with the PRR36 structure (SEQ ID NO: 3 (P113), Figure 1C). Finally, to avoid any exogenous intron interaction, the 5'UTR β-globin intron was replaced with the PRR36 5'UTR sequence. Despite the use of the contiguous endogenous intron of SEQ ID NO:3 (P113), cryptic splicing outcomes persisted, indicating the use of both canonical and cryptic splice sites within the inserted intron and in the S CDS in various combinations (Figure 1C). A second attempt involved the use of all introns from the human gene EMILIN1 (NM_007046), and a third attempt was made using a subset of introns from the human gene TTN (NC_000002.12) in the wt S CDS (data not shown). In both cases, expression was not achieved at measurable levels and extensive cryptic splicing persisted.

[0105] To assess whether the wt S CDS was driving the cryptic splicing observed in SEQ ID NO:3 (P113), two additional constructs were generated. First, 170 point mutations were introduced into the wt S CDS (wt+ss, SEQ ID NO:4 (P136), Figure 1C) to remove all cryptic splicing sites identified in the wt sequence using Alternative Splice Site Predictor (http: / / wangcomputing.com / assp / index.html) while retaining the same amino acid sequence and maintaining a similar GC% landscape. Second, the entire S CDS was codon-optimized (co, SEQ ID NO:5 (P1486), Figure 1C), a common practice for enhancing transgene expression (Figure 2B) that results in improved expression of S in an intronless context. Despite these changes to the S CDS, no or very little full-length cDNA was observed and cryptic splicing continued (Figure 1C).

[0106] Taken together, these data confirm that the addition of multiple introns, either well-defined introns from the literature and databases or commonly used introns present in widely used expression systems, does not result in improved transgene expression due to the underlying problem of alternative splicing. In other words, the assumption that a mammalian gene structure sufficient for robust gene expression can be assembled by simply inserting multiple introns and removing cryptic sites has been shown to be clearly inaccurate.

[0107] GC% landscape allowing clear definition of exons and introns Amit et al. observed that genes in low GC% genomic regions tend to have large AT-rich introns with distinct GC% gradients at intron-exon junctions (Amit et al. 2012). Whether this simply reflects an underlying bias between coding and non-coding introns in low GC% regions or is functionally significant requires experimental evaluation. To test the effect of intron-exon GC% gradients in the context of transgenes with relatively short introns, local and systematic changes were introduced into a non-functional construct, SEQ ID NO:6 (P143). This construct consists of 13 short introns from the human TTN gene inserted into the wt-ss S sequence. Transfection of SEQ ID NO:6 (P143) resulted in a variety of alternative splicing outcomes (Figure 2A) and did not result in a measurable full-sized S protein (Figure 2B). In this construct, the intron 1 / exon 2 junctions had very similar GC% profiles. The GC% of the first 60bp of exon 2 was increased from 38% to 60% (SEQ ID NO:7 (P172)) by selecting codons with the maximum number of G / C nucleotides, if possible. The splicing outcome was significantly affected by this change. From the analysis of the potentially spliced ​​mRNA produced by SEQ ID NO:6 (P143), the failure to recognize and therefore include exon 2 in the final mRNA transcript was evident. However, the GC% increase of exon 2 in SEQ ID NO:7 (P172) was now sufficient to include that region in all identified splicing outcomes (Figure 2A). Extending the same strategy over the entire length of all exons (SEQ ID NO:11 (P171)) resulted in the correct splicing of all 13 introns and further improved the expression of the S protein compared to all previous attempts at intronization as well as to an intron-less transgene with the same CDS sequence (Figure 2B). Taken together, the intron-exon GC% gradient can successfully define intron-exon boundaries in a transgene context.Such a gradient can in principle be achieved either by increasing the GC% of exons using codon optimization (as applied herein in SEQ ID NO: 11 (P171)), or by inserting introns with lower GC% into the invariant CDS sequence (as applied herein in SEQ ID NO: 6 (P143)), or a combination of both.

[0108] To further characterize what defines a functional intron-exon junction, GC% was calculated for different length segments of DNA (10-80 bp + full length of the element) measured outward from the junction for 29 adjacent intron-exon pairs from three different correctly splicing constructs (SEQ ID NO:11 (P171), SEQ ID NO:14 (P186), SEQ ID NO:25 (P237), SEQ ID NO:30 (P243), Figure 3A). The percentage of G / C nucleotides in exons, as well as the inserted introns, varied both within and between the different transgenes (20-80%), and the overall GC% range was both broad (10-52%) and overlapping with exons (Figure 3B). When adjacent introns and exons were evaluated pairwise, exons had at least equal, and usually higher, GC% compared to the preceding intron (Figure 3C). This feature was consistently present for all segment lengths starting from 30 bp, indicating that a length of 10–20 bp may be too short a measurement window for an accurate assessment of the GC% landscape.

[0109] Defining optimal exon length After resolving the important landscape requirements for correct intron and exon recognition in the transgene, the optimal number of introns can be addressed. Increasing numbers of introns were inserted into the sequence SEQ ID NO:8 (P166): 3, 7, 13, 14, 15 introns (SEQ ID NO:9 (P205), SEQ ID NO:10 (P204), SEQ ID NO:11 (P171), SEQ ID NO:12 (P231), SEQ ID NO:13 (P232), Figure 4A) and the impact on S protein expression was evaluated in a functional assay of pseudotyped S protein virus infection, where the S protein is expressed to produce infectious but replication-defective virions. The infectivity of these virus particles depends on the density and function of the S protein on their surface and is conveniently evaluated when the packaged viral genome carries a reporter. The results in Figure 4A are shown in relation to an intron-free construct. Improvements in expression were seen with the addition of several introns, gradually improving until one of the internal exons of S was reduced to 55 bp (15 intron construct).

[0110] Similar results were observed when the mCherry CDS was intronized with 1, 3, 4, 7, or 8 introns (SEQ ID NO:21 (P233) through SEQ ID NO:25 (P237), Figure 4B). The mCherry CDS is a relatively short sequence compared to S (711 bp vs. 3822 bp) and therefore fewer introns were required to achieve similar internal exon sizes. Nevertheless, expression gradually improved with more introns until the smallest internal exon was reduced to approximately 50 bp, shown as % cells expressing mCherry (Figure 4B).

[0111] Similar improvements in expression were seen when the human ACE2 CDS was intronized, most notably with the addition of 6 and 9 introns (SEQ ID NO:27 (P95), SEQ ID NO:28 (P223), SEQ ID NO:29 (P242)-SEQ ID NO:30 (P243), FIG. 4C). In this case, the exons only reached the optimal size range, and therefore no downward trend was observed.

[0112] In summary, transgene expression could be improved with internal exons between 501bp and 1146bp in size, but optimal expression results required internal exon sizes between 84bp and 372bp. These data are consistent with previous findings in human endogenous genes demonstrating that the optimal exon length for efficient splicing is between 50bp and 250bp (Movassat et al. 2019).

[0113] Minimal intron requirements To explore the type of intron sequence that allows the formation of the correct intron-exon landscape in multiple intron contexts, a series of constructs were generated with different introns embedded in the co S CDS (Figure 5A-C). The functional performance of S expressed from these constructs was assessed by the infectivity of the respective pseudotyped viruses (Figure 5D). First, 13 endogenous introns, each derived from a different human gene, were inserted into the S protein CDS (SEQ ID NO: 14 (P186), Figure 5A). These introns were selected based on their short length, low GC%, and the presence of canonical splice site sequences. Constructs containing these mixed introns resulted in equivalent S protein levels (Figure 5D), confirming that the introns do not need to be derived from the same gene to function as independent units.

[0114] A number of exogenous introns with similar criteria were then introduced into the sequence SEQ ID NO: 11 (P171) replacing the third intron (TTN intron 196). These included introns from unicellular yeast (S. cerevisiae, CMC2, intron 1, SEQ ID NO: 15 (P226)), nematode (C. elegans, rcor-1, intron 5, SEQ ID NO: 16 (P227)), fruit fly (D. melanogaster, elF4G, intron 5, SEQ ID NO: 17 (P228)), and mouse (M. musculus, Ttn, intron 125, SEQ ID NO: 18 (P229)) (Figure 5B). All the above constructs also resulted in similar expressed S protein levels, highlighting the fact that the origin of the intron sequence is an unimportant feature.

[0115] Considering the above, we then introduced two artificial intron sequences at the third intron position of SEQ ID NO:11 (P171) (SEQ ID NO:19 (P230), SEQ ID NO:20 (P241), FIG. 5C). In addition to the commonly known intron elements (splice sites, branch points, and pyrimidine tracts), the intron sequences were randomly created and guided only by the overall GC% of the intron according to the guidelines established above (FIG. 3). Both artificial introns performed equally well compared to the other constructs (FIG. 5D), indicating that the optimal GC% in addition to the known intron elements is sufficient for the intron to be correctly spliced, thus establishing the minimum requirements for a functional intron within the intronized transgene context.

[0116] Intronization results in improved expression levels in a variety of contexts The above developed rules for optimal transgene expression using multiple introns were tested in different transgene and cell line contexts (Figure 6). First, 13 introns (internal exons: 220bp to 306bp) were inserted into a cytoplasmic virus S protein CDS that naturally occurs only as an intron-less RNA sequence. Then, 7 introns were added (internal exons: 84bp to 127bp) into a synthetic sequence derived from the Discosoma sp red fluorescent gene using non-endogenous intron position sites (endogenous sites: http: / / corallimorpharia.reefgenomics.org). Finally, 9 introns (internal exons: 170bp to 372bp) were inserted into the endogenous human ACE2 CDS using a selection of its endogenous intron positions (Figure 6A).

[0117] The three intronized transgenes above were first tested against their intronless counterparts in human 293FT cells. Intronized S protein (stained with antibodies), mCherry (direct measurement of fluorescence), and ACE2 (stained with antibodies) showed improvements in both the percentage of cells expressing the protein, as well as the amount of expression per cell, shown as fold-change differences in the population median expression values ​​(Figure 6B). Improved expression was also observed in the mouse embryonic cell line JM8 (Figure 6C) and the mouse colon adenocarcinoma cell line MC38 (Figure 6D), where neither the intron nor the exon sequences were endogenous. [Table 1-1] [Table 1-2] [Table 2] [Table 3]

[0118] array Lowercase: intron Uppercase: exon Underlined capital letters: UTR SEQ ID NO: 1 (P91) [ka] [ka] [ka] [ka] [ka] [ka]

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

change

[0119] References Amit, M., M. Donyo, D. Hollander, A. Goren, E. Kim, S. Gelfman, G. Lev-Maor, D. Burstein, S. Schwartz, B. Postolsky, T. Pupko and G. Ast (2012). “Differential GC content between exons and introns establishes distinct strategies of splice-site recognition.” Cell Rep 1(5):543 - 556. Bourdon, V., A. Harvey and D. M. Lonsdale (2001). “Introns and their positions affect the translational activity of mRNA in plant cells.” EMBO Rep 2(5):394 - 398. Buchman, A. R. and P. Berg (1988). “Comparison of intron-dependent and intron-independent gene expression.” Mol Cell Biol 8(10):4395 - 4405. Callendret, B., V. Lorin, P. Charneau, P. Marianneau, H. Contamin, J. M. Betton, S. van der Werf and N. Escriou (2007). “Heterologous viral RNA export elements improve expression of severe acute respiratory syndrome (SARS) coronavirus spike protein and protective efficacy of DNA vaccines against SARS.” Virology 363(2):288 - 302. Chalfie,M.,Y.Tu,G.Euskirchen,W.W.Ward and D.C.Prasher(1994).“Green fluorescent protein as a marker for gene expression.” Science 263(5148):802-805. Chen Ling,Y.C.,Liu Xiaolin(2020).“Adenovirus vector vaccine for preventing SARS-CoV-2 infection.”CN110974950B Cockett,M.I.,C.R.Bebbington and G.T.Yarranton(1990).“High level expression of tissue inhibitor of metalloproteinases in Chinese hamster ovary cells using glutamine synthetase gene amplification.” Biotechnology(N Y) 18(7):662-667. Crane,M.M.,B.Sands,C.Battaglia,B.Johnson,S.Yun,M.Kaeberlein,R.Brent and A.Mendenhall(2019).“In vivo measurements reveal a single 5’-intron is sufficient to increase protein expression level in Caenorhabditis elegans.” Sci Rep 9(1):9192. Enenkel,B.(2017).“Artificial introns.”US9708636B2 Gray,S.J.,S.B.Foti,J.W.Schwartz,L.Bachaboina,B.Taylor-Blake,J.Coleman,M.D.Ehlers,M.J.Zylka,T.J.McCown and R.J.Samulski(2011).“Optimizing promoters for recombinant adeno-associated virus-mediated gene expression in the peripheral and central nervous system using selfcomplementary vectors.” Hum Gene Ther 22(9):1143-1153. Gromak,N.(2012).“Intronic microRNAs:a crossroad in gene regulation.” Biochem Soc Trans 40(41:759- 761. Grutzner,R.,P.Martin,C.Horn,S.Mortensen,E.J.Cram,C.W.T.Lee-Parsons,J.Stuttmann and S.Marillonnet(2021).“High-efficiency genome editing in plants mediated by a Cas9 gene containing multiple introns.” Plant Commun 2(2):100135. Gustafsson,C.,S.Govindarajan and J.Minshull(2004).“Codon bias and heterologous protein expression.” Trends Biotechnol 22(7):346-353. Hamer,D.H.and P.Leder(1979).“Splicing and the formation of stable RNA.” Cell 18(4):1299-1302. Haruyama,N.,A.Cho and A.B.Kulkarni(2009).“Overview:engineering transgenic constructs and mice.” Curr Protoc Cell Biol Chapter 19:Unit 19 10. Jin,Y.,M.Fei,S.Rosenquist,L.Jin,S.Gohil,C.Sandstrom,H.Olsson,C.Persson,A.S.Hoglund,G.Fransson,Y.Ruan,P.Aman,C.Jansson,C.Liu,R.Andersson and C.Sun(2017).“A Dual- Promoter Gene Orchestrates the Sucrose-Coordinated Synthesis of Starch and Fructan in Barley.” Mol Plant 10(12):1556-1570. Kozak,M.(1984).“Compilation and analysis of sequences upstream from the translational start site in eukaryotic mRNAs.” Nucleic Acids Res 12(2):857-872. Lacy-Hulbert,A.,R.Thomas,X.P.Li,C.E.Lilley,R.S.Coffin and J.Roes(2001).“Interruption of coding sequences by heterologous introns can enhance the functional expression of recombinant genes.” Gene Ther 8(8):649-653. Le Hir,H.,A.Nott and M.J.Moore(2003).“How introns influence and enhance eukaryotic gene expression.” Trends Biochem Sci 28(4):215-220. Lemaire,S.,N.Fontrodona,F.Aube,J.B.Claude,H.Polveche,L.Modolo,C.F.Bourgeois,F.Mortreux and D.Auboeuf(2019).“Characterizing the interplay between gene nucleotide composition bias and splicing.” Genome Biol 20(1):259. Marillonnet,S.,C.Engler,V.Klimyuk and Y.Gleba(2010).“Rna Virus-derived Plant Expression System.”EP2184363 Mordstein,C.,R.Savisaar,R.S.Young,J.Bazile,L.Talmane,J.Luft,M.Liss,M.S.Taylor,L.D.Hurst and G.Kudla(2020).“Codon Usage and Splicing Jointly Influence mRNA Localization.” Cell Syst 10(4):351-362 e358. Movassat,M.,E.Forouzmand,F.Reese and K.J.Hertel(2019).“Exon size and sequence conservation improves identification of splice-altering nucleotides.” RNA 25(12):1793-1805. Pettitt,S.J.,Q.Liang,X.Y.Rairdan,J.L.Moran,H.M.Prosser,D.R.Beier,K.C.Lloyd,A.Bradley and W.C.Skarnes(2009).”Agouti C57BL / 6N embryonic stem cells for mouse genetic resources.” Nat Methods 6(7): 493-495. Piovesan,A.,F.Antonaros,L.Vitale,P.Strippoli,M.C.Pelleri and M.Caracausi(2019). “Human proteincoding genes and gene feature statistics in 2019.” BMC Res Notes 12(1):315. Schmidt,A.,K.Tief,A.Foletti,A.Hunziker,D.Penna,E.Hummler and F.Beermann(1998).“lacZ transgenic mice to monitor gene expression in embryo and adult.” Brain Res Brain Res Protoc 3(1):54-60. Shaul,O.(2017).“How introns enhance gene expression.” Int J Biochem Cell Biol 91(Pt B):145-155. Sibley,C.R.,L.Blazquez and J.Ule(2016).“Lessons from non-canonical splicing.” Nat Rev Genet 17(7):407-421. Urlaub,G.and L.A.Chasin(1980).“Isolation of Chinese hamster cell mutants deficient in dihydrofolate reductase activity.” Proc Natl Acad Sci U S A 77(7):4216-4220. Virts,E.L.and W.C.Raschke(2001).“The role of intron sequences in high level expression from CD45 cDNA constructs.” J Biol Chem 276(23):19913-19920.

Claims

**Claim 1** A method of modifying a complementary DNA (cDNA) sequence for expression in a eukaryotic cell, comprising: providing a nucleic acid molecule comprising the cDNA sequence, wherein the cDNA sequence comprises two or more splicing consensus motifs that divide the cDNA sequence into exon regions of 50 to 1200 nucleotides, inserting a heterologous intron into the splicing consensus motif of the cDNA sequence, wherein each heterologous intron comprises a 3' region having a GC content equal to or lower than the GC content of the 5' region of the immediately downstream exon region, and thereby producing a nucleic acid molecule comprising a modified cDNA sequence for expression in a eukaryotic cell. **Claim 2** The method according to claim 1, wherein the modified cDNA sequence exhibits increased expression in a eukaryotic cell as compared to the unmodified cDNA sequence. **Claim 3** The method according to claim 1 or 2, wherein the cDNA sequence is 1000 nucleotides or longer. **Claim 4** The method according to claim 1, wherein the cDNA sequence lacks introns. **Claim 5** The method according to claim 1, wherein the method comprises inserting five or more heterologous introns into the cDNA sequence. **Claim 6** The method according to claim 1, wherein the 3' region of each heterologous intron has a GC content that is at least 8% lower than the GC content of the 5' region of the immediately downstream exon region. **Claim 7** The method according to claim 1, wherein the 3' region of each heterologous intron has a GC content that is 8% to 46% lower than the GC content of the 5' region of the immediately downstream exon region. **Claim 8** The method according to claim 1, wherein the 3' region of the heterologous intron comprises 30 nucleotides or more. **Claim 9** The method according to claim 8, wherein the 3' region of the heterologous intron consists of 30 nucleotides. **Claim 10** The method according to claim 1, wherein the 5' region of the immediately downstream exon region comprises 30 nucleotides or more. **Claim 11** The method according to claim 10, wherein the 5' region of the immediately downstream exon region consists of 30 nucleotides. **Claim 12** The method according to claim 1, wherein the two or more splicing consensus motifs divide the cDNA sequence into exon regions of 100 to 150 nucleotides. **Claim 13** The method according to claim 1, wherein the splicing consensus motif comprises the amino acid sequence (C / A / G)AGG(T / N)(T / N).

14. The method according to claim 13, wherein the splicing consensus motif comprises the amino acid sequence CAGGT T.

15. The method according to claim 1, wherein the eukaryotic cell is a higher eukaryotic cell.

16. The method according to claim 1, wherein the eukaryotic cell is a mammalian cell.

17. The method according to claim 1, wherein the eukaryotic cell is a CHO cell or a HEK cell.

18. The method according to claim 1, further comprising incorporating the recombinant nucleic acid comprising the modified cDNA sequence into an expression vector.

19. The method according to claim 18, further comprising introducing the expression vector into a eukaryotic cell.

20. The method according to claim 19, further comprising causing or enabling expression from the modified cDNA sequence to produce a gene product.

21. The method according to claim 20, further comprising isolating or purifying the gene product.

22. A recombinant nucleic acid comprising a cDNA sequence for expression in a eukaryotic cell, wherein the cDNA sequence comprises two or more heterologous introns and three or more exon regions of 50 to 1200 base pairs, and each heterologous intron comprises a 3' region having a GC content equal to or lower than the GC content of the 5' region of the immediately downstream exon region.

23. The recombinant nucleic acid according to claim 22, produced by the method according to claim 1.

24. An expression vector comprising the recombinant nucleic acid according to claim 22.

25. A eukaryotic cell comprising the recombinant nucleic acid according to claim 22.

26. A eukaryotic cell comprising the expression vector according to claim 24.