Methods of producing modified nucleic acid sequences for eliminating adverse splicing events

EP4681206A1Pending Publication Date: 2026-01-21THERMO FISHER SCI GENEART GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024714826
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-17
Filing Date
2024-03-15
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Conventional methods for optimizing nucleic acid sequences to prevent adverse splicing events, such as partial splicing in therapeutic applications, are inadequate, leading to potential harmful clinical effects due to residual splicing motifs in optimized genes.

Method used

A method for designing and optimizing nucleic acid sequences to eliminate or significantly reduce splice sequence motifs, using multi-parameter optimization approaches and silent substitutions to minimize predictive splice scores, ensuring full-length protein expression without adverse splicing events.

Benefits of technology

The method ensures the production of high-quality, full-length polypeptides or proteins by eliminating splice motifs, thereby enhancing the accuracy and quality of the protein product and reducing the likelihood of adverse splicing events, particularly in therapeutic applications like vaccines and gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000059_0001
    Figure IMGF000059_0001
  • Figure IMGF000040_0001
    Figure IMGF000040_0001
  • Figure IMGF000043_0001
    Figure IMGF000043_0001
Patent Text Reader

Abstract

The present invention relates generally to the design, modification, production, and optimization of nucleic acid sequences, preferably DNA sequences, and uses thereof for the production of polypeptides and / or proteins, wherein the nucleic acid sequences are introduced into an expression system, for example, into a vector and / or host organism or a host cell or other system for in vitro or in vivo expression, any of which can express the polynucleotide or protein. The nucleic acid sequences are free, completely or substantially, of putative splice sequence motifs that may be prone to an adverse splicing event that can disrupt the expression of the full-length polynucleotide or protein transcript. The present invention also relates to the use of the disclosed sequences in therapy. The disclosed methods may be performed in silico using a device and / or computer system programmed with computer software or executing software stored on a computer recordable medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS OF PRODUCING MODIFIED NUCLEIC ACID SEQUENCES FOR ELIMINATING ADVERSE SPLICING EVENTS

[0002] FIELD OF THE INVENTION

[0003] The present invention relates generally to the design, modification, production, and optimization of nucleic acid sequences and uses thereof, including therapeutic uses, wherein the nucleic acid sequences can produce polypeptides and / or proteins, and are free, completely or substantially, of putative splice sequence motifs.

[0004] BACKGROUND OF THE INVENTION

[0005] Nucleotide sequences contain information essential for the transcription and expression of proteins. For instance, genomic nucleotide sequences comprise splicing motifs, transcription factor binding sites, restriction enzyme binding sites, mRNA stability signals, and the like. The ability to identify splicing motifs existing in nucleic acid sequences that may potentially interfere with the expression of the corresponding full-length protein transcript in all organisms has been confounded by various constraints. Such constraints include the nature of the protein being encoded, codon usage variations, and the ability to identify and address a putative donor or acceptor splice site in order to ensure the full-length expression of an entire protein or polypeptide from the nucleic acid transcript upon translational processing. In order to recognize and eliminate these splice sequence motifs efficiently and successfully, these constraints must be addressed.

[0006] Different approaches exist for modifying or optimizing a nucleic acid sequence of interest for the expression of the corresponding protein or polypeptide on the basis of a particular amino acid sequence.

[0007] One technique for the preparation and synthesis of a protein or polypeptide sequence involves the cloning and expression of a gene sequence corresponding to the protein or polypeptide using a heterologous system. For a naturally occurring nucleic acid sequence encoding for a protein or polypeptide, e.g., a DNA sequence, one triplet of bases in the DNA sequence known as a "codon" expresses a single amino acid. An artificial, non-naturally occurring DNA sequence may also be synthesized to express a desired protein or polypeptide, and can be engineered based on alternative codon choices, where multiple codons may be selected to encode an amino acid residue based on the degeneracy of the genetic code. Both types of sequences can be used for the subsequent cloning and expression of the polypeptide or protein product in a variety of biological systems. One problem associated with such nonoptimized sequence generating approaches is that splicing motifs may either be present in the natural sequences or inadvertently introduced into the engineered sequences based on the underlying codon choice. These splicing motifs may undesirably reduce, modify, or entirely suppress the expression of the resultant protein or polypeptide, or even result in a product that can exert a deleterious impact on the host organism.

[0008] When a nucleic acid sequence is expressed in a specific organism, the choice of the codon sequence making up the translated protein should ideally be adapted to the codon scheme used by that organism. The frequency of a codon for expressing a particular amino acid, referred to as "codon usage," often differs among organisms. For instance, for a given organism, there is typically one preferred codon for a particular amino acid, with one or more alternative codons used with comparatively lower frequency for expressing that same amino acid. The selection of a codon to be included in an optimized nucleic acid sequence should thus be tailored to account for the codon usage favored by the subject organism.

[0009] Conventional methods for optimizing a nucleic acid sequence have been reported in P. S. Sarkar and Samir K. Brahmachari, Nucleic Acids Research 20 (1992) 5713 and D. M. Hoover and J. Lubkowski, Nucleic Acid Research 30 (2002), No. 10 e43. Known codon optimization approaches can also be performed using the Monte-Carlo method, which entails a random selection of codon positions in which a codon present in an initial nucleic acid sequence is replaced by an arbitrarily selected equivalent codon to generate the complete genetic sequence. Methods of this type commonly have the disadvantage that they heavily rely on the choice of so-called convergence criteria, which if not chosen correctly, will not successfully produce the intended optimized nucleic acid sequence.

[0010] One approach that improved conventional techniques for optimizing a nucleic acid sequence encoding a polypeptide or protein product is the GeneOptimizer™ software, which was developed and patented as US Patent No. 8,224,578-B2, herein incorporated by reference in its entirety. This gene optimization platform is based on the use of a predetermined amino acid sequence of the subject protein and can be implemented with low storage space requirements, minimal computing time, and additionally avoids known disadvantages associated with the use of randomized codon optimization methods.

[0011] De novo design of a DNA sequence using known techniques, which may or may not be based on an existing natural nucleic acid sequence, can present a risk of unintentionally incorporating one or more nucleic acid motifs that may substantially interfere with the intended biological function of the expressed gene. For example, in a given organism, certain designed nucleotide base motifs might result in a translated protein product with unforeseen structural or functional characteristics derived from the gene coding sequence if such a motif is used. One example of such a nucleotide motif is a c / s-active sequence motif that is a nucleic acid splice site.

[0012] RNA splicing is an essential step when expressing a gene into the resulting protein product. Through RNA splicing, a freshly synthesized precursor messenger RNA (pre-mRNA) transcript is processed and transformed into a mature messenger RNA (mRNA). During this process, introns (non-coding regions) are removed, and the remaining exon sequences (coding regions) are then joined together to create a continuous mature mRNA sequence. For nuclear-encoded genes, splicing takes place within the nucleus either during or immediately following transcription. Splicing is most often carried out in a series of reactions that are catalyzed by splicing machinery known as a spliceosome, which is a complex of small nuclear ribonucleoproteins (snRNPs) (proteins combined with small nuclear RNAs) found primarily within the nucleus of eukaryotic cells. The spliceosome removes the introns from the transcribed pre-mRNA primary transcript in a process referred to as splicing. There are two primary classes of spliceosomes that catalyze the pre-mRNA transcript splicing reaction. First, there is the "major" spliceosome class (which contain uridine-rich snRNAs, e.g., Ul, U2 etc.), commonly referred to as a "U2-type" intron spliceosome. Second, there is a "minor" spliceosome class (which contain uridine-rich snRNAs, e.g., Ull, U12 etc.), commonly referred to as a "U12-type" intron spliceosome. Within the intron non-coding regions, there exist dinucleotide splice sequence motifs required to facilitate a splicing event that ultimately excises the intron sequence from the pre- RNA transcript which is then absent from the subsequent mature mRNA transcript.

[0013] For instance, there is a highly conserved splice donor (SD) dinucleotide sequence motif, which is positioned at the 5' end of the intron sequence and borders the 3' end of the adjacent exon sequence, in addition to a highly conserved splice acceptor (SA) dinucleotide sequence motif, which is positioned at the 3' end of the intron sequence and borders the 5' end of the adjacent exon sequence. The canonical human SD sequence motif is a highly conserved GT dinucleotide sequence, although an alternative GC donor motif has been reported with a frequency of about 1% in the human genome (Sheth et al., 2006). The canonical human SA sequence motif is invariably a highly conserved AG-dinucleotide motif. The sequences surrounding these canonical splice sites form part of the splice site consensus sequence and are thus also conserved to varying degrees. For example, the surrounding "sequence context" of the highly conserved splice donor site (GT or GC motif) is located from 3 to 2 nucleotides upstream in the exon region to about 3 to 6 nucleotides downstream into the adjacent intron region (the broadest range can be denoted as "-3 to +6," where position "0" corresponds to the exonintron boundary). The noncanonical sequences around the highly conserved AG splice acceptor site are located approximately from 14 to 3 nucleotides upstream in the intron region to about 2 or 3 nucleotides downstream into the adjacent exon region (the broadest range can be denoted as "-14 to +3," where position "0" corresponds to the intron-exon boundary) (Riepe, 2021). Upstream of the AG SA sequence motif, there exists a region high in pyrimidine content e.g., C and T nucleotides), which is also referred to as a "poly-pyrimidine tract," which extends another 15-30 nucleotides into the intron.

[0014] The aforementioned patented GeneOptimizer™ software optimization algorithm can effectively process many of these dinucleotide splice sequence motifs or other sequence motifs, including restriction sites, poly(A) motifs, or cryptic splice sites. However, this and other known gene optimizing software programs have not been entirely successful in completely identifying all splice sites that may ultimately disrupt the full-length expression of a protein transcript. Thus, residual splicing of the transcript, characterized as an adverse splicing event, remains an unresolved and critical issue in gene engineering. While this problem might be tolerable in most RUO (research use only) applications, severe consequences might occur when such "optimized" genes are adopted for use in in vivo applications, such as pharmaceutical, clinical, or other therapeutic approaches where a complete, full-length expression of the protein or polypeptide product from the nucleic acid is crucial for achieving the contemplated therapeutic outcome. One exemplary in vivo application where the resulting full-length transcript of a therapeutic protein ought to be completely free from any adverse splicing events is a nucleic acid-based vaccine.

[0015] Despite attempts to strenuously avoid undesirable splicing motifs in an optimized gene product using conventional software sequence optimization programs, including avoidance of the highly conserved GT and AG dinucleotide splice motifs, the present inventors have discovered that even synthetically engineered genes optimized for mammalian expression can exhibit sporadic and residual partial splicing activity that may unfavorably impact the expression of the full-length protein transcript. For example, when the expression of the transcript results in a truncated or degraded protein product that can lead to unintended and harmful clinical effects, or even no physiological effect at all.

[0016] For instance, an optimized SARS-CoV-2 spike open reading frame in the AstraZeneca (AZ) adenovirus-based SARS-CoV-2 vaccine was reported to be susceptible to partial splicing events that appeared to result in the expression of C-terminal truncated soluble spike variants (Kowarz et al., Research Square 2021; DOI: https: / / doi.Org / 10.21203 / rs.3.rs-558954 / vl).

[0017] While partial splicing events may not be problematic in most transient experiments performed at the laboratory bench, these splicing events must be vigorously avoided in therapeutic applications, in particular gene therapy applications such as nucleic acid- or vector-based vaccines or DNA-based cancer immunization approaches, especially where thorough clinical investigations into the expression or splicing behavior of the nucleic acid sequence-based therapeutics may not be feasible.

[0018] The object of the present invention is therefore to provide improved methods for the design and optimization of the quality of polypeptides or proteins transcribed from a gene transcript, preferably engineered gene transcripts, in a manner that avoids adverse splicing events that may impede the expression of the resulting full-length protein or polypeptide transcript. The disclosed methods may be performed in silica using a computing device or computer system suitably programmed with computer software based on one or more optimization algorithms or by executing such software that is stored on a computer recordable medium, which may further render the resulting optimized nucleic acid product.

[0019] SUMMARY OF THE INVENTION

[0020] The present invention relates generally to designing, producing, modifying and optimizing nucleic acid sequences. The nucleic acid sequences can be preferably RNA sequences or DNA sequences, and even more preferably DNA sequences comprising an open reading frame. The modified nucleic acid sequences are free, completely or substantially, of active splice sequence motifs in the nucleic acid sequence that may be prone to an adverse splicing event that can disrupt the expression of the corresponding full-length protein or polypeptide transcript.

[0021] In one embodiment, the provided nucleic acid sequence is a naturally occurring sequence that can be optionally subject to a codon optimization process using, for instance, a conventional optimizing program such as the GeneOptimizer™ program, the JCat JAVA Codon Adaptation Tool (Grothe et al. JCat: a novel tool to adapt codon usage of a target gene to its potential expression host; Nucleic Acid Res., Vol. 33, Issue Suppl_2, 2005, p. W526-W531), the Gene Designer program (Villalobos et al.: Gene Designer: a synthetic biology tool for constructing artificial DNA segments, BMC Bioinformatics 7, 285 (2006)), or the ExpOptimizer program (https: / / www. novoprolabs.com / tools / codon-optimization), or the Codon Optimizer program (Condon, A. Thachuk, C: Efficient codon optimization with motif engineering, Journal of Discrete Algorithms, 2012). In another embodiment, the provided nucleic acid is an engineered, optionally non-naturally occurring artificial sequence that may or may not have been previously optimized using conventional techniques. The provided nucleic acid sequence can then be engineered according to the present invention to remove all, or substantially all, splicing motifs to thereby maximally reduce the occurrence of a splicing event when the modified nucleic acid sequence is expressed and translated in an expression system into the resultant polypeptide or protein product. In a preferred embodiment, the provided nucleic acid sequence may be engineered to remove or substantially decrease the splicing motifs using a multi-parameter optimization approach that processes all required optimization parameters simultaneously and / or on an iterative basis. The nucleic acid sequence to be optimized or modified according to the invention, or the corresponding encoded amino acid sequence, can be derived from any organism, including an animal, bacteria, a plant, a protozoan, yeast, a fungus, a pathogen including a virus, preferably a virus that is capable of gaining access to and / or reproducing in a mammalian host, even more preferably a human host. In other preferred embodiments, the nucleic acid sequence is derived from a mammal, even more preferably, from a human.

[0022] The present methods are also useful for producing the modified nucleic acid sequences disclosed herein in addition to the corresponding polypeptides and proteins, which can be accomplished by introducing the modified nucleic acid sequences into an expression system, such as a vector and / or a host organism or a host cell, or alternatively, by using other suitable approaches for the in vitro, ex vivo, and / or in vivo expression of the corresponding polypeptides and proteins. During the process of protein production based on the modified nucleic acid sequences disclosed herein, it is preferred that substantially no splicing events will occur as the pre-mRNA transcript matures into the mRNA transcript, which can cause the polypeptide sequence or protein to become truncated, degraded, or otherwise unable to become fully expressed. Especially preferred embodiments relate to mature mRNA transcripts where no such adverse splicing events occur whatsoever, since splice sequences have been completely eliminated from the modified nucleic acid sequence. Accuracy and therefore the overall quality of the produced polypeptide or protein product is thereby significantly increased and thus improved over conventional approaches.

[0023] In another aspect, the invention relates to a method of producing an optimized or modified nucleic acid sequence comprising a reduced number of active dinucleotide splice sequence motifs compared to a corresponding non-optimized nucleic acid sequence, comprising: (i) designing a modified nucleic acid sequence according to the methods described herein, and (ii) producing the modified nucleic acid sequence of step (i), wherein the producing comprises de novo synthesis, mutagenesis or a combination thereof.

[0024] As described herein, a decreased likelihood of a splicing event is based on a reduced "predictive splice score," which takes into account the appearance of consecutive nucleotide bases and their counterpart codon sequences in the modified sequence when calculating the probability of a splicing event in view of any "similarity" (homology) to the corresponding splice motif known for a particular organism. One aim of the present methods is to "maximally reduce" the predictive splice score to as close to zero as possible by eliminating most or all of the dinucleotide splice motifs in the resulting modified nucleic acid sequence. In some embodiments, the modified nucleic acid sequence is entirely devoid of residual splice sites and thus presents a "maximally reduced" or zero likelihood of splicing events, thereby rendering the modified sequence as "unbreakable."

[0025] In some embodiments, the active putative splice sequence in the provided nucleic acid sequence (also referred to as a "target nucleic acid sequence" to be modified or "Seqtarget"), is a dinucleotide splice sequence motif. In other embodiments, the active putative splice sequence comprises a "splice donor dinucleotide sequence motif" used interchangeably herein with the terms, e.g., "splice donor site," "donor splice sequence" or "(SD)," or a "splice acceptor dinucleotide splice sequence motif used interchangeably herein with the terms, e.g., "acceptor splice site," "acceptor splice sequence" or "(SA)." In preferred embodiments, the SD dinucleotide splice sequence motif is a highly conserved GT motif. In other preferred embodiments, the SA dinucleotide sequence motif is a highly conserved AG motif.

[0026] In other embodiments, the highly conserved SD GT motif is referred to herein as an "unavoidable GT motif" or, equivalently, as an "unavoidable GT dinucleotide sequence" in cases where the GT dinucleotide sequence motif is not amenable to a straightforward silent substitution using an alternative codon sequence. The present invention thus advantageously provides methods for identifying and replacing such unavoidable GT dinucleotide sequences with an alternative nucleic acid sequence that presents a lower likelihood, or no likelihood, of undergoing an adverse splicing event.

[0027] Exemplary unavoidable GT motifs are defined herein according to the following codon sequences: valine (Vai) (GTN), wherein N is any of A, T, C, or G, or in the context of a methionine (Met) (ATG) or tryptophan (Trp) (TGG) followed by a phenylalanine (Phe) (TTY), tyrosine (Tyr) (TAY), cysteine (Cys) (TGY), or tryptophan (Trp) (corresponding to ATG or TGG) followed by TTY, TAY, TGY, or TGG, wherein Y is T or C.

[0028] In another embodiment, the SD dinucleotide sequence motif is a GC motif. Exemplary GC motifs are defined herein according to the following codon sequences: alanine (GCT, GCC, GCA, GCG) or any of the following amino acid combinations, MH, MP, MQ, WH, WP or WQ.

[0029] In some embodiments, the SA dinucleotide sequence motif is a highly conserved AG motif. In other embodiments, the highly conserved splice acceptor AG motif is referred to herein as an "unavoidable AG motif" or, equivalently, as an "unavoidable AG dinucleotide sequence" in cases where the AG dinucleotide sequence motif is not amenable to a straightforward silent substitution using an alternative codon sequence. In this case, the present invention advantageously provides methods for identifying and replacing such unavoidable AG dinucleotide sequences with an alternative nucleic acid sequence presenting a reduced likelihood of undergoing a splicing event.

[0030] Exemplary unavoidable AG motifs are defined herein according to the following codon sequences: lysine (K) encoded by AAG or AAA followed by glutamic acid (E) encoded by GAG or GAA. Other exemplary unavoidable GT motifs include the amino acid combinations of EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, or #V, with the symbol "#" indicating a stop codon.

[0031] In one embodiment, the present invention relates to an in silica method for optimizing or modifying a nucleic acid sequence, including the design and production thereof, comprising a reduced number of active splice sequences compared to the original nucleic acid sequence, the method comprising:

[0032] (a) providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame, wherein Seqtarget has a length (Ntarget),

[0033] (b) defining an organism-dependent splicing model comprising:

[0034] (i) compiling a listing of sequence contexts (SeqCOn) comprising one or more intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg), each sequence in SeqCOn having a length (Neon), wherein SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (Seq con SA),

[0035] (ii) obtaining patterns (P) of length (NCOn) from SeqCOn,

[0036] (iii) providing a first quality function capable of calculating a predictive splice score for each pattern P for determining whether the pattern P indicates the presence of a putative SD sequence motif or putative SA sequence motif of Genorg based on:

[0037] (1) the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn, and / or

[0038] (2) the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn,

[0039] (iv) applying the first quality function to the pattern P obtained in (ii) to calculate their predictive splice scores;

[0040] (c) providing a splice site identifying algorithm (At) comprising:

[0041] (i) defining a maximum threshold splice score applicable for a respective Seqtarget of length N target,

[0042] (ii) providing a second quality function capable of using as a combination input: a. a pattern P and its predictive splice score calculated in (b)(iv), b. the maximum threshold splice score defined in (c)(i), and c. an input nucleic acid sequence of length NCOn, and

[0043] (iii) applying the splice site identifying algorithm Ai of (c)(ii) to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif;

[0044] (d) applying a splice site modifying algorithm (AM) comprising:

[0045] (i) selecting a sequence window (D) which contains at least one position of the open reading frame of Seqtarget, (ii) applying the third quality function defined in (c)(iii) to sequence window D to calculate a predictive splice score,

[0046] (iii) determining based on the predictive splice score whether window D comprises a putative splice sequence motif,

[0047] (iv) optionally specifying one or more silent substitutions for sequence window D of Seqtarget capable of reducing the predictive splice score;

[0048] (v) applying the one or more silent substitutions specified in (d)(iv) to window D to modify a putative splice sequence motif,

[0049] (vi) iteratively performing steps (d)(i) to (d)(v) for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

[0050] The target sequence Seqtarget may be a naturally occurring or an engineered non-naturally occurring nucleic acid sequence. Further, the Seqtarget may be further optimized for the expression of a polypeptide or protein in a target organism using an optimization program described elsewhere herein. The target sequence may be derived from an organism. In some instances, the target sequence may be derived from an animal such as a mammal. In other instances, the target sequence may be derived from a bacterium, a pathogen, a virus, a protozoan, a plant, or a fungus such as yeast. In some examples the target sequence may be derived from a human.

[0051] In one embodiment according to the methods herein, the splice donor sequence motif is a highly conserved GT motif and the splice donor sequence context having a length N comprises nucleotides positioned 5' and / or 3' of the GT dinucleotide motif. In some aspects, the splice donor sequence context comprises 5'-NNNNGTNNNN-3, 5'-NNNGTNNN-3, 5'- NNNGTNNNNN-3, 5'-NNNGTNNNNNN-3, 5'-NNNGTNNNNNNN-3, 5'-NNNGTNNNNNNNN-3, 5'-NNNGTNNNNNNNNN-3, 5'-NNNGTNNNNNNNNNN-3, 5'-NNNGTNNNNNNNNNNN-3, 5'- NNNGTNNNNNNNNNNNN-3, 5'-NNNGTNNNNNNNNNNNNN-3, 5'-NNNGTNNNNNN NNNNNNNN-3, or 5'-NNNGTNNNNN NNNNNNNNNN-3', wherein N is A, T, C, or G. In a preferred embodiment, the splice donor sequence context is 5'-NNNGTNNNN-3'. In one embodiment according to the methods herein, the splice donor sequence motif is a GC motif and the splice donor sequence context having a length N comprises nucleotides positioned 5' and / or 3' of the GC- -dinucleotide motif. In some aspects, the splice donor sequence context comprises 5'-NNNNGCNNNN-3, 5'-NNNGCNNN-3, 5'-NNNGCNNNN-3, 5'- NNNGCNNNNN-3, 5'-NNNGCNNNNNN-3, 5'-NNNGCNNNNNNN-3, 5'-NNNGCNNNNNNNN-3, 5'-NNNGCNNNNNNNNN-3, 5'-NNNGCNNNNNNNNNN-3, 5'-NNNGCNNNNNNNNNNN-3, 5'- NNNGCNNNNNNNNNNNN-3, 5'-NNNGCNNNNNNNNNNNNN-3, 5'-NNNGCNNNNNNNNN NNNNN-3, or 5'-NNNGCNNNNNN NNNNNNNNN-3', preferably 5'-NNNGCNNNNNNNN-3, wherein N is A, T, C, or G.

[0052] In one embodiment, the highly conserved splice acceptor sequence motif is an AG motif and the splice acceptor sequence context having a length N comprises nucleotides positioned 5' and / or 3' of the AG dinucleotide motif. In some aspects, the splice acceptor sequence context comprises 5'-NNNNAGNNNN-3, 5'-NNNAGNNN-3, 5'-NNNAGNNNNN-3, 5'-NNNAGNNNNNN- 3, 5'-NNNAGNNNNNNN-3, 5'-NNNAGNNNNNNNN-3, 5'-NNNAGNNNNNNNNN-3, 5'- NNNAGNNNNNNNNNN-3, 5'-NNNAGNNNNNNNNNNN-3, 5'-NNNAGNNNNNNNNNNNN-3, 5'-NNNAGNNNNNNNNNNNNN-3, 5'-NNNAGNNNNNNNNNNNNNN-3, or 5'-NNNAGNNNNN NNNNNNNNNN-3', wherein N is A, T, C, or G. In a preferred embodiment, the splice acceptor sequence context is 5'-NNNAGNNNN-3'.

[0053] In another embodiment, a set of splice donor (SD) sequence contexts (SeqCOnSD) comprises putative mammalian splice donor sequences. In preferred embodiments, the putative mammalian splice donor sequences are defined by patterns (P) such as 5'-MAGGTRAGT-3', wherein "M" corresponds to A or C, and wherein "R" corresponds to A or G.

[0054] In another embodiment, replacing one or more nucleotides 5' and / or 3' to the GT dinucleotide sequences in the SD sequence context with a silent substitution result in a substituted sequence context comprising 5'-KBHGTYBHV-3', wherein K is G orT, B is G, T or C, H is A, C or T, Y is C or T, and V is G, C or A.

[0055] In one embodiment, one or more of the method steps are performed using a suitably programmed computer. Preferably, all of the method steps herein are performed using a suitably programmed computer. In other embodiments, the aforementioned method steps can be performed without the use of a suitably programmed computer, for instance, using non-automated calculation techniques.

[0056] In one embodiment, the method further comprises optimizing, using the suitably programmed computer, the nucleic acid sequence for expression in a host cell by replacing one or more additional nucleotides with silent substitutions without modifying or otherwise altering the encoded amino acid sequence for the ultimate polypeptide or protein product. In addition to assessing the splice donor and / or acceptor dinucleotide sequences in the provided nucleic acid sequence, further optimization may comprise an adaptation of the GC content of the nucleic acid sequence, in addition to avoiding undesirable DNA sequence motifs such as splicing sites that may interfere with subsequent polypeptide or protein expression, for instance, motifs resulting in secondary structures, and / or direct or indirect DNA sequence repeats. These considerations for DNA optimization are described in US Patent No. 8,224,578-B2, and can be evaluated, for example, on the basis of a "quality function."

[0057] Adaptation of the codon usage of the gene to the codon usage of the host represents an important criterion in the present optimization methods given the different degeneracy patterns of the various codon sequences across different organisms. There are two principal methods that can be employed to score codon usage for a particular organism. The first approach involves a relative frequency or "relative adaptiveness" of codon usage which examines the frequency of a codon placed at a particular position in a sequence for expressing a specific amino acid in proportion to the frequency of the codon that most frequently expresses the same amino acid at that same site. A second approach, which is particularly useful for the methods described herein, involves a relative synonymous codon usage (RSCU) evaluation. Both approaches are standardized to the frequency of the codon(s) most used by the subject organism (cf. P. M. Sharp, W. H. Li, Nucleic Acid Research 15 (1987), 1281 to 1295). For instance, the RSCU for a codon c which expresses the amino acid A is defined by the following formula:

[0058] RSCUc = / c Ac / ( c' / c') where the sum in the denominator runs over all of the codons c' that express the amino acid Ac, and where Ac- indicates the amino acid that c encodes; and dA is the number of codons that express the amino acid A, and fcis the frequency. In order to define a criterion weight on the basis of the RSCU, the RSCU can be summed for the respective test sequence over all the codons of the test sequence or a part thereof, in particular over the number of codons of optimization positions. The difference from the criterion weight derived from the relative adaptiveness is that with this weighting each codon position is weighted with the degree of degeneracy, dAi for the amino acid at position i, so that positions at which more codons are available for selection can participate more in the criterion weight calculation than positions where only a few codons or even only a single codon are available for selection.

[0059] In a preferred embodiment, synonymous codons having a high usage frequency in a specific host can be used for nucleic acid sequence optimization based on a high codon usage adaptation index. Exemplary codon usage frequency tables for different organisms are viewable e.g., at https: / / hive.biochemistry.gwu.edu / cuts / about https: / / www.kazusa.or.jp / codon / , which was publicly available information at the filing date of this application.

[0060] According to various embodiments of the invention, one or more "quality functions" can be used to assess a sequence context, during one or more steps or iterative process runs according to the disclosed methods, and similarly applied to all or at least the majority of each iteration of the method. A quality function may be capable of computing or calculating a certain score e.g., a predictive splice score) wherein a score can be defined in such a way that either a larger value means that a sequence or sequence context is closer to or at an optimum or predefined value or more likely to match certain predefined criteria, or conversely, that a smaller value means that the sequence or sequence context is nearer to the optimum or predefined value. In some embodiments, a score can be calculated by a quality function iteratively for determining, in at least one iteration, an upper or lower limit for the value, and / or the iteration can be terminated when this value is below or above such value, as the case may be.

[0061] Where a quality function used for generating an expression-optimized nucleic acid sequence addresses a plurality of test criteria, especially when the quality function consists of a linear combination of criterion weights, a sequence context need not necessarily be assessed according to all criteria in a single iteration step. On the contrary, the assessment can be terminated as soon as it is evident that the value of the quality function is less or, i.e., less optimal than the value of the quality function of a sequence context that has already been assessed.

[0062] In one embodiment, an algorithm used for generating an expression-optimized nucleic acid sequence can comprise the following steps:

[0063] (i) assessing each nucleic acid sequence with a quality function;

[0064] (ii) ascertaining a threshold value within the values obtained by the quality function for all nucleic acid sequences generated according to the methods disclosed herein that are to be optimized.

[0065] In one embodiment, first, second and third quality functions are applied in methods described herein to identify and modify splice sequence motifs in a target sequence Seqtarget of length Ntarget. The quality functions are used to determine how similar a sequence stretch in the provided target sequence is to a highly conserved splice sequence motif used by a particular organism, for example, on the basis of a "predictive splice score." A "predictive splice score" as used herein refers to a value computed for a certain sequence that indicates the similarity with a consensus splice sequence motif as described in more detail below. For example, mapping two sequences to a value of "1" if the sequences are identical and mapping the sequences to a value of "0" if they are considered to be completely non-similar. For purpose of illustration only, for two sequences of identical length (i.e., an identical number of nucleotide bases) the numbers of identical bases are counted (this is also referred to as "matching" bases) and then divided by the length of the sequences. For IUPAC notation, each base A, C, G, or T is considered to match any IUPAC base represented by that particular IUPAC symbol.

[0066] Similarity between nucleic acid sequence regions can also be expressed in terms of homology, for instance, more than 90% similarity and / or 99% identity to a particular nucleic acid sequence, for example, an organism-specific highly conserved dinucleotide splice sequence motif, which may be present in the modified sequence. In some embodiments, the nucleic acid sequence being optimized or modified comprises no homology regions showing more than about 5%, 10%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80% similarity to an organismspecific highly conserved dinucleotide splice sequence motif. In preferred embodiments, the nucleic acid sequence being modified comprises no homology regions (0%) showing similarity to an organism-specific highly conserved dinucleotide splice sequence motif. In preferred embodiments, when the predictive splice score is "maximally reduced," the modified nucleic acid has a minimum number of homology regions in common with the highly conserved organism-specific dinucleotide splice sequence motif.

[0067] The present invention is not intended to be restricted for use in only one specific organism. A nucleic acid sequence can be optimized or modified for any organism model depending on the frequency with which that organism uses particular codon sequences for expressing an amino acid according to the codon usage scheme in that organism. In any given organism, there is usually one codon that predominates when expressing a corresponding amino acid, with one or more less preferred codons used with comparatively lower frequency. Since the optimized nucleotide sequence is envisioned for use in a particular predefined organism, the choice of the codon should be adapted to the codon usage of that organism in terms of an "organism-dependent model."

[0068] In another aspect, the present invention relates to a method for modifying a target nucleic acid sequence in a manner that reduces or even completely eliminates adverse splicing events in an organism-dependent manner, including in the design and production of these sequences. The resulting modified nucleic acid sequence product comprises a reduced overall number of dinucleotide splice sequence motifs, preferably maximally reduced to zero, compared to a corresponding non-optimized / modified nucleic acid sequence, the method comprising:

[0069] (a) providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame, wherein Seqtarget has a length (Ntarget),

[0070] (b) defining an organism-dependent splicing model comprising:

[0071] (i) compiling a listing of sequence contexts (SeqCOn) comprising one or more intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg), each sequence in SeqCOn having a length (Neon), wherein SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (Seq con SA),

[0072] (ii) obtaining patterns (P) of length (NCOn) from SeqCOn,

[0073] (iii) providing a first quality function capable of calculating a predictive splice score for each pattern P for determining whether the pattern P indicates the presence of a putative SD sequence motif or putative SA sequence motif of Genorg based on:

[0074] (1) the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn, and / or

[0075] (2) the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn,

[0076] (iv) applying the first quality function to the pattern P obtained in (ii) to calculate their predictive splice scores;

[0077] (c) providing a splice site identifying algorithm (At) comprising:

[0078] (i) defining a maximum threshold splice score applicable for a respective Seqtarget of length N target,

[0079] (ii) providing a second quality function capable of using as a combination input: d. a pattern P and its predictive splice score calculated in (b)(iv), e. the maximum threshold splice score defined in (c)(i), and f. an input nucleic acid sequence of length NCOn, and

[0080] (iii) applying the splice site identifying algorithm Ai of (c)(ii) to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif;

[0081] (d) applying a splice site modifying algorithm (AM) comprising:

[0082] (i) selecting a sequence window (D) which contains at least one position of the open reading frame of Seqtarget, (ii) applying the third quality function defined in (c)(iii) to sequence window D to calculate a predictive splice score,

[0083] (iii) determining based on the predictive splice score whether window D comprises a putative splice sequence motif,

[0084] (iv) optionally specifying one or more silent substitutions for sequence window D of Seqtarget capable of reducing the predictive splice score;

[0085] (v) applying the one or more silent substitutions specified in (d)(iv) to window D to modify a putative splice sequence motif,

[0086] (vi) iteratively performing steps (d)(i) to (d)(v) for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

[0087] As disclosed herein, the organism-dependent optimization model relies on a compiled library comprising one or more intron / exon boundary sequences for one or more genes derived from the predefined organism (Genorg), the sequence contexts collectively referred to herein as SeqCOn Each such sequence context indicated in SeqCOn, has a length NCOn, wherein NCOn is a natural number and is less than or equal to the total number of nucleotides in the entire organism genome sequence (Norg), wherein the sequence context listing comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (SeqCOnSA).

[0088] In some embodiments, the Seqcon of the putative dinucleotide splice sequence motif comprises from 2 to 100 nucleotides, including each integer value therebetween, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100. In preferred embodiments, the SeqCOn of the putative dinucleotide splice sequence motif comprises from 2 to 50 nucleotides. In other preferred embodiments, the SeqCOn comprises from 2 to 20 nucleotides. In more preferred embodiments, the SeqCOn comprises from 2 to 15 nucleotides. In most preferred embodiments, the SeqCOn comprises from 3 to 9 nucleotides.

[0089] The Seqtarget according to the disclosed methods can be derived from an organism, wherein the organism is an animal, a mammal, a bacterium, a pathogen, a virus, a protozoan, a plant, yeast, or a fungus, wherein the Seqtarget is preferably derived from a mammal, more preferably from a human. The Seqtarget comprises at least an open reading frame encoding an amino acid sequence. In some embodiments the Seqtarget may further comprise regulatory sequences at the 5' and / or 3' ends, such as e.g., a 5' UTR sequence and / or a 3' UTR sequence.

[0090] According to a preferred embodiment of the invention, step (d) is iterated on a finite series of sequence windows D in Seqtarget until all the nucleotides of the modified Seqtarget have been specified to completely eliminate all dinucleotide splice sequence motifs, some or all of which become occupied by substituted nucleotides either according to a silent substitution or, in the case of unavoidable GT or unavoidable AG motifs, based on a maximally reduced predictive splice score based on the organism-dependent codon usage scheme. For instance, in some embodiments, the highly conserved GT motif is an avoidable GT motif and preferably in step (d)(v) the avoidable GT motif is removed from the Seqtarget. In other embodiments, the splice donor dinucleotide sequence motif is a GC motif, and preferably in step (d)(v) the GC motif is removed from the Seqtarget. In yet other embodiments, in step (d)(v), the highly conserved AG motif is an avoidable AG motif and preferably in step (d)(v) the AG motif is removed from the Seqtarget.

[0091] The modification of the nucleic acid sequence according to the present invention is not generally performed on the entire provided nucleic acid sequence all at once. As described herein, in one embodiment, the modification of the nucleic acid is preferably performed on successive regions of the nucleic acid sequence by moving a respective sequence "window" (D) which contains at least one position of the open reading frame of Seqtarget in a 5' to 3' direction along the Seqtarget, moving progressively along the entire length Ntarget of the target nucleic sequence beginning from, for instance, a first position in the Seqtarget until a final position in the Seqtarget has been reached, wherein the entire sequence context has been evaluated for the presence of putative dinucleotide splice sequences. In some embodiments, this evaluation protocol can be extended to assess the entire length of the open reading frame of the provided nucleic acid sequence in step (a). In other embodiments sequence window D may be moved randomly along the sequence (i.e., not necessarily in a strict 5' to 3' direction) or two or more windows D may be evaluated or modified simultaneously.

[0092] In some cases, the silent substitutions specified in one iteration step may not be changed again in successive, subsequent iteration steps according to (d)(vi), on the contrary, the nucleotides corresponding to silent substitutions are assumed to be already optimized in the respective iteration steps. This makes it possible to take account not only of local effects on the varied positions e.g., within a given window D), but also of wider ranging correlations, e.g., in connection with the development of RNA secondary structures (e.g., assessing multiple windows D simultaneously).

[0093] In some embodiments, nucleotides replaced with silent substitutions in a given window D may be in the range from 1 to 100, from 3 to 20, or in the range from 5 to 10 nucleotides. Accordingly, it is possible to vary the positions with an acceptable usage of storage and computing time and, at the same time, achieve strong optimization of the provided target nucleic acid sequence for an intended use.

[0094] In other aspects of the invention, one or more of the positions on which the nucleotides are varied can be identical in two or more consecutive iteration steps. If the positions are connected, this means that the window D in one iteration step overlaps with the window D' of a preceding iteration step.

[0095] In some embodiments according to the present invention, for the region of each sequence context within the target nucleic acid sequence Seqtarget of length Ntarget, one iteration step may or may not include the region of a sequence context from a previous iteration step. In some instances, a window of fixed length D may be shifted along the complete sequence Ntarget during the various iterations.

[0096] If nucleotides are to be replaced in certain positions of Seqtarget they are replaced by silent substitutions that encode the same amino acids in the provided open reading frame.

[0097] In some embodiments, during the iteration step multiple criteria may be taken into account to generate an expression-optimized Seqtarget. Criteria to be taken into account for sequence optimization may include one or more of the following: codon usage in a predefined organism or target host, sequence producibility, undesired sequence elements such as repetitive sequences, splice sequences, secondary structures, or inverse complementary repeats or functional sequence motifs such as TATA boxes, termination signals, artificial recombination sites, RNA instability motifs, ribosomal entry sites, repetitive sequences, premature poly(A) sites, and the like.

[0098] If taken into account simultaneously during sequence optimization, such criteria may need to be weighted using a weighted combination.

[0099] The term "weighted combination" as used herein reflects the relative importance of one or more selection criteria, e.g., one or more optimization criteria, compared to other selection criteria and is based on the notion of "criterion weights" as described e.g., in US 8,224,578- B2. Each selection criterion has a positive or negative weight depending on the aim to maximize or minimize the score associated with the particular criterion. An overall score is obtained by adding the weighted scores of the criteria.

[0100] For example, any multi-criteria optimization problem which aims at simultaneously optimizing a number of criteria (like maximizing codon usage at certain optimization positions in a target nucleic acid sequence and minimizing the likelihood of splicing as set out elsewhere herein) can be turned into a single-criterion optimization problem by combining the criteria values with weights. For instance, if codon usage is denoted by "Pl" and the likelihood of splicing by "P2," then one can try to maximize Pl - P2 in order to maximize codon usage and minimize splicing likelihood simultaneously. (In this example, the weights according to the formula P1-P2 are 1 and -1, respectively). Similarly, one could try to maximize 3*P1 - 2*P2 (with the weights chosen as 3 and -2). In principle, the weights define by which amount one criterion is more important than the other in evaluating the resulting values. For scores that have values from the range [0,1], it is quite common to use a "convex" combination of the scores, i.e., by using weights which are also from the range [0,1] and which sum up to 1.

[0101] Splice site prediction modeling can be grouped into three main categories: (i) probability methods; (ii) alignment-based methods; and (iii) machine learning methods. Early approaches were based on probability models for splice site predictions, often relying on Markov models (Pertea et al., 2001). Alignment-based approaches use reads from RNAseq for assessing gene expression and splice site prediction (Trapnell et al., 2009; Wang et al., 2010). More recently, machine learning has become increasingly popular, given the pervasive and efficient availability of computational power. The "SpliceRover" algorithm, for instance represents one type of deep learning approach using convolutional neural networks (Zuallaert et al., 2018; Riepe et al., 2020). Moreover, because modern sequencing techniques can process high volumes of sequencing data taking into account annotation data, complex modeling can train on these data inputs.

[0102] In one embodiment, the predictive splice score of the respective putative splice sequence according to the selected prediction modeling is maximally reduced such that only a minimum of splice sequence motifs, if any, are present in the modified nucleic acid.

[0103] In another embodiment, the general method for modifying a nucleic acid sequence set forth above further comprises an additional step, wherein for each sequence context in the provided target sequence, a set of sequences corresponding to either a splice donor sequence motif or a splice acceptor sequence motif is evaluated to determine whether the sequence context comprises a putative splice sequence motif amenable to a silent substitution mutation, wherein the silent substitution is applied in the following order: (i) selecting a silent mutation that does not contain a GT or AG dinucleotide splice sequence motif, or (ii) if the dinucleotide splice sequence motif is an unavoidable GT or unavoidable AG motif, then selecting a nucleotide sequence according to a predefined organism model that contains a minimum number of GT and AG dinucleotide sequences and further, which has the lowest predictive splice score according to the method set forth above.

[0104] According to preferred embodiments of the invention, the substitution of nucleotides according to the methods disclosed herein does not result in the generation of nucleic acid sequence motifs comprising repetitive sequences, splice sequences, secondary structures, or inverse complementary repeats or other undesired functional sequence motifs such as termination signals, RNA instability motifs, premature poly(A) sites etc. in the base nucleic acid sequence that may influence the result of full-length protein and / or polypeptide expression. In another embodiment, the method further comprises optimizing the nucleic acid sequence for expression in a host by replacing one or more additional nucleotides of the Seqtarget with a silent substitution sequence in a manner to maximally reduce the predictive splice score of the Seqtarget at the same time optimizing the codon usage for the predefined organism without modifying the resulting encoded amino acid sequence. In preferred embodiments, the optimized nucleic acid sequence designed and produced according to methods disclosed herein is completely devoid of any splice donor dinucleotide sequence motifs and / or splice acceptor dinucleotide sequence motifs, or at least the presence of such splice sequence motifs is substantially reduced compared to the corresponding non-optimized nucleic acid sequence such that the likelihood of an adverse splicing event has been significantly decreased.

[0105] In another aspect, the invention relates to a method of expressing a protein or polypeptide, the method comprising:

[0106] (a) providing a nucleic acid sequence encoding the protein or polypeptide,

[0107] (b) designing an optimized or modified nucleic acid sequence according to any of the methods disclosed herein,

[0108] (c) synthesizing the modified nucleic acid sequence,

[0109] (d) producing a vector comprising the modified nucleic acid sequence,

[0110] (e) transfecting a host cell with the vector of step (d), wherein the host cell, subsequent to transfection, expresses the modified nucleic acid and produces the encoded protein or polypeptide product, and

[0111] (f) optionally isolating and purifying the protein or polypeptide produced by the transfected host cell of step (e).

[0112] In one embodiment, the optimized or modified nucleic acid sequence herein may encode a protein from an RNA virus. In another embodiment, the RNA virus can be derived from the family of Coronaviridae, which comprises the genera of coronavirus, SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2), SARS-CoV (severe acute respiratory syndrome coronavirus), MERS-CoV (Middle East respiratory syndrome (MERS) coronavirus), HIV (Human Immunodeficiency Virus), Influenza virus, Ebola virus, Flaviviridae comprising Hepatitis C virus, West Nile virus, Dengue virus, Yellow fever virus, Zika virus, or Rhabdoviruses comprising rabies virus, Paramyxoviridae comprising parainfluenza virus, Venezuelan equine encephalitis virus, Equine arteritis virus, Rotaviruses, and Enterovirus comprising Foot-and- mouth disease virus (FMDV). Coronaviruses can further include the genera of alphacoronaviruses, betacoronaviruses, gammacoronaviruses, and deltacoronaviruses. In a preferred embodiment, the nucleic acid sequence encodes a protein derived from the family Coronaviridae. In a more preferred embodiment, the nucleic acid sequence encodes a protein derived from SARS-Cov2 virus, even more preferred is a spike protein derived from SARS-Cov2 virus. Influenza viruses include the genera of influenza A, B, C and D. In a further preferred embodiment, the nucleic acid sequence encodes a protein derived from influenza A, more preferred from influenza A serotype H1N1, H2N2, H3N2, H5N1, H7N7, H1N2, H9N2, H7N2, H7N3, or H10N7. In a more preferred embodiment, the nucleic acid sequence encodes a influenza A hemagglutinin (HA) protein, wherein the HA is one of the subtypes Hl, H2, H3, H4, H5, H6, H7, H8, H9, H10, Hll, H12, H13, H14, H15, H16, H17, or H18. Alternatively, the nucleic acid sequence encodes a influenza A neuraminidase (NA) of any one of subtypes Nl, N2, N3, N4, N5, N6, N7, N8, N9, N10, or Nil.

[0113] In another aspect, the invention relates to a vector comprising a nucleic acid sequence as disclosed herein, which has been optimized by any of the presently disclosed methods. In some embodiments, the vector may be an expression vector or a viral vector. In preferred embodiments, the viral vector may be a retroviral vector, an adenoviral vector, a lentiviral vector or an AAV vector.

[0114] In another aspect, the invention relates to a host cell comprising a vector as disclosed herein. In some embodiments, the host cell may be a bacterial cell, a yeast cell, a fungal cell, a plant cell, an insect cell or a mammalian cell. In some embodiments, a mammalian host cell is preferred, and is preferably a human cell.

[0115] In one embodiment, the host cell comprising the vectors according to the invention, produces a polypeptide or protein encoded by the optimized nucleic acid sequence.

[0116] The present invention further relates to use of the disclosed modified nucleic acid sequences in therapy, wherein the nucleic acid sequence has been optimized, modified, designed, and / or produced according to any of the disclosed methods. In one embodiment, the therapy is gene therapy. In other preferred embodiments, the gene therapy is nucleic-acid based vaccine therapy such as DNA or RNA vaccination or nucleic-acid based cancer immunotherapy vaccine therapy. In some embodiments, the nucleic acid-based cancer immunotherapy vaccine therapy comprises the use of personalized neoantigens representing an artificial open reading frame of about 20-30 short, mutated protein domains specific for the individual cancer cells of a subject receiving the therapeutic intervention. In other preferred embodiments, the therapy is mRNA therapy, or a vaccine comprising a modified nucleic acid sequence according to the present invention.

[0117] In one embodiment, the invention relates to a device for modifying a target nucleic acid sequence, which is capable of being in operable connection either directly or over a network with a computer programmed with one or more algorithms that when executed, carry out the methods as disclosed herein. Each of the subroutines defined by the corresponding algorithm of the device can exist as a separate algorithm element that can optionally be associated with the same device or on one or more separate devices. In a related embodiment, the device may also be capable of being in operable connection with a computer system for designing and / or rendering the disclosed modified nucleic acid sequence. The resulting modified nucleic acid sequence that is designed and processed by the device ideally completely lacks any splice donor dinucleotide sequence motif and / or a splice acceptor dinucleotide sequence motif, or at least the presence of such splice sequence motifs is substantially reduced. In another embodiment, the method further comprises producing the modified nucleic acid sequence, wherein the producing comprises de novo synthesis, mutagenesis or a mixture of both, wherein a computer or computer system is connected to a device as described herein for carrying out simulated or actual synthesis of the resulting modified nucleic acid sequence. In another embodiment, the designing and processing steps are implemented on the same device. In yet another embodiment, the designing and processing steps are implemented on separate devices in operable connection with each other. In one embodiment, the designing and processing steps are implemented either separately or concurrently by the device.

[0118] In one embodiment, the device further comprises a machine learning module configured to determine a predictive splice score to be used by the splice site identifying algorithm Ai according to the methods disclosed herein. The predictive splice score is calculated by a first quality function for patterns P obtained from a listing of sequence contexts Seqconwherein the sequence contexts of Seqconcomprise intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg). Seqconhas a length (Ncon) and comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (SeqCOnSA). In some embodiments, the sequence contexts of SeqCOn comprise from 2 to 100 nucleotides, including each integer value therebetween, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,

[0119] 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47,

[0120] 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72,

[0121] 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,

[0122] 98, 99, and 100. In some embodiments, a sequence context comprising a putative dinucleotide splice sequence motif comprises from 2 to 50 nucleotides. In other preferred embodiments, a sequence context comprises from 2 to 20 nucleotides or from 2 to 15 nucleotides. In more preferred embodiments, a sequence context comprises from 3 to 9 nucleotides. In most preferred embodiments, the listing of putative splice donor (SD) sequence contexts comprises a sequence context of 5'-NNNGTNNNN-3' which can be used to identify all putative unavoidable GT dinucleotides that may be present in the provided nucleic acid sequence. Alternatively, or additionally, the listing of putative splice acceptor (SA) sequence contexts comprises a sequence of 5'-NNNAGNNNN-3' which can be used to identify all putative unavoidable AG dinucleotides that may be present in the provided nucleic acid sequence.

[0123] In other embodiments, the device can be configured to replace one or more of the nucleotides positioned 5' and / or 3' to the GT dinucleotides within the sequence context 5'- NNNGTNNNN-3' with silent substitutions without modifying the encoded amino acid sequence. Accordingly, the putative splice score calculated for a certain sequence context 5'- NNNGTNNNN-3' can be reduced to reduce or eliminate the likelihood of a splicing event entirely.

[0124] In other embodiments, the device can be configured to replace one or more of the nucleotides positioned 5' and / or 3' to the AG dinucleotides within the sequence context 5'- NNNAGNNNN-3' with silent substitutions without modifying the encoded amino acid sequence. Accordingly, the putative splice score calculated for a certain sequence context 5'- NNNAGNNNN-3' can be reduced to reduce or eliminate the likelihood of a splicing event entirely.

[0125] In one embodiment, the device further comprises an oligonucleotide synthesizer controlled by a computer for synthesizing the optimized nucleic acid sequence or fragments thereof.

[0126] In some embodiments, the methods according to the invention may be performed in silica using a computer system, which may be directly programed with computer software that when executed implements the presently disclosed methods, or the computer system may be adapted to implement the instant methods by executing a computer program that is stored on a computer recordable medium.

[0127] In another aspect, the invention relates to a computer program product comprising instructions encoded on a non-transitory computer-readable storage medium which, when executed by a computer, causes the computer to implement the methods, including the in silica methods, disclosed herein.

[0128] In one embodiment, the instructions can, when executed by the computer, cause the computer to render the optimized nucleic acid sequence.

[0129] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. Use of the verb 'comprise' and its conjugations does not exclude the presence of elements or steps other than those stated in a claim. The article 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The invention may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In the device claims enumerating several means, one or more of these means may be embodied by one and the same item of hardware.

[0130] BRIEF DESCRIPTION OF THE DRAWINGS

[0131] FIG. 1: Comprehensive splice consensus sequences for U2 and U12 spliceosomes in different organisms adapted from Sheth et al., "Comprehensive splice-site analysis using comparative genomics", Nucleic Acid Research, 2006, vol. 34(14):3955-3967. FIG. 1A: A. thaliana GT-AG U2 splice consensus sequence. FIG. IB: A. thaliana GC-AG U2 splice consensus sequence. FIG. 1C: A. thaliana GT-AG U12 splice consensus sequence. FIG. ID: A. thaliana AT-AC U12 splice consensus sequence. FIG. IE: H. sapiens GT-AG U2 splice consensus sequence. FIG. IF: H. sapiens GC-AG U2 splice consensus sequence. FIG. 1G: H. sapiens GT- AG U12 splice consensus sequence. FIG. 1H: H. sapiens AT-AC U12 splice consensus sequence. FIG. II: C. elegans GT-AG U2 splice consensus sequence. FIG. 1J: C. elegans GC-AG U2 splice consensus sequence. FIG. IK: C. elegans GT-AG U12 splice consensus sequence. FIG. IL: M. musculus GT-AG U2 splice consensus sequence. FIG. IM: M. musculus GC-AG U2 splice consensus sequence. FIG. IN: M. musculus GT-AG U12 splice consensus sequence. FIG. 10: M. musculus AT-AC U12 splice consensus sequence. FIG. IP: D. melanogaster GT-AG U2 splice consensus sequence. FIG. IQ: D. melanogaster GC-AG U2 splice consensus sequence. FIG. 1R: D. melanogaster GT-AG U12 splice consensus sequence. FIG. IS: D. melanogaster AT-AC U12 splice consensus sequence.

[0132] FIG. 2: Overview of nucleic acid triplets / codons encoding amino acids including start and stop codons.

[0133] FIG.3 schematically depicts an example of applying the provided quality functions in methods disclosed herein to identify splice sequence motifs in a target sequence Seqtarget of length Ntarget. In this illustrative example, a listing of putative splice donor sequence contexts (SeqConSD) with sequence contexts of length NCOn = 9 comprising a central GT dinucleotide splice sequence motif is provided (A). A first quality function is applied to calculate a predictive splice score for each pattern P of length NCOn, such as e.g., pattern "MAGGTRAGT" obtained from the plurality of sequence contexts (expressed in IUPAC notation) (B). A second quality function is then provided that is capable of calculating a predictive splice score for a given nucleotide sequence based on its similarity to a pattern P, wherein such predictive splice score is based on (i) how it matches a respective pattern P (e.g., "MAGGTRAGT"), (ii) how relevant pattern P (e.g., "MAGGTRAGT") is for splicing (as per the first quality function), and (iii) its relevance with relation to a "maximum threshold splice score" determined by a splice site identifying algorithm Ai (e.g., set to 1.0 or some other appropriate value as described below). The second quality function is then applied to a given nucleotide sequence of length Neon (e.g., "AAGGTAAGT") to calculate a respective predictive splice score for such sequence that indicates whether this sequence is "too similar" to a pattern P (e.g., "MAGGTRAGT") (C). Finally, a third quality function defined by splice site identifying algorithm Ai is applied to determine for a sequence window D within Seqtarget whether one or more sequence motifs in window D with a "high" predictive splice score (as per the second quality function) are present (D).

[0134] FIG. 4A schematically shows a computer readable medium product featuring a writable part comprising a computer program according to an embodiment of the invention.

[0135] FIG. 4B schematically shows a representation of a computer processing system according to an embodiment of the invention.

[0136] DETAILED DESCRIPTION OF THE INVENTION

[0137] Definitions

[0138] The term "base pair" as used herein is interchangeable with the standardized abbreviation of "bp" or "b" and refers to a pair of complementary nucleotide bases, for instance, in a doublestranded nucleic acid molecule.

[0139] The methods described herein for expressing the polypeptides and proteins encoded by the disclosed optimized nucleic acid sequences can be achieved using in vitro, ex vivo, and / or in vivo approaches. The term "in vitro" as used herein refers to an approach used "within the glass", i.e., in a laboratory environment using suitable labware equipment such as test tubes, flasks, Petri dishes, microtiter plates using e.g., microorganisms, cells or biological molecules isolated from their usual biological surroundings. "Ex vivo" approaches refer to experimentation or measurements performed in or on samples (such as tissue or cells) extracted from an organism in an external environment with minimal alteration of natural conditions, thereby permitting experimentation under more controlled conditions than is possible in in vivo studies. "In vivo" approaches involve living organisms or cells, preferably mammalian cells, even more preferably using human cells, and is well suited for observing the overall effects of an experiment on a living subject.

[0140] "Gene" or "genes" as used herein refers to a nucleic acid (e.g., DNA or RNA) sequence that comprises coding sequences necessary for the production of a polypeptide or a protein of interest. The polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence so long as the desired activity or functional properties (e.g., enzymatic activity, ligand binding, signal transduction, etc.) are retained. The coding portion of a sequence is also referred to as "open reading frame" or "ORF". As used herein an "open reading frame" or ORF refers to a continuous nucleic acid sequence of either DNA or RNA without any termination codon that encodes an amino acid sequence (e.g., a protein or polypeptide). Typically, the nucleic acids comprise a translation start signal or initiation codon, such as ATG or AUG, and a termination codon (TAA, TAG, TGA) defining the end of the open reading frame .The term "gene" also encompasses the coding region of a structural gene and the including sequences located adjacent to the coding region on both the 5' and 3' ends for a distance of about 1 kb on either end such that the gene corresponds to the length of the full-length mRNA. The sequences that are located 5' of the coding region and which are present on the mRNA are referred to as 5' untranslated sequences or 5'UTR. The sequences that are located 3' or downstream of the coding region and that are present on the mRNA are referred to as 3' untranslated sequences or 3'UTR. The term "gene" also encompasses both natural and artificial cDNA and genomic forms of a gene. A genomic form or clone of a gene contains the coding region interrupted with non-coding sequences termed "introns" or "intervening regions" or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA); introns may contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcript; introns therefore are absent in a messenger RNA (mRNA) transcript. The mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide.

[0141] "Expression vector," or "vector" as used herein refers to a plasmid that can introduce an optimized nucleic acid sequence into a host cell or other carrier to be transcribed and translated. The terms "vector" and "plasmid" are used interchangeably herein and refer to a polynucleotide vehicle to introduce genetic material into a cell. Vectors can be linear or circular. Vectors can integrate into a target genome of a host cell or replicate independently in a host cell. Vectors can comprise, for example, an origin of replication, a multicloning site, and / or a selectable marker. An "expression vector" typically comprises an expression cassette. Vectors and plasmids include, but are not limited to, integrating vectors, prokaryotic plasmids, eukaryotic plasmids, plant synthetic chromosomes, episomes, viral vectors, cosmids, and artificial chromosomes. In the context of the present invention, a vector can comprise a nucleic acid sequence which has been optimized according to the disclosed methods. Vectors can also include sequences encoding selectable or screenable markers, as well as polynucleotides encoding protein tags (e.g., poly-His tags, hemagglutinin tags, fluorescent protein tags, and bioluminescent tags).

[0142] As used herein, the term "expression cassette" refers to a polynucleotide construct, which can be generated recombinantly or synthetically, and comprising regulatory sequences operably linked to the polynucleotide that can facilitate expression of the polynucleotide in a host cell. For example, a regulatory sequence can facilitate transcription of the polynucleotide in a host cell, or transcription and translation of the polynucleotide in a host cell. An expression cassette can, for example, be integrated in the genome of a host cell or be present in an expression vector. The polynucleotide may be an optimized nucleic acid sequence according to the present invention.

[0143] The term "therapy" as used herein refers to an approach that uses an optimized nucleic acid according to the disclosed methods to mediate a therapeutic effect. The therapy, according to the present invention and as described herein, may comprise gene therapy, mRNA therapy, and vaccines, including nucleic acid-based therapy such as cancer immunotherapy vaccine therapy. In some embodiments, the nucleic acid sequence encodes a protein derived from the family Coronaviridae, preferably the nucleic acid sequence encodes a protein derived from SARS-Cov2 virus, and even more preferably the protein is a spike protein derived from SARS-Cov2 virus. In this context, the therapy can be DNA vaccination or RNA vaccination.

[0144] A mutation that changes a codon to a synonymous codon is referred to as a "silent mutation" and is used interchangeably with the term "silent substitution" or "silent substitution mutation." This type of mutation is a neutral mutation since the resulting amino acid is unchanged from the original amino acid. The term "silent substitution" as used herein is also interchangeable with the phrase "replacement with a synonymous codon" and refers to a replacement of one codon with a "synonymous" codon that can encode the same resultant amino acid in a manner that, for instance, does not have an observable effect on the organism's phenotype. However, whereas a silent substitution does not change the encoded amino acid, it can affect the functionality or processing of the underlying nucleic acid sequence, e.g., during transcription, splicing, mRNA transport, and translation processes and is reflected in the codon usage bias that is observed in many species. The concept of silent codon substitution is illustrated by the multiple codon options selectable for a given amino acid as shown in FIG. 3. To emphasize, mutations that are capable of causing an altered codon to produce an amino acid having similar functionality to the original unmodified codon are classified as silent where the properties of the amino acid are conserved, and the mutation does not significantly affect protein structure and / or function.

[0145] Exemplary organisms for which optimizing a nucleic acid sequence for the expression of a polypeptide or protein of interest according to the methods disclosed herein include viruses, especially vaccinia viruses; prokaryotes, for example, Escherichia coli, Caulobacter cresentus, Bacillus subtilis, Mycobacterium spec.; monocotyledonous plants, especially Oryza sativa, Zea mays, Triticum aestivum; dicotyledonous plants, especially Glycin max, Gossypium hirsutum, Nicotiana tabacum, Arabidopsis thaliana, Solanum tuberosum; yeasts, including Saccharomyces cerevisiae, Schizosaccharomyces pombe, Pichia pastoris, Pichia angusta; insects, including Spodoptera frugiperda, Drosophila spec.; mammals, including Macaca mulatto, Mus musculus, Bos taurus, Capra hircus, Ovis aries, Oryctolagus cuniculus, Rattus norvegicus, Chinese hamster ovary, and preferably Homo sapiens.

[0146] Exemplary polypeptides and proteins for which an optimized nucleic acid sequence can be generated according to the methods of invention include enzymes, such as polymerases, endonucleases, ligases, lipases, proteases, kinases, phosphatases, topoisomerases; cytokines, chemokines, transcription factors, oncogenes; proteins derived from thermophilic organisms, from cryophilic organisms, from halophilic organisms, from acidophilic organisms, from basophilic organisms; proteins with repetitive sequence elements, especially structural proteins; human antigens, especially tumor antigens, tumor markers, autoimmune antigens, diagnostic markers; viral antigens, especially from HAV, HBV, HCV, HIV, SIV, FIV, HPV, rhinoviruses, influenza viruses, herpesviruses, poliomaviruses, hendra virus, dengue virus, AAV, adenoviruses, HTLV, RSV; antigens of protozoa and / or disease-causing parasites, especially those causing malaria, leishmania; trypanosoma, toxoplasmas, amoeba; antigens of disease-causing bacteria or bacterial pathogens, especially of the genera Chlamydia, staphylococci, Klebsiella, Streptococcus, Salmonella, Listeria, Borrelia, Escherichia coli; antigens of organisms of safety level L4, especially Bacillus anthracis, Ebola virus, Marburg virus, and poxviruses.

[0147] The preceding listing of exemplary organisms, polypeptides and proteins suitable for practice in accordance with the present invention is not limited thereto but is merely intended as illustrative for implementation of the disclosed compositions, methods, and use in the described therapeutic applications.

[0148] "In silica" generally refers to a method or other task that is performed using a computer or using computer simulation. The term is used interchangeably herein with the term "computer-implemented" whereby a computer, alone or in operable connection with a computer system, can be programmed with computer software or is capable of executing software stored on a computer recordable medium to thereby perform the invention disclosed herein.

[0149] Other terms have been defined herein in the context of one or more illustrative embodiments.

[0150] Methods

[0151] The methods disclosed herein provide an optimized nucleic acid sequence by avoiding highly conserved donor and acceptor splice sequence motifs in a manner that is not restricted to the context of their broader purpose. For example, the present methods seek to reduce and preferably completely eliminate any adverse splicing events for the provided nucleic acid sequence transcript herein such that all splice donor dinucleotide sequence motifs (e.g., GT or GC dinucleotide motifs) and / or all splice acceptor dinucleotide sequence motifs (e.g., AG dinucleotide motifs) present in the nucleic acid sequence are identified and removed to the fullest extent possible. According to the presently disclosed methods, each optimized open reading frame should be protected from the occurrence of any adverse splicing events that may preclude a full-length expression of the corresponding polypeptide or protein product.

[0152] The term "spliceosome" refers to cellular machinery that carries out a splicing reaction that are normally resident inside a nucleus of a eukaryotic cell. Spliceosomes are typically a complex of small nuclear ribonucleoproteins (snRNPs), which comprise both proteins and various small nuclear RNAs (snRNAs), that remove introns from a pre-mRNA primary sequence transcript. The "major" spliceosomal-splicing pathway is sometimes referred to as "U2 dependent," based on a predominant class of introns found in the mRNA primary transcripts that are exclusively recognized by the U2 snRNP during the early stages of spliceosomal assembly. Some eukaryotes feature a second category of spliceosome, the "minor" spliceosome, which is also normally found in the cell nucleus and features a group of less abundant snRNAs, such as U12, that splices a rare class of pre-mRNA introns, denoted as U12-type. The term "U2" in the context of "U2 snRNP-dependent splicing machinery" for "U2 snRNP-dependent introns" is an essential component of the major spliceosomal complex and is based on a class of introns found in mRNA primary transcripts that are recognized exclusively by the U2 snRNP during early stages of spliceosomal assembly. U2 snRNA is implicated in intron recognition through a 7-12 nucleotide sequence between 18-40 nucleotides upstream of the 3' splice site known as the branch point sequence (BPS). Different organisms will have different consensus BPSs in the proximity of the BPS. The term "U12" in the context of "U12 snRNP-dependent splicing machinery" for "U12 snRNP-dependent introns" is an essential component of the minor spliceosomal complex and is based on a divergent class of low-abundance pre-mRNA introns. Although the U12 sequence is quite divergent from U2, the two remain functionally analogous.

[0153] The active putative splice sequence motif that is identified and optimally avoided in accordance with the calculation of the predictive splice score as described herein, may be a dinucleotide splice sequence motif, which can comprise a splice donor dinucleotide sequence motif or a splice acceptor dinucleotide sequence motif. In some embodiments, the splice donor dinucleotide sequence motif may be a highly conserved GT motif. In other embodiments, the splice acceptor dinucleotide sequence motif may be a highly conserved AG motif.

[0154] Not only GT donor dinucleotides or AG acceptor dinucleotides, but also other donor and acceptor dinucleotides that occur less frequently but are still highly conserved can be identified and replaced by the methods according to the present invention. One such example of a highly conserved donor dinucleotide splicing motif is a GC motif. Another highly conserved donor dinucleotide splicing motif is an RT motif. Additional highly conserved donor and acceptor dinucleotide motifs are also contemplated for use in the methods disclosed herein.

[0155] Splicing events catalyzed by the U2 snRNP-dependent splicing machinery in the case of U2 snRNP-dependent introns include GT donor and AG acceptor dinucleotides, GC donor and AG acceptor dinucleotides, and AT donor and AC acceptor dinucleotides. As described above, U2 snRNP-dependent introns represent the primary spliceosomal-splicing pathway for the majority of all introns.

[0156] Furthermore, splicing catalyzed by the U12 snRNP dependent splicing machinery in the case of U12 snRNP-dependent introns include AT donor and AC acceptor dinucleotides and also GT donor and AG acceptor dinucleotides. U12 snRNP-dependent introns represent the minor class of introns.

[0157] It may be the case that the dinucleotide motif, e.g., the GT donor dinucleotide, which is identified as a putative splice donor sequence motif is an unavoidable GT motif as defined herein. As disclosed herein, the term "unavoidable GT motif" refers to a GT dinucleotide sequence splice motif that cannot be replaced by a silent substitution using an alternative "synonymous" codon sequence, and is typically defined by certain codon sequence contexts such as valine (GTN), wherein N is any of A, T, C, or G; or in the context of methionine (ATG) or tryptophan (TGG) followed by a phenylalanine (TTY), tyrosine (TAY), cysteine (TGY) or tryptophan (TGG), wherein Y is T or C. Only in these "unavoidable" cases, which arithmetically account for only 7% of all amino acid positions, the broader context of the donor consensus splice site sequences must be considered for a given organism in instances where adjacent nucleotides must be altered in order to eliminate or substantially reduce a splicing event that can adversely impact the ultimate expression of the resultant polypeptide or protein product.

[0158] According to the present invention, such modifications of unavoidable GT sequences that cannot be based on a silent substitution, should not be constructed from a single "matrix", representing the frequency of each base in the putative splice donor sequences for an exemplary organism (as exemplified below). Instead, one may beneficially consider multiple donor consensus matrices reflecting corresponding consensus sequences that can be applied in this type of GT replacement to eliminate the splice site, while keeping the resultant amino acid codon transcript the same. This approach thus advantageously provides a modified or optimized nucleic acid sequence that successfully avoids residual, potentially adverse, splicing events that may compromise the expression of the resulting full-length polypeptide or protein product.

[0159] In another embodiment, the present methods can identify alternative splice donor dinucleotide sequence motifs based on, for example, a donor GC, RT orTB dinucleotide motif, wherein R refers to A or G and wherein B refers to C, G, orT, respectively. Similar to the splice donor GT dinucleotide motif, the GC, RT and TB splice donor dinucleotide motifs may each be replaced by a silent substitution. Where a silent substitution is not possible, similar to the case for an unavoidable GT splice donor motif, a predictive splice score for at least the sequence contexts 5'-NNNGCNNN-3', 5'-NNNRTNNN-3', and 5'-NNNTBNNN-3', for all such GC, RT, and TB dinucleotides, respectively, can be calculated by comparing the sequence contexts against an organism-specific set of splice donor consensus sequences.

[0160] In one exemplary implementation of the presently disclosed methods, six different donor consensus matrices are defined for humans, as depicted in FIG. 5 on p. 34 of Sibley et al, , Nat Rev Genet (2016); vol. 17(7):407-421; hereinafter: "Sibley 2016", (the disclosure of which is incorporated herein by reference), which shows various consensus splice donor sequence motifs of human introns, where the height of each letter is proportional to the frequency of the corresponding base at the given position, and bases are listed in descending order for frequency from top to bottom. The two most common human splice donor consensus sequence motifs, which are recognized by the U2-dependent spliceosome, contain the highly conserved GT dinucleotide (100%, at positions 4, 5, respectively) and occur with genomic frequencies of 53.58% and 45.10% (FIG. 5A of Sibley 2016). Approximately 99% of all human introns are spliced by the major Independent spliceosomal splicing pathway and begin with a GT splice donor sequence motif and end with an acceptor AG splice sequence motif. U2- type donor splice sites that start, alternatively, with a GC dinucleotide are referred to as "atypical" U2-type splice sites and occur with a frequency of 0.87%, 0.06%, 0.02% (FIG. 5C of Sibley 2016). In addition, introns can also be spliced by the minor U12-dependent spliceosome. The donor consensus sequences for the U12-dependent spliceosome are longer compared to the U2 donor consensus sequences and occur with a much lower frequency of 0.37% (FIG. 5B of Sibley 2016). Therefore, the methods disclosed herein may not only advantageously detect and replace GT- dinucleotide containing splice donor sites of the U2-dependent spliceosome but can also beneficially identify and process less frequent but still highly conserved "atypical" or "U12- type" splice donor dinucleotides, in addition to highly conserved AG splice acceptor dinucleotides, either alternatively or in parallel.

[0161] It may also be the case that the dinucleotide motif, e.g., the AG acceptor dinucleotide, which is identified as a putative splice acceptor sequence motif is an unavoidable AG motif. As disclosed herein, the term "unavoidable AG motif" refers to an AG dinucleotide splice sequence motif that cannot be replaced by a silent substitution using an alternative "synonymous" codon sequence and is typically defined by certain codon sequence contexts such as for the amino acid pair lysine followed glutamic acid (KE), wherein K is encoded by AAA or AAG and wherein E is encoded by GAG or GAA. Other examples include the amino acid pairs EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, #V, with "#" indicating a stop codon. In these "unavoidable" cases, the broader context of the acceptor consensus splice site sequences can be considered for a given organism in instances where adjacent nucleotides must be altered in order to eliminate or substantially reduce a splicing event that can adversely impact the ultimate expression of the resultant polypeptide or protein product.

[0162] According to the present invention, such alterations of unavoidable AG sequences that cannot be based on a silent substitution, should not be constructed from a single matrix, but instead may beneficially consider multiple donor consensus matrices reflecting corresponding consensus sequences that can be applied in this type of AG replacement to eliminate the splice site, while keeping the resultant amino acid codon transcript the same. This approach thus advantageously provides an optimized nucleic acid sequence that successfully avoids residual, potentially adverse, splicing events that may compromise the expression of the resulting full-length polypeptide or protein product.

[0163] In another embodiment, the present methods can identify alternative splice acceptor dinucleotide sequence motifs based on, for example, an acceptor AS, BG or NW, wherein S is G or C, B is C, G, or T, N is any base, and W is T or A. Similar to the AG dinucleotide splice acceptor motif, the AS, BG or NW acceptor dinucleotide motifs may each be replaced by a silent substitution. Where a silent substitution is not possible, similar to the case for an unavoidable AG splice acceptor motif, a similarity score for at least the sequence contexts 5'- NNNASNNN-3', 5'-NNNBGNNN-3' and 5'-NNNNWNNN-3", for all such AS, BG or NW dinucleotides, respectively, can be calculated by comparing the sequence context against an organism-specific set of splice acceptor consensus sequences.

[0164] In one exemplary implementation of the presently disclosed methods, six different acceptor consensus matrices are defined for humans, as depicted in FIG. 5 of Sibley 2016, which show various human consensus acceptor splice site sequences. The most common human acceptor consensus splice site sequence, which are recognized by the U2-dependent spliceosome, contain the highly conserved AG dinucleotide (100%, at positions -2 and -1, respectively) and occur with genomic frequencies of 64.55% (FIG. 5A of Sibley 2016). As described above, most (more than 99%) of human introns are spliced by the major U2-dependent spliceosome (Sibley, 2016). U2-type acceptor splice sites that start, alternatively, with a BG or NW dinucleotide are referred to as "atypical" U2-type splice sites and occur with a frequency 0.03% or 0.06%, respectively (FIG. 5C of Sibley 2016). In addition, introns can also be spliced by the minor U12-dependent spliceosome. The acceptor consensus sequences for the Independent spliceosome occur with a frequency of 0.37% (FIG. 5B of Sibley 2016).

[0165] Therefore, the methods disclosed herein may not only advantageously detect and replace AG- dinucleotide containing splice donor sites of the U2-dependent spliceosome but can also beneficially identify and process less frequent but still highly conserved "atypical" or "Ln- type" splice acceptor dinucleotides, in addition to highly conserved AG acceptor dinucleotides, either alternatively or in parallel.

[0166] Splice donor and acceptor consensus sequences may differ between different organisms. In FIG. 1 (adapted from Sheth et al., 2006), exemplary donor and acceptor consensus sequences from other species including the well-studied plant organism Arabidopsis thaliana are shown. From FIG. ID, it is apparent that the highly conserved U12-donor splice site motif from A. thaliana organism model is a motif containing seven nucleotides, whereas the splice acceptor motif is a dinucleotide, as indicated by the 100% appearance of the respective nucleotide at the position shown in the respective splice site sequence. In one embodiment, the provided nucleic acid sequence strand having an open reading frame that encodes an amino acid sequence according to the presently disclosed methods is from A. thaliana, wherein the splice donor consensus sequence comprises 5'-GGTAAG-3' (FIG. 1A). Splice site motifs are represented by the presently disclosed optimization algorithm as matrices based on a natural consensus sequence for the organism, preferably, mammalian consensus sequences, even more preferably human consensus sequences, for splice donor sequence motif and splice acceptor sequence motif sites.

[0167] One of the most common mammalian donor consensus splice sequence motifs or patterns is indicated as 5'-MAGGTRAGT-3'. The following frequency of this consensus splice sequence motif as it appears in the entire human genome, was reported in Burset et al., 2001 for each respective nucleotide position in the aforementioned consensus splice sequence motif: M(70), A(60), G(80), G(100), T(100), R(95), A(71), G(81), T(46), wherein the numbers in brackets indicate the frequency (as a percentage) of each respective nucleotide or nucleotide class in over 22199 splice donor splice consensus sequences that were examined in the human genome. IUPAC nomenclature is typically used to describe the content of subsets of nucleotides, i.e., {A,C,G,T}, thus accounting for any ambiguity introduced by possible codon variations that may exist for an amino acid as described in https: / / droog.gs.washington.edu / mdecode / images / iupac.html and listed in the following Table 1:

[0168] Table 1. IUPAC Ambiguity Code In one preferred embodiment, the donor consensus splice sequence comprises a mammalian splice consensus sequence 5'-GGCAAGT-3'. In another preferred embodiment, the splice donor consensus sequence comprises a mammalian splice donor consensus sequence defined by 5'-RTATCCT-3', which corresponds to a U12 splice donor site, and wherein R refers to A or G. In another preferred embodiment, the splice donor consensus sequence comprises mammalian splice donor consensus sequences defined by 5'-CGGTBAA-3', wherein B refers to C, G, or T.

[0169] In one embodiment, a splice donor consensus sequence comprises, in a 5' to 3' direction, from 2 to 15 nucleotide sequences, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 sequences, wherein two of the nucleotides comprised in the sequence is a highly conserved consensus splice donor GT dinucleotide motif. In a preferred embodiment, the splice donor consensus sequence comprises up to and including 9 individual nucleotides. In an even more preferred embodiment, the splice donor consensus sequence comprises a 5'-NNGTNNNN-3' motif. Similar embodiments apply to the length of the splice acceptor consensus sequence disclosed herein.

[0170] In one embodiment, determining a "predictive splice score" according to the disclosed methods can be carried out using an approach based on a probabilistic right-linear grammar. In general, grammar theory relates to the modeling of symbol strings having origins in the field of computational linguistics, which has recently extended into applications for the probabilistic modeling of nucleic acid structures, for instance, nucleic acid structures like mRNA structures. For instance, using such probabilistic context free grammars (PCFGs) in this modeling paradigm, predictions of certain nucleotide content can be made regarding future outcomes based solely on the present state of the sequence; in other words, predicting base pairings across an entire sequence is conditional on the actual state of the system being evaluated. A grammar mainly consists of a set of production rules, rewriting rules for transforming strings. Each rule specifies a replacement of a particular string (i.e., the lefthand side) with another (i.e., the right-hand side). A rule can thus be applied to each string that contains its left-hand side and produces a string in which an occurrence of that left-hand side has been replaced with its right-hand side. The probability of a base pair derivation can be represented by the product of the probabilities of the productions used in that derivation, which can be viewed as parameters of the model. The model parameters can be directly derived from frequencies of different features observed in databases of nucleic acid structures, such as a naturally occurring DNA or RNA structure, rather than by experimental determination. Since such nucleic acid sequences have been documented to preserve their structures over their primary sequence, their structure prediction can be guided by combining evolutionary information from comparative sequence analysis with biophysical knowledge about structural plausibility based on such probabilities. Search results for structural homologs using PCFG rules can be thus scored according to PCFG derivation probabilities.

[0171] For example, as disclosed herein, in human models, splice donor models can evaluate about 3 nucleotides in an exon region to 6 nucleotides of the intron region (denoted as "-3 to +6," where "0" represents the exon-intron boundary). Splice acceptor models consider a core acceptor site covering roughly 6 nucleotides of intron to 3 nucleotides of the exon region (denoted as '-6 to +3" where "0" represents the intron-exon boundary). In addition, an upstream poly-pyrimidine tract (PPT) can extend another 15-30 nucleotides into the intron region. Some conventional gene finder programs also apply branch point models, which are positioned 15-30 base pairs upstream of the core acceptor site (FIG. 5A, Sibley 2016).

[0172] The most common splice site models are position-specific 0thor 1storder Markov chains. The 0thorder Markov chains model each position in the splice site independently of one another; a probability is computed for each base of the input sequence that occurs in the corresponding position of the splice site. These probabilities are then multiplied together to produce a predictive splice score for the entire sequence. First order Markov chains are different in that they condition these probabilities on the immediately preceding base, which can then characterize any probabilistic dependency between adjacent base positions.

[0173] The example below is a representation of the most frequent splice donor consensus sequences across mammalian organisms according to probable occurrence. Table 2 presents a matrix which represents for each position in a sequence of length 9 the frequency of each observed base in the putative splice donor sequences. This matrix can be summarized by stating that "MAGGTRAGT" representing an exemplary pattern P is a splice consensus sequence motif because this consensus sequence represents the nucleotide bases that most frequently occur in the splice donor sequences in the mammalian organism. For our purposes, namely, providing a first quality function capable of calculating a predictive splice score for each obtained pattern P, it is more appropriate to use the entirety of the information provided by the matrix.

[0174] Table 2. Matrix showing the frequency of each base in the putative splice donor sequences for an exemplary organism

[0175] The right-linear grammar generates sequences having a length of 9 nucleotides by generating the first nucleotide as an "A" with a probability of 33%, a "C" with a probability of 36% etc., a second nucleotide as an "A" with a probability of 64%, a "C" with a probability of 11%, etc. This process repeats according to the probabilities provided in the matrix in an iterative fashion until 9 total nucleotides are generated. The probability value of each nucleotide occurrence in the putative splicing donor sequence indicates the importance of the corresponding nucleotide towards the prediction made, with a value of 1.0 indicating a 100% likelihood that the nucleotide appears at that position in the sequence. Notably, the "G" in position 4 and the "T" in position 5 reflect a 1.0 (100%) probability of each respective nucleotide appearing at that position in the exemplary mammalian GT-donor splice consensus sequence.

[0176] The "predictive splice score" according to this example can be calculated using a first quality function as the product of the probabilities of the nucleotides that have been generated, which is a number between 0 and 1, indicating the probability of a splicing event assigned to those predictions, with "1" indicating a 100% likelihood that a splicing event will occur and "0" indicating a 100% likelihood that a splicing event will not occur. In this way, a putative splice sequence can be assessed based on the degree of similarity of a putative sequence context, for instance, in a target sequence, to a corresponding splice consensus sequence for the organism model under evaluation, which in the above example is a mammal (human). The nucleotides in the proximity of a splice site have the highest impact on the prediction outcome. The regions around these nucleotides are also more influential than those regions positioned at the edges of the sequence context under consideration.

[0177] Any normalizing or activation function can be applied to the splice score result as a postprocessing step, e.g., by dividing through the maximum score possible to determine whether the sequence is predisposed to a high likelihood of a splicing event. This splice score can be evaluated for each relevant nucleic acid sequence. For the purpose of illustration only, using Table 2 above, the probability value calculated for a given sequence AAGGTTAGT is 33% * 64% * 81% * 100% * 100% * 3% * 70% * 78% * 48% = 0.13% which can be normalized by dividing by the maximum multiplication value that can be achieved by the above grammar (which is 36% * 64% * 81% * 100% * 100% * 60% * 70% * 78% * 48%), and the normalization then yields a value of 0.044. In another embodiment, one could also use the minimum probability observed in any of the positions, which for pattern MAGGTRAGT is 3%. The value obtained can be interpreted as a similarity score which describes the similarity of a given sequence context to the organism-derived consensus splice donor sequences such as those represented by the matrix in Table 2.

[0178] An algorithm that can apply a silent mutation (i.e., silent codon substitution) in order to replace the dinucleotide splice sequence motif will first attempt to do so to reduce the predicted splice score to a minimum, if not zero to entirely eliminate splicing activity. In cases where a silent mutation is not possible, i.e., in the case of an unavoidable GT dinucleotide motif as described herein, an additional analysis and substitution taking the SD and / or SA sequence context into account must be performed to maximally reduce the predicted splice score as far as possible such that adverse splicing events are kept to an absolute minimum.

[0179] To illustrate these principles according to the present invention, the amino acid sequence "KVS" (Lys-Val-Ser) may be encoded by the corresponding sequence AAGGTGAGT, which is conforming to the consensus pattern MAGGTRAGT resulting from Table 2. Consequently, this exemplary sequence introduces a high risk of an adverse splicing event due to the presence of a GT dinucleotide splice donor motif in the sequence at positions 4 / 5.

[0180] Accordingly, this example represents a case of an "unavoidable GT dinucleotide motif" as defined herein, since a straightforward silent substitution cannot be used to replace the GT dinucleotide splice donor motif since any codon encoding for the "V" amino acid residue necessarily begins with "GT." The set of sequences encoding for KVS can be thus exchanged by selecting one codon choice from a set of codons {AAG, AAA} for K, one codon option from {GTG, GTC, GTT, GTA} for V, and one codon option from {AGC, TCC, TCT, AGT, TCA, TCG} for S, thus leading to a codon choice from a total of 48 possible sequences {K(2) * V(4) * S(6)}, based on the codon usage for the predefined organism model (FIG. 2). In principle, each of the 48 sequences may be individually evaluated according to the corresponding predicted splice score. The codon option corresponding to the lowest splice score can be suitably selected for incorporation into the optimized nucleic acid sequence according to the present invention.

[0181] In a preferred embodiment, the set of splice donor consensus sequences comprises mammalian splice donor consensus sequences that are defined according to nucleic acid patterns P such as e.g., 5'-MAGGTRAGT-3', wherein M is A or C and wherein R is A or G. For example, where the splice consensus sequence according to the present invention is AAGGTGAGT, the replacing step may result in a substituted sequence AAAGTTTCC to thereby reduce the predictive splice score and thus reduce the likelihood of a splicing event.

[0182] In other embodiments, the occasional unavoidable GT motifs may also be analyzed by a corresponding machine learning algorithm for predicting their splicing likelihood, which advantageously allows a redesign of the nucleic acid sequence at these positions accordingly for maximum sequence optimization. This optimization approach according to the present invention may ultimately impact the overall expression rate of the polypeptide or protein product. However, the tradeoff of an improved quality of the optimized or so modified nucleic acid sequence by eliminating all dinucleotide splice sequence motifs to thereby obtain a full-length expression of the protein or polypeptide product may be a more preferred outcome in certain circumstances since the respective nucleotide substitutions render the optimized nucleic acid sequence as essentially unbreakable. Like the unavoidable GT splice donor dinucleotide sequence motif, there also exist unavoidable splice acceptor dinucleotide sequence motifs. Unavoidable AG motifs refer to an AG dinucleotide splice sequence motif that cannot be replaced by a silent substitution using an alternative "synonymous" codon sequence and is typically defined by certain codon sequence contexts such as for the amino acid pair lysine followed glutamic acid (KE), wherein K is encoded by AAA or AAG and wherein E is encoded by GAG or GAA. Other examples include the amino acid pairs EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, #V, with "#" indicating a stop codon. In these "unavoidable" cases, the broader context of the splice acceptor consensus sequences must be considered for a given organism in instances where adjacent nucleotides must be modified in order to eliminate or substantially reduce a splicing event that can adversely impact the ultimate expression of the resultant polypeptide or protein product.

[0183] In another preferred embodiment, the disclosed methods herein further comprise a step for identifying, optionally using a suitably programmed computer, all splice acceptor dinucleotide sequence motifs based on the AG dinucleotide motifs in the nucleic acid sequence.

[0184] As a further embodiment contemplated by the present invention, the present methods comprise replacing, optionally using the suitably programmed computer, the identified AG dinucleotide motifs with one or more silent substitutions, as appropriate, without modifying the encoded amino acid sequence.

[0185] According to preferred embodiments, the steps of the methods disclosed herein are iterated until all the splice donor dinucleotide sequence motifs and splice acceptor dinucleotide sequence motifs have been identified and eliminated by codon sequence replacement in the provided nucleic acid sequence, or, where such full elimination is not possible, e.g.„ in the case of unavoidable GT or unavoidable AG dinucleotide motifs, the similarity of the given sequence context of the GT dinucleotide with the organism-specific set of consensus splice donor or acceptor sequences is "maximally reduced" meaning that the likelihood of an adverse splicing event has been reduced as far as possible according to the calculated predicted splice score following the base substitution(s) in the relevant nucleic acid sequence compared to the splice sequence motif that originally appeared at that position in the corresponding non-optimized nucleic acid sequence. In another aspect, the present invention relates to a method relating to an organismdependent model method for optimizing or modifying a nucleic acid sequence as described herein, including design and production thereof, comprising a plurality of steps.

[0186] Step (a)

[0187] The first step in this exemplary embodiment is providing a nucleic acid sequence to be optimized, which can be referred to as a target nucleic acid sequence (Seqtarget). The Seqtarget comprises an open reading frame that encodes an amino acid sequence that corresponds to a resultant polypeptide or protein. The provided nucleic acid sequence can be a naturally occurring nucleic acid sequence or engineered, optionally non-naturally occurring DNA sequence. Both such types of sequences may have been previously optimized in whole or in part using conventional approaches. The target sequence may also be obtained by providing an amino acid sequence (i.e., polypeptide or protein sequence) and translating back the amino acid sequence into a nucleic acid sequence and concurrently optimizing the sequence according to methods disclosed herein. Seqtarget may be derived from an organism, wherein the organism is an animal, a mammal, a bacterium, a pathogen, a virus, a protozoan, a plant, or a fungus including yeast. In preferred embodiments Seqtarget may be derived from a mammal, more preferably from a human.

[0188] Step (b)

[0189] In step (b) an organism-dependent splicing model is defined. Optimally, the organismdependent model has the same genetic background as the provided target nucleic acid sequence to be optimized.

[0190] Step (b)(i) generally relates to defining an organism-dependent splicing model that relies on a compiled library (or listing) of sequence contexts SeqCOn comprising one or more intron / exon boundary sequences for one or more genes, preferably derived from the genome of the predefined organism (Genorg), each sequence context around each such intron / exon boundary having a length (NCOn), (wherein N is a natural number). SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (SeqCOnSA). For establishing the respective splice donor and acceptor sequence contexts, publicly available databases of annotated genomic DNA can be consulted. As an example, for the organism Homo sapiens, the well-known NCBI database https: / / www.ncbi.nlm.nih.gov / can be used to download a GenBank file for a gene (or a subset of genes) in the human genome which contains annotated mRNA regions within the gene. The GenBank file also contains the DNA sequence of the gene, so the sequences of the introns and their sequence context can be collected from that file. As an example, for the gene ACVR1B, a GenBank file containing the annotation below can be found: mRNA: join (1..139, 23570 ... 23809, 24632 ... 24880, 29274 ... 29504, 32304 ... 32471, 33497 ... 33653, 40168 ... 40298).

[0191] The positions in the above exemplary "join annotation" sequence show the sequence positions of each exon region, e.g., exon 1 at position 1...139, exon 2 at position 23750...23809, etc. Accordingly, one can derive the position of each intron region in this sequence, e.g., the intron 1 begins at position 140 and extends to position 23569. From the corresponding Genbank file, it follows that the intron starts with the sequence GTGAGTCCT, and the context left of the intron ends with GGGGTCCAG.

[0192] Using a computer or other appropriate means, the relevant nucleic acid sequence of a gene can be downloaded and the compiling of the intron sequence positions in that gene and their contexts can thus be automated. This automation process can be thus executed for a large number of genes.

[0193] There exist alternative methods of obtaining a list of exon / intron boundaries for a relevant gene for a particular organism, including, for example, on the U.S. government-sponsored National Center for Biotechnology Information (NCBI) website, which provides files containing positions and contexts of introns relative to the DNA sequence of the entire human genome, see: (https: / / www.ncbi.nlm.nih.gov / IEB / Research / Acembly / Download / Downloads.html).

[0194] Once such a listing of sequence contexts has been established for a particular gene (or a group of genes), a method of predicting a splice score of a putative splice sequence in the target nucleic acid sequence can be implemented as a first quality function that calculates a predicted splice score for patterns P of length NCOn derived from the listings of splice sequence contexts SeqCOn based on the actual or probable presence of, and thus similarity to, the respective consensus splice sequence motifs identified for that organism genome.

[0195] In one embodiment, after the respective sequence listings have been compiled, a first quality function is provided that is capable of calculating for identified pattern(s) P a predictive splice score for determining whether the pattern P indicates the presence of a putative SD sequence motif or putative SA sequence motif of Genorg.

[0196] First, whether a pattern P includes a putative splice donor / acceptor sequence motif can be calculated based on the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn. Determining whether a putative splice donor / acceptor sequence motif is present in a pattern P can also be calculated based on the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn. Other approaches useful for initially identifying a putative splice donor / acceptor sequence motif in predefined sequence contexts or derived patternscan be achieved by, for instance, training and applying a neural network such as the SpliceRover prediction tool (Zuallaert et al. 2018), or by utilizing alternative machine learning methods useful for this purpose.

[0197] Accordingly, a first quality function is provided that is capable of calculating a predictive splice score for each pattern P obtained from the listing of sequence contexts identified in the one or more genes from the predefined organism genome Genorg. Patterns P (also referred to as IUPAC pattern herein) represent consensus SD and SA sequences obtained from the plurality of individual sequences within the listings of putative splice donor sequence contexts and splice acceptor sequence contexts (SeqCOnSD and SeqCOnSA, respectively). The predictive splice score calculated by the first quality function will determine whether such pattern P of length NCon indicates the presence of a putative SD sequence motif or SA sequence motif of Genorgwherein such calculation is based on a first and / or second "probabilistic model". According to the first probabilistic model the calculation is based on the percentage of sequences of SeqConSD and / or SeqCOnSA which match the pattern P of length NCOn, (wherein "matching" means that a sequence falls within the consensus sequences represented by an IUPAC pattern P). According to the second "probabilistic model" the calculation is based on percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn. The first quality function is then applied to respective pattern(s) P to calculate their predictive splice score(s).

[0198] By way of illustration according to the present invention, continuing the nucleic acid optimizing example provided above, we identified an IUPAC pattern P "SSGTGA" that corresponds in this example to a splice donor consensus motif, in this case a human consensus motif, wherein S is G or C. The percentage of sequences within SeqCOnSD that match this IUPAC pattern is calculated according to the first probabilistic model. In addition, the probability that a uniformly and randomly generated sequence of length NCOn (e.g., a sequence of length 6, length 9, length 20, length 40 etc.) on the basis of a uniform distribution matches the IUPAC pattern SSGTGA can also be calculated according to the second probabilistic model. Both probabilistic models can be specifically evaluated for any IUPAC pattern of length 6 in this example, which is the length of the exemplary "SSGTGA" consensus sequence IUPAC pattern P. This example is for illustrative purposes only and an IUPAC pattern P of variable length NCOn is suitable for calculating a predictive splice score, although the length of the applied organism-dependent consensus sequence should be selected to allow a reliable generalization over a much larger sequence length. For example, the length of the IUPAC pattern P typically represents splice consensus sequence contexts including at least the SD or SA dinucleotide sequence motif and additional nucleotides in 5' and / or 3' orientation. In some examples, the length of pattern P can be greater than 2 or less than or equal to 15, or any whole integer amount in between, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, and 15 bases in length. In some instances, the length of pattern P may be greater than 10 or less than 40. An IUPAC pattern of variable length is also contemplated as suitable for use as a "window length," indicated as a respective dinucleotide consensus sequence splice site window (D), which can be provided in step (d)(i) herein. However, the length of a window (D) may also deviate from the length of a pattern P.

[0199] In yet another embodiment, an IUPAC pattern located at a certain position in the putative splice motifs in each respective list are analyzed and IUPAC patterns determined to have the smallest entropy compared to the entropy of a uniform distribution at those positions are selected. Alternative conventional methods suitable for this analysis are well-known to the skilled person working in the field of bioinformatics for evaluating the quality of such predictive modeling. Alternatively, in some embodiments a putative splice site in the nucleic acid sequence to be optimized may be initially identified based on available prediction tools, for example, the SpliceRover application (Zuallaert et al. 2018) or using any of the techniques reported in Riepe et al., 2021.

[0200] In another embodiment, the binary decision of whether (i) the putative splice sequence motif does contain the consensus IUPAC pattern; or (ii) the putative splice sequence motif does not contain the consensus IUPAC pattern can be replaced by similarity scoring which as described herein evaluates, on an interval between 0 and 1, whether the putative splice sequence motif contains a portion of the sequence part that is similar to the selected IUPAC pattern. This "similarity scoring" accounts for the possibility, e.g., that on a certain position in the nucleic acid sequence, an "A" (adenine) is observed in 60% of the cases while a "T" (thymine) is observed in 40% of the cases. Based on the above, the first probabilistic model would then be obtained by averaging the similarity scores for all motifs in the SeqCOnSD or SeqCOnSA listing, respectively, while the second probabilistic model would correspond to the average similarity score obtained for sequences chosen according to the uniform distribution.

[0201] Step (c)

[0202] Subsequent to step (b), in step (c) a splice site identifying algorithm (At) is defined and then applied to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif.

[0203] The splice site identifying algorithm Ai splice score comprises (i) defining a maximum threshold splice score applicable for a respective Seqtarget of length Ntarget, and (ii) providing a second quality function capable of using the following as a combination input: 1) a combination of patterns P and their predictive splice scores obtained in step (b)(iv), 2) the maximum threshold splice score defined in step (c)(i), and 3) an input nucleic acid sequence of length N con-

[0204] The splice site identifying algorithm Ai is then applied or used to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif. The maximum threshold splice score may be defined depending on the intended use of the target nucleic acid sequence and can be chosen arbitrarily. For example, where a target nucleic acid sequence is prepared for use in therapy with a requirement of no adverse splicing event to occur, the maximum threshold splice score should be chosen to allow detection of all putative splice sequence motifs in the given set of sequence listings with a minimal margin of error of x% ("a percentage of false negatives"). For example, assuming one would start with a list of 200,000 splice sequence motifs that have been observed as splice sequence contexts and assuming further that a margin of error of 1% would be acceptable, then only 2,000 of these splice sequence motifs should be reported as non-splice sequence motifs. However, in some settings an even lower margin of error (such as e.g., 0.1%, 0.001% etc.) may be required thus yielding a different maximum threshold score. The maximum threshold splice score should therefore be chosen based on the potential starting number of splice sequence motifs and the intended application (e.g., one may require that all motifs are considered but it may depend on the intended use whether this makes sense).

[0205] The term "maximally reduced" as used herein, for example, in the context of a splice sequence motif or a predictive splice score, means that the splice donor dinucleotide sequence motif and the splice acceptor dinucleotide sequence motif are substituted by way of a silent substitution, for example, either by a direct GT or AG motif replacement in the case of a straightforward "silent substitution", respectively, or by a nucleotide within the organismdependent consensus sequence comprising the unavoidable dinucleotide GT or AG motif where no further substitution of nucleotides is possible without modifying the encoded amino acid sequence. Once this optimization result has been achieved, the predictive splice score cannot be further reduced, and thus has been maximally reduced.

[0206] Step (d)

[0207] Subsequent to step (c) a splice site modifying algorithm (AM) is applied comprising the following steps: in step (i) a sequence window (D) is selected within the open reading frame of the target nucleic acid sequence (i.e., containing at least one position of the open reading frame of Seqtarget. In step (ii) the third quality function defined in in (c)(iii) is applied to sequence window D to calculate a predictive splice score for said sequence. Next, in step (iii) it is determined based on the calculated predictive splice score whether window D comprises a putative splice sequence motif. In optional step (iv) one or more silent substitutions may be specified for sequence window D of Seqtarget capable of reducing the predictive splice score. This step is considered optional, because no substitutions may need to be specified if no putative splice sequence motif has been predicted for the sequence in window D. Where it has been determined that window D comprises a putative splice sequence motif, the one or more specified silent substitutions are then applied to window D to reduce the predictive splice score for said window D without changing the encoded amino acid sequence. Preferably, the predictive splice score is maximally reduced by applying the specified silent substitutions to a sequence window D. For example, following silent nucleotide substitution, the predictive splice score for a given sequence window may be reduced to zero indicating that all putative splice sites have been eliminated from the sequence window D. In step (v) the one or more silent substitutions identified in step (iv) are applied to window D. In step (vi) previous steps (i) to (v) are then iteratively performed for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

[0208] The iteration process may be accomplished by continually moving a sequence window D of a fixed length over the entire open reading frame of Seqtarget. Alternatively, two or more sequence windows D and D' may be evaluated simultaneously starting at different positions in Seqtarget. Furthermore, a sequence window D defined in a first iteration step may overlap with a sequence window D' of a subsequent iteration step. In a preferred embodiment, step (d)(vi) is repeated until the entire open reading frame of the target sequence Seqtarget has been evaluated to maximally reduce the predictive splice score.

[0209] By way of an exemplary implementation of step (d) according to the methods disclosed herein, we refer to the exemplary optimization approach as described above, which is based on a probabilistic right-linear grammar where the probability of a base pair deviation can be represented by the product of the probabilities of the productions used in that derivation, which may be derived from the frequencies of different features as documented in various genetic sequence databases. In this way, a putative splice site sequence can be assessed based on the predictive splice or similarity score, i.e., the degree of similarity of the putative sequence to a corresponding splice consensus sequence for the organism model under evaluation, which in this case is a human model. Referring to Table 2, the probabilistic mammalian donor splice consensus pattern "MAGGTRAGT" indicates a high likelihood of a predicted splicing event according to the matrix. Thus, any deviation from this pattern P is deemed to decrease the splicing likelihood which can be calculated as the product of the probabilities of the nucleotides that have been generated. In this way, a putative splice sequence can be assessed by the calculated predictive splice score, which indicates the degree of similarity (or homology) of a given sequence context to the corresponding splice consensus patterns obtained for that organism, which is a human organism in this example.

[0210] A predictive splice score can be determined for each sequence pattern obtained from SeqCOn, the listing of sequence contexts that comprises a listing of putative splice donor sequence contexts and a listing of putative splice acceptor splice sequence contexts that have been compiled for one or more genes derived from the predefined organism genome Genorg.

[0211] A silent mutation (i.e., silent codon substitution) will be applied by the splice site modifying algorithm AM of step (d) in order to replace a putative dinucleotide splice sequence motif(s) in a given window D of Seqtarget. The algorithm AM first attempts to apply a silent codon substitution where the predictive splice score was initially indicated as a high splice score but could be maximally reduced to zero through an organism-dependent silent mutation, as determined by the third quality function according to the splice site identifying algorithm Ai of step (c)(iii). In cases of an unavoidable GT or unavoidable AG dinucleotide as described herein, an additional analysis and substitution step must be performed according to step (b)(iii) where the predicted splice score is maximally reduced as far as possible based on the second quality function of (b)(ii).

[0212] To illustrate this implementation of the presently claimed invention for identified unavoidable GT donor splice sequences as described, we again refer to the amino acid sequence "KVS" (lysine-valine-serine). The sequence KVS can be encoded by the corresponding DNA sequence AAGGTGAGT, which conforms to the above-listed consensus pattern "MAGGTRAGT" thereby indicating a putative splice donor site at the "GT" motif position and thus a high risk of a splicing event. Similar to the above, a silent mutation cannot be used to replace the GT dinucleotide motif, as any codon encoding for the V residue necessarily starts with "GT." Accordingly, the set of sequences encoding for KVS may be replaced, and thus optimized, by selecting a first codon from a set of codons {AAG, AAA} for K, a second codon from {GTG, GTC, GTT, GTA} for V, and a third codon from {AGC, TCC, TCT, AGT, TCA, TCG} for S. Each of these 48 potential codon replacement sequences may be individually evaluated according to the associated predictive splice score to identify the lowest calculated splice score, or maximally reduced splice score, such that the resulting codon can be suitably selected for incorporation into the optimized nucleic acid sequence provided by the instant methods. This exemplary process can be repeated in an iterative fashion for all sequence windows D of a given target sequence predicted splice score is maximally reduced. For purpose of illustration the above methods are further exemplified in Fig. 3.

[0213] The resulting modified nucleic acid sequence can then be additionally evaluated, rendered and / or produced as desired as described elsewhere herein, for instance, by the presently disclosed device embodiments, wherein the device may be in operable connection with one or more computer systems, either directly or over a network, wherein the device or computer system may further comprise one or more computer modules that can be suitably programmed with machine learning algorithms and / or otherwise configured to execute a computer software program product or other readable medium, which is designed to implement the methods disclosed herein.

[0214] Modified nucleic acids and their use

[0215] The present invention provides an optimized or modified nucleic acid sequence that can be designed and / or prepared and / or rendered according to the methods herein disclosed, and further relates to a vector that comprises the optimized or modified nucleic acid sequence. The invention further provides a host cell that comprises the disclosed vectors and disclosed nucleic acid sequences.

[0216] Splicing of optimized or modified nucleic acid sequences may be tested in a mammalian cell culture. For this purpose, gene variants may be cloned into suitable expression vectors such as e.g., pcDNA3.4, and transfected into Expi293™ HEK cells (both Thermo Fisher Scientific Inc.) or other suitable cells. The cells may then be harvested after 24 hours and total RNA isolated and treated with DNase I to remove any residual plasmid and genomic DNA. mRNA may then be specifically reverse transcribed into cDNA with oligo-dT primer and a target sequence-specific 5' UTR primer. Resulting DNA amplification products can be analyzed by gel electrophoresis and sequencing to quantify length variation products caused by partial splicing of the originating RNA transcripts of the designed gene variants.

[0217] Producing the optimized or modified nucleic acid sequences

[0218] The optimized or modified nucleic acid sequences disclosed herein advantageously feature a maximally reduced predictive splice score compared to the corresponding non-optimized nucleic acid sequence, wherein the optimized sequence optimally does not comprise any alterations to the encoded amino acid sequence. Methods for the production of the disclosed optimized nucleic acid sequences, either physically produced sequences or optimized sequences that are rendered or otherwise simulated using a computer medium, such as a computer system or a computer program product, are contemplated according to the present invention.

[0219] As disclosed herein, the provided nucleic acid sequences may be oligonucleotides or polynucleotides.

[0220] Nucleic acid sequence synthesis may be de novo synthesis, in situ synthesis using, for example, microarrays or multiplex gene synthesis systems in the form of programmable DNA microchips. In situ synthesis methods can include those described in, for example, WO 1998 / 41531 and the references cited therein for synthesizing polynucleotides using, e.g., phosphoramidite or other similar tools. Additional patents describing in situ nucleic acid array synthesis protocols and devices include U.S. Pub. No. 2013 / 0130321, U.S. Pub. No. 2013 / 0017977, WO 2016 / 094512, and the references cited by these documents. Additional methods for synthesizing oligonucleotides include, but are not limited to, light-directed methods utilizing masks, flow channel methods, spotting methods, pin-based methods, and methods utilizing multiple supports. An exemplary device capable of processing biological information, such as a biological sequence of interest and design thereof and rendering the biological information into a resulting biological entity is disclosed in WO 2014 / 028895, which is incorporated herein by reference in its entirety.

[0221] Vectors comprising the optimized or modified nucleic acid sequences The optimized or modified nucleic acid sequence may be present in a vector or may be cloned into a vector. Any suitable vector may be used for this purpose. For example, a vector may be a plasmid, a bacterial vector, a viral vector, a phage vector, an insect vector, a yeast vector, a mammalian vector, a BAC, a YAC, or any other like vector. In some embodiments, the type of vector that is used may be determined by the type of host cell that is chosen. Preferably, the vector is a viral vector. Suitable viral vectors can be vectors derived from retroviruses, adenoviruses, adeno-associated virus, lentiviruses, parainfluenzavirus, herpesviruses, reoviruses, paramyxoviruses, and the like. Preferably, the vector is an expression vector. Such expression vectors can be a vector comprising an optimized nucleic acid according to the invention, which may be additionally operably linked to regulatory sequences, such as promoter regions, which are capable of affecting the expression of the optimized nucleic acid sequences described herein.

[0222] Various methods are contemplated for use in the cloning process, such as for example, PCR products that may have restriction enzyme sites incorporated within, either as a result of synthesis or as a consequence of PCR amplification utilizing primers containing such sites, can be digested and cloned into a plasmid vector with compatible ends. Alternatively, selective adaptors including recognition sites compatible with the expression vector of choice can be ligated to the ends of PCR products. Other methods of nucleic acid assembly that may be used to clone one or more optimized nucleic acid sequences into a vector of choice include those described in U.S. Patent Publication Nos. 2010 / 0062495 Al; 2007 / 0292954 Al; 2003 / 0152984 AA; and 2006 / 0115850 AA and in U.S. Patents Nos. 6,083,726; 6,110,668; 5,624,827; 6,521,427; 5,869,644; 6,472,184 and 6,495,318; 8,338,091, or exonuclease-based cloning techniques as described in U.S. Patent Nos 7,575,860; 7,776,532; 8,435,736; 7,723,077; 8,815,600; 9,234,226 or 9,951,327.

[0223] In one preferred embodiment, the vector further encodes a detectable marker such as a selectable marker so that transformed cells can be selectively grown. As used herein, the term "selectable marker" refers to a gene that functions as a guide for selecting host cells containing a marker vector as described herein. Selectable markers can include, but are not limited to, fluorescent markers, luminescent markers, drug selectable markers, and the like. Fluorescent markers can include genes encoding fluorescent proteins (e.g., green fluorescent protein (GFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), red fluorescent protein (dsRFP), etc.) but not be so limited. The luminescent marker can include, but not be limited to, a gene encoding a luminescent protein such as luciferase. Drug selection markers suitable for use with the presently disclosed methods and compositions provided herein include antibiotics e.g., ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, tetracycline, chloramphenicol, and neomycin), but not limited thereto. In some embodiments, the selection can be a positive selection; that is, cells expressing the marker are isolated from the population, such that an enriched population of cells comprising the selectable marker is isolated. In other cases, the selection can be a negative selection; that is, the population being isolated from those cells, for example, resulting in an enriched population of cells that do not contain the selectable marker (such as e.g., a ccdB gene).

[0224] The disclosed vectors containing the optimized or modified nucleic acid sequence of interest may be transferred into a host cell using well-known methods, depending on the type of cellular host. For example, calcium chloride transfection is commonly utilized for prokaryotic cells, whereas calcium phosphate treatment, lipofection or electroporation are exemplary procedures that may be used for other cellular hosts. Other methods used to transform mammalian cells include the use of viral infection, polybrene, protoplast fusion, liposomes, cationic transfection procedures, and microinjection.

[0225] Once the vector has been incorporated into an appropriate host cell, the host cell may be maintained under conditions suitable for high level expression of the optimized nucleic acid sequence comprised in the vector; the expressed polypeptides may be collected and purified.

[0226] Host cells

[0227] Any suitable host cell type may be used with the disclosed methods, for example, prokaryotic, eukaryotic, bacterial, yeast, insect, and mammalian cells. For example, host cells may be bacterial cells comprising Escherichia coli, Bacillus subtilis, Mycobacterium spp., M. tuberculosis, or other suitable bacterial cells; yeast cells comprising Saccharomyces spp., Picchia spp., Candida spp., or other suitable yeast species, e.g.,, S. cerevisiae, C. albicans, S. pombe, and others; xenopus cells; mouse cells; monkey cells; human cells; mammalian cells comprising CHO, COS, BHK, HEK 293; fungal cells such as Neurospora crassa; insect cells comprising SF9 cells and Drosophila cells; worm cells such as Caenorhabditis spp.; plant cells; or other suitable cells, including for example, transgenic or other recombinant cell lines. In a preferred embodiment, the host cell is a mammalian cell, and more preferred a human cell.

[0228] In one embodiment, the optimized or modified nucleic acid sequence may be used for a vaccine. For such use, the optimized nucleic acid sequence can encode for a specific polypeptide including a pathogenic polypeptide, a cancer-specific antigen or fragments or variants thereof. Also disclosed are methods for preparing a vaccine, comprising synthesizing an optimized nucleic acid according to the invention and a formulation of a vaccine composition.

[0229] As used herein, the term "DNA vaccine" or "RNA vaccine" refers to a vaccine comprising a DNA or RNA molecule, respectively. The vaccine may comprise, however, other substances and components that may be required or which are otherwise advantageous in cases where the vaccine is administered to an individual, for example, pharmaceutical excipients and / or other adjuvants typically used in these preparations.

[0230] Formulation of the nucleic acid vaccine compositions may require the use of stabilizers, such as those described in European Patent Nos. EP0353108 or EP0869814, with the use of suitable adjuvants such as aluminum salts, phosphates and hydroxides; Freund adjuvant; N- acetylmuramyl-L-alanyl-D-isoglutamyl-L-alanine-2-[l,2-dipalmitoyl-sn-glycero-3-(hydroxy- phosphoryloxy) (Sanchez-Pescador et al., J. Immu., 141, 1720-1727 1988); molecules derived from Quillaja saponaria, as the Stimulon® (Aquila, US) the Iscoms® adjuvant (CSL ltd, US); all molecules from cholesterol or analogy thereof, as the DC-Chol® (Targeted Genetics); the glycolipid Bay R1005® (Bayer, DE); antigens from Leishmania brasiliensis as the LelF (technical name) available a Corixa Corp. (US); polymers from the polyphosphazenes family, as the Adjumer (technical name) available at the "Virus Research Institute " (US). The optimized nucleic acid can be formulated with or without the other components disclosed herein in liposomes, nanosomes, as a monoplex or polyplex formulation, or can be coupled to fatty acids derivatives, including cholesterol molecules.

[0231] DNA vaccines In one embodiment, the optimized nucleic acid sequences according to the present invention can be used for the production of DNA vaccines. Such vaccines may comprise the optimized nucleic acid sequence as part of a DNA vector (including viral vectors such as e.g., adenoviral vectors) or in the form of plasmid DNA, wherein the optimized nucleic acid encodes for a specific antigen including a pathogenic antigen.

[0232] Coronavirus-based vaccines

[0233] In one embodiment, the optimized nucleic acid sequences according to the invention may encode for a specific polypeptide or protein derived from the family Coronaviridae, preferably the nucleic acid sequence encodes a protein derived from SARS-Cov2 virus, and even more preferably the protein is a spike protein derived from SARS-Cov2 virus.

[0234] As described herein, the present invention also provides optimized nucleic acid sequences encoding proteins and polypeptides derived from members of the family Coronaviridae, a large family of single-stranded positive sense RNA viruses. Coronaviruses can further include the genera of alphacoronaviruses, betacoronaviruses, gammacoronaviruses, and deltacoronaviruses.

[0235] The term "coronavirus" or "CoV" refers to any virus of the coronavirus family, including but not limited to SARS-CoV-2, MERS-CoV, and SARS-CoV. SARS-CoV-2 refers to a novel coronavirus which has been identified as the cause of a viral outbreak that was first-identified in Wuhan, China, which rapidly spread throughout the world. SARS-CoV-2 has also been referred to as 2019-nCoV and / or the Wuhan coronavirus. This coronavirus has been reported to bind to a human host cell angiotensin-converting enzyme 2 (ACE2) receptor via a viral spike protein. The spike protein also binds to and is cleaved by TMPRSS2, which activates the spike protein for membrane fusion and subsequent entry of the virus into a cell.

[0236] The term "spike protein", "CoV-S", also called "S" or "S protein" refers to the spike protein element characteristic to a coronavirus, and can also refer to a specific class of S proteins, for instance, SARS-CoV-2-S, MERS-CoV S, and SARS-CoV S. The SARS-CoV-2-spike protein is a 1273 amino acid type I membrane glycoprotein which assembles into trimers that constitute the spikes or peplomers on the surface of the enveloped coronavirus particle. The protein has two essential functions, host receptor binding and membrane fusion, which are attributed to the N-terminal (SI) and C-terminal (S2) halves of the S protein. CoV-S binds to its cognate receptor via a receptor binding domain (RBD) present in the SI subunit.

[0237] In one embodiment, the optimized or modified nucleic acid sequence according to the present invention encodes the spike protein of SARS-CoV-2. The wild-type sequence of SARS- CoV-2 is shown below as SEQ ID NO:1 and can be found under https: / / www.ncbi.nlm.nih.gov / nuccore / NC_045512.2?report=fasta&from=21563&to= 25384, a webpage that is publicly accessible at the filing date of this application. Through the use of presently disclosed sequence optimization methods as applied to a spike protein sequence in a vaccine formulation, instead of a wild-type nucleic acid sequence or even a codon-optimized wildtype nucleic acid variant, fewer adverse splicing events will occur in the host cells of a subject receiving a vaccine therapy as disclosed herein due to the maximally reduced number of dinucleotide splice site motifs in the optimized sequence. Without being bound to theory, the present methods result in a reduced probability of adverse splicing events due to the complete, or nearly complete, elimination of dinucleotide donor / acceptor splice sequence motifs in the resulting optimized nucleic acid sequences that is incorporated into the vaccine preparation. Accordingly, the absence of splicing events leads to a reduction in adverse clinical indications that can be caused by misfolded or truncated polypeptides or proteins once the optimized nucleic acid sequences are expressed in a cell or tissue.

[0238] Not only the optimized nucleic acid sequences for the spike protein of SARS-CoV-2 for use in the production of a vaccine are contemplated herein, but also any other viral protein or polypeptide derivative suitable for use in a vaccine including antigens used in Hepatitis B virus (HBV), Human papilloma virus (HPV), influenza virus or other virus-based vaccines. The present invention additionally provides a method for generating, enhancing, or modulating a protective and / or therapeutic immune response to infection by a member of the family Coronaviridae, including SARS-CoV-2, in a vertebrate, comprising administering to a vertebrate in thereof a vaccine as disclosed herein. The vaccine may be used for the prevention or treatment of a coronavirus infection, i.e., for reducing the incidence of coronavirus infection by a member of the family Coronaviridae, including SARS-CoV-2. The presently disclosed vaccines may also be useful for therapeutic vaccination of a subject infected with a coronavirus, or suspected of being infected with a coronavirus, i.e., by a member of the family Coronaviridae, including SARS-CoV-2. Therapy

[0239] As used herein, "therapy" refers to any modality of therapeutic intervention using the presently disclosed optimized nucleic acid sequences as an active component of the therapeutic agent for treating a disease, disorder, or infection e.g., a coronavirus infection). As used herein, "therapy" can refer to gene therapy, RNA-based therapeutics, DNA-based therapeutics, DNA vaccines, and / or RNA vaccines. As used herein, "gene therapy" refers to the transfer of nucleic acid molecules, such as DNA or RNA, to certain cells, target cells of mammals, particularly humans, having a disease, disorder, or infection for which such treatment is sought. DNA or RNA is introduced into the selected target cells in such a way that the heterologous DNA is expressed, or the RNA is translated, thereby producing the therapeutic product. In particular, gene therapy can be used for disease, disorder, or infection treatment purposes. Non-limiting examples of RNA-based therapeutics include mRNA, antisense RNA and oligonucleotides, ribozymes, aptamers, interfering RNAs (RNAi), Dicersubstrate dsRNA, small hairpin RNA (shRNA), asymmetrical interfering RNA (aiRNA), microRNA (miRNA). Non-limiting examples of DNA-based therapeutics include minicircle DNA, minigene, viral DNA (e.g., Lentiviral or AAV genome), viral synthetic DNA vectors, closed- ended linear duplex DNA (ceDNA / CELiD), plasmids, bacmids, doggybone (dbDNA™) DNA vectors, minimalistic immunological -defined gene expression (MIDGE)-vector, nonviral ministring DNA vector (linear-covalently closed DNA vector), or dumbbell-shaped DNA minimal vector ("dumbbell DNA"). DNA or RNA vaccines refer to polynucleotide vaccines.

[0240] The therapy may be an immunotherapy, preferably cancer immunotherapy. The presently disclosed optimized nucleic acid sequences may also be used as the active component of a cancer vaccine. Accordingly, the one or more optimized nucleic acid sequences as disclosed herein are contemplated for use in cancer immunotherapy.

[0241] Device and computer-implemented embodiments

[0242] The present invention provides for a device for optimizing or modifying a nucleic acid sequence according to the methods disclosed herein, wherein the optimized or modified nucleic acid sequence lacks or substantially lacks a splice donor dinucleotide sequence motif and / or a splice acceptor dinucleotide sequence motif. The device can be in operable connection with a computer system useful for designing and / or rendering the optimized or modified nucleic acid sequence. The device may also optimize a nucleic acid sequence for a reduced number or preferably completely eliminated number of splice sequence motifs on the basis of the amino acid sequence of the polypeptide or protein encoded by the nucleic acid sequence, which can operate in connection with a computer processing system for designing the optimized sequence that lacks or substantially lacks a splice donor dinucleotide sequence motif and / or splice acceptor dinucleotide sequence motif.

[0243] In one embodiment, the device for optimizing or modifying a nucleotide sequence according to the invention can provide for the expression of a protein on the basis of the amino acid sequence of the protein that is encoded by the optimized or modified nucleic acid sequence, which is capable of being in operable connection either directly or over a network with a computer programmed with one or more algorithms, wherein the one or more algorithms are stored, for example, in the same or different modules on the device, comprising:

[0244] (a) means for providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame, wherein Seqtarget has a length (Ntarget),

[0245] (b) an algorithm for defining an organism-dependent splicing model comprising:

[0246] (i) compiling a listing of sequence contexts (SeqCOn) comprising one or more intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg), each sequence in SeqCOn having a length (Neon), wherein SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (Seq con SA),

[0247] (ii) an algorithm for obtaining patterns (P) of length (Ncon) from SeqCOn,

[0248] (iii) means for providing a first quality function capable of calculating a predictive splice score for each pattern P for whether the pattern P indicates the presence of a putative SD sequence motif or putative SA sequence motif of Genorgbased on:

[0249] (1) the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn, and / or (2) the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn,

[0250] (iv) an algorithm for applying the first quality function to the pattern P obtained in

[0251] (ii) to calculate their predictive splice scores;

[0252] (c) a splice site identifying algorithm (At) for:

[0253] (i) defining a maximum threshold splice score applicable for a respective Seqtarget of length N target,

[0254] (ii) providing a second quality function capable of using as a combination input: a. a pattern P and its predictive splice score calculated in (b)(iv), b. optionally, the maximum threshold splice score defined in (c)(i), and c. an input nucleic acid sequence of length NCOn, and

[0255] (iii) an algorithm for applying the splice site identifying algorithm Ai of (c)(ii) to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif;

[0256] (d) a splice site modifying algorithm (AM) for:

[0257] (i) selecting a sequence window (D) which contains at least one position of the open reading frame of Seqtarget,

[0258] (ii) applying the third quality function defined in (c)(iii) to sequence window D to calculate a predictive splice score,

[0259] (iii) determining based on the predictive splice score whether window D comprises a putative splice sequence motif,

[0260] (iv) optionally specifying one or more silent substitutions for sequence window D of Seqtarget capable of reducing the predictive splice score;

[0261] (v) applying the one or more silent substitutions specified in (d)(iv) to window D to modify a putative splice sequence motif, (vi) iteratively performing steps (d)(i) to (d)(v) for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

[0262] In a preferred aspect, each of the subroutines defined by the corresponding algorithms of the device can exist as a separate algorithm element that can optionally be associated with the same device or on one or more separate devices. The device may further comprise a machine learning module which is configured to determine a predictive splice score to be used by the splice site identifying algorithm Ai according to the methods described above.

[0263] In another aspect, the invention relates to the production of optimized or modified nucleic acid sequences having maximally reduced splice sequence motifs, wherein the method is performed in a device as disclosed herein, and wherein the computer may be in operable connection with the device for implementing the design and / or rendering of the optimized nucleic acid sequence, and wherein after having designed the optimized nucleic acid sequence, transfers the data relating to the optimized or modified nucleic acid sequence to an oligonucleotide synthesizer to thereby cause the oligonucleotide synthesizer to perform the synthesis of the optimized nucleic acid sequence or fragments thereof. In cases where the optimized nucleic acid sequence exceeds a certain length e.g., more than 200 bp, more than 500 bp), the oligonucleotide synthesizer may perform the synthesis of multiple shorter parts (or "building blocks") together comprising the optimized or modified nucleic acid sequence, which are subsequently assembled into the full-length sequence using any of the assembly methods disclosed elsewhere herein. The device may therefore further comprise an assembly element for assembling oligonucleotides or building blocks synthesized with the oligonucleotide synthesizer.

[0264] The device according to the invention may comprise one or more processing modules for implementing the methods disclosed herein. For example, the device herein may include an oligonucleotide synthesizer that is controlled by a computer to thereby synthesize the optimized nucleic acid sequence or parts thereof, and optionally further include an assembly element for assembling the parts into the full-length nucleic acid sequence. In one embodiment, the optimized or modified nucleic acid sequence can be synthesized either automatically or through an appropriate command from the user.

[0265] The presently disclosed invention may also be implemented by a computer, computer software, a computer system and / or a computer program product encoded on a non- transitory computer-readable storage medium.

[0266] One exemplary device according to the present invention relates to the device being configured to be in operable connection with a computer system comprising a receiving element for receiving nucleic acid sequence data relating to the optimized nucleic acid provided by the optimization methods disclosed herein, which may be transmitted from a transmitting element that can be in direct operable connection with the receiving element or located remotely; and an assembly element that can be connected to the receiving element and is further configured to produce the optimized nucleic acid sequence, wherein the assembly element comprises or is in operable connection with building block components for assembling the optimized nucleic acid sequence, including all relevant reagent components, such that the respective components may be synthesized within the system in an automated fashion. In another embodiment, following the design of the optimized nucleic acid sequence, the networked computer and device can transfer data relating to the optimized nucleic acid sequence to an oligonucleotide synthesizer to thereby cause the oligonucleotide synthesizer to perform the synthesis of the optimized nucleic acid sequence. The aforementioned system elements can be arranged on the same or different computer networks. The assembly element can further convert the nucleic acid sequence into a biological product such as a polypeptide or protein product. In some embodiments, the computer system is entirely automated.

[0267] In one embodiment, the disclosed methods may be executed using software, which comprises instructions for causing a computer processing system to perform the methods (FIG. 4A, FIG. 4B). The software may only include those steps taken by a particular sub-entity of the system. The software may be stored in a suitable storage medium, such as a hard disk, a floppy, a memory, an optical disc, etc. The software may be sent as a signal along a wire, or wireless, or using a data network, e.g., the Internet. The software may be made available for download and / or for remote usage on a server.

[0268] It will be appreciated that the invention also extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the invention into practice. The program may be in the form of source code, object code, a code intermediate source, and object code such as partially compiled form, or in any other form suitable for use in the implementation of the method according to the invention.

[0269] One embodiment relating to a computer program product comprises computer executable instructions corresponding to each of the processing steps of at least one of the methods disclosed herein. These instructions may be subdivided into subroutines and / or be stored in one or more files that may be linked statically or dynamically. Another embodiment relating to a computer program product comprises computer executable instructions corresponding to each of the means of at least one of the systems described herein.

[0270] In another embodiment, the invention relates to a computer program product comprising instructions encoded on a non-transitory computer-readable storage medium which, when executed by a computer, the computer program product can cause the computer to implement the methods disclosed herein, optionally in an iterative fashion. In addition, the computer program product comprises instructions, that when executed by the computer, can cause the computer to render the optimized nucleic acid sequence.

[0271] FIG. 4A shows a computer readable medium 400 having a writable part 410 comprising a computer program 420, the computer program 420 comprising instructions for causing a computer system to perform a computer implemented method, according to the presently disclosed embodiments. The computer program 420 may be embodied on the computer readable medium 400 as physical marks or by means of magnetization of the computer readable medium 400. However, any other suitable embodiment is also envisioned. Furthermore, it will be appreciated that, although the computer readable medium 400 is shown here as an optical disc, the computer readable medium 400 may be any suitable computer readable medium, such as a hard disk, solid state memory, flash memory, etc., and may be non-recordable or recordable. The computer program 420 comprises instructions for causing a computer processing system to perform the computer implemented method.

[0272] FIG. 4B shows a schematic representation of a computer processing system 450 according to an embodiment of a device. The computer system comprises one or more integrated circuits 460. The architecture of the one or more integrated circuits 460 is schematically shown in FIG. 4B. Circuit 460 comprises a processing unit 470, e.g., a CPU, for running computer program components to execute a method according to an embodiment and / or implement its modules or units. Circuit 460 comprises a memory 472 for storing programming code, data, etc. Part of the memory 472 may be read-only. Circuit 560 may comprise a communication element 476, e.g., an antenna, connectors or both, and the like. Circuit 560 may comprise a dedicated integrated circuit 474 for performing part or all of the processing defined in the method. Processor 470, memory 472, dedicated IC 474 and communication element 476 may be connected to each other via an interconnect 480, e.g., a bus. The computer system 450 may be arranged for contact and / or contact-less communication, using an antenna and / or connectors, respectively.

[0273] For example, in one embodiment, the device may comprise a processor circuit and a memory circuit, the processor being arranged to execute software stored in the memory circuit. For example, the processor circuit may be an Intel Core i7 processor, ARM Cortex-R8, etc. In an embodiment, the processor circuit may be ARM Cortex M0. The memory circuit may be an ROM circuit, or a non-volatile memory, e.g., a flash memory. The memory circuit may be a volatile memory, e.g., an SRAM memory. In the latter case, the device may comprise a nonvolatile software interface, e.g., a hard drive, a network interface, etc., arranged for providing the software.

[0274] In another embodiment, the disclosed computer system 450 may comprise one or more modules for identifying a splice donor dinucleotide sequence that may be a GT dinucleotide motif, a module for identifying a splice acceptor dinucleotide sequence motif that may be a AG dinucleotide motif, a module for determining whether the donor (GT) and / or acceptor (AG) dinucleotide motifs can be replaced by a silent substitution that does not alter the core amino acid sequence, and / or a module for calculating which nucleotides in proximity to the unavoidable GT dinucleotide motifs can be replaced by a silent substitution or an alternative substitution in order to reduce the predictive splice score of the nucleotide sequence. An exemplary computer system showing the automated conversion of biological information, such as a biological sequence of interest, into a final biological entity using a multiple computer module concept is shown in WO 2014 / 028895, which is incorporated herein by reference.

[0275] The aforementioned computer modules can be implemented on a single device or alternatively, on a plurality of different devices that are in operable connection with one another.

[0276] The methods disclosed herein can be combined with the use of conventional gene optimizing programs, for example, the GeneOptimizer™ software as described herein. Additional sequence optimization programs or algorithms that may be combined with the methods disclosed herein include those described in EP3782057, U.S. Patent No. 8,326,547, WO 2020 / 024917 or CN106951726A. The method thus further comprises optimizing, using a suitably programmed computer, the provided nucleic acid sequence for expression in a host cell by replacing one or more additional nucleotides with silent substitutions without changing the encoded amino acid sequence.

[0277] Advantages of the disclosed embodiments

[0278] The methods disclosed herein relate to an expanded optimization approach for providing optimized or modified nucleic acid sequences that are rendered unbreakable in terms of being resistant to splicing events, such as an adverse splicing event, thereby introducing an extra measure of protection when using these sequences when expressing a resultant polypeptide or protein product where full-length expression is critical. This is the case in particularfor pharmaceutical applications, where preferably all donor / acceptor dinucleotide splicing motifs should be entirely avoided, or at least maximally reduced, in the optimized nucleic acid sequence. For example, in applications involving nucleic acid- or vector-based vaccines, or in immunotherapy or gene therapy strategies. In applications where dependable expression of a full-length polypeptide or protein is more desirable than maximum expression yield, the optimized genes provided by the disclosed methods represent a significant milestone in creating reliable therapeutics based on engineered and optimized nucleic acid sequences. The application of the disclosed optimization strategies is implementable in the aforementioned GeneOptimizer™ computer program and is thus combinable with all such methods of optimizing nucleic acid sequences described herein.

[0279] EXAMPLES

[0280] EXAMPLE 1: Substitutions of donor and acceptor dinucleotide splice sequence motifs

[0281] In the following EXAMPLE 1A, all splice donor dinucleotide motifs indicated as "GT" and all splice acceptor dinucleotide motifs indicated as "AG" that are present in a provided nucleic acid sequence can be removed, and thus optimized, based on one or more silent codon substitutions.

[0282] The sequence 5'- CTG TTG CAG TAG GGC -3' encodes for the amino acid sequence LLQYG which is present in the spike protein derived from the SARS-CoV-2 Coronavirus. Notably, this sequence contains a "GT" motif at positions 3 / 4 as well as an "AG" at positions 8 / 9 (underlined, respectively). By applying one or more straightforward silent codon substitutions, all occurrences of GT and AG can be successfully removed, while still maintaining the same amino acid sequence upon expression, as follows: (1) apply the silent codon substitution TTG --> CTG (for the second "L" amino acid residue); and (2) apply the silent substitution CAG --> CAA (for the "Q" amino acid residue). These two silent substitutions result in the sequence 5'-CTG CTG CAA TAG GGC -3' (resulting replacement codons are indicated in bold), which similarly encodes for the amino acid sequence LLQYG, but beneficially omits all GT or AG dinucleotide slicing motifs that may potentially create an adverse splice event resulting in truncation or other interference of the full-length expression of the intended polypeptide / protein product. Accordingly, in this illustrative example the predicted splice score has been maximally reduced to zero.

[0283] The following EXAMPLE IB illustrates when the splice acceptor dinucleotide AG cannot be removed and replaced using a silent substitution, thus creating an "unavoidable" AG situation:

[0284] Consider any amino acid sequence which contains the amino acid pair "KE." As shown in FIG. 2, the K amino acid residue can be encoded by AAG or AAA, and residue E can be encoded by GAG or GAA. The sequence 5' - AAG GAG - 3' thus encodes for the amino acid sequence KE. In order to remove at least the AG motif in the codon encoding for the amino acid K, one has to apply the silent mutation AAG --> AAA. Similarly, in order to remove the AG in the codon GAG encoding for E, one has to apply the silent mutation GAG --> GAA. This results in the DNA sequence 5' - AAA GAA - 3' which now contains only one AG motif at positions 3 / 4 (underlined). In this example, the splice score has therefore been maximally reduced. By applying silent substitutions to replace the original dinucleotide motif AG, such splicing motif would still exist in the sequence but is generated at a different position.

[0285] The following "AA" amino acid pair combinations may result in unavoidable AG motifs: EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, #V (with "#" indicating a stop codon). The following "AA" combinations may also result in unavoidable GT motifs: V, MC, MF, MW, MY, M#, WC, WF, WW, WY, W#. The "AA" motif is exemplary and non-limiting; other codon combinations may also result in unavoidable splice donor or splice acceptor dinucleotide motifs.

[0286] A third EXAMPLE 1C illustrates the situation where either the donor dinucleotide splice sequence motif GT or the acceptor dinucleotide splice sequence motif AG can be removed by silent substitutions, however, not at the same time:

[0287] Consider the amino acid sequence "WS": The amino acid residue "W" can be encoded only by the codon sequence TGG. However, residue "S" can be encoded by AGC, TCC, TCT, AGT, TCA, TCG (codon options shown in FIG. 2). In order to avoid a GT donor splice motif at the border between the "W" codon and the "S" codon, one must choose a codon for S which starts with an "A" nucleotide. However, there are only two such codon options coding for S, and both begin with "AG", the acceptor splice motif. Accordingly, in this case, the simultaneous removal of both motifs by silent substitutions is not possible.

[0288] EXAMPLE 2: Optimized nucleic acid sequence with maximally reduced splice sites

[0289] SEQ ID NO: 1 represents the wild-type amino acid sequence of the SARS-CoV-2 Coronavirus spike protein and was used to design a series of optimized nucleic acid sequences according to the methods disclosed herein.

[0290] SEQ ID NO: 1. (Wild-type SARS-CoV-2 Coronavirus spike protein) MFVFLVLLPLVSSQCVNLTTRTQLPPAYTNSFTRGVYYPDKVFRSSVLHSTQDLFLPFFSNVTWFHAIHVS GTNGTKRFDNPVLPFNDGVYFASTEKSNIIRGWIFGTTLDSKTQSLLIVNNATNVVIKVCEFQFCNDPFLG VYYHKNNKSWMESEFRVYSSANNCTFEYVSQPFLMDLEGKQGNFKNLREFVFKNIDGYFKIYSKHTPINL VRDLPQGFSALEPLVDLPIGINITRFQTLLALHRSYLTPGDSSSGWTAGAAAYYVGYLQPRTFLLKYNENGT ITDAVDCALDPLSETKCTLKSFTVEKGIYQTSNFRVQPTESIVRFPNITNLCPFGEVFNATRFASVYAWNRK RISNCVADYSVLYNSASFSTFKCYGVSPTKLNDLCFTNVYADSFVIRGDEVRQIAPGQTGKIADYNYKLPD DFTGCVIAWNSNNLDSKVGGNYNYLYRLFRKSNLKPFERDISTEIYQAGSTPCNGVEGFNCYFPLQSYGF QPTNGVGYQPYRVVVLSFELLHAPATVCGPKKSTNLVKNKCVNFNFNGLTGTGVLTESNKKFLPFQQFG RDIADTTDAVRDPQTLEILDITPCSFGGVSVITPGTNTSNQVAVLYQDVNCTEVPVAIHADQLTPTWRVY STGSNVFQTRAGCLIGAEHVNNSYECDIPIGAGICASYQTQTNSPRRARSVASQSIIAYTMSLGAENSVAY SNNSIAIPTNFTISVTTEILPVSMTKTSVDCTMYICGDSTECSNLLLQYGSFCTQLNRALTGIAVEQDKNTQ EVFAQVKQIYKTPPIKDFGGFNFSQILPDPSKPSKRSFIEDLLFNKVTLADAGFIKQYGDCLGDIAARDLICA QKFNGLTVLPPLLTDEMIAQYTSALLAGTITSGWTFGAGAALQIPFAMQMAYRFNGIGVTQNVLYENQK LIANQFNSAIGKIQDSLSSTASALGKLQDVVNQNAQALNTLVKQLSSNFGAISSVLNDILSRLDKVEAEVQI DRLITGRLQSLQTYVTQQLIRAAEIRASANLAATKMSECVLGQSKRVDFCGKGYHLMSFPQSAPHGVVFL HVTYVPAQEKNFTTAPAICHDGKAHFPREGVFVSNGTHWFVTQRNFYEPQIITTDNTFVSGNCDVVIGIV NNTVYDPLQPELDSFKEELDKYFKNHTSPDVDLGDISGINASVVNIQKEIDRLNEVAKNLNESLIDLQELGK YEQYIKWPWYIWLGFIAGLIAIVMVTIMLCCMTSCCSCLKGCCSCGSCCKFDEDDSEPVLKGVKLHYT*

[0291] Three different DNA sequences (SEQ ID NOs: 2-4) all of which encode SEQ ID NO: 1 were designed by back-translation of the underlying amino acid sequence. The SpliceRover tool (Riepe et al.; Version July 18th, 2021; Zuallaert et al., 2018) was used for predicting splicing donor and splicing acceptor site motifs and the critical threshold for reported splice donors and acceptors was set at 0.1.

[0292] The sequences SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4 were each designed using the multi-parameter GeneOptimizer™ software tool based on the following criteria:

[0293] SEQ ID NO: 2 was engineered based on optimization criteria for mammalian expression taking into account codon usage adaptation, avoidance of any secondary structures, inhibitory motifs, balancing GC content etc., except that no motifs and / or constraints were used that target the removal of the mammalian splice sites that may be present in these sequences as potential splice site motifs. SEQ ID NO: 3 was engineered using the optimization criteria as specified for SEQ ID NO: 2 but also defining as an extra constraint the removal of as many occurrences of "GT" (putative splice donor motifs) and "AG" (putative splice acceptor motifs) as possible. Accordingly, by allowing only silent codon substitutions in this instance, not all "GT" and "AG" could be removed.

[0294] SEQ ID NO: 4 was engineered using the optimization criteria as specified for SEQ ID NO: 3 but in addition defining a number of extra constraints and motifs specifically targeted to remove all putative splice donor motifs and putative splice acceptor motifs according to the presently disclosed methods.

[0295] For each of these sequences, the putative GT splice donor motifs are indicated in bold, and the putative AG splice acceptor motifs are indicated in italics and underlining. For all three designed nucleic acid sequences, the SpliceRover tool was used to predict the splice score for the remaining splice donor and splice acceptor sites as shown in Table 3 below.

[0296] Interestingly, a non-negligible splice score was predicted by the SpliceRover in connection with a GT at position 3305. The prediction score at that specific position could not be decreased significantly by silent mutations in a small window around position 3305. Therefore, silent substitutions within a longer window that spanned positions 3280 to 3321 were tested, which yielded SEQ ID NO: 4 where the reported splicing prediction score was close to the threshold value of 0.1.

[0297] SEQ ID NO: 2

[0298] ATGTTCGTGTTCCTG GTG CTG CTG CCTCTG GTG A GCA GCCA GTG CGTG A ACCTG ACC ACCA GA AC AC A GCTG CCTCCAGCCT ACACC AACA GCTTT ACCA GA GG CGTGTACTACCCCG ACAA GGTGTTCA GATC CAGCGTGCTGCACTCTACCCAGGACCTGTTCCTGCCTTTCTTCAGCAACGTGACCTGGTTCCACGCCA TCCACGTGAGCGGCACCAATGGCACCAAGAGATTCGACAACCCCGTGCTGCCCTTCAACGACGGCG TGTATTTTG CCA GCACCGA GAA GTCCAAC ATC ATCA GA GGCTG G ATCTTCGG CACCAC ACTGG ACA G CAA GACCCA GA GCCTG CTG ATCGTG AACAACGCCACCAACGTGGTG ATC AA GGTGTGCGA GTTCCA GTTCTGCAACGACCCCTTCCTGGGAGTGTACTATCACAAGAACAACAAGAGCTGGATGGAAAGCGA GTTCCGG GTGTACA GCA GCGCCAAC AACTGC ACCTTCG A GTACGTGA GCCA GCCTTTCCTG ATG G AC CTGGAAGGCAAGCAGGGCAACTTCAAGAACCTGCGCGAGTTCGTGTTTAAGAACATCGACGGCTAC TTCAAGATCTACAGCAAGCACACCCCTATCAACCTGGTGCGGGATCTGCCTCAGGGCTTCTCTGCTCT G G AACCCCTG GTG G ATCTG CCCATCG GC ATC AACATC ACCCG GTTTCAGACACTG CTGG CCCTGC AC AGAAGCTACCTGACACCTGGCGATAGCAGCTCTGGATGGACAGCTGGCGCCGCTGCCTACTATGTG G G ATACCTGCA GCCTCG G ACCTTCCTG CTG AA GTACAACGA GAACG GC ACC ATCACCG ACG CCGTG GATTGTGCTCTGGATCCTCTGAGCGAGACAAAGTGCACCCTGAAGTCCTTCACCGTGGAAAAGGGC ATCTACCA GACCA GCAACTTCCGG GTG CA GCCC ACCG A GA GCATCGTG A GATTCCCC AATATCACCA ATCTGTGCCCCTTCGGCGAGGTGTTCAATGCCACCAGATTCGCCAGCGTGTACGCCTGGAACCGGAA GA GAATCA GCAACTGCGTGGCCG ACTACTCCGTG CTGTAC AACTCCGCCA GCTTCA GC ACCTTCAA G TGCTACGGCGTGAGCCCCACCAAGCTGAACGATCTGTGCTTCACCAATGTGTACGCCGACAGCTTCG TGATCAGAGGCGACGAGGTGAGACAGATCGCCCCTGGACAGACAGGCAAGATCGCCGACTACAAC TACAAGCTGCCCGACGACTTCACCGGCTGTGTGATCGCCTGGAATAGCAACAACCTGGACTCCAAG GTCGGCGGCAACTACAATTACCTGTACCGGCTGTTCCGGAAGTCCAATCTGAAGCCCTTCGAGCGG GACATCTCCACCGAGATCTATCAGGCCGGCAGCACCCCTTGTAACGGCGTGGAAGGCTTCAACTGCT ACTTCCCACTGCAGAGCTACGGCTTTCAGCCCACAAATGGCGTGGGCTACCAGCCTTACAGAGTGGT GGTGCTGAGCTTCGAGCTGCTGCATGCTCCTGCCACAGTGTGCGGCCCTAAGAAAAGCACAAACCT GGTGAAGAACAAGTGCGTCAACTTCAATTTCAACGGCCTGACCGGCACCGGCGTGCTGACAGAGAG CAACAAGAAGTTCCTGCCATTCCAGCAGTTCGGCCGGGATATCGCCGATACCACAGACGCCGTTAGA GATCCCCAGACACTGGAAATCCTGGACATCACCCCTTGCAGCTTTGGCGGCGTGTCTGTGATCACCC CTGGCACCAATACCAGCAACCAGGTGGCTGTGCTGTACCAGGACGTGAACTGTACCGAAGTGCCCG TGGCCATTCACGCCGATCAGCTGACACCTACATGGCGGGTGTACTCCACCGGCTCCAACGTGTTTCA GACCA GA GCCG G ATGTCTG ATCGG A GCCG A GC ACGTG AAC AATA GCTACG AGTGCG AC ATCCCCAT CGGCGCTGGCATCTGTGCCAGCTACCAGACACAGACAAACAGCCCCAGACGGGCCAGATCTGTGGC CA GCCA GA GC ATC ATTG CCT AC AC A ATGTCTCTG G G CG CCGAGAACAGCGTG G CCTACAGC AACAA CTCTATCGCTATCCCCACCAACTTCACCATCAGCGTGACCACAGAGATCCTGCCTGTGAGCATGACCA AGACCAGCGTGGACTGCACCATGTACATCTGCGGCGATTCCACCGAGTGCTCCAACCTGCTGCTGCA GTACGGCAGCTTCTGCACCCAGCTGAATAGAGCCCTGACAGGGATCGCCGTGGAACAGGACAAGA ACACCC AA GA GGTGTTCGCCCA GGTG AA GCA GATCTAC AA GACCCCTCCTATC AA GG ACTTCGG CG G CTTTAACTTCAGCCAGATTCTG CCCG ATCCTAGC AAGCCCAGCAAGCG G AGCTTC ATCGAGG ACCT GCTGTTCAACAAGGTGACCCTGGCCGATGCCGGCTTCATCAAGCAGTATGGCGATTGCCTGGGCGA CATTGCCGCCAGGGATCTGATTTGCGCCCAGAAGTTTAACGGACTGACAGTGCTGCCTCCTCTGCTG ACCGATGAGATGATCGCCCAGTACACATCTGCCCTGCTGGCCGGCACAATCACAAGCGGCTGGACA

[0299] TTTGGAGCTGGCGCTGCCCTGCAGATCCCTTTCGCTATGCAGATGGCCTACCGGTTCAATGGCATCG GA GTG ACCCA GAACGTG CTGT ATGA GAACCA GAA GCTG ATCGCCAACCA GTTC AACA GCG CC ATCG G CAAGATCCAGG ACAGCCTGAGCAGCACAGC AAGCG CCCTGG G AAAACTG CAGG ACGTG GTG AAC CAGAACGCCCAGGCTCTGAACACACTGGTGAAACAGCTGTCTAGCAACTTCGGCGCCATCAGCTCTG TGCTGAACGACATCCTGAGCCGCCTGGACAAGGTGGAAGCCGAGGTGCAGATCGACAGACTGATC ACCGGAAGGCTGCAGTCCCTGCAGACCTACGTTACCCAGCAGCTGATCAGAGCCGCCGAGATTAGA GCCTCTGCCAATCTGG CCGCCACCAAGATGTCTG AGTGTGTG CTG GG CCAGAGCAAGAGAGTG G AC TTTTGCGGCAAGGGCTACCACCTGATGAGCTTCCCTCAGTCTGCTCCTCACGGCGTGGTGTTTCTGCA CGTG ACATACGTGCCCG CTCAA GA GM GAATTTC ACCACCG CTCCAGCC ATCTG CCACG ACGG C AAA GCCC ACTTTCCTA GA GAA GG CGTGTTCGTG GC AACGG AACCCACTG GTTCGTG ACCCA GCGG AAC TTCTACGAGCCCCAGATCATCACCACCGACAACACCTTCGTGTCCGGCAACTGTGATGTGGTGATCG GCATTGTGAACAATACCGTGTACGACCCTCTGCAGCCCGAGCTGGACAGCTTCAAAGAGGAACTGG ACAAGTACTTTAAGAACCACACAAGCCCCGACGTGGACCTGGGAGATATCAGCGGCATCAATGCCA GCGTG GTG AATATCCA GAAA GA GATCG ACCG GCTG AACG A GGTGG CCAA GAATCTG AACG A GA GC CTG ATCG ACCTG C AAG A ACTG G G AAA GTACG A GCA GTAC ATC AA GTG G CCCTG GT AC ATCTG GCTG GGCTTTATCGCCGGACTGATTGCCATCGTGATGGTGACAATCATGCTGTGTTGCATGACCAGCTGCT GTA GCTG CCTG AA GGG CTGTTGTA GCTGTG GCTCCTGCTG CAA GTTCG ACGA GG ACG ATTCTG A GC CCGTGCTGAAGGGCGTGAAACTGCACTACACCTGA

[0300] SEQ ID NO: 3

[0301] ATGTTCGTCTTTCTGGTGCTGCTGCCCCTGGTTTCCTCTCAATGCGTGAACCTGACCACACGGACCCA ACTGCCGCCTGCCTACACCAATTCTTTCACACGGGGCGTCTACTACCCCGACAAGGTCTTTCGATCCT CCGTGCTGCACTCTACCCAGGACCTCTTCCTGCCATTCTTCTCCAACGTGACCTGGTTCCACGCCATCC ACGTTTCCGGCACCAATGGCACCAAACGCTTCGACAACCCCGTGCTGCCCTTCAACGACGGCGTCTA TTTTGCCTCCACCGAAAAATCCAACATCATCCGCGGCTGGATCTTCGGCACCACACTGGACTCCAAAA CACAATCCCTGCTGATCGTGAACAACGCCACCAACGTGGTGATCAAGGTCTGCGAATTCCAATTCTG CAACGACCCCTTCCTGGGCGTTTACTATCACAAAAACAACAAATCCTGGATGGAATCCGAATTCCGG GTCTACTCCTCCGCCAACAACTGCACCTTCGAATATGTCTCCCAACCTTTCCTGATGGACCTGGAAGG CAAACAGGGCAACTTCAAAAACCTGCGGGAATTCGTCTTTAAAAACATCGACGGCTACTTCAAAATC TACTCCAAACACACGCCCATCAACCTGGTGCGGGATCTGCCTCAGGGCTTCTCTGCTCTGGAACCCCT GGTGGATCTGCCCATCGGCATCAACATCACCCGCTTTCAAACCCTGCTGGCCCTGCACCGCTCTTACC TTACACCTGGCGATTCCTCTTCCGGATGGACTGCTGGCGCCGCTGCCTACTATGTGGGCTATCTTCAA CCCCGGACCTTCCTGCTGAAATACAACGAAAACGGCACCATCACCGACGCCGTGGACTGCGCTCTTG ATCCCCTCTCCGAAACAAAATGCACCCTGAAATCCTTCACCGTGGAAAAGGGCATCTACCAAACCTCC AACTTCCGGGTGCAACCCACCGAATCCATCGTGCGCTTCCCCAATATCACCAATCTCTGCCCCTTCGG CG AGGTCTTTAATGCCACCCGCTTCG CCTCTGTCTACG CCTG G AATCG G AAACG G ATCTCC AACTGC GTGGCCGACTACTCCGTGCTCTACAACTCCGCCTCCTTCTCCACCTTCAAATGCTACGGCGTTTCCCCT ACCAAACTGAACGACCTCTGCTTCACCAACGTCTACGCCGACTCCTTTGTGATCCGGGGCGACGAGG TTCGGCAAATCGCTCCTGGACAAACCGGCAAAATCGCCGACTACAACTACAAACTGCCCGACGACTT CACCG GCTG CGTG ATCGCTTG G AACTCCAAC AACCTG G ATTCCAAGGTCG GCGG CAACTAC AATTAC CTCTACCGGCTCTTTCGGAAATCCAACCTGAAACCTTTCGAACGGGACATCTCTACCGAAATCTACCA GGCCGGCTCTACCCCTTGCAACGGCGTGGAAGGCTTCAACTGCTACTTCCCTCTGCAATCCTACGGCT TCCAACCTACC AATG GCGTG GG CTACC AACCTTACCG CGTG GTGGTG CTCTCCTTCG AACTG CTGC A TG CTCCTG CC ACCGTCTG CG GCCCTAAAAAATCC ACCAATCTG GTG AAAAACAAATG CGTCAACTTC AATTTCAACGGCCTGACCGGCACCGGCGTGCTGACCGAATCTAACAAAAAATTCCTGCCTTTCCAAC AATTCGGCCGGGATATCGCCGACACCACTGATGCTGTTCGGGACCCTCAAACACTGGAAATTCTGGA C ATCACCCCTTG CTCATTCG GCGG CGTTTCTGTG ATCACCCCTGG CACC AACACCTCTAACCAGGTG G CCGTGCTTTACCAGGACGTGAACTGCACCGAAGTGCCCGTGGCCATTCACGCCGATCAACTGACACC TACCTGGCGCGTCTATTCCACCGGCTCCAACGTCTTTCAAACACGCGCCGGATGCCTGATCGGGGCC GAACACGTGAACAATTCCTACGAATGCGACATCCCCATCGGCGCTGGAATCTGCGCCTCCTATCAAA CCCAAAC AAACTCCCCTCGG CG CG CTCG CTCTGTGG CCTCTCAATCTATTATCGCCTACAC AATG AGC CTGGGCGCCGAAAACTCCGTGGCCTACTCTAACAACTCTATCGCTATCCCCACCAACTTCACCATCTC CGTGACCACCGAAATCCTGCCTGTCTCCATGACCAAAACCTCCGTGGATTGCACCATGTACATCTGC GGCGATTCCACCGAATGCTCCAACCTGCTGCTGCAATACGGCTCCTTCTGCACCCAACTGAATCGGG CCCTGACCGGAATTGCCGTGGAACAGGACAAAAACACCCAAGAGGTCTTTGCCCAGGTGAAACAAA TCTACAAAACCCCGCCTATCAAGGACTTCGGCGGCTTTAACTTCTCCCAAATTCTGCCCGATCCTTCG AAACCCTCCAAACGCTCCTTCATCGAGGACCTGCTTTTCAACAAGGTGACCCTGGCCGACGCCGGCT TCATCAAACAATACGGCGATTGCCTGGGCGACATTGCCGCTCGGGATCTGATCTGCGCCCAAAAATT CAATGGACTGACTGTGCTGCCTCCTCTGCTGACTGACGAAATGATCGCCCAATACACCTCCGCTCTGC TGGCCGGCACAATCACATCTGGCTGGACATTTGGCGCTGGCGCTGCCCTGCAAATCCCCTTCGCTAT GCAAATGGCCTACCGCTTCAACGGCATCGGGGTGACCCAAAATGTGCTCTACGAAAATCAAAAACT

[0302] G ATCG CC AATC AATTCAATTCCG CC ATCGG G AAAATCCAGG ACTCCCTCTCCTCTACCG CATCCG CTC TGGGCAAACTGCAGGATGTGGTGAATCAAAACGCTCAGGCCCTGAACACCCTGGTCAAACAACTCT CCTCCAATTTTGGCGCCATCTCCTCTGTGCTGAACGACATCCTCTCTCGGCTGGACAAGGTGGAAGC CGAGGTGCAAATCGACCGGCTGATTACCGGACGGCTGCAATCACTGCAAACCTACGTGACACAACA ACTGATCCGGGCTGCCGAAATTCGGGCCTCTGCTAATCTGGCCGCCACCAAAATGAGCGAATGCGT

[0303] GCTGGGCCAATCCAAACGCGTGGACTTTTGCGGCAAGGGCTACCACCTGATGAGCTTCCCTCAATCT

[0304] GCCCCTCACGGCGTGGTTTTTCTGCACGTGACATACGTGCCCGCGCAAGAAAAAAACTTTACCACTG

[0305] CTCCCGCCATCTGCCACGACGGCAAAGCTCATTTTCCCCGCGAGGGCGTCTTTGTCTCTAATGGCACC

[0306] CACTGGTTCGTGACCCAACGGAACTTCTACGAACCCCAAATCATCACCACCGACAACACCTTTGTCTC

[0307] CGGCAACTGCGACGTCGTGATCGGCATTGTGAACAATACCGTCTACGACCCACTGCAACCCGAACTG

[0308] GATTCCTTCAAAGAGGAACTGGACAAATATTTCAAAAATCACACATCTCCCGACGTGGACCTGGGCG

[0309] ATATCTCTGGCATCAATGCCTCCGTGGTGAACATCCAAAAAGAAATCGATCGCCTGAACGAGGTGGC

[0310] CAAAAATCTGAACGAATCCCTGATCGACCTGCAAGAACTGGGGAAATACGAACAATATATCAAATG

[0311] GCCCTGGTATATCTGGCTCGGCTTTATCGCCGGCCTGATTGCCATCGTGATGGTGACAATCATGCTTT

[0312] GCTGCATGACCTCTTGCTGCTCTTGCCTGAAGGGATGCTGCTCCTGCGGATCCTGCTGCAAATTCGAC

[0313] GAGGACGACTCCGAACCTGTCCTGAAGGGCGTGAAACTGCACTACACCTGA

[0314] SEQ ID NO: 4

[0315] ATGTTCGTCTTTCTCGTCCTGCTGCCTCTCGTTTCCTCTCAATGCGTCAATTTGACGACTCGGACTCAA

[0316] CTGCCTCCGGCCTATACAAATTCCTTCACTCGGGGCGTCTATTATCCCGATAAAGTCTTTCGCTCCTCC

[0317] GTCCTGCATTCGACGCAAGATCTCTTCCTGCCATTCTTCTCCAATGTGACATGGTTCCATGCCATCCAT

[0318] GTCTCCGGAACGAATGGGACAAAACGATTCGACAATCCCGTTCTGCCATTCAACGACGGCGTTTATT

[0319] TCGCCTCGACGGAAAAATCCAATATCATCCGCGGATGGATCTTCGGGACAACTCTGGACTCCAAAAC

[0320] TCAATCTCTGCTGATCGTCAACAACGCGACGAATGTCGTCATCAAAGTCTGCGAATTCCAATTCTGCA

[0321] ACGATCCTTTCCTGGGCGTTTATTATCATAAAAACAACAAATCCTGGATGGAATCCGAATTCCGCGTC

[0322] TATTCCTCCGCCAATAATTGCACATTCGAATATGTCTCGCAACCATTCCTGATGGATCTCGAGGGGAA

[0323] ACAAGGAAATTTCAAAAATCTGCGGGAATTCGTTTTCAAAAATATCGACGGCTATTTCAAAATCTATT

[0324] CGAAACACACGCCCATCAATCTCGTCCGGGATCTGCCCCAAGGATTCTCTGCTCTCGAGCCCCTCGTC

[0325] GATCTGCCCATCGGCATCAATATAACTCGATTCCAAACTCTGCTGGCCCTCCATCGATCCTATTTGACT

[0326] CCCGGCGATTCCTCCTCCGGATGGACTGCTGGCGCCGCCGCATATTATGTCGGCTATCTTCAACCCCG

[0327] GACATTCCTGCTGAAATATAATGAAAACGGGACGATAACGGATGCCGTCGACTGCGCTCTGGATCC

[0328] CCTCTCCGAAACAAAATGCACTCTGAAATCCTTCACTGTCGAAAAAGGAATCTATCAAACGAGCAAT

[0329] TTCCGCGTTCAACCGACGGAATCCATCGTTCGATTCCCCAATATAACGAATCTCTGCCCATTCGGGGA

[0330] AG7TTTCAACGCGACACGATTCGCCTCCGTCTATGCTTGGAATCGGAAACGGATCTCCAATTGCGTC

[0331] GCCGATTATTCCGTCCTCTATAATTCCGCCTCCTTCTCGACATTCAAATGCTATGGCGTTTCCCCGACG AAACTGAACGATCTCTGCTTCACAAATGTCTATGCCGATTCCTTCGTCATCCGGGGCGACGAAG7TC GGCAAATCGCTCCTGGCCAAACTGGCAAAATTGCCGATTATAATTATAAACTGCCCGATGATTTCAC GGGATGCGTCATTGCTTGGAACTCCAATAATCTCGACTCCAAAGTCGGCGGCAACTATAATTATCTCT ATCGCCTCTTCCGGAAATCCAATCTGAAACCATTTGAACGGGATATCTCGACTGAAATCTATCAAGCT GGCTCGACTCCTTGCAACGGCGTTGAGGGATTCAACTGCTATTTCCCTCTGCAATCCTATGGATTTCA ACCGACAAATGGCGTCGGCTATCAACCCTATCGCGTCGTCGTCCTCTCCTTTGAACTGCTGCATGCTC CCGCAACTGTCTGCGGCCCCAAAAAATCAACTAATCTCGTCAAAAACAAATGCGTTAATTTCAATTTC AATGGATTGACTGGAACGGGCGTCTTGACTGAATCCAATAAAAAATTTCTCCCATTTCAACAATTCG GCCGGGATATCGCCGATACAACTGATGCCGTTCGGGATCCTCAAACGCTCGAAATCCTCGATATAAC GCCATGCTCCTTCGGCGGCGTTTCCGTCATAACTCCCGGAACAAATACGAGCAATCAAGTCGCCGTT CTCTATCAAGATGTCAACTGCACTGAAGTCCCCGTCGCCATTCATGCTGATCAATTGACTCCAACATG GCGCGTTTATTCGACGGGCTCCAATGTCTTTCAAACTCGCGCCGGATGCCTCATCGGGGCCGAGCAT GTCAACAACTCCTATGAATGCGATATTCCCATTGGCGCCGGAATCTGCGCCTCCTATCAAACTCAAAC A A ATTCCCCTCG G CGCGCTCGCTCCGTCG CCTCTC A ATCT ATC ATTG CTT AT AC A ATG AGCTTG G G CG CCGAAAACTCTGTCGCCTATTCCAATAATTCTATCGCGATCCCGACGAATTTCACGATCTCCGTGACA ACTGAAATTCTGCCCGTTTCCATGACGAAAACGAGCGTCGATTGCACGATGTATATCTGCGGCGACT CAACGGAATGCTCCAATCTGCTGCTGCAATATGGCTCCTTCTGCACTCAACTGAATCGGGCTCTGACT GGGATCGCCGTTGAACAAGACAAAAACACGCAAGAAGTCTTTGCCCAAG7TAAACAAATCTATAAA ACGCCTCCGATCAAGGATTTCGGCGGATTCAATTTTTCCCAAATTCTCCCCGATCCTTCGAAACCCTCT AAACG CTCCTTCATCG AGG ATCTG CTTTTC AACAAAGTG ACTCTG GCCG ATG CCG G ATTC ATTAAACA ATATGGCGACTGCCTCGGGGATATTGCTGCTCGGGATCTGATCTGCGCCCAAAAATTCAATGGATTG

[0332] ACTGTCCTGCCTCCTCTGCTGACGGATGAAATGATCGCCCAATATACTTCTGCCCTCCTGGCCGGAAC TATAACATCTGGATGGACATTCGGCGCCGGGGCCGCTCTGCAAATTCCATTCGCTATGCAAATGGCC TATCGATTCAATGGCATCGGCGTGACTCAAAATGTCCTCTATGAAAATCAAAAACTGATCGCCAATC AATTCAACTCCG CC ATCG GC AAAATCCAAGACTCCCTCTCCTCAACTG CTTCTG CTCTG GG AAAACTG CAAGATGTCGTTAATCAAAACGCGCAAGCTCTGAACACGCTCGTTAAACAACTCTCCTCCAATTTCGG GGCCATCTCCTCTGTCCTGAACGATATCCTCTCTCGGCTCGACAAAGTCGAAGCTGAAGTCCAAATC GATCGGCTCATAACTGGCCGCCTCCAATCTCTCCAAACATATGTGACTCAACAACTGATCCGGGCCG CCGAAATTCGGGCCTCTGCTAATCTGGCCGCAACAAAAATGAGCGAATGCGTCCTGGGCCAATCTAA ACGCGTCGATTTCTGCGGCAAAGGATATCATCTGATGAGCTTCCCTCAATCTGCCCCTCATGGCGTC GTTTTTCTG CATGTG ACTTATGTCCCCG CG CAAGAAAAAAATTTC ACAACTG CTCCCGCCATCTGCC A TGATGGCAAAGCTCATTTTCCCCGCGAGGGCGTGTTTGTTTCAAATGGCACACACTGGTTTGTCACA CAAAGGAATTTCTATGAGCCCCAAATCATAACGACGGATAATACTTTTGTCTCCGGCAACTGCGATG

[0333] TTGTTATCGGCATTGTCAACAATACTGTCTATGATCCTCTGCAACCGGAACTGGACTCCTTCAAAGAA

[0334] GAACTGGACAAATATTTCAAAAATCATACGAGCCCCGATGTCGATCTGGGCGATATCTCTGGCATCA

[0335] ATGCCTCCGTCGTCAATATCCAAAAAGAAATTGATCGCCTCAACGAAGTCGCCAAAAATCTCAACGA

[0336] ATCCCTCATCGATCTCCAAGAACTGGGGAAATATGAACAATATATCAAATGGCCATGGTATATCTGG

[0337] CTCGGATTCATTGCTGGACTGATTGCCATCGTCATGGTGACAATCATGCTTTGCTGCATGACAAGCT

[0338] GCTGCTCTTGCCTCAAAGGATGCTGCTCCTGCGGCTCCTGCTGCAAATTTGATGAGGACGACTCTGA

[0339] GCCCGTCCTGAAGGGCGTTAAACTGCATTATACTTGA

[0340] Table 3. Splicing prediction scores provided by SpliceRover using a score threshold value set at 0.1*:

[0341] * Note that a threshold value of 0.1 was chosen in this case because below that value a very large number of splicing site artefacts is likely to be reported.

[0342] For SEQ ID NO: 2 SpliceRover reported 27 human splice donors, the two highest threshold scores at positions 1218 (score 0.9859) and position 387 (score 0.9595) as well as 15 human splice acceptors with the highest scores at positions 2701 (score 0.6519) and position 1930 (score 0.2614).

[0343] For SEQ ID NO: 3 SpliceRover reported 14 human splice donors, the two highest threshold scores at the positions 1332 (score 0.9430) and position 387 (score 0.9316) as well as 8 human splice acceptors, the two highest threshold scores at positions 1837 (score 0.7371) and at position 2956 (score 0.5089).

[0344] For SEQ ID NO: 4 SpliceRover reported only one human splice donor at position 3305 with a threshold score of 0.1163 as well as a human splice acceptor at position 2474 with a threshold score of 0.1011. Example 2 thus demonstrates that, using the methods disclosed herein, an already optimized nucleic acid sequence can be additionally modified to further maximally reduce the number of GT splice donor dinucleotide sequence motifs and / or AG splice acceptor dinucleotide sequence motifs. The reduction of such splicing motifs means that the optimized or modified nucleic acid sequence is less likely to undergo an adverse splice event that may detrimentally impact the full-length expression, and possibly the physiological function, of the corresponding polypeptide or protein product.

Claims

Claims1. A method for generating a modified nucleic acid sequence comprising:(a) providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame, wherein Seqtarget has a length (Ntarget),(b) defining an organism-dependent splicing model comprising:(i) compiling a listing of sequence contexts (SeqCOn) comprising one or more intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg), each sequence in SeqCOn having a length (Neon), wherein SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (Seq con SA),(ii) obtaining patterns (P) of length (Ncon) from SeqCOn,(iii) providing a first quality function capable of calculating a predictive splice score for each pattern P for determining whether the pattern P indicates the presence of a putative SD sequence motif or putative SA sequence motif of Genorg based on:(1) the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn, and / or(2) the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn,(iv) applying the first quality function to the pattern P obtained in (ii) to calculate their predictive splice scores;(c) providing a splice site identifying algorithm (At) comprising:(i) defining a maximum threshold splice score applicable for a respective Seqtarget of length N target,(ii) providing a second quality function capable of using as a combination input:a. a pattern P and its predictive splice score calculated in (b)(iv), b. optionally, the maximum threshold splice score defined in (c)(i), and c. an input nucleic acid sequence of length NCOn, and(iii) applying the splice site identifying algorithm Ai of (c)(ii) to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif;(d) applying a splice site modifying algorithm (AM) comprising:(i) selecting a sequence window (D) which contains at least one position of the open reading frame of Seqtarget,(ii) applying the third quality function defined in (c)(iii) to sequence window D to calculate a predictive splice score,(iii) determining based on the predictive splice score whether window D comprises a putative splice sequence motif,(iv) optionally specifying one or more silent substitutions for sequence window D of Seqtarget capable of reducing the predictive splice score;(v) applying the one or more silent substitutions specified in (d)(iv) to window D to modify a putative splice sequence motif, and(vi) iteratively performing steps (d)(i) to (d)(v) for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

2. The method of claim 1, wherein Seqtarget of (a) is a naturally occurring or an engineered, optionally non-naturally occurring nucleic acid sequence.

3. The method of claim 2, wherein the Seqtarget is further optimized for the expression of a polypeptide or protein using an optimization program, optionally wherein the optimization program is a GeneOptimizer™ program.

4. The method of any of claims 1 to 3, wherein the Seqtarget is derived from an organism, wherein the organism is an animal, a mammal, a bacterium, a pathogen, a virus, a protozoan, a plant, or a fungus including yeast, wherein the Seqtarget is preferably derived from a mammal, more preferably from a human.

5. The method of any preceding claim, wherein the combination of patterns P used as input in step (c)(ii) a. comprises patterns P with a high predictive splice score.

6. The method of any preceding claim, wherein the SD sequence motif is a highly conserved GT motif.

7. The method of claim 6, wherein the highly conserved GT motif is an unavoidable GT motif encoding for a corresponding amino acid sequence:(i) valine (GTN), wherein N is any of A, T, C, or G, or(ii) methionine (ATG) or tryptophan (TGG) followed by a phenylalanine (TTY), tyrosine (TAY), cysteine (TGY) or tryptophan (TGG), wherein Y is T or C.

8. The method of claim 6, wherein the highly conserved GT motif is an avoidable GT motif and wherein step (d)(v) removes the avoidable GT motif from the Seqtarget.

9. The method of any of claims 1 to 6, wherein the SD sequence motif is a GC motif, and preferably wherein step (d)(v) removes the GC motif from the Seqtarget.

10. The method of claim 9, wherein the highly conserved GC motif encodes for a corresponding amino acid sequence:(i) alanine (GCT, GCC, GCA, GCG), or(ii) any of the following amino acid combinations: MH, MP, MQ, WH, WP and WQ.

11. The method of any preceding claim, wherein the SA sequence motif is a highly conserved AG motif.

12. The method of claim 11, wherein the highly conserved AG motif is an unavoidable AG motif encoding for a corresponding amino acid sequence:(i) lysine (K)-glutamic acid (E), wherein K is encoded by AAG or AAA and E is encoded by GAG or GAA, or(ii) any of the following amino acid combinations: EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, or #V, with "#" indicating a stop codon.

13. The method of claim 11, wherein the highly conserved AG motif is an avoidable AG motif and wherein step (d)(v) removes the avoidable AG motif from the Seqtarget.

14. The method of any preceding claim, wherein the predictive splice score calculated in step (d)(ii) for one or more windows D is maximally reduced.

15. The method of claim 14, wherein step (d)(iv) further comprises specifying zero silent substitutions for the sequence window D when the predictive splice score is calculated as maximally reduced.

16. The method of any preceding claim, wherein sequences within SeqCOn comprise from2 to 100 nucleotides, including each integer value therebetween, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30,31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53,54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76,77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100.

17. The method of claim 16, wherein sequences within SeqCOn comprise from 2 to 50 nucleotides.

18. The method of claim 16, wherein sequences within SeqCOn comprise from 2 to 20 nucleotides.

19. The method of claim 16, wherein sequences within SeqCOn comprises from 2 to 15 nucleotides.

20. The method of claim 16, wherein sequences within SeqCOn comprise from 3 to 9 nucleotides.

21. The method of any preceding claim, wherein the method further comprises optimizing Seqtarget for subsequent expression, wherein the encoded amino acid sequence is not modified.

22. The method of any preceding claim, wherein step (d) is performed for successive regions across the Seqtarget for identifying the presence of one or more putative dinucleotide splice sequence motifs by advancing the respective sequence window D in a 5' to 3' direction over the entire length N target-23. The method of any one of claims 1-22, wherein the modified sequence is for use in therapy.

24. The method of claim 23, wherein the therapy comprises gene therapy, vector-based therapy, vaccination nucleic acid-based cancer immunotherapy vaccine therapy, or DNA vaccination.

25. An in silica method for generating a modified nucleic acid sequence, comprising:(a) providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame encoding an amino acid sequence;(b) identifying, using a suitably programmed computer, a SD sequence motif comprising a GT dinucleotide sequence in the Seqtarget;(c) determining, using the suitably programmed computer,(i) whether the GT dinucleotide sequences identified in (b) can be replaced with at least one silent substitution without modifying the encoded amino acid sequence, and(ii) whether the GT dinucleotide sequences identified in (b) are unavoidable GT dinucleotides that cannot be replaced with at least one silent substitution without modifying the encoded amino acid sequence;(d) determining, using the suitably programmed computer, a similarity score for a SD sequence context of length (N), preferably 5'-NNNGTNNNN-3', for each GT dinucleotide sequence identified in (c)(ii) by comparing the sequence context in the Seqtarget against a species-specific set of SD consensus sequences of the same length; and(e) replacing, using the suitably programmed computer,(i) the GT dinucleotide sequences identified in (c)(i) with one or more silent substitutions without modifying the encoded amino acid sequence, and(ii) one or more of the nucleotides 5' and / or 3' of the unavoidable GT dinucleotide sequences within the SD sequence context for all GT dinucleotide sequences identified in (c)(ii) with silent substitutions, without modifying the encoded amino acid sequence to reduce the similarity of the SD sequence context with the species-specific set of SD consensus sequences, to thereby generate the modified nucleic acid sequence.

26. The in silica method of claim 25, wherein the method further comprises:(f) identifying, using a suitably programmed computer, a SA dinucleotide sequence motif comprising an AG dinucleotide sequence in the Seqtarget;(g) determining, using the suitably programmed computer,(i) whether the AG dinucleotide sequences identified in (f) can be replaced with at least one silent substitution without modifying the encoded amino acid sequence, and(ii) whether the AG dinucleotide sequences identified in (f) are unavoidable AG dinucleotides that cannot be replaced with at least one silent substitution without modifying the encoded amino acid sequence;(h) determining, using the suitably programmed computer, a similarity score for a SA sequence context of length (N), preferably 5'-NNNAGNNNN-3', for each AG dinucleotide identified in (g)(ii) by comparing the sequence context in theSeqtarget against a species-specific set of SA consensus sequences of the same length; and(i) replacing, using the suitably programmed computer,(i) the AG dinucleotide sequences identified in (g)(i) with one or more silent substitutions without modifying the encoded amino acid sequence, and(ii) one or more of the nucleotides 5' and / or 3' of the unavoidable AG dinucleotide sequences within the SA sequence context for all AG dinucleotide sequences identified in (g)(ii) with silent substitutions, without modifying the encoded amino acid sequence for reducing the similarity of the SA sequence context 5'- with the species-specific set of SA consensus sequences, to thereby generate the modified nucleic acid sequence.

27. The in silica method of claim 25 or 26, wherein in step (c)(ii) the unavoidable GT dinucleotides are determined as non-replaceable by silent substitutions and encode for a corresponding amino acid sequence:(i) valine (GTN), wherein N is any of A, T, C, or G, or(ii) methionine (ATG) or tryptophan (TGG) followed by a phenylalanine (TTY), tyrosine (TAY), cysteine (TGY) or tryptophan (TGG), wherein Y is T or C.

28. The in silica method of claim 25 or 26, wherein in step (g)(ii) the unavoidable AG dinucleotides are determined as non-replaceable by silent substitutions and encode for a corresponding amino acid sequence:(i) lysine (K)-glutamic acid (E), wherein K is encoded by AAG or AAA and E is encoded by GAG or GAA, or(ii) any of the following amino acid combinations: EA, ED, EE, EG, EV, KA, KD, KE, KG, KV, QA, QD, QE, QG, QV, #A, #D, #E, #G, or #V, with "#" indicating a stop codon.

29. The in silica method of claims 25 to 28, wherein the set of donor consensus splice sequences comprises mammalian donor consensus splice sequences defined by 5'- MAGGTRAGT-3 wherein M is A or C, and wherein R is A or G.

30. The in silica method of any of claims 25 to 29, wherein replacing the one or more nucleotides 5' and / or 3' to the GT dinucleotides with a silent substitution results in a substituted sequence context 5'-KBHGTYBHV-3', wherein K is G or T, wherein B is G, T or C, wherein H is A, C or T, wherein Y is C or T, or wherein V is G, C or A.

31. The in silica method of any of claims 25 to 30, wherein the method further comprises optimizing, using the suitably programmed computer, Seqtarget for expression in a designated host cell by replacing one or more additional nucleotides with one or more silent substitutions without modifying the encoded amino acid sequence.

32. The in silica method of any of claims 25 to 31, wherein the method further comprises producing the modified nucleic acid sequence, wherein the producing comprises de novo synthesis, mutagenesis or a mixture of both, optionally wherein the computer is connected to a suitably programmed device to carry out nucleic acid synthesis.

33. The in silica method of claims 25 to 32, wherein the produced nucleic acid sequence lacks or substantially lacks a SD sequence motif and / or SA sequence motif.

34. The in silica method of claims 25 to 33, wherein the substitution of nucleotides does not result in a sequence motif repetitive sequence, secondary structure or inverse complementary repeat capable of modifying the expression of the resultant polypeptide or protein.

35. A modified nucleic acid sequence obtained according to the method of any of the preceding claims, and wherein the modified nucleic acid sequence lacks or substantially lacks a SD sequence motif and / or a SA sequence motif.

36. The modified nucleic acid sequence of claim 35, wherein the nucleic acid sequence encodes a protein derived from the family Coronaviridae, preferably the nucleic acid sequence encodes a protein derived from SARS-Cov2 virus, and even more preferably the protein is a spike protein derived from SARS-Cov2 virus.

37. A vector comprising the modified nucleic acid sequence of claims 35 or 36.

38. The vector of claim 37, wherein the vector is an expression vector.

39. The vector of claim 37 or 38, wherein the vector is a viral vector.

40. The vector of claim 39, wherein the viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector or an AAV vector.

41. A host cell comprising the vector of any of claims 39 to 40.

42. The host cell of claim 41, wherein the host cell is a bacterial cell, a yeast cell, a fungal cell, a plant cell, an insect cell or a mammalian cell, wherein the mammalian cell is preferably a human cell.

43. The host cell of claim 41 or42, wherein the host cell produces a polypeptide or protein encoded by the nucleic acid sequence.

44. A method of producing a modified nucleic acid sequence comprising a reduced number of splice sequence motifs compared to a corresponding non-modified nucleic acid sequence, comprising:(i) generating a modified nucleic acid sequence according to any of claims 1 to 33, and(ii) producing the modified nucleic acid sequence of step (i), wherein the producing comprises de novo synthesis, mutagenesis or a combination thereof.

45. The method of claim 44, wherein the generating step (i) and the processing step (ii) are implemented by a device, either separately or concurrently.

46. The method of claim 45, wherein the device, subsequent to generating the modified nucleic acid sequence, transfers modified nucleic acid sequence data to an oligonucleotide synthesizer, wherein the oligonucleotide synthesizer synthesizes the modified nucleic acid sequence.

47. A modified nucleic acid sequence according to claims 35 or 36 for use in therapy.

48. The modified nucleic acid according to claim 47, wherein the therapy comprises gene therapy, nucleic acid-based cancer immunotherapy vaccine therapy, DNA vaccination, and RNA vaccination.

49. The modified nucleic acid according to claim 47 or 48, wherein the sequence encodes a protein derived from the family Coronaviridae, preferably the nucleic acid sequence encodes a protein derived from SARS-Cov2 virus, and even more preferably the protein is a spike protein derived from SARS-Cov2 virus, and wherein the therapy is DNA or RNA vaccination.

50. A DNA vaccine comprising an optimized nucleic acid according to claim 35 or 36.

51. Use of a modified nucleic acid sequence according to claim 35 or 36 for the production of a vaccine.

52. A method of expressing a protein or polypeptide, the method comprising:(a) providing a nucleic acid sequence encoding the protein or polypeptide;(b) generating a modified nucleic acid sequence according to any of claims 1-34;(c) synthesizing the modified nucleic acid sequence;(d) producing a vector comprising the modified nucleic acid sequence;(e) transfecting a host cell with the vector of step (d), wherein the host cell, subsequent to transfection, expresses the modified nucleic acid and produces the encoded protein or polypeptide product; and(f) optionally isolating and purifying the protein or polypeptide produced by the transfected host cell of step (e).

53. A computer program product comprising instructions encoded on a non-transitory computer-readable storage medium which, when executed by a computer, causes the computer to implement the methods of claims 1 to 24 and / or the in silica method of claims 25 to 34.

54. The computer program product of claim 53, wherein the instructions, when executed by the computer, causes the computer to render the modified nucleic acid sequence.

55. A computer system comprising means, for example, the computer program product of claims 53 or 54, that when executed generates and optionally synthesizes a modified nucleic acid sequence according to the methods of claims 1 to 24 and / or the in silica method of claims 25 to 34.

56. A device for generating a modified nucleic acid sequence, which is capable of being in operable connection either directly or over a network with a computer programmed with one or more algorithms, wherein the one or more algorithms are preferably in the same or different modules on the device, comprising:(a) means for providing a target nucleic acid sequence (Seqtarget) comprising an open reading frame, wherein Seqtarget has a length (Ntarget),(b) an algorithm for defining an organism-dependent splicing model comprising:(i) compiling a listing of sequence contexts (SeqCOn) comprising one or more intron / exon boundary sequences for one or more genes derived from a predefined organism genome (Genorg), each sequence in SeqCOn having a length (NCOn), wherein SeqCOn comprises a listing of putative splice donor (SD) sequence contexts (SeqCOnSD) and a listing of putative splice acceptor (SA) sequence contexts (SeqCOnSA),(ii) an algorithm for obtaining patterns (P) of length (NCOn) from SeqCOn,(iii) means for providing a first quality function capable of calculating a predictive splice score for each pattern P for determining whether such pattern P indicates the presence of a putative SD sequence motif or SA sequence motif of Genorgbased on:(1) the percentage of sequences within SeqCOnSD and / or SeqCOnSA which match the pattern P of length NCOn, and / or(2) the percentage at which pattern P matches a uniformly and randomly generated sequence of length NCOn,(iv) an algorithm for applying the first quality function to the pattern P obtained in (ii) to calculate their predictive splice scores;(c) a splice site identifying algorithm (At) for:(i) defining a maximum threshold splice score applicable for a respective Seqtarget of length N target,(ii) providing a second quality function capable of using as a combination input: a. a pattern P and its predictive splice score calculated in (b)(iv), b. optionally, the maximum threshold splice score defined in (c)(i), and c. an input nucleic acid sequence of length NCOn, and(iii) an algorithm for applying the splice site identifying algorithm Ai of (c)(ii) to define a third quality function capable of computing for each nucleic acid sequence of length (N) a predictive splice score, wherein the predictive splice score indicates the presence of a putative splice sequence motif;(d) a splice site modifying algorithm (AM) for:(i) selecting a sequence window (D) which contains at least one position of the open reading frame of Seqtarget,(ii) applying the third quality function defined in (c)( iii ) to sequence window D to calculate a predictive splice score,(iii) determining based on the predictive splice score whether window D comprises a putative splice sequence motif,(iv) optionally specifying one or more silent substitutions for sequence window D of Seqtarget capable of reducing the predictive splice score;(v) applying the one or more silent substitutions specified in (d)(iv) to window D to modify a putative splice sequence motif, and(vi) iteratively performing steps (d)(i) to (d)(v) for a finite series of sequence windows D in the open reading frame of Seqtarget for reducing the predictive splice score, to thereby generate the modified nucleic acid sequence.

57. The device of claim 56, wherein each of the subroutines defined by the corresponding algorithms of the device can exist as a separate algorithm element that can optionally be associated with the same device or on one or more separate devices.

58. The device of claim 56 or 57, wherein the device further comprises a machine learning module which is configured to determine a predictive splice score to be used by the splice site identifying algorithm Ai according to claim 1.

59. The device of any of claims 56 to 58, wherein the device further comprises an oligonucleotide synthesizer controlled by the computer for synthesizing the modified nucleic acid sequence.