Artificial nucleic acid molecule

By designing artificial nucleic acid molecules with highly efficient 5'-UTR and 3'-UTR elements, the problems of DNA insertion mutations and RNA instability have been solved, enabling stable expression and efficient translation in gene therapy and vaccines, thus improving the efficacy of treatment and immune responses.

CN114381470BActive Publication Date: 2026-08-25CUREVAC SE
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202210036661.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-08-28
Filing Date
2016-08-22
Publication Date
2026-08-25
Estimated Expiration
2036-08-22

AI Technical Summary

Technical Problem

In existing gene therapy and gene vaccination, DNA, as a nucleic acid molecule, carries the risk of insertion into the genome, leading to mutations and antibody production. RNA, on the other hand, is unstable and easily degraded, making it difficult to achieve stable expression and efficient translation.

Method used

Artificial nucleic acid molecules containing highly efficient 5'-UTR and 3'-UTR elements were designed, and the nucleotide sequence was optimized to increase the G/C content. The 5'-cap structure was combined to improve the stability and translation efficiency of the mRNA.

Benefits of technology

This technology enables the stable expression and efficient translation of artificial nucleic acid molecules in vivo for gene therapy and gene vaccination, reducing the risk of mutations and antibody production and improving the effectiveness of treatment and immune responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003468667140000311
    Figure BDA0003468667140000311
  • Figure BDA0003468667140000521
    Figure BDA0003468667140000521
  • Figure BDA0003468667140000522
    Figure BDA0003468667140000522
Patent Text Reader

Abstract

The present invention relates to artificial nucleic acid molecules comprising at least one open reading frame and at least one 3'-untranslated region element (3'-UTR element) and / or at least one 5'-untranslated region element (5'-UTR element), wherein said artificial nucleic acid molecule is characterized by a high translation efficiency. Said translation efficiency is at least partially due to said 5'-UTR element or said 3'-UTR element, or due to both said 5'-UTR element and said 3'-UTR element. The present invention further relates to the use of said artificial nucleic acid molecules in gene therapy and / or gene vaccination. Furthermore, novel 3'-UTR elements and 5'-UTR elements are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201680049973.1, filed on August 22, 2016, entitled "Artificial Nucleic Acid Molecule".

[0002] This invention relates to artificial nucleic acid molecules comprising a read frame, a 3'-UTR element and / or a 5'-UTR element, and optionally a poly(A) sequence and / or a polyadenylation signal. The invention also relates to vectors comprising 3'-UTR elements and / or 5'-UTR elements, cells comprising said artificial nucleic acid molecule or said vector, pharmaceutical compositions comprising said artificial nucleic acid molecule or said vector, and kits comprising said artificial nucleic acid molecule, said vector, and / or said pharmaceutical composition, preferably for use in gene therapy and / or gene vaccination.

[0003] Gene therapy and gene vaccination are among the most promising and rapidly developing methods in modern medicine. They can provide highly specific and personalized options for the treatment of a wide range of diseases. In particular, not only genetic diseases, but also autoimmune diseases, cancer or tumor-related diseases, and inflammatory diseases can be treated with these methods. Moreover, it is envisioned that these methods can prevent the early onset of such diseases.

[0004] The core concept behind gene therapy is the appropriate regulation of impaired gene expression associated with the pathological state of a specific disease. Pathologically altered gene expression can lead to the lack or overproduction of key gene products, such as signaling factors like hormones, housekeeping factors, metabolic enzymes, and structural proteins. Altered gene expression can result not only from misregulation of transcription and / or translation but also from mutations within the ORF encoding a specific protein. Pathological mutations can be caused by, for example, chromosomal aberrations or by more specific mutations such as point or frameshift mutations, all of which result in restricted function and potentially complete loss of function of the gene product. However, misregulation of transcription or translation can also occur if the mutation affects a gene encoding a protein involved in the cellular transcription or translation mechanisms. Such mutations can lead to pathological upregulation or downregulation of a gene that is functional itself. Genes encoding gene products that perform this regulatory function can be, for example, transcription factors, signal receptors, messenger proteins, etc. However, the loss of function of such genes encoding regulatory proteins can, in some cases, be reversed by artificially introducing other factors that further act downstream of the impaired gene product. This type of gene defect can also be compensated for through gene therapy that replaces the affected gene itself.

[0005] Genetic vaccination allows for the induction of the desired immune response against selected antigens, such as characteristic components of bacterial surfaces, viral particles, tumor antigens, etc. Generally speaking, vaccination is one of the key achievements of modern medicine. However, currently, only a limited number of diseases have effective vaccines available. Therefore, infections that cannot be prevented by vaccination still affect millions of people each year.

[0006] Vaccines are generally categorized into "first-generation," "second-generation," and "third-generation" vaccines. First-generation vaccines are typically whole-biofilm vaccines. They are based on live or attenuated or killed pathogens, such as viruses and bacteria. A major drawback of live and attenuated vaccines is the risk of reversion to life-threatening variants. Therefore, even when attenuated, the pathogen itself can still carry unpredictable risks. Killed pathogens may not generate a specific immune response as effectively as desired. To minimize these risks, "second-generation" vaccines were developed. These are typically subunit vaccines, which consist of defined antigenic or recombinant protein components derived from the pathogen.

[0007] Gene vaccines, also known as vaccines administered via gene delivery, are often understood as "third-generation" vaccines. They typically consist of genetically modified nucleic acid molecules that allow the expression of peptides or protein (antigen) fragments characteristic of pathogens or tumor antigens within the body. After being administered to a patient, the gene vaccine is expressed by target cells. The expression of the administered nucleic acid leads to the production of the encoded protein. If these proteins are recognized as foreign substances by the patient's immune system, an immune response is triggered.

[0008] As can be seen from the above, both gene therapy and gene vaccination are essentially based on administering nucleic acid molecules to a patient and subsequently transcribing and / or translating the encoded genetic information. Alternatively, gene vaccination or gene therapy may also include methods that involve isolating specific somatic cells from the patient to be treated, subsequently transfecting these cells in vitro, and then administering the treated cells back to the patient.

[0009] DNA and RNA can be used as nucleic acid molecules administered within the scope of gene therapy or gene vaccination. DNA is known to be relatively stable and easy to manipulate. However, the use of DNA carries the risk of unwanted insertion of the administered DNA fragment into the patient's genome, potentially leading to mutational events such as loss of gene function. As a further risk, undesirable anti-DNA antibodies may arise. Another drawback is the limited expression levels of encoded peptides or proteins that can be obtained after DNA administration, because the DNA must enter the nucleus to be transcribed, and the resulting mRNA can then be translated. Among other reasons, the expression level of the administered DNA will depend on the presence of specific transcription factors that regulate DNA transcription. In the absence of such factors, DNA transcription will not produce a satisfactory amount of RNA. Consequently, the level of translated peptides or proteins obtained is limited.

[0010] By using RNA to replace DNA in gene therapy or gene-based vaccines, the risks of unwanted genome integration and the generation of anti-DNA antibodies are minimized or avoided. However, RNA is considered a rather unstable molecule that can be easily degraded by ubiquitous RNases.

[0011] Typically, RNA degradation promotes regulation of RNA half-life. This effect is thought to be and demonstrated to be a fine-tuning of eukaryotic gene expression regulation (Friedel et al., 2009. Conserved principles of mammalian transcriptional regulation revealed by RNA half-life, Nucleic Acid Research 37(17):1-12). Therefore, each naturally occurring mRNA has its own half-life, dependent on the gene from which it originates and in which cell type it is expressed. This promotes regulation of the expression level of that gene. Unstable RNA is important for timely transient gene expression at different points. However, durable RNA may be associated with the accumulation of different proteins or the continuous expression of genes. In vivo, the half-life of mRNA can also depend on environmental factors, such as hormone treatment, as shown for example, insulin-like growth factor I, actin, and albumin mRNA (Johnson et al., Newly synthesized RNA: Simultaneous measurement in intact cells of transcription rates and RNA stability of insulin-like growth factor I, actin, and albumin in growth hormone-stimulated hepatocytes, Proc. Natl. Acad. Sci., Vol. 88, pp. 5287-5291, 1991).

[0012] For gene therapy and gene vaccines, stable RNA is typically required. This is partly due to the fact that products encoded by RNA sequences usually need to accumulate in the body. Furthermore, RNA must maintain its structural and functional integrity when prepared into suitable dosage forms during its storage and when administered. Therefore, efforts are being made to provide stable RNA molecules for gene therapy or gene vaccines that prevent them from undergoing premature degradation or decay.

[0013] It has been reported that the G / C content of nucleic acid molecules can affect their stability. Therefore, nucleic acids containing increased amounts of guanine (G) and / or cytosine (C) residues can be functionally more stable than those containing large amounts of adenine (A) and thymine (T) or uracil (U) nucleotides. In this regard, WO02 / 098443 provides a pharmaceutical composition containing mRNA stabilized by sequence modifications in the coding region. This sequence modification utilizes the degeneracy of the genetic code. Therefore, codons containing less favorable nucleotide combinations (less favorable in terms of RNA stability) can be replaced by alternative codons without altering the encoded amino acid sequence. This RNA stabilization method is limited by providing specific nucleotide sequences for individual RNA molecules that do not allow for the retention of desired amino acid sequence intervals. Furthermore, this method is limited to the coding region of the RNA.

[0014] As alternatives to mRNA stabilization, naturally occurring eukaryotic mRNA molecules have been found to contain specific stabilizing elements. These can be, for example, contained in the so-called untranslated region (UTR) at its 5′ end (5′-UTR) and / or at its 3′ end (3′-UTR), as well as other structural features such as a 5′-cap structure or a 3′-poly(A) tail. Both the 5′-UTR and 3′-UTR are typically transcribed from genomic DNA and are therefore elements of the premature mRNA. During mRNA processing, specific structural features of mature mRNA, such as the 5′-cap and 3′-poly(A) tail (also known as the poly(A) tail or poly(A) sequence), are typically added to the transcribed (premature) mRNA.

[0015] A 3′-polyadenylated tail is typically a monotonous adenosine nucleotide sequence added to the 3′ end of transcribed mRNA. It can contain up to approximately 400 adenosine nucleotides. The length of this 3′-polyadenylated tail has been found to be a potentially key factor in the stability of individual mRNAs.

[0016] Furthermore, it has been shown that the 3′-UTR of α-globin mRNA may be a well-known important factor in the stability of α-globin mRNA (Rodgers et al., Regulated α-globin mRNA decay is a cytoplasmic event proceeding through 3′-to-5′ exosome-dependent decapping, RNA, 8, pp. 1526-1537, 2002). The 3′-UTR of α-globin mRNA is clearly involved in the formation of a specific nucleoprotein complex (α-complex), and its presence is associated with the in vitro stability of mRNA (Wang et al., An mRNA stability complex functions with poly(A)-binding protein to stabilize mRNA in vitro, Molecular and Cellular Biology, Vol. 19, No. 7, July 1999, pp. 4552-4560).

[0017] Interesting regulatory functions have been further revealed in the UTRs of ribosomal protein mRNAs: while the 5'-UTR of ribosomal protein mRNA controls the translation of growth-related mRNAs, the strictness of this regulation is conferred by individual 3'-UTRs within the ribosomal protein mRNA (Ledda et al., Effect of the 3'-UTR length on the translational regulation of 5'-terminal oligopyrimidine mRNAs, Gene, Vol. 344, 2005, pp. 213-220). This mechanism promotes the specific expression of ribosomal proteins that are typically transcribed in a constant manner, thus some ribosomal protein mRNAs, such as ribosomal protein S9 or ribosomal protein L32, are referred to as housekeeping genes (Janovick-Guretzky et al., Housekeeping Gene Expression in Bovine Liver is Affected by Physiological State, Feed Intake, and Dietary Treatment, J. Dairy Sci., Vol. 90, 2007, pp. 2246-2252). The growth-related expression patterns of ribosomal proteins are therefore primarily due to the regulation of translation levels.

[0018] WO 2014 / 164253 A1 describes some specific nucleic acid molecules with 5'-UTR and / or 3'-UTR, but does not provide details on the translation efficiency of such molecules.

[0019] Regardless of factors affecting mRNA stability, the efficient translation of nucleic acid molecules delivered by target cells or tissues is crucial for any method using nucleic acid molecules for gene therapy or gene vaccination. As the examples cited above illustrate, along with stability regulation, the translation of most mRNAs is also regulated by structural features such as the UTR, 5′-cap, and 3′-polyadenylated tail. In this case, the length of the polyadenylated tail has been reported to play a significant role in translation efficiency. However, stabilizing the 3′ element may also have a detrimental effect on translation.

[0020] The object of this invention is to provide nucleic acid molecules suitable for use in gene therapy and / or gene vaccination. In particular, the object of this invention is to provide mRNA species that are stable against early degradation or decay without exhibiting significant functional loss in translation efficiency. Another object of this invention is to provide artificial nucleic acid molecules, preferably mRNAs, characterized by high translation efficiency. A particular object of this invention is to provide mRNAs in which the translation (ribosome production of the individual encoded proteins) efficiency is enhanced, for example, relative to a reference nucleic acid molecule (reference mRNA). Another object of this invention is to provide nucleic acid molecules encoding such superior mRNA species, which are suitable for use in gene therapy and / or gene vaccination. A further object of this invention is to provide pharmaceutical compositions for gene therapy and / or gene vaccination. In summary, the object of this invention is to provide improved nucleic acid species that overcome the deficiencies of the prior art discussed above through cost savings and a direct approach.

[0021] The fundamental objective of this invention is achieved through the claimed subject matter. Specifically, the inventors have identified UTR elements (5'-UTR elements and 3'-UTR elements) that provide high translation efficiency. They are all characterized by translation efficiency higher than that of previously known nucleic acid molecules, a common aspect of the preferred nucleic acid molecules of this invention. A method for determining high translation efficiency relative to a reference nucleic acid molecule is also provided.

[0022] This invention was supported by the U.S. government, and was awarded contract number HR0011-11-3-0001 by DARPA. The U.S. government owns certain rights to this invention.

[0023] For clarity and readability, the following definitions are provided. Any technical features mentioned in these definitions can be understood based on each embodiment of the invention. Other definitions and interpretations may be specifically provided within the scope of these embodiments.

[0024] Adaptive immune response: Adaptive immune responses are typically understood as antigen-specific responses of the immune system. Antigen specificity allows for the generation of responses that adapt to specific pathogens or pathogen-infected cells. The ability to initiate these adaptive responses is usually maintained in the body by "memory cells." If a pathogen infects the body more than once, these specific memory cells are used to quickly clear it. In this case, the first step in an adaptive immune response is to activate naïve cells capable of inducing antigen-specific immune responses through antigen-presenting cells. Antigen-specific T cells, or different immune cells, are cells that present antigens. This occurs in lymphoid tissues and organs through which naive T cells constantly pass. Three cell types that can serve as antigen-presenting cells are dendritic cells, macrophages, and B cells. Each of these cells has a different function in evoking an immune response. Dendritic cells can take up antigens through phagocytosis and pinocytosis and can be stimulated by contact with, for example, foreign antigens, thus migrating to local lymphoid tissues where they differentiate into mature dendritic cells. Macrophages take up particulate antigens such as bacteria and express MHC molecules through induction by infectious agents or other suitable stimuli. The unique ability of B cells to bind and internalize soluble protein antigens through their receptors can also be important for inducing T cells. MHC molecules are typically responsible for presenting antigens to T cells. Presentation of antigens to MHC molecules leads to T cell activation, which induces their proliferation and differentiation into armed effector T cells. The most important functions of effector T cells are to kill infected cells via CD8+ cytotoxic T cells, activate macrophages via Th1 cells (together constituting cell-mediated immunity), and activate B cells via Th2 and Th1 cells to produce different types of antibodies, thereby driving humoral immune responses. T cells recognize antigens through their T cell receptors. They do not directly recognize and bind antigens, but rather recognize short peptide fragments, such as fragments of protein antigens derived from pathogens, such as so-called epitopes, which bind to MHC molecules on the surface of other cells.

[0025] Adaptive immune system: The adaptive immune system is essentially dedicated to eliminating or preventing the growth of pathogens. It typically modulates the adaptive immune response by providing the vertebrate immune system with the ability to recognize and remember specific pathogens (to generate immunity) and to produce a stronger attack on each subsequent encounter with the pathogen. This system is highly adaptive due to somatic hypermutation (an accelerated process of somatic mutation) and V(D)J recombination (irreversible genetic recombination of segments of antigen receptor genes). This mechanism allows a small number of genes to produce a large number of different antigen receptors, which are then uniquely expressed on each individual's lymphocytes. Because gene rearrangements result in irreversible changes to the DNA of individual cells, all subsequent offspring of such cells will inherit genes encoding the same receptor-specific genes, including memory B cells and memory T cells, which are key to durable specific immunity.

[0026] Adjuvant / Adjuvant Components Adjuvants, or adjuvant components, are typically pharmaceutical and / or immunomodulatory agents that can modify (e.g., enhance) the effects of other agents (such as drugs or vaccines). The term will be interpreted broadly and refers to broad-spectrum substances. Typically, these substances can increase the immunogenicity of an antigen. For example, adjuvants can be recognized by the innate immune system and, for example, can elicit an innate immune response. "Adjuvants" typically do not elicit an adaptive immune response. To this extent, "adjuvants" are not considered antigens. Their mode of action differs from the effects elicited by antigens that cause an adaptive immune response.

[0027] antigen: In the context of this invention, "antigen" typically refers to a substance that can be recognized by the immune system, preferably the adaptive immune system, and is capable of inducing an antigen-specific immune response, for example, through the formation of antibodies and / or antigen-specific T cells as part of an adaptive immune response. Typically, an antigen can be or may comprise a peptide or protein that can be presented to T cells by the MHC. In the sense of this invention, an antigen can be the translation product of a provided nucleic acid molecule (preferably mRNA as defined herein). In this case, fragments, variants, and derivatives of peptides and proteins containing at least one epitope are also understood as antigens. In the context of this invention, tumor antigens and pathogen antigens as defined herein are particularly preferred.

[0028] In other words, an artificial nucleic acid molecule can be understood as a non-natural nucleic acid molecule. Such a nucleic acid molecule may be non-natural due to its individual sequence (which does not naturally exist) and / or due to other modifications (e.g., structural modifications of nucleotides that do not naturally exist). An artificial nucleic acid molecule can be a DNA molecule, an RNA molecule, or a hybrid molecule containing both DNA and RNA portions. Typically, artificial nucleic acid molecules can be designed and / or generated using genetic engineering methods to conform to a desired artificial nucleotide sequence (heterologous sequence). In this case, the artificial sequence is usually a sequence that may not naturally exist, i.e., it differs from the wild-type sequence by at least one nucleotide. The term "wild-type" can be understood as a naturally occurring sequence. When any particular "artificial nucleic acid molecule" is described herein as being "based on" any particular wild-type nucleic acid molecule, the artificial nucleic acid molecule differs from the wild-type artificial nucleic acid molecule by at least one nucleotide. Furthermore, the term "artificial nucleic acid molecule" is not limited to meaning "a single molecule," but is typically understood to encompass all identical molecules. Therefore, it can refer to multiple identical molecules contained in equal portions.

[0029] Bicistronic RNA, Polycistronic RNA Bicistronic or polycistronic RNA is typically RNA that can have two (bicistronic) or more (polycistronic) open reading frames (ORFs), preferably mRNA. In this case, the ORF is a codon sequence that can be translated into a peptide or protein.

[0030] Carrier / Polymerizing Carrier: The carriers in this invention are typically compounds that facilitate the transport and / or recombination of another compound (cargo). Polymerized carriers are typically carriers formed from polymers. The carrier can be associated with the cargo through covalent or non-covalent interactions. The carrier can transport nucleic acids (e.g., RNA or DNA) to target cells. In some embodiments, the carrier can be a cationic component.

[0031] cationic components The term "cationic component" typically refers to a charged molecule that is positively charged (cationic) at pH values ​​typically from 1 to 9, preferably or below 9 (e.g., from 5 to 9), from 8 to 8 (e.g., from 5 to 8), from 7 to 7 (e.g., from 5 to 7), and most preferably at physiological pH (e.g., from 7.3 to 7.4). Therefore, a cationic component can be any positively charged compound or polymer, preferably a cationic peptide or protein that is positively charged under physiological conditions, especially under in vivo physiological conditions. A "cationic peptide or protein" may contain at least one positively charged amino acid, or more than one positively charged amino acid, such as those selected from Arg, His, Lys, or Orn. ​​Therefore, a "polycationic" component also exists within a range exhibiting more than one positive charge under given conditions.

[0032] 5'- Cap: The 5′-cap is an entity, typically a modified nucleotide entity, that is usually “capped” at the 5′ end of mature mRNA. The 5′-cap can typically be formed from modified nucleotides, especially guanine nucleotide derivatives. Preferably, the 5′-cap is linked to the 5′ end via a 5′-5′-triphosphate bond. The 5′-cap can be methylated, for example, m7GpppN, where N is the terminal 5′ nucleotide of the nucleic acid carrying the 5′-cap, typically the 5′ end of RNA. Further examples of 5′ cap structures include glycerol (inverted deoxy-debasing residue (partial)), 4′,5′ methylene nucleotide, 1-(β-D-erythrofuranosyl) nucleotide, 4′-thionucleotide, carbocyclic nucleotide, 1,5-dehydrated hexitol nucleotide, L-nucleotide, α-nucleotide, modified base nucleotide, threo-pentafuranosyl nucleotide, acyclic 3′,4′-open nucleotide, acyclic 3,4-dihydroxybutyl nucleotide, acyclic 3,5-dihydroxypentayl nucleotide, 3′-3′-inverted nucleotide moiety, 3′-3′-inverted debasing moiety, 3′-2′-inverted nucleotide moiety, 3′-2′-inverted debasing moiety, 1,4-butanediol phosphate, 3′-aminophosphate, hexyl phosphate, aminohexyl phosphate, 3′-phosphate, 3′-thiophosphate, dithiophosphate, or bridged or non-bridged methyl phosphate moiety.

[0033] Cellular immunity / cellular immune response:Cellular immunity typically involves the activation of macrophages, natural killer (NK) cells, and antigen-specific cytotoxic T-lymphocytes, and the release of various cytokines in response to antigens. More generally, cellular immunity is not based on antibodies, but on the activation of immune system cells. Typically, a cellular immune response can be characterized, for example, by the activation of antigen-specific cytotoxic T-lymphocytes capable of inducing apoptosis in cells (e.g., specific immune cells such as dendritic cells or other cells), displaying epitopes of foreign antigens on their surface. Such cells can be virus-infected, intracellularly bacterial-infected, or cancer cells displaying tumor antigens. A further feature can be the activation of macrophages and natural killer cells, enabling them to destroy pathogens and stimulate the secretion of various cytokines that influence the function of other cells involved in adaptive and innate immune responses.

[0034] DNA: DNA is the common abbreviation for deoxyribonucleic acid. It is a nucleic acid molecule, a polymer composed of nucleotides. These nucleotides are typically monomers of deoxyadenosine monophosphate, deoxythymidine monophosphate, deoxyguanosine monophosphate, and deoxycytidine monophosphate, which themselves consist of a sugar moiety (deoxyribose), a base moiety, and a phosphate moiety, polymerized through a characteristic backbone structure. This backbone structure is typically formed by a phosphodiester bond between the sugar moiety (deoxyribose) of the nucleotide of the first adjacent monomer and the phosphate moiety of the second adjacent monomer. The specific sequence of the monomers, i.e., the sequence of bases linked to the sugar / phosphate backbone, is called the DNA sequence. DNA can be single-stranded or double-stranded. In the double-stranded form, the nucleotides of the first strand typically hybridize with the nucleotides of the second strand, for example, through A / T base pairing and G / C base pairing.

[0035] Epitope:Epitopes (also known as "antigenic determinants") can be distinguished between T-cell epitopes and B-cell epitopes. T-cell epitopes or portions of proteins within the scope of this invention may comprise fragments, preferably having a length of about 6 to about 20 or even more amino acids, such as fragments processed and presented by class I MHC molecules, preferably having a length of about 8 to about 10 amino acids, for example 8, 9, or 10 (or even 11 or 12 amino acids), or fragments processed and presented by class II MHC molecules, preferably having a length of about 13 or more amino acids, for example 13, 14, 15, 16, 17, 18, 19, 20 or even more amino acids, wherein these fragments may be selected from any portion of the amino acid sequence. These fragments are typically recognized by T cells in the form of a complex consisting of a peptide fragment and an MHC molecule, i.e., the fragments are typically not recognized in their native form. B-cell epitopes are typically fragments located on the outer surface of (natural) protein or peptide antigens as defined herein, preferably having 5 to 15 amino acids, more preferably 5 to 12 amino acids, and even more preferably 6 to 9 amino acids, which can be recognized by antibodies in their natural form.

[0036] Furthermore, such epitopes of proteins or peptides can be selected from any variant of such proteins or peptides mentioned herein. In this case, the antigenic determinant can be a conformational epitope or a discontinuous epitope consisting of protein or peptide fragments whose amino acid sequences are discontinuous but aggregated in a three-dimensional structure, as defined herein, or a continuous or linear epitope consisting of a single polypeptide chain.

[0037] Sequence fragment: Sequence fragments can typically be shorter portions of the full-length sequence of, for example, a nucleic acid molecule or an amino acid sequence. Therefore, typically, fragments consist of a sequence identical to a corresponding segment within the full-length sequence. Preferred sequence fragments within the scope of this invention consist of entities (e.g., nucleotides or amino acids) of a continuous segment corresponding to a continuous segment entity in the molecule from which the fragment originates, representing at least 5%, 10%, 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 60%, even more preferably at least 70%, and most preferably at least 80% of the total (i.e., full-length) molecule from which the fragment originates.

[0038] G / C modified:G / C-modified nucleic acids can typically be nucleic acids, preferably artificial nucleic acid molecules as defined herein, based on a modified wild-type sequence preferably containing an increased number of guanosine and / or cytosine nucleotides compared to the wild-type sequence. This increased number can be achieved by replacing a codon containing adenosine or thymidine nucleotides with a codon containing a guanosine or cytosine nucleotide. If the enriched G / C content occurs in the coding region of DNA or RNA, it utilizes the degeneracy of the genetic code. Therefore, the codon substitution preferably does not change the encoded amino acid residues but only increases the G / C content of the nucleic acid molecule.

[0039] Gene therapy: Gene therapy can typically be understood as treating a patient’s body or isolated components of a patient’s body, such as isolated tissues / cells, with nucleic acids encoding peptides or proteins. It typically includes at least one of the following steps: a) directly administering nucleic acids (preferably artificial nucleic acid molecules as defined herein) to the patient via any route of administration or in vitro to isolated cells / tissues of the patient, resulting in in vivo / in vitro or in vitro transfection of the patient’s cells; b) transcribing and / or translating the introduced nucleic acid molecules; and optionally c) if the nucleic acids are not directly administered to the patient, then re-administering the isolated, transfected cells to the patient.

[0040] Genetic vaccination: Genetic vaccination can typically be understood as vaccination by administering a nucleic acid molecule encoding an antigen or immunogen or a fragment thereof. The nucleic acid molecule can be administered to the subject's body or to isolated cells of the subject. When certain cells in the body or isolated cells are transfected, the antigen or immunogen can be expressed by those cells and subsequently presented to the immune system, eliciting an adaptive (i.e., antigen-specific) immune response. Therefore, genetic vaccination typically includes at least one of the following steps: a) administering a nucleic acid (preferably an artificial nucleic acid molecule as defined herein) to the subject, preferably a patient, or to isolated cells of the subject (preferably a patient), which generally results in in vivo or in vitro transfection of the subject's cells; b) transcribing and / or translating the introduced nucleic acid molecule; and optionally c) if the nucleic acid was not directly administered to the patient, then re-administering the isolated, transfected cells to the subject, preferably a patient.

[0041] Heterogeneous sequences: Two sequences are typically understood as 'heterologous' if they do not originate from the same gene. That is, although heterologous sequences may originate from the same organism, they do not naturally exist in the same nucleic acid molecules, such as in the same mRNA.

[0042] Humoral immunity / humoral immune responseHumoral immunity typically refers to antibody production and optionally to the additional processes that accompany antibody production. Humoral immune responses can be typically characterized by, for example, Th2 activation and cytokine production, germinal center formation and allotype switching, affinity maturation, and memory cell production. Humoral immunity can also typically refer to the effector functions of antibodies, including pathogen and toxin neutralization, classical complement activation, and opsonin-mediated phagocytosis and pathogen clearance.

[0043] Immunogen: In Within the scope of this invention, an immunogen can typically be understood as a compound capable of stimulating an immune response. Preferably, an immunogen is a peptide, polypeptide, or protein. In a particularly preferred embodiment, for the purposes of this invention, an immunogen is the translational product of a provided nucleic acid molecule (preferably an artificial nucleic acid molecule as defined herein). Typically, an immunogen elicits at least an adaptive immune response.

[0044] Immunostimulatory composition: Within the scope of this invention, an immunostimulatory composition can be typically understood as a composition containing at least one component capable of inducing an immune response or a component from which the component capable of inducing an immune response is derived. Such an immune response may preferably be an innate immune response or a combination of adaptive and innate immune responses. Preferably, the immunostimulatory composition within the scope of this invention contains at least one artificial nucleic acid molecule, more preferably RNA, such as an mRNA molecule. The immunostimulatory component, such as mRNA, may be complexed with a suitable carrier. Therefore, the immunostimulatory composition may comprise an mRNA / carrier complex. Furthermore, the immunostimulatory composition may comprise an adjuvant and / or a suitable carrier for the immunostimulatory component (e.g., mRNA).

[0045] Immune response Immune responses can typically be specific responses of the adaptive immune system to a specific antigen (so-called specific or adaptive immune responses) or non-specific responses of the innate immune system (so-called non-specific or innate immune responses), or a combination thereof.

[0046] Immune system: The immune system protects an organism from infection. If a pathogen successfully crosses the organism's physical barriers and enters the organism, the innate immune system provides an immediate, but nonspecific, response. If the pathogen evades this innate response, vertebrates possess a second layer of protection: the adaptive immune system. Here, the immune system alters its response during infection to improve its recognition of the pathogen. After the pathogen is eliminated, this improved response is then retained as immune memory, allowing the adaptive immune system to launch a faster and stronger attack each time it encounters the pathogen. Accordingly, the immune system comprises both the innate and adaptive immune systems. Each of these typically contains what are called humoral and cellular components.

[0047] Immune stimulating RNA: Within the scope of this invention, immunostimulatory RNA (isRNA) can typically be RNA capable of inducing an innate immune response. It generally does not have a read frame and therefore does not provide peptide-antigens or immunogens, but evokes an immune response, for example, by binding to a specific type of Toll-like receptor (TLR) or other suitable receptor. However, of course, mRNAs with a read frame and encoding peptides / proteins can also induce an innate immune response and, therefore, can be immunostimulatory RNAs.

[0048] Innate immune system:The innate immune system, also known as the non-specific immune system, typically comprises cells and mechanisms that protect the host from infection by other organisms in a non-specific manner. This means that the cells of the innate system can recognize and respond to pathogens in a general way, but unlike the adaptive immune system, it does not confer durable or protective immunity to the host. The innate immune system can be activated, for example, by: Toll-like receptor (TLR) ligands or other helpers such as lipopolysaccharide, TNF-α, CD40 ligands, or cytokines, monokines, lymphokines, interleukins or chemokines, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, IL-10, IL-11, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, IL-19, IL-20, IL-21, IL-22, IL-23, IL-24, IL-25, IL-26, IL-27, IL-28, IL-29, IL-30, IL-31, IL-32, I... L-33, IFN-α, IFN-β, IFN-γ, GM-CSF, G-CSF, M-CSF, LT-β, ​​TNF-α, growth factors, and hGH, ligands of human Toll-like receptors TLR1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, and TLR10, ligands of mouse Toll-like receptors TLR1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TLR10, TLR11, TLR12, or TLR13, ligands of NOD-like receptors, ligands of RIG-I-like receptors, immunostimulatory nucleic acids, immunostimulatory RNA (isRNA), CpG-DNA, antibacterial agents, or antiviral agents. The pharmaceutical compositions of this invention may contain one or more of these substances. Typically, the response of the innate immune system includes: recruiting immune cells to the site of infection by producing chemical factors, including specialized chemical mediators called cytokines; activating the complement cascade; identifying and removing foreign substances from organs, tissues, blood, and lymph by specialized leukocytes; activating the adaptive immune system; and / or acting as physical and chemical barriers against infectious agents.

[0049] Cloning site:A cloning site is typically understood as a nucleic acid molecular fragment suitable for insertion into a nucleic acid sequence (e.g., a nucleic acid sequence containing a read frame). Insertion can be performed by any molecular biology method known to those skilled in the art, such as restriction enzyme digestion and ligation. A cloning site typically contains one or more restriction enzyme recognition sites (restriction sites). These one or more restriction sites can be recognized by restriction enzymes that cleave DNA at these sites. Cloning sites containing one or more restriction sites may also be called multiple cloning sites (MCS) or multiple adapters.

[0050] Nucleic acid molecules Nucleic acid molecules are molecules containing nucleic acid components, preferably molecules composed of nucleic acid components. The term nucleic acid molecule preferably refers to DNA or RNA molecules. It is preferably used synonymously with the term "polynucleotide". Preferably, a nucleic acid molecule is a polymer comprising or composed of nucleotide monomers covalently linked together by phosphodiester bonds of a sugar / phosphate ester backbone. The term "nucleic acid molecule" also includes modified nucleic acid molecules, such as base-modified, sugar-modified, or backbone-modified DNA or RNA molecules.

[0051] Readable box: The open reading frame (ORF) within the scope of this invention can typically be a sequence of nucleotide triplets that can be translated into a peptide or protein. The ORF preferably contains a start codon at its 5′ end, which is typically a combination of three consecutive nucleotides encoding the amino acid methionine (ATG), and usually presents as a contiguous region of multiple 3 nucleotides in length. Typically, this is the only stop codon of the ORF. Therefore, the ORF within the scope of this invention is preferably a nucleotide sequence consisting of a number of nucleotides divisible by three, starting with a start codon (e.g., ATG) and preferably ending with a stop codon (e.g., TAA, TGA, or TAG). The ORF can be isolated or can be integrated into a longer nucleic acid sequence, such as a vector or mRNA. The ORF can also be referred to as a “protein-coding region”.

[0052] peptides: Peptides, or polypeptides, are typically polymers of amino acid monomers linked by peptide bonds. They typically contain fewer than 50 monomer units. Nevertheless, the term peptide does not preclude the existence of molecules with more than 50 monomer units. Long peptides, also known as polypeptides, typically contain 50 to 600 monomer units.

[0053] Pharmaceutical effective dose The pharmaceutically effective amount within the scope of this invention is typically understood as, for example, an amount sufficient to induce a pharmaceutical effect (such as an immune response) under pathological conditions, alter the pathological level of expressed peptides or proteins, or replace the amount of missing gene products.

[0054] protein:Proteins typically contain more than one peptide or polypeptide. Proteins are typically folded into a 3-dimensional form, which may be required for the protein to perform its biological functions.

[0055] Poly(A) sequence A polyadenylated sequence, also known as a polyadenylated tail or 3′-polyadenylated tail, is typically understood as an adenosine nucleotide sequence of, for example, up to about 400 adenosine nucleotides, such as from about 20 to about 400, preferably from about 50 to about 400, more preferably from about 50 to about 300, even more preferably from about 50 to about 250, and most preferably from about 60 to about 250 adenosine nucleotides. The polyadenylated sequence is typically located at the 3′ end of mRNA. Within the scope of this invention, the polyadenylated sequence can be located, for example, within mRNA via transcription by a vector or within any other nucleic acid molecule, such as, for example, within a vector, for example, within a vector used as a template for generating RNA, preferably in mRNA.

[0056] Polyadenylation Polyadenylation is typically understood as the addition of polyadenylated sequences to nucleic acid molecules, such as RNA molecules, for example, pre-mature mRNA. Polyadenylation can be caused by so-called... Polyadenylation signaling Induction. This signal is preferably located within a nucleotide segment at the 3' end of the nucleic acid molecule to be polyadenylated, such as an RNA molecule. The polyadenylation signal typically comprises a hexamer of adenine and uracil / thymine nucleotides, preferably the hexamer sequence AAUAAA. Other sequences, preferably hexamer sequences, are also possible. Polyadenylation typically occurs during the processing of pre-mRNA (also known as pre-mRNA). Typically, RNA maturation (from pre-mRNA to mature mRNA) involves the polyadenylation step.

[0057] Restrictive sites: A restriction site, also known as a 'restriction enzyme recognition site', is a nucleotide sequence recognized by a restriction enzyme. Restriction sites are typically short, preferably palindromic nucleotide sequences, such as sequences containing 4 to 8 nucleotides. Restriction sites are preferably specifically recognized by restriction enzymes. Restriction enzymes typically cleave the nucleotide sequence containing the restriction site. In double-stranded nucleotide sequences, such as double-stranded DNA sequences, restriction enzymes typically cleave both strands of the nucleotide sequence.

[0058] RNA, mRNA:RNA is the common abbreviation for ribonucleic acid. It is a nucleic acid molecule, a polymer composed of nucleotides. These nucleotides are typically adenosine monophosphate, uridine monophosphate, guanosine monophosphate, and cytidine monophosphate monomers linked together along a so-called backbone. The backbone is formed by phosphodiester bonds between the sugar, ribose, of the first adjacent monomer and the phosphate moiety of the second adjacent monomer. The specific sequence of monomers is called the RNA sequence. RNA is usually obtained by transcription, for example, of a DNA sequence within a cell. In eukaryotic cells, transcription typically takes place in the nucleus or mitochondria. Typically, transcription of DNA usually produces so-called pre-mature RNA, which must be processed into so-called messenger RNA, usually abbreviated as mRNA. For example, in eukaryotes, the processing of pre-mature RNA involves a variety of different post-transcriptional modifications, such as splicing, 5′-capping, polyadenylation, and export from the nucleus or mitochondria. The sum of these processes is also called RNA maturation. Mature messenger RNA typically provides a nucleotide sequence that can be translated into a specific peptide or protein amino acid sequence. Typically, mature mRNA contains a 5′-cap, a 5′-UTR, a reading frame, a 3′-UTR, and a polyadenylated sequence. In addition to messenger RNA, there are several non-coding RNA types that can participate in regulating transcription and / or translation.

[0059] Nucleic acid molecular sequence: A nucleic acid molecule sequence is typically understood as a specific and unique sequence, that is, a sequence of its nucleotides. A protein or peptide sequence is typically understood as that sequence, that is, a sequence of its amino acids.

[0060] Sequence identity: Two or more sequences are considered identical if they consist of nucleotides or amino acids of the same length and sequence. The percentage of identity typically describes the degree to which two sequences are identical, that is, it typically describes the percentage of nucleotides corresponding to nucleotides at the same sequence position as a reference sequence. To determine the degree of identity, the sequences to be compared are considered to be of the same length, i.e., the longest sequence length of the sequences to be compared. This means that a first sequence consisting of 8 nucleotides is 80% identical to a second sequence consisting of 10 nucleotides containing the first sequence. In other words, within the scope of this invention, sequence identity preferably refers to the percentage of nucleotides at the same position in two or more sequences of the same length. Gaps are generally considered to be positions of dissimilarity, regardless of their actual position in the alignment.

[0061] Stable nucleic acid molecules:Stable nucleic acid molecules are those modified to be more stable than unmodified nucleic acid molecules due to degradation or breakdown caused by environmental factors or enzymatic digestion, such as degradation by exonucleases or endonucleases. Preferably, DNA or RNA molecules. Preferably, within the scope of this invention, stable nucleic acid molecules are stable in cells, such as prokaryotic or eukaryotic cells, particularly mammalian cells, such as human cells. The stabilizing effect can also occur extracellularly, for example in buffer solutions, or during the manufacture of pharmaceutical compositions containing said stabilized nucleic acid molecules.

[0062] Transfection: The term "transfection" refers to the introduction of nucleic acid molecules, such as DNA or RNA (e.g., mRNA), into cells, preferably eukaryotic cells. Within the scope of this invention, the term "transfection" includes any method known to those skilled in the art for introducing nucleic acid molecules into cells, preferably eukaryotic cells, such as mammalian cells. Such methods include, for example, electroporation, lipid transfection based on cationic lipids and / or liposomes, calcium phosphate precipitation, nanoparticle-based transfection, virus-based transfection, or transfection based on cationic polymers (such as DEAE-glucan or polyethyleneimine), etc. Preferably, the introduction is non-viral.

[0063] Translation efficiency (or translation efficiency): When used herein, this term generally refers to nucleic acid molecules (e.g., mRNA) that contain a readability frame (ORF). Translation efficiency is experimentally measurable. Translation efficiency is typically measured by determining the amount of protein translated by the ORF. For experimental measurements of translation efficiency, the ORF preferably encodes a reporter protein or any other quantifiable protein. However, it is not intended to be limited to specific theories, and it should be understood that high translation efficiency is generally provided by specific UTR elements (specific 5'-UTR elements or specific 3'-UTR elements). Thus, in the context of this invention, the term translation efficiency is specifically used for nucleic acid molecules that, in addition to the ORF, contain at least one 5'-UTR element and / or at least one 3'-UTR element, preferably as defined herein. Although the ORF suitably encodes a reporter protein or any other quantifiable protein for experimental quantification of translation efficiency, the invention is not limited to such purposes; therefore, at least one 5'-UTR element and / or at least one 3'-UTR element (which provides high translation efficiency) of the present invention can be included in nucleic acid molecules containing an ORF that does not encode a reporter protein.

[0064] Translation efficiency is a relative term. Therefore, the translation efficiencies of multiple (e.g., two or more) nucleic acid molecules can be determined and compared, for example, by experimentally quantifying proteins encoded by ORFs. This can be performed under standard conditions, for example, in the standard assays described herein. Therefore, the translation efficiency determined under standard conditions is an objective characteristic. Preferably, the translation efficiencies of two nucleic acid molecules are determined and compared. The first of the two nucleic acid molecules may be referred to as the "test nucleic acid molecule" or "test construct," and the second of the two nucleic acid molecules may be referred to as the "reference nucleic acid molecule" or "reference construct." The test nucleic acid molecule may be an artificial nucleic acid molecule as described in this invention. For this purpose, the reference nucleic acid molecule and the test nucleic acid molecule share the same ORF (identical nucleic acid sequence); and preferably, the nucleic acid sequence of the test nucleic acid molecule is the same as that of the reference nucleic acid molecule, the difference being the UTR element being tested, i.e., the 5'-UTR element or the 3'-UTR element; in other words, preferably, the test nucleic acid molecule and the reference nucleic acid molecule differ from each other only in that the 5'-UTR element or the 3'-UTR element has a different nucleic acid sequence; making the 5'-UTR element or the 3'-UTR element the only structural feature that distinguishes the test nucleic acid molecule from the reference nucleic acid molecule.

[0065] In the assays described herein, test nucleic acid molecules and reference nucleic acid molecules are transfected into mammalian cells, and translation efficiency is determined. A test nucleic acid molecule is considered to be characterized by high translation efficiency when its translation efficiency is higher than that of the reference nucleic acid molecule. In this case, the 5'-UTR element or 3'-UTR element (for the test nucleic acid molecule) that distinguishes it from the reference nucleic acid molecule is considered to provide high translation efficiency. In other words, high translation efficiency is considered to be "provided" by specific 5'-UTR elements or 3'-UTR elements present in the test nucleic acid molecule but not in the reference nucleic acid molecule; and a test nucleic acid molecule (or artificial nucleic acid molecule) containing such a 5'-UTR or 3'-UTR and at least one ORF is considered to be "characterized" by high translation efficiency.

[0066] vaccine: A vaccine is typically understood as a prophylactic or therapeutic substance that provides at least one antigen, preferably an immunogen. The antigen or immunogen can be derived from any substance suitable for vaccination. For example, the antigen or immunogen can be derived from pathogens, such as bacterial or viral particles, or from tumor or cancerous tissue. The antigen or immunogen stimulates the body's adaptive immune system to provide an adaptive immune response.

[0067] Carrier:The term "vector" refers to a nucleic acid molecule, preferably an artificial nucleic acid molecule. Within the scope of this invention, a vector is suitable for binding or carrying a desired nucleic acid sequence, such as a nucleic acid sequence containing a read frame. Such vectors can be storage vectors, expression vectors, cloning vectors, transfer vectors, etc. A storage vector is a vector that allows convenient storage of nucleic acid molecules (e.g., mRNA molecules). Therefore, the vector may contain a sequence corresponding to, for example, a desired mRNA sequence or a portion thereof, such as a sequence corresponding to the read frame and the 3′UTR and / or 5′UTR of the mRNA. An expression vector can be used to produce an expression product, such as RNA (e.g., mRNA), or a peptide, polypeptide, or protein. For example, an expression vector may contain a sequence required for transcription of a segment of the vector, such as a promoter sequence, for example, an RNA polymerase promoter sequence. A cloning vector is typically a vector containing a cloning site that can be used to incorporate a nucleic acid sequence into the vector. A cloning vector can be, for example, a plasmid vector or a phage vector. A transfer vector can be a vector suitable for transferring nucleic acid molecules into cells or organisms, for example, a viral vector. Within the scope of this invention, the vector can be, for example, an RNA vector or a DNA vector. Preferably, the vector is a DNA molecule. Preferably, in the sense of this application, the vector includes a cloning site, a selection marker (such as an antibiotic resistance factor), and a sequence suitable for vector amplification, such as an origin of replication. Preferably, the vector within the scope of this application is a plasmid vector.

[0068] Excipients (vehicle): Excipients are typically understood as materials suitable for the storage, transport, and / or administration of compounds, such as pharmaceutically active compounds. For example, they can be physiologically acceptable liquids suitable for the storage, transport, and / or administration of pharmaceutically active compounds.

[0069] 3'-Untranslated Region (3'-UTR):Typically, the term "3'-UTR" refers to a portion of an artificial nucleic acid molecule located at the 3' (i.e., "downstream") of the open reading frame (ORF) and not translated into a protein. Generally, the 3'-UTR is a portion of mRNA between the protein-coding region (ORF or coding sequence (CDS)) and the polyadenylated nucleotide sequence. In the context of this invention, the term 3'-UTR may also include elements that are not encoded in the template, are transcribed from the RNA, but are added post-transcriptionally during maturation, such as polyadenylated nucleotide sequences. The 3'-UTR of mRNA is not translated into an amino acid sequence. The 3'-UTR sequence is typically encoded by a gene that is transcribed into its respective mRNA during gene expression. The genome sequence is first transcribed into pre-mature mRNA containing optional introns. The pre-mature mRNA is then further processed into mature mRNA during maturation. The maturation process includes the following steps: 5′ capping, splicing of pre-mature mRNA to remove optional introns and 3′ end modifications (such as polyadenylation of the 3′ end of the pre-mature mRNA and optional endonuclease / or exonuclease cleavage). Within the scope of this invention, a 3′-UTR corresponds to a mature mRNA sequence located between a protein-coding stop codon (preferably immediately following the 3′ end of the protein-coding stop codon) and the polyadenylated sequence of the mRNA. The term "corresponds to" means that the 3′-UTR sequence can be an RNA sequence as defined in the mRNA sequence used to define the 3′-UTR sequence, or a DNA sequence corresponding to this RNA sequence. Within the scope of this invention, the term "gene's 3′-UTR" refers to the sequence corresponding to the 3′-UTR of a mature mRNA derived from that gene, i.e., mRNA obtained through gene transcription and maturation of pre-mature mRNA. The term "gene's 3′-UTR" includes both the DNA and RNA sequences of the 3′-UTR (both sense and antisense strands, and both mature and immature strands). Preferably, the 3′UTR has a length of more than 20, 30, 40 or 50 nucleotides.

[0070] 5'-Untranslated Region (5'-UTR):Generally, the term "5'-UTR" refers to a portion of an artificial nucleic acid molecule located at the 5' (i.e., "upstream") of the reading frame and not translated into a protein. A 5'-UTR is generally understood as a specific segment of messenger RNA (mRNA) located at the 5' of the reading frame of the mRNA. Typically, a 5'-UTR begins at a transcription start site and terminates one nucleotide before the start codon of the reading frame. Preferably, the 5'-UTR has a length of more than 20, 30, 40, or 50 nucleotides. A 5'-UTR may contain elements for controlling gene expression, also known as regulatory elements. These regulatory elements may be, for example, ribosome binding sites. A 5'-UTR can be post-transcribed, for example, by adding a 5'-cap. The 5'-UTR of mRNA is not translated into an amino acid sequence. The 5'-UTR sequence is typically encoded by a gene transcribed into individual mRNAs during gene expression. The genomic sequence is first transcribed into pre-mature mRNA, which contains optional introns. The pre-mature mRNA is then further processed into mature mRNA during maturation. The maturation process includes the following steps: 5′ capping, splicing of pre-mature mRNA to remove optional introns and 3′ end modifications (such as polyadenylation of the 3′ end of the pre-mature mRNA and optional endonuclease / exonuclease cleavage, etc.). Within the scope of this invention, the 5′-UTR corresponds to the mature mRNA sequence located between the start codon and, for example, the 5′-cap. Preferably, the 5′-UTR corresponds to a sequence extending from the nucleotide on the 3′ side of the 5′ cap, more preferably from the nucleotide immediately adjacent to the 5′ cap, toward the nucleotide on the 5′ side of the start codon in the protein-coding region, and more preferably toward the nucleotide immediately adjacent to the 5′ side of the start codon in the protein-coding region. The nucleotide immediately adjacent to the 3′ side of the 5′ cap of the mature mRNA typically corresponds to the transcription start site. The term "corresponds to" means that the 5′-UTR sequence can be an RNA sequence as defined in the mRNA sequence used to define the 5′-UTR sequence, or a DNA sequence corresponding to this RNA sequence. Within the scope of this invention, the term "5′-UTR of a gene" refers to the sequence corresponding to the 5′-UTR of a mature mRNA derived from that gene, i.e., mRNA obtained through gene transcription and maturation of pre-mature mRNA. The term "5′-UTR of a gene" includes both the DNA and RNA sequences of the 5′-UTR (both the sense and antisense strands, and both the mature and immature strands).

[0071] 5′ terminal oligopyrimidine bundle (TOP):A 5′ terminal oligopyrimidine bundle (TOP) is typically a pyrimidine nucleotide located in the 5′ terminal region of a nucleic acid molecule, such as the 5′ terminal region of some mRNA molecules or the 5′ terminal region of a functional entity, such as the transcriptome of some genes. This sequence begins with a cytidine, which usually corresponds to the transcription start site, and is followed by a pyrimidine nucleotide bundle typically of about 3 to 30. For example, a TOP can contain 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or even more nucleotides. The pyrimidine sequence and the resulting 5′ TOP terminate at a nucleotide on the 5′ side of the first purine nucleotide downstream of the TOP. Messenger RNA containing a 5′ terminal oligopyrimidine bundle is often referred to as TOP mRNA. Therefore, the gene providing this messenger RNA is called a TOP gene. For example, TOP sequences have been found in genes and mRNAs encoding peptide elongation factors and ribosomal proteins.

[0072] TOP motif: Within the scope of this invention, the TOP motif is a nucleic acid sequence corresponding to the 5′ TOP as defined above. Therefore, the TOP motif within the scope of this invention is preferably a pyrimidine nucleotide segment having a length of 3-30 nucleotides. Preferably, the TOP motif consists of at least 3 pyrimidine nucleotides, more preferably at least 4 pyrimidine nucleotides, more preferably at least 5 pyrimidine nucleotides, more preferably at least 6 nucleotides, more preferably at least 7 nucleotides, and most preferably at least 8 pyrimidine nucleotides, wherein the pyrimidine nucleotide segment preferably begins with a cytosine nucleotide at its 5′ end. In TOP genes and TOP mRNA, the TOP motif preferably begins at a transcription start site at its 5′ end and terminates at a nucleotide 5′ to the first purine residue in the gene or mRNA. The TOP motif in the context of this invention is preferably located at the 5′ end of a sequence representing a 5′-UTR or at the 5′ end of a sequence encoding a 5′-UTR. Therefore, preferably, in the sense of the present invention, a sequence of three or more pyrimidine nucleotides is referred to as a "TOP motif" if the sequence is located at the 5' end of its respective sequence (such as an artificial nucleic acid molecule, a 5'-UTR element of an artificial nucleic acid molecule, or a nucleic acid sequence derived from the 5'-UTR of the TOP gene as described herein). In other words, a sequence of three or more pyrimidine nucleotides located not at the 5' end of the 5'-UTR or 5'-UTR element, but at any position within the 5'-UTR or 5'-UTR element, is preferably not referred to as a "TOP motif".

[0073] TOP genes:TOP genes are typically characterized by the presence of a 5′-terminal oligopyrimidine bundle. Furthermore, most TOP genes are characterized by growth-related translational regulation. However, tissue-specific translational regulation of TOP genes is also known. As defined above, the 5′-UTR of a TOP gene corresponds to the 5′-UTR sequence of the mature mRNA derived from the TOP gene, which preferably extends from the nucleotides on the 3′ side of the 5′-cap to the nucleotides on the 5′ side of the start codon. The 5′-UTR of a TOP gene typically does not contain any start codon and preferably has no upstream AUG (uAUG) or upstream open reading frame (uORF). Herein, upstream AUG and upstream open reading frame are typically understood as the AUG and open reading frame present on the 5′ side of the start codon (AUG) of the reading frame that should be translated. The 5′-UTR of a TOP gene is generally quite short. The length of the 5′-UTR of a TOP gene can vary from 20 nucleotides to up to 500 nucleotides, and is typically less than about 200 nucleotides, preferably less than about 150 nucleotides, and more preferably less than about 100 nucleotides. In the context of this invention, a typical 5'-UTR of the TOP gene is a nucleic acid sequence extending from the nucleotide at position 5 to the nucleotide immediately following the 5' side of the start codon (e.g., ATG) in SEQ ID Nos. 1-1363 of patent application WO2013 / 143700, the disclosure of which is incorporated herein by reference. In this case, a particularly preferred fragment of the 5'-UTR of the TOP gene is the 5'-UTR of the TOP gene lacking the 5' TOP motif. The term "5'-UTR of the TOP gene" or "5'-TOP UTR" preferably refers to the 5'-UTR of the naturally occurring TOP gene. A preferred example is shown in SEQ ID NO: 208 (the 5'-UTR of human macroribosomal protein 32 lacking the 5' terminal oligopyrimidine bundle); which corresponds to SEQ ID NO. 1368 of patent application WO2013 / 143700.

[0074] wild typeFor example, wild-type nucleic acid molecules: The term "wild-type" can be understood as a naturally occurring sequence. Wild-type nucleic acid molecules can generally be understood as naturally occurring nucleic acid molecules, such as DNA or RNA. In other words, artificial nucleic acid molecules can be understood as natural nucleic acid molecules. Such nucleic acid molecules may be natural due to their respective sequences (which are naturally occurring) and / or due to other naturally occurring modifications (e.g., structural modifications of nucleotides). Wild-type nucleic acid molecules can be DNA molecules, RNA molecules, or hybrid molecules containing both DNA and RNA portions. Here, the term "wild-type" means any molecule that is naturally occurring and does not need to be reflected in a publicly accessible sequence library (such as GenBank). The National Institutes of Health (NIH) provides a publicly accessible library of annotated, publicly available nucleotide sequences ("GenBank", accessible via the NCBI Entrez search system: http: / / www.ncbi.nlm.nih.gov, (Nucleic Acids Research, 2013; 41(D1):D36-42), which includes publicly available wild-type sequences. Each GenBank record is assigned a unique, immutable identifier called an accession number, which appears on the accession line of the GenBank record; and changes in sequence information are traced by the integer continuation of the accession number appearing on the version line of the GenBank record. Furthermore, the term "wild-type nucleic acid molecule" is not limited to referring to "a single molecule," but is typically understood to encompass all identical molecules. Therefore, it can refer to multiple identical molecules contained in an equal subset.

[0075] Detailed Explanation

[0076] This invention relates to artificial nucleic acid molecules comprising at least one read frame and at least one 3'-untranslated region element (3'-UTR element) and / or at least one 5'-untranslated region element (5'-UTR element), wherein the artificial nucleic acid molecules are characterized by high translation efficiency. The translation efficiency is at least partially attributable to the 5'-UTR element or the 3'-UTR element, or to both the 5'-UTR element and the 3'-UTR element. The invention also relates to the use of such artificial nucleic acid molecules in gene therapy and / or gene vaccination. Furthermore, novel 3'-UTR elements and 5'-UTR elements are provided.

[0077] In a first aspect, the present invention relates to an artificial nucleic acid molecule comprising:

[0078] a. At least one readable bounding box (ORF); and

[0079] b. At least one 3'-untranslated region element (3'-UTR element) and / or at least one 5'-untranslated region element (5'-UTR element), wherein the artificial nucleic acid molecule is characterized by high translation efficiency.

[0080] For example, the artificial nucleic acid molecule of the present invention may differ from other (e.g., wild-type or artificial) nucleic acid molecules in that at least one 3'-UTR element (preferably one 3'-UTR element) is replaced by at least one different 3'-UTR element (preferably one 3'-UTR element), or at least one 5'-UTR element (preferably one 5'-UTR element) is replaced by at least one different 5'-UTR element (preferably one 5'-UTR element).

[0081] Typically, replacing one (or at least one) 3'-UTR element or one (or at least one) 5'-UTR element (in the starting or reference nucleic acid molecule) generally means that the sequence of at least some nucleotides of the 3'-UTR element or the 5'-UTR element is replaced by a different nucleic acid sequence. Preferably, when replacing a 5'-UTR element, the sequential nucleotide sequence of the 5'-UTR element is replaced by a different sequential nucleotide sequence (element); and when replacing a 3'-UTR element, the sequential nucleotide sequence of the 3'-UTR element is replaced by a different sequential nucleotide sequence (element). The sequential nucleotide sequences of the initial (replaced) element and the sequential nucleotide sequences of the replaced element can each be of any length independently, such as 1 to greater than 500 nucleotides, 20 to 500 nucleotides, 40 to 400 nucleotides, 60 to 300 nucleotides, 80 to 200 nucleotides, or 100 to 150 nucleotides. When replacing a (5'- or 3'-)UTR element, this does not necessarily mean replacing all nucleotides on the 5' side of the start codon or all nucleotides on the 3' side of the stop codon. In a preferred embodiment, a consecutive sequence is replaced, which is a subset of all nucleotides on the 5' side of the start codon and all nucleotides on the 3' side of the stop codon, respectively. Examples are shown... Figure 1BIn (5') and 1C(3'), known 5'-UTR elements or known 3'-UTR elements present in the starting or reference sequence can be precisely replaced. For example, when (e.g., the reference) nucleic acid contains a 5'-UTR element corresponding to the sequence of SEQ ID NO:208; the artificial nucleic acid molecule of the present invention can be prepared by precisely replacing the entire continuous sequence of SEQ ID NO:208 with different consecutive 5'-UTR elements (i.e., 5'-UTR elements that provide high translation efficiency) as described in the present invention. Alternatively, it is possible that known UTR elements are not precisely replaced, for example, a continuous nucleic acid sequence containing additional nucleotides is not precisely replaced, i.e., in addition to the replacement of known UTR elements, or, for example, only a portion of the known UTR elements (typically the major portion, i.e., more than 90% of the length) is replaced. Although these possibilities have been exemplified for the case of 5'-UTR elements, the same possibilities exist for 3'-UTR elements. In other words, the replaced UTR element does not need to be limited by the precise boundaries of the known UTR element. These possibilities are further distributed by way of example in Figure 1B and Figure 1C Examples in the text, respectively with Figure 1A Comparisons are made. Similarly, the replaced UTR element does not need to be constrained by the exact boundaries of the known UTR element.

[0082] Preferably, the artificial nucleic acid molecule of the present invention does not contain the 3'-UTR (element) and / or 5'-UTR (element) of ribosomal protein S6, RPL36AL, rps16 or ribosomal protein L9. More preferably, the artificial nucleic acid molecule of the present invention does not contain the 3'-UTR (element) and / or 5'-UTR (element) of ribosomal protein S6, RPL36AL, rps16 or ribosomal protein L9, and the read frame of the artificial nucleic acid molecule of the present invention does not encode GFP protein. Even more preferably, the artificial nucleic acid molecule of the present invention does not contain the 3'-UTR (element) and / or 5'-UTR (element) of ribosomal protein S6, RPL36AL, rps16 or ribosomal protein L9, and the read frame of the artificial nucleic acid molecule of the present invention does not encode, for example, a reporter protein selected from the group consisting of: globulin (especially β-globulin), luciferase protein, GFP protein, glucuronidase protein (especially β-glucuronidase) or variants thereof, for example, variants showing at least 70% sequence identity with globulin, luciferase protein, GFP protein, or glucuronidase protein.

[0083] The term "3'-UTR element" refers to a nucleic acid sequence that contains or is composed of the following nucleic acid sequences: nucleic acid sequences derived from a 3'-UTR or variants or fragments derived from a 3'-UTR. "3'-UTR element" preferably refers to a nucleic acid sequence contained in the 3'-UTR of an artificial nucleic acid sequence (such as artificial mRNA). Therefore, in the sense of the invention, preferably, a 3'-UTR element can be contained in the 3'-UTR of mRNA (preferably artificial mRNA), or a 3'-UTR element can be contained in the 3'-UTR of its respective transcription template. Preferably, a 3'-UTR element is a nucleic acid sequence corresponding to the 3'-UTR of mRNA, preferably an artificial mRNA (such as mRNA obtained through a vector construct modified by transcription). Preferably, in the sense of the invention, a 3'-UTR element is a nucleotide sequence that functions as a 3'-UTR or encodes a 3'-UTR that performs the function of a 3'-UTR.

[0084] Therefore, the term "5'-UTR element" refers to a nucleic acid sequence that contains or is composed of the following nucleic acid sequences: nucleic acid sequences derived from a 5'-UTR or variants or fragments derived from a 5'-UTR. "5'-UTR element" preferably refers to a nucleic acid sequence contained in the 5'-UTR of an artificial nucleic acid sequence (such as artificial mRNA). Therefore, in the sense of the invention, preferably, a 5'-UTR element can be contained in the 5'-UTR of mRNA (preferably artificial mRNA), or a 5'-UTR element can be contained in the 5'-UTR of its respective transcription template. Preferably, a 5'-UTR element is a nucleic acid sequence corresponding to the 5'-UTR of mRNA, preferably an artificial mRNA (such as mRNA obtained through a vector construct modified by transcription). Preferably, in the sense of the invention, a 5'-UTR element is a nucleotide sequence that functions as a 5'-UTR or encodes a 5'-UTR that performs the function of a 5'-UTR.

[0085] The 3'-UTR element and / or 5'-UTR element in the artificial nucleic acid molecule of the present invention provide high translation efficiency for the artificial nucleic acid molecule. Therefore, the artificial nucleic acid molecule of the present invention may particularly include:

[0086] —Provides the artificial nucleic acid molecule with a 3'-UTR element that has high translation efficiency.

[0087] —Provides the artificial nucleic acid molecule with a 5'-UTR element that has high translation efficiency.

[0088] —Provides the artificial nucleic acid molecule with a 3'-UTR element and a 5'-UTR element that provide high translation efficiency.

[0089] Preferably, the artificial nucleic acid molecule of the present invention comprises a 3'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule and / or a 5'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule.

[0090] Preferably, the artificial nucleic acid molecule of the present invention comprises at least one 3'-UTR element and at least one 5'-UTR element, that is, at least one 3'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule and at least one 5'-UTR element that extends and provides high translation efficiency for the artificial nucleic acid molecule.

[0091] As detailed below, the at least one 3'-UTR element providing high translation efficiency for the artificial nucleic acid molecule or the at least one 5'-UTR element providing high translation efficiency for the artificial nucleic acid molecule can be selected from naturally occurring (preferably heterologous) 3'-UTR elements and 5'-UTR elements (both naturally occurring UTR elements or wild-type UTR elements), and from artificial 3'-UTR elements and artificial 5'-UTR elements (both artificial UTR elements). Wild-type UTR elements can be selected from the group including wild-type UTR elements published in references and publicly available databases (such as GenBank (NCBI)) and previously unpublished wild-type UTR elements. The latter can be identified by sequencing mRNA present in cells, preferably mammalian cells. Using this method, the inventors have identified some previously unpublished wild-type UTR elements, and this invention provides UTR elements of this type. The term artificial UTR element is not particularly limited and refers to any nucleic acid sequence not found in nature, i.e., different from wild-type UTR elements. However, in a preferred embodiment, the artificial UTR element used in this invention is a nucleic acid sequence exhibiting a specific degree (e.g., 10-99.9%, 20-99%, 30-98%, 40-97%, 50-96%, 60-95%, 70-90%) sequence identity with the wild-type UTR element. In a preferred embodiment, the artificial UTR used in this invention is identical to the wild-type UTR, except that one, two, three, four, five, or more than five nucleotides have been replaced by the same number of nucleotides (e.g., one nucleotide is replaced by one nucleotide). Preferably, the nucleotide replacement is a replacement of the respective complementary nucleotide. Preferred artificial UTR elements correspond to wild-type UTR elements, differing in that: (i) some or all of the ATG triplets (if present) in the wild-type 5'-UTR element are converted to the TAG triplet; and / or (ii) the cleavage site for the specific restriction enzyme (if present) in the wild-type 5'-UTR element or wild-type 3'-UTR element is removed by replacing one nucleotide within the cleavage site of the specific restriction enzyme with a complementary nucleotide, thereby removing the cleavage site of the specific restriction enzyme. The latter is typically required when (e.g., wild-type) UTR elements contain the cleavage site for the specific restriction enzyme and when the specific restriction enzyme is used (planned for) in subsequent cloning steps. Since such internal cleavage of the 5'-UTR element and 3'-UTR element is undesirable, artificial UTR elements with the restriction cleavage site of the specific restriction enzyme removed can be produced. Such substitutions can be performed by any suitable method known to those skilled in the art, for example, by PCR using modified primers.

[0092] Optionally, the artificial nucleic acid molecule of the present invention comprises at least one 3'-UTR element and at least one 5'-UTR element, that is, at least one 3'-UTR element that is extended and / or increased from the protein of the artificial nucleic acid molecule and at least one 5'-UTR element that is extended and / or increased from the protein of the artificial nucleic acid molecule.

[0093] Protein production; assays to determine whether protein production is prolonged and / or increased.

[0094] "Extending and / or increasing protein production from the artificial nucleic acid molecule" generally refers to the amount of protein produced from the artificial nucleic acid molecule of the present invention having individual 3'-UTR elements and / or 5'-UTR elements, compared to the amount of protein produced from individual reference nucleic acids lacking 3'-UTR and / or 5'-UTR or containing reference 3'-UTR and / or reference 5'-UTR (such as 3'-UTR and / or 5'-UTR naturally present in combination with ORF).

[0095] In particular, compared to individual nucleic acids lacking 3'-UTR and / or 5'-UTR or containing a reference 3'-UTR and / or 5'-UTR (such as 3'- and / or 5'-UTR naturally present in combination with ORF), at least one 3'-UTR element and / or 5'-UTR element of the artificial nucleic acid molecule of the present invention is extended from the protein generated from the artificial nucleic acid molecule of the present invention (e.g., from the mRNA of the present invention).

[0096] In particular, compared to individual nucleic acids lacking 3'- and / or 5'-UTR or containing a reference 3'- and / or 5'-UTR (such as the 3'- and / or 5'-UTR naturally present in combination with ORF), at least one 3'-UTR element and / or 5'-UTR element of the artificial nucleic acid molecule of the present invention increases protein production from the artificial nucleic acid molecule of the present invention (e.g., from the mRNA of the present invention), especially protein expression and / or total protein production.

[0097] Preferably, the at least one 3'-UTR element and / or the at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention does not negatively affect the translation efficiency of the nucleic acid compared to the translation efficiency of individual nucleic acids lacking a 3'-UTR and / or a 5'-UTR or containing a reference 3'-UTR and / or a reference 5'-UTR (such as a 3'-UTR and / or a 5'-UTR naturally present in combination with an ORF). Alternatively, the translation efficiency is enhanced by the 3'-UTR and / or 5'-UTR compared to the translation efficiency of proteins encoded by individual ORFs in their natural state.

[0098] As used herein, the terms “various nucleic acid molecules” or “reference nucleic acid molecules” mean that, except for different 3'-UTRs and / or 5'-UTRs, a reference nucleic acid molecule is equivalent to, preferably identical to, the artificial nucleic acid molecule of the present invention containing 3'-UTR elements and / or 5'-UTR elements.

[0099] To evaluate the in vivo or in vitro protein production of the artificial nucleic acid molecules of the present invention as defined herein (i.e., in vitro involving (“living”) cells and / or tissues (including tissues of a living subject); cells particularly include cell lines, primary cells, tissues, or cells in a subject, preferably mammalian cells, such as human and mouse cells, and particularly preferably human cell lines HeLa, HEPG2, and U-937 and mouse cell lines NIH3T3, JAWSII, and L929, with primary cells being particularly preferred, and in a particularly preferred embodiment, human skin fibroblasts (HDF)), the expression of the encoded protein is determined after the artificial nucleic acid molecules of the present invention are injected / transfected into target cells / tissues, and compared with protein expression induced by a reference nucleic acid. Quantitative methods for determining protein expression are known in the art (e.g., Western blotting, FACS, ELISA, mass spectrometry). In this case, it is particularly useful to determine the expression of reporter proteins such as luciferase, green fluorescent protein (GFP), or secreted alkaline phosphatase (SEAP). Therefore, the artificial nucleic acid or reference nucleic acid described in this invention is introduced into target tissues or cells, preferably in mammalian expression systems, such as mammalian cells, for example, HDF, L929, HepG2, and / or HeLa cells. Several hours or days (e.g., 6, 12, 24, 48, or 72 hours) after expression initiation or introduction of the nucleic acid molecule, target cell samples are collected and measured and / or lysed by FACS. The lysate can then be used to detect the expressed protein using various methods, such as Western blotting, FACS, ELISA, mass spectrometry, or by fluorescence or luminescence measurements (and thereby determine the efficiency of protein expression).

[0100] Therefore, if the protein expression of the artificial nucleic acid molecule described in this invention is introduced separately into the target tissue / cell at specific time points (e.g., 6, 12, 24, 48, or 72 hours after expression initiation or introduction of the nucleic acid molecule), compared to protein expression from a reference nucleic acid molecule, and tissue / cell samples are collected after the specific time points, protein lysis products are prepared according to a specific protocol tailored to a specific detection method (e.g., Western blotting, ELISA, fluorescence or luminescence measurement, etc., as known in the art), and the protein is detected by a selected detection method. As an alternative to measuring the amount of protein expressed in the cell lysis products—or, in addition to measuring the amount of protein in the cell lysis products before the cells are collected by lysis or using equal fractions in parallel—the protein amount can also be determined by using FACS analysis.

[0101] The term "prolonged" protein production from artificial nucleic acid molecules such as artificial mRNA preferably means, more preferably in mammalian expression systems such as HDF, L929, HEP2G, or HeLa cells, prolonged protein production from artificial nucleic acid molecules such as artificial mRNA (e.g., containing or lacking a reference 3'- and / or 5'-UTR). Therefore, protein production from artificial nucleic acid molecules such as artificial mRNA can be observed for a longer period compared to protein production from reference nucleic acid molecules. In other words, the amount of protein produced from artificial nucleic acid molecules such as artificial mRNA measured at a later time point, such as 48 or 72 hours post-transfection, is greater than the amount of protein produced from reference nucleic acid molecules such as reference mRNA at the corresponding later time point. The "later time point" can be, for example, any time beyond the initial expression, such as 24 hours after transfection of the nucleic acid molecule, for example, the initial expression, i.e., 36, 48, 60, 72, or 96 hours post-transfection. In addition, for the same nucleic acid, the amount of protein produced at a later time point can be standardized relative to the amount produced at an earlier (reference) time point. For example, the amount of protein at a later time point can be expressed as a percentage of the amount of protein 24 hours after transfection.

[0102] Preferably, the effect of extending protein production is determined by the following steps: (i) measuring, for example, the amount of protein obtained over time by expressing a reporter protein such as luciferase, preferably in mammalian expression systems such as HDF, L929, HEP2G, or HeLa cells; (ii) determining the amount of protein observed at a “reference” time point t1, for example, after t1 = 24 hours, and setting this amount of protein as 100%; and (iii) determining the amount of protein observed at one or more later time points t2, t3, etc., for example, after transfection, t2 = 48 hours and t3 = 72 hours, and calculating the relative amount of protein observed at the later time points as a percentage of the amount of protein at time point t1. For example, a protein expressed at “80” at t1, at “20” at t2, and at “10” at t3, would have a relative amount of 25% at t2 and 12.5% ​​at t3. These relative amounts at later time points can then be compared in step (iv) with the relative protein amounts at the corresponding time points for nucleic acid molecules that are respectively lacking 3'- and / or 5'-UTR or respectively containing reference 3'- and / or 5'-UTR. By comparing the relative protein amounts produced from the artificial nucleic acid molecules of the present invention with the relative protein amounts produced from reference nucleic acid molecules (i.e., nucleic acid molecules that are respectively lacking 3'- and / or 5'-UTR or respectively containing reference 3'- and / or 5'-UTR), the factors causing the prolongation of protein production from the artificial nucleic acid molecules of the present invention compared to protein production from reference nucleic acid molecules can be determined.

[0103] Preferably, compared to protein production from reference nucleic acid molecules that are, for example, lacking 3'- and / or 5'-UTRs or containing reference 3'- and / or 5'-UTRs respectively, the at least one 3'- and / or 5'-UTR element of the artificial nucleic acid molecule of the present invention is extended by at least 1.2 times, preferably at least 1.5 times, more preferably at least 2 times, and even more preferably at least 2.5 times, from protein production from said artificial nucleic acid molecule at a later time point as described above, compared to the (relative) amount of protein produced from reference nucleic acid molecules (which, for example, lack 3'- and / or 5'-UTRs or contain reference 3'- and / or 5'-UTRs respectively).

[0104] Alternatively, the effect of prolonging protein production can also be determined by the following steps: (i) measuring, for example, the amount of protein obtained over time by expressing a reporter protein such as luciferase, encoded in mammalian expression systems such as HDF, L929, HEP2G, or HeLa cells; (ii) determining time points where the protein amount is lower than, for example, at 1, 2, 3, 4, 5, or 6 hours after initiation of expression, or at 1, 2, 3, 4, 5, or 6 hours after transfection with an artificial nucleic acid molecule; and (iii) comparing the time points where the protein amount is lower than that observed at 1, 2, 3, 4, 5, or 6 hours after initiation of expression with said time points determined for nucleic acid molecules that are respectively lacking 3'- and / or 5'-UTR or respectively containing a reference 3'- and / or 5'-UTR.

[0105] For example, compared to protein production from a reference nucleic acid molecule (such as a reference mRNA), in mammalian expression systems, such as mammalian cells, for example in HDF, L929, HEP2G, or HeLa cells, protein production from an artificial nucleic acid molecule such as an artificial mRNA—in amounts observed at least during the initial phase of expression, such as after the initiation of expression, such as 1, 2, 3, 4, 5, or 6 hours after transfection of the nucleic acid molecule—is prolonged by at least about 5 hours, preferably at least about 10 hours, and more preferably at least about 24 hours. Therefore, compared to reference nucleic acid molecules that are respectively lacking 3'- and / or 5'-UTR or respectively containing reference 3'- and / or 5'-UTR, the artificial nucleic acid molecules of the present invention preferably allow for a prolonged protein production amount observed at least during the initial phase of expression, such as after the initiation of expression, such as 1, 2, 3, 4, 5, or 6 hours after transfection—in amounts observed at least 5 hours, preferably at least about 10 hours, and more preferably at least about 24 hours.

[0106] In a preferred embodiment, the protein production period from the artificial nucleic acid molecule of the present invention is extended by at least 1.2 times, preferably at least 1.5 times, more preferably at least 2 times, and even more preferably at least 2.5 times, compared to protein production from reference nucleic acid molecules that are respectively lacking 3'- and / or 5'-UTR or respectively containing reference 3'- and / or 5'-UTR.

[0107] Preferably, to achieve this prolonged protein production effect, the total amount of protein produced from the artificial nucleic acid molecule of the present invention within a time interval of, for example, 48 or 72 hours, corresponds at least to the amount of protein produced from a reference nucleic acid molecule that is deficient in 3'- and / or 5'-UTR or contains reference 3'- and / or 5'-UTRs (such as the 3'-UTRs and / or 5'-UTRs present in the ORF of the natural and artificial nucleic acid molecule). Thus, the present invention provides an artificial nucleic acid molecule that allows for prolonged protein production in mammalian expression systems, such as mammalian cells, for example in HDF, L929, HEP2G, or HeLa cells, as described above, wherein the total amount of protein produced from said artificial nucleic acid molecule (e.g., within a time interval of 48 or 72 hours) is at least, for example, the total amount of protein produced within said time interval from a reference nucleic acid molecule that is deficient in 3'- and / or 5'-UTR or contains reference 3'- and / or 5'-UTRs (such as the 3'- and / or 5'-UTRs present in the ORF of the natural and artificial nucleic acid molecule).

[0108] Furthermore, the term "extended protein expression" also includes "stabilized protein expression," whereby "stabilized protein expression" preferably means that, when compared with a reference nucleic acid molecule (e.g., mRNA containing reference 3'- and / or 5'-UTR or mRNA lacking 3'- and / or 5'-UTR respectively), there is a more uniform protein production from the artificial nucleic acid molecule described in this invention over a predetermined period, such as 24 hours, more preferably 48 hours, and even more preferably 72 hours.

[0109] Therefore, for example in mammalian systems, the protein production level of the artificial nucleic acid molecule containing 3'- and / or 5'-UTR elements described in this invention, such as the mRNA described in this invention, preferably does not decrease to the level observed for a reference nucleic acid molecule, such as the reference mRNA described above. To assess the extent to which protein production decreases from a specific nucleic acid molecule, for example, the amount of protein (encoded by the respective ORFs) observed 24 hours after initial expression, such as 24 hours after transfection of the artificial nucleic acid molecule described in this invention into cells, such as mammalian cells, can be compared to the amount of protein observed 48 hours after initial expression, such as 48 hours after transfection. Therefore, the ratio of the amount of protein encoded by the ORF of the artificial nucleic acid molecule of the present invention, such as the amount of reporter protein, for example, luciferase, observed at a later time point, for example, 48 hours after initiation of expression (e.g., post-transfection), to the amount of protein observed at an earlier time point, for example, 24 hours after initiation of expression (e.g., post-transfection), is preferably higher than the corresponding ratio for reference nucleic acid molecules that contain reference 3'- and / or 5'-UTRs or lack 3'- and / or 5'-UTRs, respectively (including the same time point).

[0110] Preferably, the ratio of the amount of protein encoded by the ORF of the artificial nucleic acid molecule of the present invention, such as the amount of a reporter protein, for example, luciferase, observed at a later time point, for example, 48 hours after initiation of expression (e.g., post-transfection), to the amount of protein observed at an earlier time point, for example, 24 hours after initiation of expression (e.g., post-transfection), is preferably at least 0.2, more preferably at least about 0.3, even more preferably at least about 0.4, even more preferably at least about 0.5, and particularly preferably at least about 0.7. For each reference nucleic acid molecule, for example, mRNA containing reference 3'- and / or 5'-UTR or lacking 3'- and / or 5'-UTR respectively, the ratio may be, for example, between about 0.05 and about 0.35.

[0111] Therefore, the present invention provides an artificial nucleic acid molecule comprising an ORF and 3'- and / or 5'-UTR elements as described above, wherein, preferably in a mammalian expression system, such as in mammalian cells, for example in HDF cells or HeLa cells, the ratio of the amount of protein observed 48 hours after initiation of expression, for example, the amount of luciferase, to the amount of protein observed 24 hours after initiation of expression is preferably at least 0.2, more preferably at least about 0.3, more preferably at least about 0.4, even more preferably at least about 0.5, even more preferably at least about 0.6, and particularly preferably at least about 0.7. Thus, preferably, for example, the total amount of protein produced from the artificial nucleic acid molecule within a 48-hour time interval corresponds at least, for example, to the total amount of protein produced within said time interval from a reference nucleic acid molecule that is lacking 3'- and / or 5'-UTR or contains reference 3'- and / or 5'-UTRs (such as the 3'-UTRs and / or 5'-UTRs present in the ORF of the natural and artificial nucleic acid molecules).

[0112] Preferably, the present invention provides an artificial nucleic acid molecule comprising ORF and 3'-UTR elements and / or 5'-UTR elements as described above, wherein preferably in a mammalian expression system, such as in mammalian cells, for example in HeLa cells or HDF cells, the ratio of the amount of protein observed 72 hours after initiation of expression, for example, the amount of luciferase, to the amount of protein observed 24 hours after initiation of expression is preferably greater than about 0.05, more preferably greater than about 0.1, more preferably greater than about 0.2, and even more preferably greater than about 0.3, wherein preferably, for example, the total amount of protein produced from the artificial nucleic acid molecule within a 72-hour time interval, at least for example, within said time interval, from reference nucleic acid molecules respectively lacking 3'- and / or 5'-UTR or respectively containing reference 3'- and / or 5'-UTR (such as 3'- and / or 5'-UTR present in the ORF of the natural and artificial nucleic acid molecules).

[0113] In the context of this invention, "increased protein expression" or "enhanced protein expression" preferably means the increased / enhanced protein expression or the total amount of increased / enhanced protein expression at a time point after initial expression, compared to the expression induced by a reference nucleic acid molecule. Thus, the protein levels observed at a certain time point after initial expression, for example, after transfection with the artificial nucleic acid molecule described in this invention, or for example, after transfection with the mRNA described in this invention, such as 6, 12, 24, 48, or 72 hours after transfection, are preferably higher than the protein levels observed at the same time point after initial expression, for example, after transfection with a reference nucleic acid molecule, such as a reference mRNA containing reference 3'- and / or 5'-UTRs or lacking 3'- and / or 5'-UTRs respectively. In a preferred embodiment, the maximum amount of protein expressed from the artificial nucleic acid molecule (as determined, for example, by protein activity or quality) increases with respect to the amount of protein expressed from the reference nucleic acid containing reference 3'- and / or 5'-UTRs or lacking 3'- and / or 5'-UTRs respectively. Preferably, for example, peak expression levels are reached within 48 hours, more preferably within 24 hours, and even more preferably within 12 hours after transfection.

[0114] Preferably, the term "increased" or "enhanced" total protein production of the artificial nucleic acid molecules of the present invention refers, preferably in mammalian expression systems, such as mammalian cells, for example in HDF, L929, HEP2G, or HeLa cells, to increased / enhanced protein production over time intervals, such as 48 hours or 72 hours, compared to reference nucleic acid molecules that are deficient in 3'- and / or 5'-UTR or contain reference 3'- and / or 5'-UTR, respectively. According to a preferred embodiment, when using the artificial nucleic acid molecules of the present invention, the accumulation of expressed protein increases over time.

[0115] The total protein amount at a specific time point can be determined by the following steps: (i) collecting tissues or cells at multiple time points after the introduction of the artificial nucleic acid molecule (e.g., 6, 12, 24, 48, and 72 hours after initiation of expression or after introduction of the nucleic acid molecule), and determining the protein amount at each time point as described above. To calculate the accumulated protein amount, mathematical methods for determining the total protein amount can be used, such as the area under the curve (AUC), which can be determined according to the following formula:

[0116]

[0117] To calculate the area under the curve for total protein, the integral of the expression curve equation is calculated from each endpoint (a and b).

[0118] Therefore, "total protein production" preferably refers to the area under the curve (AUC) that represents the relative time of protein production.

[0119] Preferably, compared to protein production from a reference nucleic acid molecule lacking 3'- and / or 5'-UTR respectively, at least one 3'- or 5'-UTR element of the present invention increases protein production from the artificial nucleic acid molecule by at least 1.5 times, preferably at least 2 times, and more preferably at least 2.5 times. In other words, compared to the (relative) amount of protein produced from a reference nucleic acid molecule (which, for example, lacks 3'- and / or 5'-UTR respectively or contains reference 3'- and / or 5'-UTR respectively) at a corresponding later time point, at a certain time point, for example after the initiation of expression, such as 48 hours or 72 hours after transfection, the total amount of protein produced from the artificial nucleic acid molecule of the present invention increases by at least 1.5 times, preferably at least 2 times, and more preferably at least 2.5 times.

[0120] The elongation effect and efficiency of mRNA and / or protein of variants, fragments and / or variant fragments of 3'-UTR and / or 5'-UTR, as well as the elongation effect and efficiency of mRNA and / or protein of at least one 3'-UTR element and / or at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention, and the elongation effect and efficiency of mRNA and / or protein of at least one 3'-UTR element, and / or the increase effect and efficiency of protein production, can be determined by any method known to those skilled in the art that is suitable for this purpose.

[0121] For example, artificial mRNA molecules can be generated that contain a coding sequence / reading frame (ORF) of a reporter protein (such as luciferase) and the 3'-UTR element described in this invention, i.e., its elongation and / or addition from the protein of the artificial mRNA molecule. Furthermore, such inventive mRNA molecules may also contain the 5'-UTR element described in this invention, i.e., its elongation and / or addition from the protein of the artificial mRNA molecule, or may not contain a 5'-UTR element or may contain a 5'-UTR element that is not described in this invention, such as a reference 5'-UTR. Therefore, artificial mRNA molecules can be generated that contain a coding sequence / reading frame (ORF) of a reporter protein (such as luciferase) and the 5'-UTR element described in this invention, i.e., its elongation and / or addition from the protein of the artificial mRNA molecule. Furthermore, such inventive mRNA molecules may also contain the 3'-UTR element described in this invention, i.e., its elongation and / or addition from the protein of the artificial mRNA molecule, or may not contain a 3'-UTR element or may contain a 3'-UTR element that is not described in this invention, such as a reference 3'-UTR.

[0122] According to the present invention, mRNA can be generated, for example, by in vitro transcription using various vectors such as plasmid vectors, said vectors containing, for example, a T7 promoter and a sequence encoding the respective mRNA sequence. The generated mRNA molecules can be transfected into cells using a transfection method suitable for mRNA transfection, for example, they can be lipid-transfected into mammalian cells, such as HeLa cells or HDF cells, and the samples can be analyzed at certain time points after transfection, for example, 6 hours, 24 hours, 48 ​​hours, and 72 hours post-transfection. The amount of mRNA and / or protein in said samples can be analyzed using methods known to those skilled in the art. For example, the amount of reporter mRNA present in cells at the sample time points can be determined by quantitative PCR. The amount of reporter protein encoded by each mRNA can be determined, for example, by Western blotting, ELISA assay, FACS analysis, or reporter assays such as luciferase assay, depending on the reporter protein used. The effect of stabilizing and / or prolonging protein expression can be analyzed, for example, by determining the ratio of protein levels observed at 48 hours post-transfection to protein levels observed at 24 hours post-transfection. The closer the value is to 1, the more stable the protein expression is during that period. The measurements can, of course, be performed at 72 hours or longer, and the ratio of the protein level observed 72 hours after transfection to the protein level observed 24 hours after transfection can be determined, thereby determining the stability of protein expression.

[0123] Translation efficiency; determination of translation efficiency

[0124] As detailed above, translation efficiency is a characteristic of nucleic acid molecules (e.g., mRNA) that contain a readability frame (ORF) and is typically provided by specific UTR elements (specific 5'-UTR elements or specific 3'-UTR elements) contained in the nucleic acid molecule. Translation efficiency is usually determined by the amount of protein translated from the ORF. In this regard, the present invention provides both general methods and specific assays. For example, by experimentally quantifying the protein encoded by the ORF, the translation efficiency of multiple (e.g., two or more) nucleic acid molecules can be compared.

[0125] Typically, cell-based methods are used to assess translation efficiency. As defined herein, the translation efficiency of the artificial nucleic acid molecules of the present invention is determined experimentally in vivo or in vitro (i.e., in vitro refers to (“living”) cells and / or tissues, including tissues of a living subject; cells include cells in a specific cell line, primary cells, tissue, or subject, preferably mammalian cells, such as human and mouse cells, particularly preferably human cell lines HeLa, HEPG2, and U-937 and mouse cell lines NIH3T3, JAWSII, and L929, further preferably primary cells, with a particularly preferred embodiment being human skin fibroblasts (HDF)). The expression of the encoded protein is determined after the artificial nucleic acid molecules of the present invention are injected / transfected into target cells / tissues and compared with protein expression induced by a reference nucleic acid. Quantitative methods for determining protein expression are known in the art (e.g., Western blotting, FACS, ELISA, mass spectrometry). In this context, it is particularly useful to determine the expression of reporter proteins, such as luciferase, green fluorescent protein (GFP), or secreted alkaline phosphatase (SEAP). Therefore, the artificial nucleic acid or reference nucleic acid described in this invention is introduced into target tissues or cells, preferably in mammalian expression systems, such as mammalian cells, for example, HDF, L929, HepG2, and / or HeLa cells. Several hours or days (e.g., 6, 12, 24, 48, or 72 hours) after expression initiation or introduction of the nucleic acid molecule, target cell samples are collected and measured and / or lysed by FACS. The lysate can then be used to detect the expressed protein using various methods, such as Western blotting, FACS, ELISA, mass spectrometry, or by fluorescence or luminescence measurements (and thereby determine the efficiency of protein expression).

[0126] Therefore, if the protein expression of the artificial nucleic acid molecule described in this invention is introduced separately into the target tissue / cell at specific time points (e.g., 6, 12, 24, 48, or 72 hours after expression initiation or introduction of the nucleic acid molecule), compared to protein expression from a reference nucleic acid molecule, and tissue / cell samples are collected after the specific time points, protein lysis products are prepared according to a specific protocol tailored to a specific detection method (e.g., Western blotting, ELISA, fluorescence or luminescence measurement, etc., as known in the art), and the protein is detected by a selected detection method. As an alternative to measuring the amount of protein expressed in the cell lysis products—or, in addition to measuring the amount of protein in the cell lysis products before the cells are collected by lysis or using equal fractions in parallel—the protein amount can also be determined by using FACS analysis.

[0127] Unless the context otherwise indicates, the other general aspects correspond to those described above regarding protein production.

[0128] Preferably, the at least one 3'-untranslated region element (3'-UTR element) and / or the at least one 5'-untranslated region element (5'-UTR element) provide high translation efficiency for the artificial nucleic acid molecule of the present invention. "Provides" means that in the absence of the at least one 3'-untranslated region element (3'-UTR element) and / or the at least one 5'-untranslated region element (5'-UTR element), the translation efficiency is low, as defined herein.

[0129] The term "high translation efficiency" can be used to compare the translation efficiency of the artificial nucleic acid molecule of the present invention with that of a reference nucleic acid molecule. The term "reference nucleic acid molecule" as used herein means—except for different 3'-UTR and / or 5'-UTR—that the reference nucleic acid molecule is equivalent, preferably identical, to the artificial nucleic acid molecule of the present invention containing said 3'-UTR element and / or said 5'-UTR element. Thus, the translation efficiency of the artificial nucleic acid molecule can be compared with that of the reference nucleic acid molecule. Various methods for comparing translation efficiencies are provided herein.

[0130] Unless the context otherwise indicates, all aspects described above regarding the determination of protein production also apply to methods for comparing translation efficiency. Since methods for comparing translation efficiency involve comparing two nucleic acid molecules, two parallel experiments are performed, the only difference between these experiments being the nature of the nucleic acid molecules. In other words, all other conditions, such as, for example, growth medium, temperature, and pH, are identical between the experiments. Based on common sense and the guidance provided herein, those skilled in the art routinely choose appropriate conditions, requiring that the same conditions be chosen for both parallel experiments. Typically, one of the parallel experiments involves a reference nucleic acid molecule (which may be wild-type or artificial), while the other involves an artificial nucleic acid molecule (also called a test nucleic acid molecule). This method allows comparison of whether the test nucleic acid molecule is characterized by high translation efficiency; that is, whether it is an artificial nucleic acid molecule as described in this invention. Specific aspects of the methods for comparing translation efficiency are as follows:

[0131] In this method, the reference nucleic acid molecule contains at least one open reading frame (ORF) identical to at least one open reading frame (ORF) of the artificial nucleic acid molecule; and the reference nucleic acid molecule does not contain at least one 3'-UTR element and / or at least one 5'-UTR element of the artificial nucleic acid molecule. The translation efficiency of the artificial nucleic acid molecule is compared with that of the reference nucleic acid molecule by a method including the following steps:

[0132] (i) Transfect mammalian cells with the artificial nucleic acid molecule and measure the expression level of the protein encoded by the ORF of the artificial nucleic acid molecule at a specific time point after transfection (e.g., 24 or 48 hours).

[0133] (ii) Transfect mammalian cells with the reference nucleic acid molecule and measure the expression level of the protein encoded by the ORF of the reference nucleic acid molecule at the same time point after transfection.

[0134] (iii) Calculate the ratio of the amount of protein expressed from the artificial nucleic acid molecule to the amount of protein expressed from the reference nucleic acid molecule.

[0135] The ratio calculated in (iii) is ≥1, preferably >1.

[0136] Therefore, by comparing the relative amount of protein produced from the test nucleic acid molecule described in this invention with the relative amount of protein produced from the reference nucleic acid molecule, the factor by which the translation efficiency of the artificial nucleic acid molecule described in this invention is increased compared to the translation efficiency of the reference nucleic acid molecule can be determined.

[0137] "Post-transfection" is used to indicate the time point at which cells have been transfected with the test nucleic acid molecule and the reference nucleic acid molecule.

[0138] Preferably, compared to protein generation from a reference nucleic acid molecule containing reference 3'- and / or 5'-UTRs respectively, at least one 3'- and / or 5'-UTR element in the artificial nucleic acid molecule of the present invention increases the translation efficiency of the artificial nucleic acid molecule by at least 1.2 times, preferably at least 1.5 times, more preferably at least 2 times, and even more preferably at least 2.5 times (i.e., the ratio calculated in (iii) is at least 1.2, preferably at least 1.5, more preferably at least 2, and even more preferably at least 2.5). In other words, for the same point in time, the increase in translation efficiency compared to the translation efficiency of the reference nucleic acid molecule is at least 1.2 times, preferably at least 1.5 times, more preferably at least 2, and even more preferably at least 2.5 times.

[0139] The high translation efficiency is attributed to the 5'-UTR or 3'-UTR that differs from that of the reference nucleic acid molecule.

[0140] In a preferred embodiment of the method for comparing translation efficiency according to the present invention, (i) the reference nucleic acid molecule contains a 3'-UTR element that is not present in the artificial nucleic acid molecule, or (ii) the reference nucleic acid molecule contains a 5'-UTR element that is not present in the artificial nucleic acid molecule. In a more preferred embodiment, Smart Light, (i) the reference nucleic acid molecule contains a 3'-UTR element derived from the albumin gene (preferably derived from the 3'-UTR of the sequence corresponding to SEQ ID NO:207), or (ii) the reference nucleic acid molecule contains a 5'-UTR element derived from the TOP gene (preferably the 5'-UTR of human macroribosomal protein 32 lacking the 5' end oligopyrimidine bundle, preferably the 5'-UTR of the sequence corresponding to SEQ ID NO:208).

[0141] In an alternative preferred embodiment, (i) the reference nucleic acid molecule does not contain any 3'-untranslated region element (3'-UTR element), or (ii) the reference nucleic acid molecule does not contain any 5'-untranslated region element (5'-UTR element).

[0142] In a preferred embodiment, the cells used in the step of transfecting the artificial nucleic acid molecule and the reference nucleic acid molecule are mammalian cells, preferably selected from the group consisting of HDF, L929, HEP2G, and HeLa cells. More preferably, the method is performed in parallel in more than one type of mammalian cell, such as in any combination of two, preferably any three, more preferably all four of the above cell lines. When the method is performed in parallel in more than one type of mammalian cell, the test nucleic acid molecule is considered to have high translation efficiency when the ratio calculated in step (iii) in at least one of the said mammalian cell types is ≥1, preferably >1. However, preferably, the test nucleic acid molecule is considered to have high translation efficiency when the ratio calculated in step (iii) in at least two, such as at least three, such as four types of mammalian cells used is ≥1, preferably >1.

[0143] The expression levels of the ORF-encoded proteins of the artificial nucleic acid molecule and the reference nucleic acid molecule are measured at specific time points after transfection (e.g., 12h, 24h, 36h, 48h, or 72h). Preferably, the expression levels of the ORF-encoded proteins are measured 24 hours after transfection.

[0144] The present invention also provides particularly preferred reference constructs that can be used in the methods of the present invention. Particularly preferred is the mRNA of SEQ ID NO:205 (…). Figure 1AIn this case, it is further preferred that the reference nucleic acid molecule differs from the artificial nucleic acid molecule only in that at least one 3'-UTR element is replaced by at least one additional 3'-UTR element (preferably one 3'-UTR element), or at least one 5'-UTR element is replaced by at least one additional 5'-UTR element (preferably one 5'-UTR element). In each case, this allows for direct detection of whether the additional 3'-UTR element or the additional 5'-UTR element, respectively, provides high translation efficiency.

[0145] Particularly preferably, the method for determining translation efficiency further has the following characteristics:

[0146] (a) In the transfection step, the cells are mammalian cells selected from the group consisting of HDF, L929, HEP2G, and HeLa cells; and

[0147] (b) The measurement was taken 24 hours after transfection; and

[0148] (c) The reference nucleic acid molecule is the mRNA of SEQ ID NO:205. Figure 1A );and

[0149] Preferably, (d) the reference nucleic acid molecule differs from the artificial nucleic acid molecule only in that (i) at least one 3'-UTR element (preferably one 3'-UTR element) is replaced by at least one other 3'-UTR element (preferably one 3'-UTR element), or (ii) at least one 5'-UTR element (preferably one 5'-UTR element) is replaced by at least one other 5'-UTR element (preferably one 5'-UTR element); in other words, except for at least one 5'-UTR element and Figure 1A The 5'-UTR shown is different, or at least one 3'-UTR element is different. Figure 1A Except for the 3'-UTR shown, the test nucleic acid molecule is the same as the reference nucleic acid molecule. Figure 1B This shows an example of 5'-UTR element replacement; Figure 1C This shows an example of 3'-UTR element replacement.

[0150] When all of (a) through (d) are achieved, the method is also referred to as a "standard determination of translation efficiency," or simply "determination of translation efficiency." Particularly preferably, the artificial nucleic acid molecules of the present invention are characterized by high translation efficiency determined by this standard determination of translation efficiency. This standard determination allows for the direct identification of particularly useful artificial nucleic acid molecules.

[0151] In Example 3 of this invention, a standard assay was used. Thus, the inventors discovered that a group of novel artificial nucleic acids share the advantageous characteristic of high translation efficiency, namely, their translation efficiency is higher than... Figure 1A The translation efficiency of the reference nucleic acid molecule is described in Example 3.5. This is a very surprising and advantageous finding, especially given the background that the reference nucleic acid molecule contains a 3'-UTR element derived from the albumin gene (corresponding to the 3'-UTR of the DNA sequence described in SEQ ID NO: 207) and a 5'-UTR element derived from the TOP gene (particularly the 5'-UTR of human macroribosomal protein 32 lacking the 5'-terminal oligopyrimidine bundle (corresponding to the 5'-UTR of the sequence described in SEQ ID NO: 208)). Such 3'-UTR elements and 5'-UTR elements are known to be advantageous in cases of protein translation in cells (for the 3'-UTR derived from the albumin gene, see WO2013 / 143700; for the 5'-TOP-UTR, see WO2013 / 143700). The present invention provides further improvements.

[0152] In some embodiments, the ORF of the artificial nucleic acid molecule encodes a quantifiable protein, preferably a reporter protein. These embodiments are well-suited for comparing the translation efficiency methods described in this invention. In other embodiments, the ORF of the artificial nucleic acid molecule does not encode a quantifiable protein, preferably not a reporter protein. For example, when using an ORF of an artificial nucleic acid molecule encoding a quantifiable protein to determine a specific 5'-UTR or 3'-UTR that provides high translation efficiency by the methods of this invention, the 5'-UTR or 3'-UTR can be combined with any ORF. For example, although in standard assays for determining translation efficiency, the ORF encodes a reporter protein, the artificial nucleic acid of this invention is not limited to the ORF used in standard assays.

[0153] These methods can also be used in combination; that is, artificial nucleic acid molecules that have been identified as having high translation efficiency characteristics by the methods in this paper can themselves be used as “reference constructs” for the next round of methods in this paper. This allows for the detection of other artificial nucleic acid molecules that release translation efficiency characteristics that are even higher (or whether the UTR contained in other artificial nucleic acid molecules provides even higher translation efficiency). This allows for the identification of even further improved artificial nucleic acid molecules and UTRs.

[0154] Stable mRNA

[0155] In some embodiments, the at least one 3'-UTR element and / or the at least one 5'-UTR element in the artificial nucleic acid molecule of the present invention are derived from a stable mRNA. Thus, "derived from" a stable mRNA means that the at least one 3'-UTR element and / or the at least one 5'-UTR element has at least 50%, preferably at least 60%, preferably at least 70%, more preferably at least 75%, more preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, even more preferably at least 95%, and particularly preferably at least 98% sequence identity with the 3'-UTR element and / or 5'-UTR element of the stable mRNA. Preferably, the stable mRNA is a naturally occurring mRNA, and thus, the 3'-UTR element and / or 5'-UTR element of the stable mRNA refers to the 3'-UTR and / or 5'-UTR of a naturally occurring mRNA, or fragments or variants thereof. Furthermore, 3'-UTR elements and / or 5'-UTR elements derived from stable mRNA preferably refer to those modified compared to naturally occurring 3'-UTR elements and / or 5'-UTR elements, for example, to increase RNA stability, or even further and / or lengthen and / or increase protein-generated 3'-UTR elements and / or 5'-UTR elements. This does not mean that said modification is preferred, for example, that it does not disrupt RNA stability compared to naturally occurring (unmodified) 3'-UTR elements and / or 5'-UTR elements. In particular, as used herein, the term mRNA refers to the mRNA molecule; however, it can also refer to the types of mRNA as defined herein.

[0156] Preferably, the stability of the mRNA, i.e., the mRNA degradation and / or half-life, is assessed under standard conditions, such as standard conditions (standard culture medium, incubation, etc.) for a particular cell line in use.

[0157] As used herein, the term "stable mRNA" generally refers to mRNA with slow mRNA degradation. Thus, "stable mRNA" typically has a long half-life. The half-life of mRNA is the time required for 50% degradation of the mRNA molecule present in vivo or in vitro. Therefore, the stability of mRNA is typically assessed in vivo or in vitro. In vitro, in particular, refers to ("live") cells and / or tissues, including tissues of a live subject. Cells specifically include cell lines, primary cells, tissues, or cells in a subject. In specific embodiments, cell types that allow cell culture may be suitable for the present invention. Mammalian cells, such as human and mouse cells, are particularly preferred. In particularly preferred embodiments, human cell lines HeLa, HEPG2, and U-937, and mouse cell lines NIH3T3, JAWSII, and L929 are used. Furthermore, primary cells are particularly preferred, and in particularly preferred embodiments, human skin fibroblasts (HDF) may be used. Alternatively, tissues of a subject may also be used.

[0158] Preferably, the half-life of the “stable mRNA” is at least 5 hours, at least 6 hours, at least 7 hours, at least 8 hours, at least 9 hours, at least 10 hours, at least 11 hours, at least 12 hours, at least 13 hours, at least 14 hours, and / or at least 15 hours. The half-life of the mRNA under study can be determined by various methods known to those skilled in the art. Typically, the half-life of the mRNA under study is determined by determining a degradation constant, thereby inferring a generally ideal in vivo (or in vitro as defined above) state in which transcription of the mRNA under study can be completely “shut down” (or at least tuned to an undetectable level). Under this ideal state, it is generally presumed that mRNA degradation follows first-order kinetics. Therefore, mRNA degradation can generally be described by the following equation:

[0159] A(t) = A0 * e -λt

[0160] A0 is the amount (or concentration) of the mRNA under study at time 0, i.e., before degradation begins, A(t) is the amount (or concentration) of the mRNA under study at time t during the degradation process, and λ is the degradation constant. Therefore, if the amount (or concentration) of the mRNA under study at time 0 (A0) and the amount (or concentration) of the mRNA under study at a certain time t during the degradation process (A(t) and t) are known, the degradation constant λ can be calculated. Based on the degradation constant λ, the half-life t can be calculated using the following equation. 1 / 2 :

[0161] t 1 / 2 =ln2 / λ.

[0162] Because according to the definition at t 1 / 2A(t) / A0 = 1 / 2. Therefore, in order to assess the half-life of the mRNA under study, the amount or concentration of mRNA is usually determined during RNA degradation in vivo (or in vitro as defined above).

[0163] Various methods, known to those skilled in the art, can be used to determine the amount or concentration of mRNA during RNA degradation in vivo (or in vitro as defined above). Non-limiting examples of such methods include, for example, general inhibition of transcription using transcriptional inhibitors such as actinomycin D; specific initiation of transient transcription using inducible promoters, such as the c-fos serum-inducible promoter system and the Tet-off regulated promoter system; and, for example, kinetic labeling techniques using 4-thiouridine (4sU), 5-ethynyluridine (EU), or 5'-bromouridine (BrU), such as pulse labeling. Further details and preferred embodiments of how to determine the amount or concentration of mRNA during RNA degradation are outlined below in the case of methods for identifying the 3'-UTR element and / or at least one 5'-UTR element described in this invention. Various descriptions and preferred embodiments of how to determine the amount or concentration of mRNA during RNA degradation also apply herein.

[0164] Preferably, the "stable mRNA" in the sense of this invention has a slower mRNA degradation compared to average mRNA, preferably assessed in vivo (or in vitro as defined above). For example, "average mRNA degradation" can be determined by studying a variety of mRNA species, preferably 100, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10000, at least 11000, at least 12000, at least 13000, at least 1400. The assessment can be performed on mRNA degradation of at least 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, or 30,000 mRNA species. Particularly preferred is the assessment of the entire transcriptome, or as many mRNA species as possible within the transcriptome. This can be achieved, for example, by using microarrays that provide full transcriptome coverage.

[0165] As used herein, “mRNA species” corresponds to a genomic transcription unit, which typically corresponds to a gene. Thus, for example, different transcripts can exist within a single “mRNA species” due to mRNA processing. For example, mRNA species can be represented by dots on a microarray. Therefore, microarrays provide a useful tool for determining, for example, the amount of multiple mRNA species at a given time point during mRNA degradation. However, other techniques known to those skilled in the art can also be used, such as RNA-seq, quantitative PCR, etc.

[0166] In this invention, it is particularly preferred that the stable mRNA is characterized by a ratio of the amount of mRNA at the second time point to the amount of mRNA at the first time point of at least 0.5 (50%), at least 0.6 (60%), at least 0.7 (70%), at least 0.75 (75%), at least 0.8 (80%), at least 0.85 (85%), at least 0.9 (90%), or at least 0.95 (95%) of the mRNA degradation. Thus, the second time point is later than the first time point during the degradation process.

[0167] Preferably, a first time point is chosen such that only mRNA undergoing degradation is considered, i.e., mRNA that has emerged—e.g., during transcription. For example, if a kinetic labeling technique, such as pulse labeling, is used, a first time point is preferably chosen such that the introduction of the label into the mRNA is completed, i.e., continuous introduction of the label into the mRNA does not occur. Thus, if kinetic labeling is used, the first time point can be at least 10 min, at least 20 min, at least 30 min, at least 40 min, at least 50 min, at least 60 min, at least 70 min, at least 80 min, or at least 90 min after the cells have been incubated with the label.

[0168] For example, the first time point may preferably be the termination of transcription in the presence of an inducible promoter (e.g., by a transcription inhibitor), after the termination of promoter induction or after the termination pulse or label supply, for example, 0 to 6 hours after labeling ends. More preferably, the first time point may be the termination of transcription in the presence of an inducible promoter (e.g., by a transcription inhibitor), after the termination of promoter induction or after the termination pulse or label supply, for example, 30 minutes to 5 hours after labeling ends, even more preferably 1 hour to 4 hours, and particularly preferably about 3 hours.

[0169] Preferably, the second time point is chosen as late as possible during mRNA degradation. However, if multiple mRNA species are considered, it is preferable to choose a second time point such that a considerable amount of the various mRNA species, preferably at least 10% of the mRNA species, are still present in detectable amounts, i.e., in amounts greater than 0. Preferably, the second time point is at least 5 h, at least 6 h, at least 7 h, at least 8 h, at least 9 h, at least 10 h, at least 11 h, at least 12 h, at least 13 h, at least 14 h, or at least 15 h after the end of transcription or the experimental labeling process.

[0170] Therefore, the time interval between the first time point and the second time point is preferably as large as possible within the aforementioned limits. Thus, the time interval between the first time point and the second time point is preferably at least 4 hours, at least 5 hours, at least 6 hours, at least 7 hours, at least 8 hours, at least 9 hours, at least 10 hours, at least 11 hours, or at least 12 hours.

[0171] Furthermore, it is possible that the at least one 3'-UTR element and / or the at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention can be identified by the methods described herein for identifying the 3'-UTR elements and / or 5'-UTR elements of the present invention. Particularly preferred is that the at least one 3'-UTR element and / or the at least one 5'-UTR element in the artificial nucleic acid molecule of the present invention can be identified by the methods described herein for identifying 3'-UTR elements and / or 5'-UTR elements that provide high translation efficiency for the artificial nucleic acid molecule.

[0172] Preferred embodiments of the artificial nucleic acid molecule of the present invention

[0173] Preferably, the at least one 3'-UTR element and / or at least one 5'-UTR element in the artificial nucleic acid molecule of the present invention comprises or is composed of the following nucleic acid sequences: the nucleic acid sequences are derived from the 3'-UTR and / or 5'-UTR of eukaryotic protein-coding genes, preferably from the 3'-UTR and / or 5'-UTR of vertebrate protein-coding genes, more preferably from the 3'-UTR and / or 5'-UTR of mammalian protein-coding genes (e.g., from mouse and human protein-coding genes), even more preferably from the 3'-UTR and / or 5'-UTR of primate or rodent protein-coding genes, especially the 3'-UTR and / or 5'-UTR of human or mouse protein-coding genes.

[0174] Generally, it is to be understood that at least one 3'-UTR element in the artificial nucleic acid molecule of the present invention comprises or is composed of the following nucleic acid sequence: preferably a nucleic acid sequence derived from a naturally occurring (natural) 3'-UTR, and at least one 5'-UTR element in the artificial nucleic acid molecule of the present invention comprises or is composed of the following nucleic acid sequence: preferably a nucleic acid sequence derived from a naturally occurring (natural) 5'-UTR.

[0175] Preferably, the at least one read frame is heterologous with at least one 3'-UTR element and / or with at least one 5'-UTR element. In this case, the term "heterologous" means that the two sequence elements contained in the artificial nucleic acid molecule, such as the read frame and the 3'-UTR element and / or the read frame and the 5'-UTR element, are not naturally (in their natural state) combined in this way. They are typically recombinant. Preferably, the 3'-UTR element and / or the 5'-UTR element originates from a gene different from the read frame. For example, the ORF can originate from a gene different from the 3'-UTR element and / or at least one 5'-UTR element, such as encoding a different protein or the same protein but belonging to a different class. That is, the read frame originates from a gene different from the gene from which the 3'-UTR element originates and / or from which the at least one 5'-UTR element originates. In a preferred embodiment, the ORF does not encode human or plant (e.g., Arabidopsis) ribosomal proteins, preferably not human ribosomal protein S6 (RPS6), human ribosomal protein L36a-like (RPL36AL), or Arabidopsis ribosomal protein S16 (RPS16). In a further preferred embodiment, the read frame (ORF) does not encode ribosomal protein S6 (RPS6), ribosomal protein L36a-like (RPL36AL), or ribosomal protein S16 (RPS16).

[0176] In a specific embodiment, it is preferred that the read frame (ORF) does not encode a reporter protein, for example, a reporter protein selected from the group consisting of: globulins (especially β-globulins), luciferase proteins, GFP proteins, or variants thereof, for example, variants exhibiting at least 70% sequence identity with globulins, luciferase proteins, or GFP proteins. Thus, it is particularly preferred that the ORF does not encode a GFP protein. It is also particularly preferred that the ORF does not encode a reporter gene or is not derived from a reporter gene, wherein the reporter gene is preferably not selected from the group consisting of: globulins (especially β-globulins), luciferase proteins, β-glucuronidase (GUS), and GFP proteins, or variants thereof, preferably not selected from EGFP, or variants of any of the above genes, said variants generally exhibiting at least 70% sequence identity with any of these reporter genes, preferably globulins, luciferase proteins, or GFP proteins.

[0177] Even more preferably, the 3'-UTR element and / or 5'-UTR element are heterologous to any other element contained in the artificial nucleic acid as defined herein. For example, if the artificial nucleic acid of the present invention contains a 3'-UTR element from a given gene, it preferably does not contain any other nucleic acid sequence, especially not a functional nucleic acid sequence (e.g., coding or regulatory sequence element) from the same gene, including its regulatory sequence at the 5' and 3' ends of the ORF of the gene. Thus, for example, if the artificial nucleic acid of the present invention contains a 5'-UTR element from a given gene, it preferably does not contain any other nucleic acid sequence, especially a functional nucleic acid sequence (e.g., coding or regulatory sequence element) from the same gene, including its regulatory sequence at the 5' and 3' ends of the ORF of the gene.

[0178] Furthermore, preferably, the artificial nucleic acid of the present invention comprises at least one read frame, at least one 3'-UTR (element), and at least one 5'-UTR (element), wherein the at least one 3'-UTR (element) is the 3'-UTR element of the present invention and / or the at least one 5'-UTR (element) is the 5'-UTR element of the present invention. In the preferred artificial nucleic acid of the present invention comprising at least one read frame, at least one 3'-UTR (element), and at least one 5'-UTR (element), it is particularly preferred that each of the at least one read frame, at least one 3'-UTR (element), and at least one 5'-UTR (element) is heterologous, i.e., at least one 3'-UTR (element) and at least one 5'-UTR (element), as well as the read frame and the 3'-UTR (element) or 5'-UTR (element), are not naturally occurring (in their natural state) in this combination. This means that the artificial nucleic acid molecule comprises an ORF, a 3'-UTR (element), and a 5'-UTR (element), which are heterologous to each other, for example, because each of them originates from a different gene (and their 5' and 3'UTRs), and they are recombinant. In another preferred embodiment, the 3'-UTR (element) is not derived from the 3'-UTR (element) of a viral gene or is not of viral origin.

[0179] Preferably, at least one 3'-UTR element and / or at least one 5'-UTR element are functionally linked to the ORF. This means that, preferably, the 3'-UTR element and / or at least one 5'-UTR element are associated with the ORF so that it can perform functions such as enhancing or stabilizing the expression of the encoded peptide or protein, or stabilizing the artificial nucleic acid molecule. Preferably, the ORF and the 3'-UTR element are associated in a 5'→3' direction and / or the 5'-UTR element and the ORF are associated in a 5'→3' direction. Thus, preferably, the artificial nucleic acid molecule typically comprises the structure 5'-[5'-UTR element]-(optional)-linker-ORF-(optional)-linker-[3'-UTR element]-3', wherein the artificial nucleic acid molecule may contain only the 5'-UTR element without the 3'-UTR element, only the 3'-UTR element without the 5'-UTR element, or contain both the 3'-UTR element and the 5'-UTR element. Furthermore, the linker may or may not be present. For example, a linker can be one or more nucleotides, such as a chain of 1-50 or 1-20 nucleotides, for example, containing one or more restriction enzyme recognition sites (restriction sites) or consisting of one or more restriction enzyme recognition sites (restriction sites).

[0180] Preferably, the at least one 3'-UTR element and / or the at least one 5'-UTR element comprises or consists of the following nucleic acid sequences: nucleic acid sequences of the 3'-UTR and / or 5'-UTR of transcripts of genes selected from the group consisting of: ZNF460, TGM2, IL7R, BGN, TK1, RAB3B, CBX6, FZD2, COL8A1, NDUFS7, PHGDH, PLK2, TSPO, PTGS1, FBXO32, NID2, ATP5D, EXOSC4, NOL9, UBB4B, VPS18, ORMDL2, FSCN1, TMEM33, TUBA4A, EMP3, TMEM201, CRIP2, BRAT1, SERPINH1, CD9, DPYSL2, CDK9, TFRC, PSMB3 5'-UTR, FASN, PSMB6, PRSS56, KPNA6, SFT2D2, PARD6B, LPP, SPARC, SCAND1, VASN, SLC26A1, LCLAT1, FBXL18 , SLC35F6, RAB3D, MAP1B, VMA21, CYBA, SEZ6L2, PCOLCE, VTN, ALDH16A1, RAVER1, KPNA6, SERINC5, JUP, CPN2 , CRIP2, EPT1, PNPO, SSSCA1, POLR2L, LIN7C, UQCR10, PYCRL, AMN, MAP1S, NDUFS7, PHGDH, TSPO, ATP5D, EXOS C4, TUBB4B, TUBA4A, EMP3, CRIP2, BRAT1, CD9, CDK9, PSMB3, PSMB6, PRSS56, SCAND1, AMN, CYBA, PCOLCE, MAP 1S, VTN, ALDH16A1 (all preferred for humans) and Dpysl2, Ccnd1, Acox2, Cbx6, Ubc, Ldlr, Nutt22, Pcyox1l, Ankrd1, Tmem37, Tspyl4, Slc7a3, Cst6, Aacs, Nosip, Itga7, Ccnd2, Ebp, Sf3b5, Fasn, Hmgcs1, Osr1, Lmnb1, Vma21, Kif20a, C dca8, Slc7a1, Ubqln2, Prps2, Shmt2, Aurkb, Fignl1, Cad, Anln, Slfn9, Ncaph, Pole, Uhrf1, Gja1, Fam64a, Kif2c, Tspan10, Scand1, Gpr84, Fads3, Cers6, Cxcr4, Gprc5c, Fen1, Cspg4, Mrpl34, Comtd1, Armc6, Emr4,Atp5d, 1110001J03Rik, Csf2ra, Aarsd1, Kif22, Cth, Tpgs1, Ccl17, Alkbh7, Ms4a8a, Acox2, Ubc, Slpi, Pcyox1l, Igf2bp1, Tmem37, Slc7a3, Cst6, Ebp, Sf3b5, Plk1, Cdca8, Kif22, Cad, Cth, Pole, Kif2c, Scand1, Gpr84, Tpgs1, Ccl17, Alkbh7, Ms4a8a, Mrpl34, Comtd1, Armc6, Atp5d, 1110001J03Rik, Nudt22, Aarsd1 (all preferably mouse-derived).

[0181] Preferably, at least one 3'-UTR element and / or at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention comprises a "functional fragment", "functional variant" or "functional fragment of a variant" of the 3'-UTR and / or 5'-UTR of the gene transcript, or is composed of a "functional fragment", "functional variant" or "functional fragment of a variant" of the 3'-UTR and / or 5'-UTR of the gene transcript.

[0182] Preferably, the at least one 5'-UTR element comprises a nucleic acid sequence derived from the 5'-UTR of a gene transcript selected from the group consisting of: ZNF460-5'-UTR, TGM2-5'-UTR, IL7R-5'-UTR, BGN-5'-UTR, TK1-5'-UTR, RAB3B-5'-UTR, CBX6-5'-UTR, FZD2-5'-UTR, COL8A1-5'-UTR, NDUFS7-5'-UTR, PHGDH-5'-UTR, PLK2-5'-UTR, TSPO-5'-UTR, PTGS1-5'-UTR, FBXO32-5'-UTR, NID2-5'- UTR, ATP5D-5'-UTR, EXOSC4-5'-UTR, NOL9-5'-UTR, UBB4B-5'-UTR, VPS18-5'-UTR, ORMDL2-5'-UTR, FSCN1-5'-UTR, TMEM33-5'-UTR, TUBA4A-5'-UTR , EMP3-5'-UTR, TMEM201-5'-UTR, CRIP2-5'-UTR, BRAT1-5'-UTR, SERPINH1-5'-UTR, CD9-5'-UTR, DPYSL2-5'-UTR, CDK9-5'-UTR, TFRC-5'-UTR, PSMB3 5'-UTR, FASN-5'-UTR, PSMB6-5'-UTR, PRSS56-5'-UTR, KPNA6-5'-UTR, SFT2D2-5'-UTR, PARD6B-5'-UTR, LPP-5'-UTR, SPARC-5'-UTR, SCAND1-5'-UT R, VASN-5'-UTR, SLC26A1-5'-UTR, LCLAT1-5'-UTR, FBXL18-5'-UTR, SLC35F6-5'-UTR, RAB3D-5'-UTR, MAP1B-5'-UTR, VMA21-5'-UTR, CYBA-5'-UTR, S EZ6L2-5'-UTR, PCOLCE-5'-UTR, VTN-5'-UTR, ALDH16A1-5'-UTR, RAVER1-5'-UTR, KPNA6-5'-UTR, SERINC5-5'-UTR, JUP-5'-UTR, CPN2-5'-UTR, CRIP2 -5'-UTR, EPT1-5'-UTR, PNPO-5'-UTR, SSSCA1-5'-UTR, POLR2L-5'-UTR, LIN7C-5'-UTR, UQCR10-5'-UTR, PYCRL-5'-UTR, AMN-5'-UTR, MAP1S-5'-UTR,(All are preferably human) and Dpysl2-5'-UTR, Ccnd1-5'-UTR, Acox2-5'-UTR, Cbx6-5'-UTR, Ubc-5'-UTR, Ldlr-5'-UTR, Nudt22-5'-UTR, Pcyox1l-5'-UTR, Ankrd1-5'-UTR, Tmem37-5'-UTR, Tspyl4-5'-UTR, Slc7a3-5'-UTR, Cst6-5'-UTR, Aacs-5'-UTR, Nosip-5'-UTR, Itga7-5'-UTR, Ccnd2-5'-UTR, Ebp-5'-UTR, Sf3b5-5'-UTR, Fasn-5'-UTR, Hmgcs1-5'-UTR, Osr1-5'-UTR, Lmnb1-5'-UTR, Vma21-5'-UTR, Kif20a-5'-UTR, Cdca8-5'-UTR, Slc7a1-5'-UTR, Ubqln2-5'-UTR, Prps2-5'-UTR, Shmt2-5'-UTR, Aurkb-5'-UTR, Fignl1-5'-UTR, Cad-5'-UTR, Anln-5'-UTR, Slfn9-5'-UTR, Ncaph-5'-UTR, Pole-5'-UTR, Uhrf1-5'-UTR, Gja1-5'-UTR, Fam64a-5'-UTR, Kif2c-5'-UTR, Tspan10-5'-UTR, Scand1-5'-UTR, Gpr84-5'-UTR, Fads3-5'-UTR, Cers6-5'-UTR, Cxcr4-5'-UTR, Gprc5c-5'-UTR, Fen1-5'-UTR, Cspg4-5'-UTR, Mrpl34-5'-UTR, Comtd1-5'-UTR, Armc6-5'-UTR, Emr4-5'-UTR, Atp5d-5'-UTR, 1110001J03Rik-5'-UTR, Csf2ra-5'-UTR, Aarsd1-5'-UTR, Kif22-5'-UTR, Cth-5'-UTR, Tpgs1-5'-UTR, Ccl17-5'-UTR, Alkbh7-5'-UTR, Ms4a8a-5'-UTR (all are preferably mouse).

[0183] Preferably, the at least one 3'-UTR element comprises a nucleic acid sequence of a 3'-UTR derived from a gene transcript selected from the group consisting of: NDUFS7-3'-UTR, PHGDH-3'-UTR, TSPO-3'-UTR, ATP5D-3'-UTR, EXOSC4-3'-UTR, TUBB4B-3'-UTR, TUBA4A-3'-UTR, EMP3-3'-UTR, CRIP2-3 '-UTR, BRAT1-3'-UTR, CD9-3'-UTR, CDK9-3'-UTR, PSMB3-3'-UTR, PSMB6-3'-UTR, PRSS56-3'-UTR, SCAND1-3'-UTR, AMN-3'-UTR, CYBA-3'-UTR, PCOLCE-3'-UTR, MAP1S-3'-UTR, VTN-3'-UTR, ALDH16A1-3'-UTR (all preferred for humans) and Acox2-3'-UTR, Ub c-3'-UTR, Slpi-3'-UTR, Pcyox1l-3'-UTR, Igf2bp1-3'-UTR, Tmem37-3'-UTR, Slc7a3-3'-UTR, Cst6-3'-UTR, Ebp-3'- UTR, Sf3b5-3'-UTR, Plk1-3'-UTR, Cdca8-3'-UTR, Kif22-3'-UTR, Cad-3'-UTR, Cth-3'-UTR, Pole-3'-UTR, Kif2c-3'-U TR, Scand1-3'-UTR, Gpr84-3'-UTR, Tpgs1-3'-UTR, Ccl17-3'-UTR, Alkbh7-3'-UTR, Ms4a8a-3'-UTR, Mrpl34-3'-UTR, Comtd1-3'-UTR, Armc6-3'-UTR, Atp5d-3'-UTR, 1110001J03Rik-3'-UTR, Nudt22-3'-UTR, Aarsd1-3'-UTR (all preferably mouse-derived).

[0184] More preferably, the artificial nucleic acid of the present invention comprises (i) at least one preferred 5'-UTR and (i) at least one preferred 3'-UTR.

[0185] In a particularly preferred embodiment, the at least one 5'-UTR element comprises a nucleic acid sequence derived from the 5'-UTR of a gene transcript selected from the group consisting of: ZNF460-5'-UTR, TGM2-5'-UTR, IL7R-5'-UTR, COL8A1-5'-UTR, NDUFS7-5'-UTR, PLK2-5'-UTR, FBXO32-5'-UTR, ATP5D-5'-UTR, TUBB4B-5'-UTR, ORMDL2-5'-UTR, FSCN1-5'-UTR, CD9-5'-UTR, PYSL2-5'-UTR, PSMB3-5'-UTR, PSMB6-5'-UTR. -UTR, KPNA6-5'-UTR, SFT2D2-5'-UTR, LCLAT1-5'-UTR, FBXL18-5'-UTR, SLC35F6-5'-UTR, VMA21-5'-UTR, SEZ6L2-5'-UTR, PCOLCE-5'-UTR, VTN-5'-U TR, ALDH16A1-5'-UTR, KPNA6-5'-UTR, JUP-5'-UTR, CPN2-5'-UTR, PNPO-5'-UTR, SSSCA1-5'-UTR, POLR2L-5'-UTR, LIN7C-5'-UTR, UQCR10-5'-UTR, PYC RL-5'-UTR, AMN-5'-UTR, MAP1S-5'-UTR (all human), Dpysl2-5'-UTR, Acox2-5'-UTR, Ubc-5'-UTR, Nudt22-5'-UTR, Pcyox1l-5'-UTR, Ankrd1-5'-UTR, Ts pyl4-5'-UTR, Slc7a3-5'-UTR, Aacs-5'-UTR, Nosip-5'-UTR, Itga7-5'-UTR, Ccnd2-5'-UTR, Ebp-5'-UTR, Sf3b5-5'-UTR, Fasn-5'-UTR, Hmgcs1-5'-UT R, Osr1-5'-UTR, Lmnb1-5'-UTR, Vma21-5'-UTR, Kif20a-5'-UTR, Cdca8-5'-UTR, Slc7a1-5'-UTR, Ubqln2-5'-UTR, Prps2-5'-UTR, Shmt2-5'-UTR, Fign l1-5'-UTR, Cad-5'-UTR, Anln-5'-UTR, Slfn9-5'-UTR, Ncaph-5'-UTR, Pole-5'-UTR, Uhrf1-5'-UTR, Gja1-5'-UTR, Fam64a-5'-UTR, Tspan10-5'-UTR,Scand1-5'-UTR, Gpr84-5'-UTR, Cers6-5'-UTR, Cxcr4-5'-UTR, Gprc5c-5'-UTR, Fen1-5'-UTR, Cspg4-5'-UTR, Mrpl34-5'-UTR, Comtd1-5'-UTR, Armc6-5'-UTR, Emr4-5'-UTR, Atp5d-5'-UTR, Csf2ra-5'-UTR, Aarsd1-5'-UTR, Cth-5'-UTR, Tpgs1-5'-UTR, Ccl17-5'-UTR, Alkbh7-5'-UTR, Ms4a8a-5'-UTR (all mouse). This indicates that such UTR elements contribute to high translation efficiency.

[0186] In a particularly preferred embodiment, the at least one 3'-UTR element comprises a nucleic acid sequence derived from the 3'-UTR of a gene transcript selected from the group consisting of: NDUFS7-3'-UTR, PHGDH-3'-UTR, TSPO-3'-UTR, ATP5D-3'-UTR, EXOSC4-3'-UTR, TUBB4B-3'-UTR, TUBA4A-3'-UTR, EMP3-3'-UTR, CRIP2-3'-UTR, BRAT1-3'-UTR. CD9-3'-UTR, CDK9-3'-UTR, PSMB3-3'-UTR, PSMB6-3'-UTR, PRSS56-3'-UTR, SCAND1-3'-UTR, AMN-3'-UTR, CYBA-3'-UTR, PCOLCE-3'-UTR, MAP1S-3'-UTR, VTN-3'-UTR, ALDH16A1-3'-UTR (all preferred for humans) and Acox2-3'-UTR, Ubc-3'-UTR, Slpi -3'-UTR, Pcyox1l-3'-UTR, Igf2bp1-3'-UTR, Tmem37-3'-UTR, Slc7a3-3'-UTR, Cst6-3'-UTR, Ebp-3'-UTR, Sf3b5- 3'-UTR, Plk1-3'-UTR, Cdca8-3'-UTR, Kif22-3'-UTR, Cad-3'-UTR, Cth-3'-UTR, Pole-3'-UTR, Kif2c-3'-UTR, Scan d1-3'-UTR, Gpr84-3'-UTR, Tpgs1-3'-UTR, Ccl17-3'-UTR, Alkbh7-3'-UTR, Ms4a8a-3'-UTR, Mrpl34-3'-UTR, Comtd1-3'-UTR, Armc6-3'-UTR, Atp5d-3'-UTR, 1110001J03Rik-3'-UTR, Nudt22-3'-UTR, Aarsd1-3'-UTR (all preferably mouse-derived). This indicates that such UTR elements contribute to high translation efficiency.

[0187] A particularly preferred embodiment is combinable such that the artificial nucleic acid of the present invention comprises (i) at least one particularly preferred 5'-UTR and (i) at least one particularly preferred 3'-UTR.

[0188] The phrase “nucleic acid sequence of the 3'-UTR and / or 5'-UTR of a gene transcript” preferably refers to a nucleic acid sequence based on the 3'-UTR and / or 5'-UTR sequence of a gene transcript or a fragment or portion thereof (preferably a naturally occurring gene or a fragment or portion thereof). In this case, the term “naturally occurring” is used synonymously with the term “wild-type.” The phrase includes sequences corresponding to the entire 3'-UTR sequence and / or the entire 5'-UTR sequence, i.e., the full-length 3'-UTR and / or 5'-UTR sequence of the gene transcript, and sequences corresponding to fragments of the 3'-UTR and / or 5'-UTR sequence of the gene transcript. Preferably, the fragment of the 3'-UTR and / or 5'-UTR of the gene transcript consists of a continuous nucleotide segment corresponding to a continuous nucleotide segment in the full-length 3'-UTR and / or 5'-UTR of the gene transcript, representing at least 5%, 10%, 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90% of the full-length 3'-UTR and / or 5'-UTR of the gene transcript. In the context of this invention, the fragment is preferably a functional fragment as described herein. Preferably, the fragment preserves the regulatory function of the translation of the ORF linked to the 3'-UTR and / or 5'-UTR or a fragment thereof.

[0189] Within the scope of the 3'-UTR and / or 5'-UTR of a gene transcript, the terms "variants of the 3'-UTR and / or 5'-UTR of a gene transcript" and "variants thereof" refer to naturally occurring variants of the 3'-UTR and / or 5'-UTR of a gene transcript, preferably variants of the 3'-UTR and / or 5'-UTR of a vertebrate gene transcript, more preferably variants of the 3'-UTR and / or 5'-UTR of a mammalian gene transcript, and even more preferably variants of the 3'-UTR and / or 5'-UTR of a primate gene, particularly variants of the 3'-UTR and / or 5'-UTR of a human gene transcript as described above. The variants may be modified 3'-UTRs and / or 5'-UTRs of the gene transcript. For example, variants of the 3'-UTR and / or 5'-UTR may exhibit one or more nucleotide deletions, insertions, additions, and / or substitutions compared to the naturally occurring 3'-UTR and / or 5'-UTR from which the variant originates. Preferably, the variants of the 3'-UTR and / or 5'-UTR of the gene transcript are at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, even more preferably at least 80%, even more preferably at least 90%, and most preferably at least 95% identical to the naturally occurring 3'-UTR and / or 5'-UTR from which the variant originates. Preferably, the variant is a functional variant as described herein.

[0190] The phrase “nucleic acid sequence of a variant of the 3'-UTR and / or 5'-UTR of a gene transcript” preferably refers to a nucleic acid sequence based on a variant of the 3'-UTR and / or 5'-UTR of a gene transcript as described above, or a fragment or portion thereof. This phrase includes the entire sequence corresponding to a variant of the 3'-UTR and / or 5'-UTR of the gene transcript, i.e., the full-length variant 3'-UTR sequence and / or the full-length variant 5'-UTR sequence of the gene transcript, and the sequence of a fragment corresponding to a variant 3'-UTR sequence and / or a fragment of a variant 5'-UTR sequence of the gene transcript. Preferably, the fragment of the variant of the 3'-UTR and / or 5'-UTR of the gene transcript consists of a continuous nucleotide segment corresponding to a continuous nucleotide segment in the full-length variant of the 3'-UTR and / or 5'-UTR of the gene transcript, representing at least 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90% of the full-length variant of the 3'-UTR and / or 5'-UTR of the gene transcript. In the context of this invention, the fragment of the variant is preferably a functional fragment of the variant described herein.

[0191] In the context of this invention, the terms "functional variant," "functional fragment," and "functional fragment of a variant" (also referred to as "functional variant fragment") mean a fragment of the 3'-UTR and / or 5'-UTR of a gene transcript, a variant of the 3'-UTR and / or 5'-UTR, or a fragment of a variant of the 3'-UTR and / or 5'-UTR that satisfies at least one function, preferably more than one function, of the naturally occurring 3'-UTR and / or 5'-UTR of the transcript of the gene from which the variant, fragment, or variant fragment originates. The function may, for example, stabilize mRNA and / or enhance, stabilize, and / or prolong protein production from mRNA and / or increase protein expression or total protein production from mRNA (preferably in mammalian cells, such as human cells). Preferably, the function of the 3'-UTR and / or 5'-UTR relates to the translation of proteins encoded by the ORF. More preferably, the function includes enhancing the translation efficiency of the ORF linked to the 3'-UTR and / or 5'-UTR or a fragment or variant thereof. Particularly preferred is that, in the case of the present invention, variants, fragments, and variant fragments satisfy the following requirements: compared to mRNAs containing a reference 3'-UTR and / or a reference 5'-UTR or lacking a 3'-UTR and / or 5'-UTR, preferably in mammalian cells, such as human cells, stabilizing the function of the mRNA; and / or enhancing, stabilizing, and / or prolonging the protein production function from the mRNA compared to mRNAs containing a reference 3'-UTR and / or a reference 5'-UTR or lacking a 3'-UTR and / or 5'-UTR; and / or increasing the protein production function from the mRNA compared to mRNAs containing a reference 3'-UTR and / or a reference 5'-UTR or lacking a 3'-UTR and / or 5'-UTR. The reference 3'-UTR and / or the reference 5'-UTR can be, for example, a 3'-UTR and / or 5'-UTR that is naturally present in combination with an ORF. Furthermore, functional variants, functional fragments, or functional variant fragments of the 3'-UTR and / or 5'-UTR of the gene transcript preferably do not significantly reduce the translation efficiency of mRNAs containing the 3'-UTR and / or 5'-UTR, compared to the wild-type 3'-UTR and / or wild-type 5'-UTR from which the variant, fragment, or variant fragment originates. In the context of this invention, a particularly preferred function of the "functional fragment," "functional variant," or "functional fragment of a variant" of the 3'-UTR and / or 5'-UTR of the gene transcript is produced by expressing an enhancing, stabilizing, and / or elongating protein in mRNA carrying the functional fragment, functional variant, or functional fragment of a variant as described above.

[0192] Preferably, in terms of the translation efficiency exhibited by the naturally occurring 3'-UTR and / or 5'-UTR of the transcript of the gene from which the variant, fragment, or variant fragment originates, the efficiency of one or more functions exhibited by the functional variant, functional fragment, or functional variant fragment, such as providing translation efficiency, increases by at least 5%, more preferably at least 10%, more preferably at least 20%, more preferably at least 30%, more preferably at least 40%, more preferably at least 50%, more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90%.

[0193] In the context of this invention, the fragments of the 3'-UTR and / or 5'-UTR of the gene transcript or variants of the 3'-UTR and / or 5'-UTR of the gene transcript preferably have a length of at least about 3 nucleotides, preferably at least about 5 nucleotides, more preferably at least about 10, 15, 20, 25 or 30 nucleotides, even more preferably at least about 50 nucleotides, and most preferably at least about 70 nucleotides. Preferably, the fragments of the 3'-UTR and / or 5'-UTR of the gene transcript or variants of the gene transcript are functional fragments as described above. In a preferred embodiment, the 3'-UTR and / or 5'-UTR of the gene transcript or its fragments or variants have a length of 3 to about 500 nucleotides, preferably 5 to about 150 nucleotides, more preferably 10 to 100 nucleotides, even more preferably 15 to 90 nucleotides, and most preferably 20 to 70 nucleotides. Typically, 5'-UTR elements and / or 3'-UTR elements are characterized by fewer than 500, 400, 300, 200, 150, or fewer than 100 nucleotides.

[0194] 5'-UTR element

[0195] Preferably, at least one 5'-UTR element comprises or is composed of such a nucleic acid sequence having at least about 1, 2, 3, 4, 5, 10, 15, 20, 30, or 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, or wherein at least one 5'-UTR element comprises or is composed of such a fragment of a nucleic acid sequence having at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, with nucleic acid sequences selected from the group consisting of SEQ ID NO: 1-151 or corresponding RNA sequences.

[0196] In the sequences shown in detail below, SEQ ID NO:1-136 can be considered as wild-type 5'-UTR sequences, and SEQ ID NO:136-151 can be considered as artificial 5'-UTR sequences.

[0197] SEQ ID NO:1–Homo sapiens ZNF460 5'-UTR

[0198] NM_006635.3

[0199]

[0200] SEQ ID NO:2-Homo sapiens TGM2 5'-UTR

[0201] NM_004613.2

[0202]

[0203] SEQ ID NO:3-Homo sapiens IL7R 5'-UTR

[0204] NM_002185.3

[0205]

[0206] SEQ ID NO:4-Homo sapiens BGN 5'-UTR

[0207] NM_001711.4

[0208]

[0209] SEQ ID NO:5-Homo sapiens TK1 5'-UTR

[0210] NM_003258.4

[0211]

[0212] SEQ ID NO:6-Homo sapiens RAB3B 5'-UTR

[0213] NM_002867.3

[0214]

[0215] SEQ ID NO:7-Homo sapiens CBX6 5'-UTR

[0216] NM_014292.3

[0217]

[0218] SEQ ID NO:8-Homo sapiens FZD2 5'-UTR

[0219] NM_001466.3

[0220]

[0221] SEQ ID NO:9-Homo sapiens COL8A1 5'-UTR

[0222] NM_001850.4

[0223]

[0224] SEQ ID NO:10-Homo sapiens NDUFS7 5'-UTR

[0225] NM_024407.4

[0226]

[0227] SEQ ID NO:11-Homo sapiens PHGDH 5'-UTR

[0228] NM_006623.3

[0229]

[0230] SEQ ID NO:12-Homo sapiens PLK2 5'-UTR

[0231] NM_006622.3

[0232]

[0233] SEQ ID NO:13-Homo sapiens TSPO 5'-UTR

[0234] NM_000714.5

[0235]

[0236] SEQ ID NO:14-Homo sapiens PTGS1 5'-UTR

[0237] NM_000962.3

[0238]

[0239] SEQ ID NO:15-Homo sapiens FBXO32 5'-UTR

[0240] NM_058229.3

[0241]

[0242] SEQ ID NO:16-Homo sapiens NID2 5'-UTR

[0243] NM_007361.3

[0244]

[0245] SEQ ID NO:17-Homo sapiens ATP5D 5'-UTR

[0246] NM_001687.4

[0247]

[0248] SEQ ID NO:18-Homo sapiens EXOSC4 5'-UTR

[0249] NM_019037.2

[0250]

[0251] SEQ ID NO:19-Homo sapiens NOL9 5'-UTR

[0252] NM_024654.4

[0253]

[0254] SEQ ID NO:20-Homo sapiens TUBB4B 5'-UTR

[0255] NM_006088.5

[0256]

[0257] SEQ ID NO:21-Homo sapiens VPS18 5'-UTR

[0258] NM_020857.2

[0259]

[0260] SEQ ID NO:22-Homo sapiens ORMDL2 5'-UTR

[0261] NM_014182

[0262]

[0263] SEQ ID NO:23-Homo sapiens FSCN1 5'-UTR

[0264] NM_003088.3

[0265]

[0266] SEQ ID NO:24-Homo sapiens TMEM33 5'-UTR

[0267] NM_018126.2

[0268]

[0269] SEQ ID NO:25-Homo sapiens TUBA4A5'-UTR

[0270] NM_006000.2

[0271]

[0272] SEQ ID NO:26-Homo sapiens EMP3 5'-UTR

[0273] NM_001425.2

[0274]

[0275] SEQ ID NO:27-Homo sapiens TMEM201 5'-UTR

[0276] NM_001010866.3

[0277]

[0278] SEQ ID NO:28-Homo sapiens CRIP2 5'-UTR

[0279] NM_001312.3

[0280]

[0281] SEQ ID NO:29-Homo sapiens BRAT1 5'-UTR

[0282] NM_152743.3

[0283]

[0284] SEQ ID NO:30-Homo sapiens SERPINH1 5'-UTR

[0285] NM_001235.3

[0286]

[0287] SEQ ID NO:31-Homo sapiens CD9 5'-UTR

[0288] NM_001769.3

[0289]

[0290] SEQ ID NO:32-Homo sapiens DPYSL2 5'-UTR

[0291] NM_001386.5

[0292]

[0293] SEQ ID NO:33-Homo sapiens CDK9 5'-UTR

[0294] NM_001261.3

[0295]

[0296] SEQ ID NO:34-Homo sapiens SSCA1 5'-UTR

[0297] NM_006396.1

[0298]

[0299] SEQ ID NO:35-Homo sapiens POLR2L 5'-UTR

[0300] NM_021128.4

[0301]

[0302] SEQ ID NO:36-Homo sapiens LIN7C 5'-UTR

[0303] NM_018362

[0304]

[0305] SEQ ID NO:37-Homo sapiens UQCR10 5'-UTR

[0306] NM_001003684.1

[0307]

[0308] SEQ ID NO:38-Homo sapiens TFRC 5'-UTR

[0309] NM_001128148.1

[0310]

[0311] SEQ ID NO:39-Homo sapiens PSMB3 5'-UTR

[0312] NM_002795.2

[0313]

[0314] SEQ ID NO:40-Homo sapiens FASN 5'-UTR

[0315] NM_004104.4

[0316]

[0317] SEQ ID NO:41-Homo sapiens PSMB6 5'-UTR

[0318] NM_002798.2

[0319]

[0320] SEQ ID NO:42-Homo sapiens PYCRL 5'-UTR

[0321] NM_023078.3

[0322]

[0323] SEQ ID NO:43-Homo sapiens PRSS56 5'-UTR

[0324] NM_001195129.1

[0325]

[0326] SEQ ID NO:44-Homo sapiens KPNA6 5'-UTR

[0327] NM_012316.4

[0328]

[0329] SEQ ID NO:45-Homo sapiens SFT2D2 5'-UTR

[0330] NM_199344.2

[0331]

[0332] SEQ ID NO:46-Homo sapiens PARD6B 5'-UTR

[0333] NM_032521.2

[0334]

[0335] SEQ ID NO:47-Homo sapiens LPP 5'-UTR

[0336]

[0337] SEQ ID NO:48-Homo sapiens SPARC 5'-UTR

[0338] NM_003118.3

[0339]

[0340] SEQ ID NO:49-Homo sapiens SCAND1 5'-UTR

[0341]

[0342] SEQ ID NO:50-Homo sapiens VASN 5'-UTR

[0343] NM_138440.2

[0344]

[0345] SEQ ID NO:51-Homo sapiens SLC26A1 5'-UTR

[0346] NM_022042.3

[0347]

[0348] SEQ ID NO:52-Homo sapiens LCLAT1 5'-UTR

[0349] NM_182551.3

[0350]

[0351] SEQ ID NO:53-Homo sapiens FBXL18 5'-UTR

[0352] NM_024963.4

[0353]

[0354] SEQ ID NO:54-Homo sapiens SLC35F6 5'-UTR

[0355] NM_017877.3

[0356]

[0357] SEQ ID NO:55-Homo sapiens RAB3D 5'-UTR

[0358] NM_004283.3

[0359]

[0360] SEQ ID NO:56-Homo sapiens MAP1B 5'-UTR

[0361] NM_005909.3

[0362]

[0363] SEQ ID NO:57-Homo sapiens VMA21 5'-UTR

[0364] NM_001017980.3

[0365]

[0366] SEQ ID NO:58-Homo sapiens AMN 5'-UTR

[0367] NM_030943.3

[0368]

[0369] SEQ ID NO:59-Homo sapiens CYBA5'-UTR

[0370] NM_000101.3

[0371]

[0372] SEQ ID NO:60-Homo sapiens SEZ6L2 5'-UTR

[0373] NM_012410.3

[0374]

[0375] SEQ ID NO:61-Homo sapiens PCOLCE 5'-UTR

[0376] NM_002593.3

[0377]

[0378] SEQ ID NO:62-Homo sapiens MAP1S 5'-UTR

[0379] NM_018174.4

[0380]

[0381] SEQ ID NO:63-Homo sapiens VTN 5'-UTR

[0382] NM_000638.3

[0383]

[0384] SEQ ID NO:64-Homo sapiens ALDH16A1 5'-UTR

[0385] NM_001145396.1

[0386]

[0387] SEQ ID NO:65-Homo sapiens RAVER1 5'-UTR

[0388] NM_133452.2

[0389]

[0390] SEQ ID NO:66-Homo sapiens KPNA6 5'-UTR (identical to SEQ ID NO:44)

[0391] NM_012316.4

[0392]

[0393] SEQ ID NO:67 - Homo sapiens SERINC5 5'-UTR

[0394] NM_178276.5

[0395]

[0396] SEQ ID NO:68 - Homo sapiens JUP 5'-UTR

[0397] NM_002230.2

[0398]

[0399] SEQ ID NO:69 - Homo sapiens CPN2 5'-UTR

[0400] NM_001080513.2

[0401]

[0402] SEQ ID NO:70 - Homo sapiens CRIP2 5'-UTR

[0403] NM_001312.3

[0404]

[0405] SEQ ID NO:71 - Homo sapiens EPT1 5'-UTR

[0406] NM_033505.2

[0407]

[0408] SEQ ID NO:72 - Homo sapiens PNPO 5'-UTR

[0409] NM_018129.3

[0410]

[0411] SEQ ID NO:73 – Mus musculus Dpysl2 5'-UTR

[0412] NM_009955.3

[0413]

[0414] SEQ ID NO:74 – Mus musculus Ccnd1 5'-UTR

[0415] NM_007631.2

[0416]

[0417] SEQ ID NO:75–House Mouse Acox2 5'-UTR

[0418] NM_001161667.1

[0419]

[0420] SEQ ID NO:76–House Mouse Cbx6(Npcd)5'-UTR

[0421] NM_001013360.2

[0422]

[0423] SEQ ID NO:77–House Mouse Ubc 5'-UTR

[0424]

[0425] SEQ ID NO:78–House Mouse Ldlr 5'-UTR

[0426] NM_001252659.1

[0427]

[0428] SEQ ID NO:79–House Mouse Nudt22 5'UTR

[0429]

[0430] SEQ ID NO:80–Pcyox1l 5'-UTR

[0431] NM_172832.4

[0432]

[0433] SEQ ID NO:81–Ankrd1 5'-UTR

[0434] NM_013468.3

[0435]

[0436] SEQ ID NO:82–Tmem37 5'-UTR

[0437] NM_019432.2

[0438]

[0439] SEQ ID NO:83–Tspyl4 5'-UTR of the house mouse

[0440] NM_030203.2

[0441]

[0442] SEQ ID NO:84–House Mouse Slc7a3 5'-UTR

[0443]

[0444] SEQ ID NO:85–House Mouse Cst6 5'-UTR

[0445]

[0446] SEQ ID NO:86–Aacs 5'-UTR

[0447] NM_030210.1

[0448]

[0449] SEQ ID NO:87–Nosip 5'-UTR

[0450] NM_001163684.1

[0451]

[0452] SEQ ID NO:88–House Mouse Itga7 5'-UTR

[0453] NM_008398.2

[0454]

[0455] SEQ ID NO:89–House Mouse Ccnd2 5'-UTR

[0456] NM_009829.3

[0457]

[0458] SEQ ID NO:90–House Mouse Ebp 5'-UTR

[0459]

[0460] SEQ ID NO:91–House Mouse Sf3b5 5'-UTR

[0461] NM_009829.3

[0462]

[0463] SEQ ID NO:92–House Mouse Fasn 5'-UTR

[0464]

[0465] SEQ ID NO:93–House Mouse Hmgcs1 5'-UTR

[0466]

[0467] SEQ ID NO:94–House Mouse Osr1 5'-UTR

[0468] NM_011859.3

[0469]

[0470] SEQ ID NO:95–Lmnb1 5'-UTR

[0471] NM_011121.3

[0472]

[0473] SEQ ID NO:96–House Mouse Aarsd1 5'-UTR

[0474] NM_144829.1

[0475]

[0476] SEQ ID NO:97–Vma21 5'-UTR

[0477] NM_001081356.2

[0478]

[0479] SEQ ID NO:98–Kif20a 5'-UTR

[0480]

[0481] SEQ ID NO:99–Cdca8 5'-UTR

[0482] NM_026560.4

[0483]

[0484] SEQ ID NO:100–House Mouse Slc7a1 5'-UTR

[0485] NM_007513.4

[0486]

[0487]

[0488] SEQ ID NO:101–House Mouse Ubqln2 5'-UTR

[0489] NM_018798.2

[0490]

[0491] SEQ ID NO:102–House Mouse Prps2 5'-UTR

[0492] NM_026662.4

[0493]

[0494] SEQ ID NO:103–Shmt2 5'-UTR

[0495] NM_026662.4

[0496]

[0497] SEQ ID NO:104–Kif22 5'-UTR

[0498] NM_145588.1

[0499]

[0500] SEQ ID NO:105–House Mouse Aurkb 5'-UTR

[0501] NM_011496.1

[0502]

[0503] SEQ ID NO:106–Fignl1 5'-UTR

[0504]

[0505] SEQ ID NO:107–Cad 5'-UTR

[0506] NM_023525.2

[0507]

[0508] SEQ ID NO:108–House Mouse Anln 5'-UTR

[0509] NM_028390.3

[0510]

[0511] SEQ ID NO:109–House Mouse Slfn9 5'-UTR

[0512]

[0513] SEQ ID NO:110–Ncaph 5'-UTR

[0514] NM_144818.3

[0515]

[0516] SEQ ID NO:111–House Mouse Cth 5'-UTR

[0517] NM_145953.2

[0518]

[0519] SEQ ID NO:112–Pole 5'-UTR

[0520] NM_011132.2

[0521]

[0522] SEQ ID NO:113–House Mouse Uhrf1 5'-UTR

[0523] NM_001111079.1

[0524]

[0525] SEQ ID NO:114–Gja1 5'-UTR

[0526] NM_010288.3

[0527]

[0528] SEQ ID NO:115–House Mouse Fam64a 5'-UTR

[0529]

[0530] SEQ ID NO:116–Kif2c 5'-UTR

[0531] NM_134471.4

[0532]

[0533] SEQ ID NO:117–House Mouse Tspan10 5'-UTR

[0534] NM_145363.2

[0535]

[0536] SEQ ID NO:118–Scand1 5'-UTR

[0537]

[0538] SEQ ID NO:119–House Mouse Gpr84 5'-UTR

[0539] NM_030720.1

[0540]

[0541] SEQ ID NO:120–House Mouse Tpgs1 5'-UTR

[0542] BC138516.1

[0543]

[0544] SEQ ID NO:121–House Mouse Ccl17 5'-UTR

[0545] NM_011332.3

[0546]

[0547] SEQ ID NO:122–House Mouse Fads3 5'-UTR

[0548] NM_021890.3

[0549]

[0550] SEQ ID NO:123–Cers6 5'-UTR

[0551] NM_172856.3

[0552]

[0553] SEQ ID NO:124–House Mouse Alkbh7 5'-UTR

[0554] NM_027372.1

[0555]

[0556] SEQ ID NO:125–House Mouse Ms4a8a 5'-UTR

[0557]

[0558] SEQ ID NO:126–House Mouse Cxcr4 5'-UTR

[0559] NM_009911.3

[0560]

[0561] SEQ ID NO:127–House Mouse Gprc5c 5'-UTR

[0562]

[0563] SEQ ID NO:128–House Mouse Fen1 5'-UTR

[0564] NM_007999.4

[0565]

[0566] SEQ ID NO:129–House Mouse Cspg4 5'-UTR

[0567] NM_139001.2

[0568]

[0569] SEQ ID NO:130–Mrpl34 5'-UTR

[0570] NM_053162.2

[0571]

[0572] SEQ ID NO:131–House Mouse Comtd1 5'-UTR

[0573]

[0574] SEQ ID NO:132–Armc6 5'-UTR

[0575] NM_133972.2

[0576]

[0577] SEQ ID NO:133–House Mouse Emr4 5'-UTR

[0578] NM_139138.3

[0579]

[0580] SEQ ID NO:134–House Mouse Atp5d 5'-UTR

[0581] NM_025313.2

[0582]

[0583] SEQ ID NO:135–House Mouse 1110001J03Rik 5'-UTR

[0584]

[0585] SEQ ID NO:136–Csf2ra 5'-UTR

[0586]

[0587] Some of these wild-type 5'-UTR elements differ from publicly available 5'-UTR sequences identified in NCBI GenBank, as shown in Table 1 (Example 1). This table reflects the inventors' sequencing results.

[0588] Therefore, the present invention provides a 5'-UTR element selected from the group consisting of: SEQ ID NO:45, SEQ ID NO:47, SEQ ID NO:49, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:98, SEQ ID NO:106, SEQ ID NO:109, SEQ ID NO:115, SEQ ID NO:118, SEQ ID NO:125, SEQ ID NO:127, SEQ ID NO:131, SEQ ID NO:135, and can be used in any aspect of the present invention.

[0589] Examples of artificial 5'-UTR elements used in this invention

[0590] In some embodiments, the 5'-UTR element described in this invention differs from the wild-type 5'-UTR element. Such 5'-UTR elements are referred to as "artificial 5'-UTR elements".

[0591] Preferably, the artificial 5'-UTR element exhibits a degree of sequence identity with the wild-type 5'-UTR element, wherein the sequence identity is (a) less than 100% and (B) greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90%, or greater than 95%, and the wild-type 5'-UTR element is selected from the group consisting of SEQ ID NOs: 1-136. Typically, the artificial 5'-UTR element differs from the wild-type 5'-UTR element upon which it is based, wherein at least one nucleotide, such as two, three, four, five, six, seven, eight, nine, ten, or more than ten nucleotides, is exchanged. For example, the nucleotide exchange may be recommended in cases where the wild-type 5'-UTR element contains nucleotide elements considered unfavorable. For example, in some embodiments, the nucleotide elements considered unfavorable are selected from: (i) internal ATG triplets (i.e., ATG triplets other than the start codon of the read frame of the nucleic acid of the present invention) or (ii) restriction enzyme recognition sites (cleavage sites), particularly restriction enzyme recognition sites (cleavage sites) recognized (cleaved) by the restriction enzymes of the process used to prepare (clone) the artificial nucleic acid of the present invention. Therefore, specific bases can be specifically introduced (for each wild-type base, by exchange, preferably substitution) so that the artificial 5'-UTR element does not contain the nucleotide elements considered unfavorable. In a specific embodiment, the artificial 5'-UTR element is selected from the group consisting of SEQ ID NOs:137-151:

[0592] SEQ ID NO:137 – Artificial sequence (based on SEQ ID NO:3)

[0593]

[0594] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that AAGCTT is... TAGCTT Replacement, the above is underlined for emphasis]

[0595] SEQ ID NO:138 – Artificial sequence (based on SEQ ID NO:9)

[0596]

[0597] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is preferably... TAG Replacement, the above is underlined for emphasis]

[0598] SEQ ID NO:139 – Artificial sequence (based on SEQ ID NO:14)

[0599]

[0600] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is preferably... TAG Replacement, the above is underlined for emphasis]

[0601] SEQ ID NO:140 – Artificial sequence (based on SEQ ID NO:17)

[0602]

[0603] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that GACGTC is... GTCGTC Replacement, the above is underlined for emphasis]

[0604] SEQ ID NO:141 – Artificial sequence (based on SEQ ID NO:18)

[0605]

[0606] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is... TAG Replacement, the above is underlined for emphasis]

[0607] SEQ ID NO:142 – Artificial sequence (based on SEQ ID NO:29)

[0608]

[0609] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that the two ATGs are respectively... TAG Replace the two underlines above; they indicate emphasis.

[0610] SEQ ID NO:143 – Artificial sequence (based on SEQ ID NO:31)

[0611]

[0612] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is... TAG Replacement, the above is underlined for emphasis]

[0613] SEQ ID NO:144 – Artificial sequence (based on SEQ ID NO:42)

[0614]

[0615] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that AGATCT is... TGATCT Replacement, the above is underlined for emphasis]

[0616] SEQ ID NO:145 – Artificial sequence (based on SEQ ID NO:43)

[0617]

[0618] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that (i) ATG is... TAG Replace, and (ii) GAATTC is CAATTC Replacement, the above are underlined for emphasis.

[0619] SEQ ID NO:146 – Artificial sequence (based on SEQ ID NO:52)

[0620]

[0621] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that the three ATGs are respectively... TAG Replacement, the above are underlined for emphasis.

[0622] SEQ ID NO:147 – Artificial sequence (based on SEQ ID NO:109)

[0623]

[0624] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that GAATTC is... CAATTC Replacement, the above is underlined for emphasis]

[0625] SEQ ID NO:148 – Artificial sequence (based on SEQ ID NO:110)

[0626]

[0627] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that GACGTC is... CACGTC Replacement, the above is underlined for emphasis]

[0628] SEQ ID NO:149 – Artificial sequence (based on SEQ ID NO:119)

[0629]

[0630] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is... TAG Replacement, the above is underlined for emphasis]

[0631] SEQ ID NO:150 – Artificial sequence (based on SEQ ID NO:120)

[0632]

[0633] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is... TAG Replacement, the above is underlined for emphasis]

[0634] SEQ ID NO:151 – Artificial sequence (based on SEQ ID NO:136)

[0635]

[0636] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that ATG is... TAG Replacement, the above is underlined for emphasis]

[0637] 3'-UTR element

[0638] Preferably, the at least one 3'-UTR element comprises or is composed of a nucleic acid sequence in which the nucleic acid sequence has at least about 1, 2, 3, 4, 5, 10, 15, 20, 30, or 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, respectively; or wherein the at least one 3'-UTR element comprises or is composed of a fragment of a nucleic acid sequence in which the fragment of the nucleic acid sequence has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, respectively.

[0639] In the sequences shown in detail below, SEQ ID NO:152-203 can be considered as wild-type 3'-UTR sequences, and SEQ ID NO:204 can be considered as artificial 5'-UTR sequences.

[0640] SEQ ID NO:152-Homo sapiens NDUFS7 3'-UTR

[0641] NM_024407.4

[0642]

[0643] SEQ ID NO:153-Homo sapiens PHGDH 3'-UTR

[0644] NM_006623.3

[0645]

[0646] SEQ ID NO:154-Homo sapiens TSPO 3'-UTR

[0647] NM_000714.5

[0648]

[0649] SEQ ID NO:155-Homo sapiens ATP5D 3'-UTR

[0650] NM_001687.4

[0651]

[0652] SEQ ID NO:156-Homo sapiens EXOSC4 3'-UTR

[0653] NM_019037.2

[0654]

[0655] SEQ ID NO:157-Homo sapiens TUBB4B 3'-UTR

[0656] NM_006088.5

[0657]

[0658] SEQ ID NO:158-Homo sapiens TUBA4A3'-UTR

[0659] NM_006000.1

[0660]

[0661] SEQ ID NO:159-Homo sapiens EMP3 3'-UTR

[0662]

[0663] SEQ ID NO:160-Homo sapiens CRIP2 3'-UTR

[0664] NM_001312.3

[0665]

[0666] SEQ ID NO:161-Homo sapiens BRAT1 3'-UTR

[0667] NM_152743.3

[0668]

[0669] SEQ ID NO:162-Homo sapiens CD9 3'-UTR

[0670] NM_001769.3

[0671]

[0672] SEQ ID NO:163-Homo sapiens CDK9 3'-UTR

[0673] NM_001261.3

[0674]

[0675] SEQ ID NO:164-Homo sapiens PSMB3 3'-UTR

[0676] NM_002795.2

[0677]

[0678] SEQ ID NO:165-Homo sapiens PSMB6 3'-UTR

[0679] NM_002798.2

[0680]

[0681] SEQ ID NO:166-Homo sapiens PRSS56 3'-UTR

[0682] NM_001195129.1

[0683]

[0684] SEQ ID NO:167-Homo sapiens SCAND1 3'-UTR

[0685] NM_016558

[0686]

[0687] SEQ ID NO:168-Homo sapiens AMN 3'-UTR

[0688] NM_030943.3

[0689]

[0690] SEQ ID NO:169-Homo sapiens CYBA 3'-UTR

[0691] NM_000101.3

[0692]

[0693] SEQ ID NO:170-Homo sapiens PCOLCE 3'-UTR

[0694] NM_002593.3

[0695]

[0696] SEQ ID NO:171-Homo sapiens MAP1S 3'-UTR

[0697] NM_018174.4

[0698]

[0699] SEQ ID NO:172-Homo sapiens VTN 3'-UTR

[0700] NM_000638.3

[0701]

[0702] SEQ ID NO:173-Homo sapiens ALDH16A1 3'-UTR

[0703] NM_001145396.1

[0704]

[0705] SEQ ID NO:174–House Mouse Acox2 3'-UTR

[0706] NM_001161667.1

[0707]

[0708] SEQ ID NO:175–House Mouse Ubc 3'-UTR

[0709] BC006680.1

[0710]

[0711] SEQ ID NO:176–Slpi 3'-UTR of the house mouse

[0712] NM_011414.3

[0713]

[0714] SEQ ID NO:177–Nudt22 3'-UTR

[0715] NM_026675.2

[0716]

[0717] SEQ ID NO:178–Pcyox1l 3'-UTR

[0718] NM_172832.4

[0719]

[0720] SEQ ID NO:179–House Mouse Igf2bp1 3'-UTR

[0721]

[0722] SEQ ID NO:180–Tmem37 3'-UTR

[0723]

[0724] SEQ ID NO:181–House Mouse Slc7a3 3'-UTR(240)

[0725] NM_007515

[0726]

[0727] SEQ ID NO:182–House Mouse Cst6 3'-UTR

[0728]

[0729] SEQ ID NO:183–House Mouse Ebp 3'-UTR

[0730]

[0731] SEQ ID NO:184–House Mouse Sf3b5 3'-UTR

[0732] NM_009829.3

[0733]

[0734] SEQ ID NO:185–House Mouse Plk1 3'-UTR

[0735] NM_011121.3

[0736]

[0737] SEQ ID NO:186–House Mouse Aarsd1 3'-UTR

[0738] NM_144829.1

[0739]

[0740] SEQ ID NO:187–House Mouse Cdca8 3'-UTR

[0741] NM_026560.4

[0742]

[0743] SEQ ID NO:188–Kif22 3'-UTR

[0744] NM_145588.1

[0745]

[0746] SEQ ID NO:189–Cad 3'-UTR

[0747] NM_023525.2

[0748]

[0749] SEQ ID NO:190–House Mouse Cth 3'-UTR

[0750] NM_145953.2

[0751]

[0752] SEQ ID NO:191–Pole 3'-UTR

[0753] NM_011132.2

[0754]

[0755] SEQ ID NO:192–Kif2c 3'-UTR

[0756] NM_134471.4

[0757]

[0758] SEQ ID NO:193–Scand1 3'-UTR

[0759] NM_020255.3

[0760]

[0761] SEQ ID NO:194–House Mouse Gpr84 3'-UTR

[0762] NM_030720.1

[0763]

[0764] SEQ ID NO:195–House Mouse Tpgs1 3'-UTR

[0765] NM_148934.2

[0766]

[0767] SEQ ID NO:196–House Mouse Ccl17 3'-UTR

[0768] NM_011332.3

[0769]

[0770] SEQ ID NO:197–House Mouse Alkbh7 3'-UTR

[0771]

[0772] SEQ ID NO:198–House Mouse Ms4a8a 3'-UTR

[0773] NM_022430.2

[0774]

[0775] SEQ ID NO:199–Mrpl34 3'-UTR

[0776] NM_053162.2

[0777]

[0778] SEQ ID NO:200–House Mouse Comtd1 3'-UTR

[0779]

[0780] SEQ ID NO:201–Armc6 3'-UTR

[0781] NM_133972.2

[0782]

[0783] SEQ ID NO:202–House Mouse Atp5d 3'-UTR

[0784] NM_025313.2

[0785]

[0786] SEQ ID NO:203–House Mouse 1110001J03Rik 3'-UTR

[0787] NM_025363.3

[0788]

[0789] Some of these wild-type 3'-UTR elements differ from publicly available 3'-UTR sequences identified by NCBI GenBank, see Table 2 (Example 1). This table reflects the inventors' sequencing results.

[0790] Therefore, the present invention provides a 3'-UTR element selected from the group consisting of: SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:197, SEQ ID NO:200, and can be used in any aspect of the present invention.

[0791] Examples of artificial 3'-UTR elements used in this invention

[0792] In some embodiments, the 3'-UTR element described in this invention differs from the wild-type 3'-UTR element. Such 3'-UTR elements are referred to as "artificial 3'-UTR elements".

[0793] Preferably, the artificial 3'-UTR element exhibits a certain degree of sequence identity with the wild-type 3'-UTR element, wherein the sequence identity is (a) less than 100% and (B) greater than 10%, greater than 20%, greater than 30%, greater than 40%, greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90%, or greater than 95%, and the wild-type 3'-UTR element is selected from the group consisting of SEQ ID NOs: 152-203. Typically, the artificial 3'-UTR element differs from the wild-type 3'-UTR element upon which it is based, wherein at least one nucleotide, such as two, three, four, five, six, seven, eight, nine, ten, or more than ten nucleotides, is exchanged. For example, the nucleotide exchange may be recommended in cases where the wild-type 3'-UTR element contains nucleotide elements considered unfavorable. For example, in some embodiments, the nucleotide elements considered unfavorable are restriction enzyme recognition sites (cleavage sites), particularly those recognized (cleavable) by restriction enzymes in the process of preparing (cloning) the artificial nucleic acid of the present invention. Therefore, specific bases can be specifically introduced (for each wild-type base, by exchange, preferably substitution) so that the artificial 3'-UTR element does not contain the nucleotide elements considered unfavorable. In a specific embodiment, the artificial 3'-UTR element is SEQ ID NO:204:

[0794] SEQ ID NO:204 – Artificial sequence (based on SEQ ID NO:192)

[0795]

[0796] This sequence corresponds to the wild-type sequence that forms its basis, the difference being that AGATCT is... TGATCT Replacement, the above is underlined for emphasis]

[0797] Novel 5'-UTR element and novel 3'-UTR element

[0798] This invention also provides novel 5'-UTR elements and novel 3'-UTR elements, namely, 5'-UTR elements and 3'-UTR elements discovered by the inventors that are expressed in human cells and mouse cells, respectively, but not known in public databases (see Example 1). In some embodiments of this invention, any 5'-UTR selected from the group consisting of SEQ ID NO:45, SEQ ID NO:47, and SEQ ID NO:49 may be preferred. In some embodiments of this invention, any 5'-UTR selected from the group consisting of SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:98, SEQ ID NO:106, SEQ ID NO:109, SEQ ID NO:115, SEQ ID NO:118, SEQ ID NO:125, SEQ ID NO:127, SEQ ID NO:131, and SEQ ID NO:135 may be preferred. In some embodiments of the present invention, any 3'-UTR selected from the group consisting of SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:197, and SEQ ID NO:200 may be preferred.

[0799] Description of 5'-UTR element and preferred 3'-UTR element

[0800] Preferably, at least one 3'-UTR element of the artificial nucleic acid molecule of the present invention comprises or is composed of a nucleic acid sequence that has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with the 3'-UTR of a gene selected from the group consisting of: NDUFS7-3'-UTR, PHGDH- 3'-UTR, TSPO-3'-UTR, ATP5D-3'-UTR, EXOSC4-3'-UTR, TUBB4B-3'-UTR, TUBA4A-3'-UTR, EMP3-3'-UTR, CRIP2-3'-UTR , BRAT1-3'-UTR, PSMB3-3'-UTR, PSMB6-3'-UTR, SCAND1-3'-UTR, AMN-3'-UTR, CYBA-3'-UTR, PCOLCE-3'-UTR, MAP1S-3 '-UTR, VTN-3'-UTR, ALDH16A1-3'-UTR (all human), Acox2-3'-UTR, Ubc-3'-UTR, Slpi-3'-UTR, Igf2bp1-3'-UTR, Tmem37-3'-UTR, Slc7a3-3'-UTR, Cst6-3'-UTR, Ebp-3'-UTR, Sf3b5-3'-UTR, Cdca8-3'-UTR, Kif22-3'-UTR, Cad-3'-UTR, Pole- 3'-UTR, Kif2c-3'-UTR, Scand1-3'-UTR, Gpr84-3'-UTR, Tpgs1-3'-UTR, Ccl17-3'-UTR, Alkbh7-3'-UTR, Ms4a8a-3'-UTR, Mrpl34-3'-UTR, Comtd1-3'-UTR, Armc6-3'-UTR, Atp5d-3'-UTR, 1110001J03Rik-3'-UTR, Nudt22-3'-UTR (all from mice).More preferably, at least one 3'-UTR element of the artificial nucleic acid molecule of the present invention comprises or is composed of a nucleic acid sequence that has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with a sequence selected from the group consisting of the following or corresponding RNA sequences: SEQ ID NO:152, SEQ ID NO:153, SEQ ID NO:154, SEQ ID NO:155, SEQ ID NO:156, SEQ ID NO:157, SEQ ID NO:158, SEQ ID NO:159, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171 ...9, SEQ ID NO:169, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO NO:172, SEQ ID NO:173; SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:184, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO: 189, SEQ ID NO: 191, SEQ ID NO: 204 (or SEQ ID NO: 192), SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO:201, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:177.

[0801] Preferably, at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention comprises or is composed of a nucleic acid sequence that has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with the 5'-UTR sequences of the transcripts of the following: ZNF460-5'-UTR, TGM2-5'-UTR, IL7R-5'-UTR, COL8A1-5'-UTR, NDUFS7-5'-UTR, PLK2-5'-UTR, FBXO32 -5'-UTR, ATP5D-5'-UTR, TUBB4B-5'-UTR, ORMDL2-5'-UTR, FSCN1-5'-UTR, CD9-5'-UTR, PYSL2-5'-UTR, PSMB3-5'-UTR, PSMB6-5'-UTR, KPNA6-5'-UTR, SFT2D2-5'-UTR, LCLAT1-5'-UTR, FBXL18-5'-UTR, SLC35F6-5'-UTR, VMA21-5'-UTR, SEZ6L2-5'-UTR, PCOLCE-5'-UTR, VTN-5'-UTR, ALDH16A1-5'-UTR, KPNA6-5'-UTR, JUP-5'-UTR, CPN2-5'-UTR, PNPO-5'-UTR, SSSCA1-5'-UTR, POLR2L-5'-UTR, LIN7C-5'-UTR, UQCR10-5'-UTR, PYCRL-5'-UTR, AMN-5'-UT R, MAP1S-5'-UTR (all human), Dpysl2-5'-UTR, Acox2-5'-UTR, Ubc-5'-UTR, Nudt22-5'-UTR, Pcyox1l-5'-UTR, Ankrd1-5'-UTR, Tspyl4-5'-UTR, Slc7a3-5' -UTR, Aacs-5'-UTR, Nosip-5'-UTR, Itga7-5'-UTR, Ccnd2-5'-UTR, Ebp-5'-UTR, Sf3b5-5'-UTR, Fasn-5'-UTR, Hmgcs1-5'-UTR, Osr1-5'-UTR, Lmnb1-5 '-UTR, Vma21-5'-UTR, Kif20a-5'-UTR, Cdca8-5'-UTR, Slc7a1-5'-UTR, Ubqln2-5'-UTR, Prps2-5'-UTR, Shmt2-5'-UTR, Fignl1-5'-UTR, Cad-5'-UTR,Anln-5'-UTR, Slfn9-5'-UTR, Ncaph-5'-UTR, Pole-5'-UTR, Uhrf1-5'-UTR, Gja1-5'-UTR, Fam64a-5'-UTR, T span10-5'-UTR, Scand1-5'-UTR, Gpr84-5'-UTR, Cers6-5'-UTR, Cxcr4-5'-UTR, Gprc5c-5'-UTR, Fen1-5'-UT R, Cspg4-5'-UTR, Mrpl34-5'-UTR, Comtd1-5'-UTR, Armc6-5'-UTR, Emr4-5'-UTR, Atp5d-5'-UTR, Csf2ra-5'-UTR, Aarsd1-5'-UTR, Cth-5'-UTR, Tpgs1-5'-UTR, Ccl17-5'-UTR, Alkbh7-5'-UTR, Ms4a8a-5'-UTR (all from mice). Most preferably, at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention comprises or is composed of a nucleic acid sequence having at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with the sequences of the following: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:39, SEQ ID NO:41, ...0, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:1 NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:66, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:72, SEQ ID NO:34, SEQ ID NO:35,SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:144 (or SEQ ID NO:42), SEQ ID NO:58, SEQ ID NO:62; SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147 (or SEQ ID NO:109), SEQ ID NO:148 (or SEQ ID NO:110), SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149 (or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:151 (or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150 (based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125, or the corresponding RNA sequences.,

[0802] At least one 3'-UTR element of the artificial nucleic acid molecule of the present invention may also comprise or consist of a fragment of such a nucleic acid sequence: the fragment of the nucleic acid sequence has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with the nucleic acid sequence of the 3'-UTR of the transcript of the gene, such as the sequence of SEQ ID NOs: 152 to 204, wherein the fragment is preferably a functional fragment or a functional variant fragment as described above. The fragment preferably has a length of at least about 3 nucleotides, preferably at least about 5 nucleotides, more preferably at least about 10, 15, 20, 25 or 30 nucleotides, even more preferably at least about 50 nucleotides, and most preferably at least about 70 nucleotides. In a preferred embodiment, the fragment or a variant thereof has a length of 3 to about 500 nucleotides, preferably 5 to about 150 nucleotides, more preferably 10 to 100 nucleotides, even more preferably 15 to 90 nucleotides, and most preferably 20 to 70 nucleotides.Preferably, the variant, fragment, or variant fragment is a functional variant, functional fragment, or functional variant fragment of the 3'-UTR, to produce a protein extension efficiency of at least 30%, preferably at least 40%, more preferably at least 50%, more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90% from the protein produced by the artificial nucleic acid molecule of the present invention: SEQ ID NO:152, SEQ ID NO:153, SEQ ID NO:154, SEQ ID NO:155, SEQ ID NO:156, SEQ ID NO:157, SEQ ID NO:158, SEQ ID NO:159, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171 ...9, SEQ ID NO:169, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:169, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO: NO:172, SEQ ID NO:173; SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:184, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO: 189, SEQ ID NO: 191, SEQ ID NO: 204 (or SEQ ID NO: 192), SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO:201, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:177.

[0803] At least one 5'-UTR element of the artificial nucleic acid molecule of the present invention may also comprise or be composed of a fragment of such a nucleic acid sequence, wherein the nucleic acid sequence fragment has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99%, and most preferably 100% identity with the nucleic acid sequence of the 5'-UTR of the transcript of the gene, such as the sequence of SEQ ID NO: 1 to 151. The fragment is preferably a functional fragment or a functional variant fragment as described above. The fragment preferably exhibits a length of at least about 3 nucleotides, preferably at least about 5 nucleotides, more preferably at least about 10, 15, 20, 25, or 30 nucleotides, even more preferably at least about 50 nucleotides, and most preferably at least about 70 nucleotides. In a preferred embodiment, the fragment or a variant thereof has a length of 3 to about 500 nucleotides, preferably 5 to about 150 nucleotides, more preferably 10 to 100 nucleotides, even more preferably 15 to 90 nucleotides, and most preferably 20 to 70 nucleotides. Preferably, the variant, fragment, or variant fragment is a functional variant, functional fragment, or functional variant fragment of the 5'-UTR, increasing the protein production efficiency of the artificial nucleic acid molecule comprising a nucleic acid sequence selected from the group consisting of at least 30%, preferably at least 40%, more preferably at least 50%, more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90% of the protein production efficiency from the artificial nucleic acid molecule of the present invention: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:137 ...47, SEQ ID NO:140, SEQ ID NO:140, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:66,SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:72, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:144 (or SEQ ID NO:42), SEQ ID NO:58, SEQ ID NO:62; SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147 (or SEQ ID NO:109), SEQ ID NO:148 (or SEQ ID NO:110), SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149 (or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:151 (or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150 (based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125.,

[0804] Further optimized implementation scheme

[0805] Preferably, the at least one 3'-UTR element and / or the at least one 5'-UTR element of the artificial nucleic acid molecule of the present invention presents a length of at least about 3 nucleotides, preferably at least about 5 nucleotides, more preferably at least about 10, 15, 20, 25 or 30 nucleotides, even more preferably at least about 50 nucleotides, and most preferably at least about 70 nucleotides. The upper limit of the length of the at least one 3'-UTR element and / or the at least one 5'-UTR element may be 500 nucleotides or less, for example 400, 300, 200, 150 or 100 nucleotides. For other embodiments, the upper limit may be selected in the range of 50 to 100 nucleotides. For example, the fragment or a variant thereof may present a length of 3 to about 500 nucleotides, preferably 5 to about 150 nucleotides, more preferably 10 to 100 nucleotides, even more preferably 15 to 90 nucleotides, and most preferably 20 to 70 nucleotides.

[0806] Preferably, the UTR element contained in the artificial nucleic acid molecule comprises (or consists of) a continuous sequence shorter than any of the sequences selected from SEQ ID NOs:1-204 provided herein, but still having a degree of identity of 70% or more, 80% or more, 90% or more, or 95% or more with respect to any of the sequences selected from SEQ ID NOs:1-204. In this case, preferably, the substitution UTR element comprises a continuous sequence identical in its entire length to any of the sequences selected from SEQ ID NOs:1-204; however, the substitution UTR element is shorter, for example, having a length of 70% or more, 80% or more, 90% or more, or 95% or more of the full length of another identical sequence selected from SEQ ID NOs:1-204.

[0807] Furthermore, the artificial nucleic acid molecule of the present invention may contain more than one 3'-UTR element and / or more than one 5'-UTR element (as described above). For example, the artificial nucleic acid molecule of the present invention may contain one, two, three, four or more 3'-UTR elements, and / or one, two, three, four or more 5'-UTR elements, wherein the individual 3'-UTR elements may be the same or they may be different, and similarly, the individual 5'-UTR elements may be the same or they may be different. For example, the artificial nucleic acid molecule of the present invention may contain two substantially identical 3'-UTR elements as described above, for example, the two 3'-UTR elements comprising or composed of nucleic acid sequences derived from the 3'-UTR of a gene transcript, such as sequences derived from SEQ ID NO: 152 to 204, or fragments or variants of the 3'-UTR of a gene transcript as described above, its functional variants, its functional fragments, or functional variant fragments thereof. Therefore, for example, the artificial nucleic acid molecule of the present invention may comprise two substantially identical 5'-UTR elements as described above, such as two 5'-UTR elements comprising or composed of such nucleic acid sequences: said nucleic acid sequences are derived from the 5'-UTR of a gene transcript, such as sequences derived from SEQ ID NO: 1 to 151, or fragments or variants of the 5'-UTR of a gene transcript, its functional variants, its functional fragments, or functional variant fragments thereof, as described above.

[0808] Surprisingly, the inventors have discovered that artificial nucleic acid molecules comprising the 3'-UTR element as described above and / or the 5'-UTR element as described above can represent or provide mRNA molecules that offer high translation efficiency. Thus, the 3'-UTR element and / or the 5'-UTR element described herein can improve the translation efficiency of mRNA molecules.

[0809] In particular, the artificial nucleic acid molecule of the present invention may comprise (i) at least one 3'-UTR element and at least one 5'-UTR element that provide high translation efficiency; (ii) at least one 3'-UTR element that provides high translation efficiency, but not a 5'-UTR element that provides high translation efficiency; or (iii) at least one 5'-UTR element that provides high translation efficiency, but not a 3'-UTR element that provides high translation efficiency.

[0810] However, particularly in cases (ii) and (iii), but possibly also in case (i), the artificial nucleic acid molecule of the present invention may also contain one or more “other 3'-UTR elements and / or 5'-UTR elements,” that is, 3'-UTR elements and / or 5'-UTR elements that do not meet the requirements described above. For example, an artificial nucleic acid molecule of the present invention containing the 3'-UTR element of the present invention (i.e., the 3'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule) may additionally contain any additional 3'-UTR and / or any additional 5'-UTR, especially additional 5'-UTRs, such as 5'-TOP UTRs, or any other 5'-UTR or 5'-UTR element. Similarly, for example, an artificial nucleic acid molecule of the present invention that includes the 5'-UTR element described in the present invention (i.e., the 5'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule) may additionally include any additional 3'-UTR and / or any additional 5'-UTR, especially additional 3'-UTRs, such as 3'-UTRs derived from the albumin gene 3'-UTR, particularly preferably including the 3'-UTR of the sequence of SEQ ID NO: 206 or 207, especially SEQ ID NO: 207, or any other 3'-UTR or 3'-UTR element.

[0811] If, in addition to at least one 5'-UTR element and / or at least one 3'-UTR element of the present invention providing high translation efficiency, additional 3'-UTR elements and / or additional 5'-UTR elements are present in the artificial nucleic acid molecule of the present invention, then the additional 5'-UTR elements and / or the additional 3'-UTR elements can interact with the 3'-UTR elements and / or the 5'-UTR elements of the present invention, thereby supporting the translation efficiency effects of the 3'-UTR elements and / or the 5'-UTR elements of the present invention, respectively. The additional 3'-UTR elements and / or 5'-UTR elements can further support stability and translation efficiency. Furthermore, if both (the 3'-UTR elements and the 5'-UTR elements of the present invention) are present in the artificial nucleic acid molecule of the present invention, the translation efficiency effects of the 5'-UTR elements and the 3'-UTR elements of the present invention preferably lead to high translation efficiency in a synergistic manner.

[0812] Preferably, the additional 3'-UTR comprises or is composed of a nucleic acid sequence derived from the 3'-UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene, or a variant derived from the 3'-UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene, which is based on SEQ ID No. 1369-1390 of patent application WO2013 / 143700, the disclosure of which is incorporated herein by reference. In a particularly preferred embodiment, the additional 3'-UTR comprises or consists of the nucleic acid sequence of the 3'-UTR of the human albumin gene derived from the albumin gene, preferably the vertebrate albumin gene, more preferably the mammalian albumin gene, and most preferably the human albumin gene of SEQ ID NO: 206:

[0813] SEQ ID NO:206:

[0814]

[0815] (Human albumin 3'-UTR; corresponding to SEQ ID No:1369 of patent application WO 2013 / 143700).

[0816] In another particularly preferred embodiment, the additional 3'-UTR comprises or is composed of the following nucleic acid sequences: said nucleic acid sequences are derived from α-globin genes, preferably vertebrate α- or β-globin genes, more preferably mammalian α- or β-globin genes, more preferably the 3'-UTR of the human α- or β-globin gene of patent application WO 2013 / 143700 (3'-UTR of Homo sapiens hemoglobin, α1(HBA1)) of SEQ ID No. 1370 of patent application WO 2013 / 143700 or the 3'-UTR of the human α- or β-globin gene of patent application WO 2013 / 143700 (3'-UTR of Homo sapiens hemoglobin, α2(HBA2)) of SEQ ID No. 1371 of patent application WO 2013 / 143700 and / or the 3'-UTR of the human α- or β-globin gene of patent application WO 2013 / 143700 (3'-UTR of Homo sapiens hemoglobin, β(HBB)).

[0817] For example, the additional 3'-UTR may comprise or consist of the central α-complex-binding portion of the 3'-UTR of the α-globin gene of SEQ ID No. 1393 in patent application WO 2013 / 143700.

[0818] In this case, it is particularly preferred that the nucleic acid molecule of the present invention contains an additional 3'-UTR element derived from the nucleic acid or fragments thereof, homologs or variants thereof, according to SEQ ID No. 1369-1390 of patent application WO2013 / 143700.

[0819] The most preferred additional 3'-UTR contains a nucleic acid sequence of a fragment derived from the human albumin gene of SEQ ID NO:207:

[0820] SEQ ID NO:207:

[0821]

[0822] (Albumin 7 3'-UTR; corresponding to SEQ ID No: 1376 of patent application WO 2013 / 143700)

[0823] In this case, it is particularly preferred that the additional 3'-UTR of the artificial nucleic acid molecule of the present invention comprises or is composed of the nucleic acid sequence of SEQ ID NO:207 or the corresponding RNA sequence.

[0824] In some embodiments, the additional 3'-UTR comprises or consists of the following nucleic acid sequences: said nucleic acid sequences are derived from ribosomal protein-coding genes, preferably as described in international patent applications WO 2015 / 101414 or WO2015 / 101415, the disclosures of which are incorporated herein by reference.

[0825] The additional 3'-UTR may also contain or be composed of nucleic acid sequences derived from ribosomal protein-coding genes, wherein the additional 3'-UTR may be derived from ribosomal protein-coding genes including, but not limited to, ribosomal protein L9 (RPL9), ribosomal protein L3 (RPL3), ribosomal protein L4 (RPL4), ribosomal protein L5 (RPL5), ribosomal protein L6 (RPL6), ribosomal protein L7 (RPL7), ribosomal protein L7a (RPL7A), ribosomal protein L11 (RPL11), ribosomal protein L12 (RPL12), ribosomal protein L13 (RPL13), and ribosomal protein L23 (RPL23). Ribosomal protein L18 (RPL18), ribosomal protein L18a (RPL18A), ribosomal protein L19 (RPL19), ribosomal protein L21 (RPL21), ribosomal protein L22 (RPL22), ribosomal protein L23a (RPL23A), ribosomal protein L17 (RPL17), ribosomal protein L24 (RPL24), ribosomal protein L26 (RPL26), ribosomal protein L27 (RPL27), ribosomal protein L30 (RPL30), ribosomal protein L27a (RPL27A), ribosomal protein L28 (RPL28), ribosomal protein L29 (RPL29), ribosomal protein L31 (RPL31 ...21 (RPL31), ribosomal protein L28 (RPL28), ribosomal protein L29 (RPL29), ribosomal protein L21 (RPL31), ribosomal protein L28 (RPL28), ribosomal protein L29 (RPL29), ribosomal protein L21 (RPL31), ribosomal protein L28 (R Ribosomal protein L32 (RPL32), ribosomal protein L35a (RPL35A), ribosomal protein L37 (RPL37), ribosomal protein L37a (RPL37A), ribosomal protein L38 (RPL38), ribosomal protein L39 (RPL39), large ribosomal protein P0 (RPLP0), large ribosomal protein P1 (RPLP1), large ribosomal protein P2 (RPLP2), ribosomal protein S3 (RPS3), ribosomal protein S3A (RPS3A), ribosomal protein S4, X-linked (RPS4X), ribosomal protein S4, Y-linked 1 (RPS4Y1), ribosomal protein S5 (RPS5), ribosomal protein S6 (RPS6), ribosomal protein S32 (RPL32), ribosomal protein L35a (RPL35A), ribosomal protein L37 (RPL37), ribosomal protein L37a (RPL37A), ribosomal protein L38 (RPL38), ribosomal protein L39 (RPL39), large ribosomal protein P0 (RPLP0), large ribosomal protein P1 (RPLP1), large ribosomal protein P2 (RPLP2), ribosomal protein S3 (RPS3), ribosomal protein S3A (RPS3A), ribosomal protein S4, X-linked (RPS4X), ribosomal protein S4, Y-linked (RPS4Y1), ribosomal protein S5 (RPS5), ribosomal protein S6 (RPS6), ribosomal protein S32 (RPL32), ribosomal protein S35a (RPL35A), ribosomal protein L37 (RPL37), ribosomal Ribosomal protein S7 (RPS7), ribosomal protein S8 (RPS8), ribosomal protein S9 (RPS9), ribosomal protein S10 (RPS10), ribosomal protein S11 (RPS11), ribosomal protein S12 (RPS12), ribosomal protein S13 (RPS13), ribosomal protein S15 (RPS15), ribosomal protein S15a (RPS15A), ribosomal protein S16 (RPS16), ribosomal protein S19 (RPS19), ribosomal protein S20 (RPS20), ribosomal protein S21 (RPS21), ribosomal protein S23 (RPS23), ribosomal protein S25 (RPS25), ribosomal protein S26 (RPS26),Ribosomal protein S27 (RPS27), ribosomal protein S27a (RPS27a), ribosomal protein S28 (RPS28), ribosomal protein S29 (RPS29), ribosomal protein L15 (RPL15), ribosomal protein S2 (RPS2), ribosomal protein L14 (RPL14), ribosomal protein S14 (RPS14), ribosomal protein L10 (RPL10), ribosomal protein L10a (RPL10A), ribosomal protein L35 (RPL35), ribosomal protein L1 3a (RPL13A), ribosomal protein L36 (RPL36), ribosomal protein L36a (RPL36A), ribosomal protein L41 (RPL41), ribosomal protein S18 (RPS18), ribosomal protein S24 (RPS24), ribosomal protein L8 (RPL8), ribosomal protein L34 (RPL34), ribosomal protein S17 (RPS17), ribosomal protein SA (RPSA), ubiquitin A-52 residue ribosomal protein fusion product 1 (UBA52), widely expressed Fi nkel-Biskis-Reilly murine sarcoma virus (FBR-MuSV) (FAU), ribosomal protein L22-like 1 (RPL22L1), ribosomal protein S17 (RPS17), ribosomal protein L39-like (RPL39L), ribosomal protein L10-like (RPL10L), ribosomal protein L36a-like (RPL36AL), ribosomal protein L3-like (RPL3L), ribosomal protein S27-like (RPS27L), ribosomal protein L26-like 1 (RP The following ribosomal proteins are identified: L26L1, L7-like 1 (RPL7L1), L13a pseudogene (RPL13AP), L37a pseudogene 8 (RPL37AP8), S10 pseudogene 5 (RPS10P5), S26 pseudogene 11 (RPS26P11), L39 pseudogene 5 (RPL39P5), P0 pseudogene 6 (RPLP0P6), and L36 pseudogene 14 (RPL36P14).

[0826] Preferably, the additional 5'-UTR contains or is composed of a nucleic acid sequence derived from the 5'-UTR of the TOP gene or a fragment, homologue, or variant of the 5'-UTR of the TOP gene.

[0827] Particularly preferred is that the 5'-UTR element does not contain the TOP motif or 5'TOP as defined above. In particular, it is preferred that the 5'-UTR of the TOP gene is the 5'-UTR of the TOP gene lacking the TOP motif.

[0828] The 5'-UTR nucleic acid sequence derived from the TOP gene originates from the eukaryotic TOP gene, preferably from plant or animal TOP genes, more preferably from chordate TOP genes, even more preferably from vertebrate TOP genes, and most preferably from mammalian TOP genes, such as human TOP genes.

[0829] For example, the additional 5'-UTR is preferably selected from 5'-UTR elements comprising or composed of nucleic acid sequences derived from the group consisting of: SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700 (the disclosure of which is incorporated herein by reference), homologues of SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700, variants thereof, or preferably corresponding RNA sequences. The term “homologous to SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700” refers to sequences from species other than Homo sapiens that are homologous to the sequences of SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700.

[0830] In a preferred embodiment, the additional 5'-UTR comprises or is composed of a nucleic acid sequence derived from homologues of SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700, variants of which, or the corresponding RNA sequence, extends from nucleotide position 5 (i.e., the nucleotide at position 5 in the sequence) to the 5' nucleotide position immediately following the start codon (located at the 3' end of the sequence), for example, the nucleotide sequence immediately following the 5' nucleotide position of the ATG sequence. Particularly preferred is that the additional 5'-UTR is derived from homologues of SEQ ID NOs.1-1363, SEQ ID NO.1395, SEQ ID NO.1421 and SEQ ID NO.1422 of patent application WO2013 / 143700, variants thereof, or nucleic acid sequences of the corresponding RNA sequence extending from the 3' nucleotide position immediately following the 5' TOP to the 5' nucleotide position immediately following the start codon (located at the 3' end of the sequence), for example, a nucleic acid sequence immediately following the 5' nucleotide position of the ATG sequence.

[0831] In a particularly preferred embodiment, the additional 5'-UTR comprises or is composed of a nucleic acid sequence derived from or a variant of the 5'-UTR of the TOP gene encoding a ribosomal protein. For example, the 5'-UTR element comprises or is composed of a nucleic acid sequence derived from SEQ ID NO. 143700 according to patent application WO2013 / 143700. NOs: 170, 232, 244, 259, 1284, 1285, 1286, 1287, 1288, 1289, 1290, 1291, 1292, 1293, 1294, 1295, 1296, 1297, 1298, 1299, 1300, 1301, 130 2, 1303, 1304, 1305, 1306, 1307, 1308, 1309, 1310, 1311, 1312, 1313, 1314, 1315, 1316, 1317, 1318, 1319, 1320, 1321, 1322, 1323, 1324, 1 The 5'-UTR of any of the following nucleic acid sequences: 325, 1326, 1327, 1328, 1329, 1330, 1331, 1332, 1333, 1334, 1335, 1336, 1337, 1338, 1339, 1340, 1341, 1342, 1343, 1344, 1346, 1347, 1348, 1349, 1350, 1351, 1352, 1353, 1354, 1355, 1356, 1357, 1358, 1359, or 1360; the corresponding RNA sequence, its homologue, or a variant thereof, preferably lacking the 5'-TOP motif, as described herein. As described above, the sequence extending from position 5 to the 5' nucleotide immediately following ATG (located at the 3' end of the sequence) corresponds to the 5'-UTR of the sequence.

[0832] Preferably, the additional 5'-UTR comprises or is composed of a nucleic acid sequence derived from the 5'-UTR of the TOP gene encoding ribosomal macroprotein (RPL) or a homologue or variant of the 5'-UTR of the TOP gene encoding ribosomal macroprotein (RPL). For example, the 5'-UTR element comprises or is composed of a nucleic acid sequence derived from the 5'-UTR of the nucleic acid sequence of any one of SEQ ID NOs: SEQ ID NOs: 67, 259, 1284-1318, 1344, 1346, 1348-1354, 1357, 1358, 1421 and 1422 according to patent application WO2013 / 143700, the corresponding RNA sequence, its homologue, or a variant thereof, preferably lacking the 5'TOP motif, as described herein.

[0833] In a particularly preferred embodiment, the 5'-UTR element comprises or is composed of a nucleic acid sequence derived from the 5'-UTR of the macroribosomal protein 32 gene (preferably derived from the vertebrate macroribosomal protein 32 (L32) gene, more preferably derived from the mammalian macroribosomal protein 32 (L32) gene, and most preferably derived from the human macroribosomal protein 32 (L32) gene), or a variant derived from the 5'-UTR of the macroribosomal protein 32 gene (preferably derived from the vertebrate macroribosomal protein 32 (L32) gene, more preferably derived from the mammalian macroribosomal protein 32 (L32) gene, and most preferably derived from the human macroribosomal protein 32 (L32) gene), wherein preferably the additional 5'-UTR does not contain the 5' TOP of the gene.

[0834] Therefore, in a particularly preferred embodiment, the additional 5'-UTR comprises or is composed of a nucleic acid sequence that has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99% identity with the nucleic acid sequence of SEQ ID NO:208 (5'-UTR of human macroribosomal protein 32 lacking a 5' terminal oligopyrimidine bundle: GGCGCTGCCTACGGAGGTGGCAGCCATCTCCTTCTCGGCATC (SEQ ID NO:208); corresponding to SEQ ID NO.1368 of patent application WO2013 / 143700), or preferably the corresponding RNA sequence. Alternatively, the additional 5'-UTR comprises or is composed of a fragment of such a nucleic acid sequence that has at least 40% identity with the nucleic acid sequence of SEQ ID NO:208 (5'-UTR of human macroribosomal protein 32 lacking a 5' terminal oligopyrimidine bundle: GGCGCTGCCTACGGAGGTGGCAGCCATCTCCTTCTCGGCATC (SEQ ID NO:208); or preferably at least about 99% identity with the nucleic acid sequence of SEQ ID NO:2013 / 143700). The nucleic acid sequence of NO:208, or more preferably the corresponding RNA sequence, has at least about 40%, preferably at least about 50%, preferably at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%, even more preferably at least about 99% identity, wherein, preferably, the fragment is, as described above, a continuous nucleotide segment representing at least 20% of the full-length 5'-UTR. Preferably, the fragment has a length of at least about 20 nucleotides or more, preferably at least about 30 nucleotides or more, more preferably at least about 40 nucleotides or more. Preferably, the fragment is a functional fragment as described herein.

[0835] In some embodiments, the artificial nucleic acid molecule includes an additional 5'-UTR containing or composed of a nucleic acid sequence derived from the 5'-UTR of a vertebrate TOP gene (e.g., a mammalian TOP gene, such as the human TOP gene): RPSA, RPS2, RPS3, RPS3A, RPS4, RPS5, RPS6, RPS7, RPS8, RPS9, RPS10, RPS11, RPS12, RPS13, RPS14, RPS15, RPS15A, RPS16, RPS17, RPS18, RPS19, RPS20, RPS21, RPS23, RPS24. , RPS25, RPS26, RPS27, RPS27A, RPS28, RPS29, RPS30, RPL3, RPL4, RPL5, RPL6, RPL7, RPL7A, RPL8, RPL9, RPL10, RPL10A, RPL11, RPL12, RPL13, RPL 13A, RPL14, RPL15, RPL17, RPL18, RPL18A, RPL19, RPL21, RPL22, RPL23, RPL23A, RPL24, RPL26, RPL27, RPL27A, RPL28, RPL29, RPL30, RPL31, RPL32 , RPL34, RPL35, RPL35A, RPL36, RPL36A, RPL37, RPL37A, RPL38, RPL39, RPL40, RPL41, RPLP0, RPLP1, RPLP2, RPLP3, RPLP0, RPLP1, RPLP2, EEF1A1, EEF1B2, EEF1D, EEF1G, EEF2, EIF3E, EIF3F, EIF3H, EIF2S3, EIF3C, EIF3K, EIF3EIP, EIF4A2, PABPC1, HNRNPA1, TPT1, TUBB1, UBA52, NPM1, ATP5G2 GNB2L1, NME2, UQCRB or their homologues or variants, wherein preferably the additional 5'-UTR does not contain the TOP motif or the 5'TOP of the gene, and wherein optionally the additional 5'-UTR begins at its 5'-end with a nucleotide at position 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 downstream of the 5'-terminal oligopyrimidine bundle (TOP), and wherein further optionally, the additional 5'-UTR derived from the TOP gene terminates at its 3'-end with a nucleotide at position 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 upstream of the start codon (A(U / T)G) of the gene from which it originates.

[0836] In some embodiments, the additional 5'-UTR contains or is composed of a nucleic acid sequence derived from the 5'-UTR described in International Patent Application WO 2016 / 107877, the disclosure of which is incorporated herein by reference.

[0837] The artificial nucleic acid molecule described in this invention can be RNA, such as mRNA or viral RNA or a replicon, DNA, such as DNA plasmid or viral DNA, or can be a modified RNA or DNA molecule. It can be provided as a double-stranded molecule with a sense strand and an antisense strand, for example, a DNA molecule with a sense strand and an antisense strand.

[0838] The artificial nucleic acid molecule of the present invention may also optionally include a 5'-cap. The optional 5'-cap is preferably located at the 5' of the ORF, more preferably at the 5' of at least one 5'-UTR or any other 5'-UTR within the artificial nucleic acid molecule of the present invention.

[0839] Preferably, the artificial nucleic acid molecule of the present invention further comprises a poly(A) sequence and / or a polyadenylation signal. Preferably, the optional polyadenylation sequence is located at the 3' end of at least one 3'-UTR element or any other 3'-UTR, more preferably the optional polyadenylation sequence is linked to the 3'-end of the 3'-UTR element. The linking can be direct or indirect, for example, through a nucleotide segment of 2, 4, 6, 8, 10, 20, etc., such as through a linker containing, for example, 1-50, preferably 1-20 nucleotides comprising one or more restriction sites or consisting of one or more restriction sites. However, even if the artificial nucleic acid molecule of the present invention does not contain a 3'-UTR, for example if it only contains at least one 5'-UTR element, it preferably still contains a polyadenylation sequence and / or a polyadenylation signal.

[0840] In one embodiment, the optional polyadenylation signal is located downstream of the 3' end of the 3'-UTR element. Preferably, the polyadenylation signal comprises the consensual sequence NN(U / T)ANA, where N = A or U, preferably AA(U / T)AAA or A(U / T)(U / T)AAA. This consensual sequence can be recognized by most animal and bacterial cell systems, for example by polyadenylation factors, such as cleavage / polyadenylation-specific factors (CPSF) that interact with CstF, PAP, PAB2, CFI, and / or CFII. Preferably, the polyadenylation signal, preferably the consensual sequence NNUANA, is located less than about 50 bases downstream of the 3'-end of the 3'-UTR element or ORF (if the 3'-UTR element is absent), more preferably less than about 30 bases, and most preferably less than about 25 bases, for example, at 21 bases.

[0841] Transcription of the artificial nucleic acid molecule described in this invention (e.g., an artificial DNA molecule containing a polyadenylation signal downstream of a 3'-UTR element (or ORF)) will produce mature pre-RNA containing a polyadenylation signal downstream of its 3'-UTR element (or ORF).

[0842] Subsequently, using a suitable transcription system will result in the attachment of a polyadenylated sequence to the pre-mature RNA. For example, the artificial nucleic acid molecule of the present invention can be a DNA molecule containing a 3'-UTR element and a polyadenylation signal as described above, which can cause polyadenylation of RNA during transcription of the DNA molecule. Thus, the resulting RNA can contain a combination of the 3'-UTR element followed by a polyadenylated sequence of the present invention.

[0843] Possible transcription systems include in vitro transcription systems or cellular transcription systems. Therefore, transcription of the artificial nucleic acid molecules described in this invention, such as transcription of artificial nucleic acid molecules containing a read frame, a 3'-UTR element, and / or a 5'-UTR element and optionally a polyadenylation signal, can produce mRNA molecules containing a read frame, a 3'-UTR element, and optionally a polyadenylation sequence.

[0844] Therefore, the present invention also provides an artificial nucleic acid molecule, which is an mRNA molecule comprising a read frame, a 3'-UTR element as described above and / or a 5'-UTR element as described above and optionally a polyadenylate sequence.

[0845] In another embodiment, the 3'-UTR of the artificial nucleic acid molecule of the present invention does not contain a polyadenylation signal or a polyadenylation sequence. More preferably, the artificial nucleic acid molecule of the present invention does not contain a polyadenylation signal or a polyadenylation sequence. More preferably, the 3'-UTR of the artificial nucleic acid molecule, or the artificial nucleic acid molecule of the present invention itself, does not contain a polyadenylation signal, particularly it does not contain the polyadenylation signal AAU / TAAA.

[0846] In a preferred embodiment, the present invention provides an artificial nucleic acid molecule, which is an artificial RNA molecule, comprising a read frame and an RNA sequence corresponding to a DNA sequence selected from the group consisting of the following sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:5 ... ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 66, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 72, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 144 (or SEQ ID NO: 42), SEQ ID NO: 58, SEQ ID NO: 62;SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147 (or SEQ ID NO:109), SEQ ID NO:148 (or SEQ ID NO:110), SEQ SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149 (or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:151 (or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150 (based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125, or fragments thereof as described above. In addition, corresponding artificial DNA molecules are also provided.

[0847] In another preferred embodiment, the present invention provides an artificial nucleic acid molecule, which is an artificial DNA molecule, comprising a read frame and a sequence selected from the group consisting of the following sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:52, ...2, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:52, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, NO: 63, SEQ ID NO: 64, SEQ ID NO: 66, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 72, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 144 (or SEQ ID NO: 42), SEQ ID NO: 58, SEQ ID NO: 62;SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ IDNO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ IDNO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147(or SEQ ID NO:109), SEQ ID NO:148(or SEQ ID NO:110), SEQ ID NO:112, SEQ IDNO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149(or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ IDNO:134, SEQ ID NO:151(or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150(based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125.;

[0848] Therefore, the present invention provides an artificial nucleic acid molecule that can be used as a template for an RNA molecule (preferably an mRNA molecule) characterized by high translation efficiency. In other words, the artificial nucleic acid molecule can be DNA that can be used as a template for the production of mRNA. The available mRNA can be translated accordingly to produce a desired peptide or protein encoded by a read frame. If the artificial nucleic acid molecule is DNA, it can be used, for example, as a double-stranded storage form for the continuous and repetitive in vitro or in vivo production of mRNA. Thus, in vitro specifically refers to (“living”) cells and / or tissues, including tissues of a living subject. Cells specifically include cell lines, primary cells, and cells in tissues or subjects. In specific embodiments, cell types that allow cell culture can be suitable for the present invention. Mammalian cells, such as human cells and mouse cells, are particularly preferred. In particularly preferred embodiments, human cell lines HeLa, HEPG2, and U-937 and mouse cell lines NIH3T3, JAWSII, and L929 are used. Furthermore, primary cells are particularly preferred, and in particularly preferred embodiments, human skin fibroblasts (HDF) can be used. The preferred group of cell lines includes

[0849] Alternatively, tissue from the test subject can also be used.

[0850] In one embodiment, the artificial nucleic acid molecule of the present invention further comprises a polyadenylated sequence. For example, a DNA molecule comprising an ORF (optionally followed by a 3'UTR) may contain a thymine nucleotide that can be transcribed into a polyadenylated sequence in the resulting mRNA. The length of the polyadenylated sequence can vary. For example, the polyadenylated sequence may have a length of about 20 adenine nucleotides to up to about 300 adenine nucleotides, preferably about 40 to about 200 adenine nucleotides, more preferably from about 50 to about 100 adenine nucleotides, such as about 60, 70, 80, 90 or 100 adenine nucleotides. Most preferably, the nucleic acid of the present invention comprises a polyadenylated sequence of about 60 to about 70 nucleotides, most preferably 64 adenine nucleotides.

[0851] Artificial RNA molecules can also be obtained in vitro through conventional chemical synthesis methods without the need for transcription from DNA precursors.

[0852] In a particularly preferred embodiment, the artificial nucleic acid molecule of the present invention is an RNA molecule, preferably an mRNA molecule containing a read frame, a 3'-UTR element as described above, and a polyadenylated sequence in the 5'-to-3'-direction, or containing a 5'-UTR element, a read frame, and a polyadenylated sequence as described above in the 5'-to-3'-direction.

[0853] In a preferred embodiment, the read frame is derived from a gene that is different from the gene from which the 3'-UTR element and / or 5'-UTR element of the artificial nucleic acid of the present invention originates. In some other preferred embodiments, the read frame does not encode genes selected from the group consisting of: ZNF460, TGM2, IL7R, BGN, TK1, RAB3B, CBX6, FZD2, COL8A1, NDUFS7, PHGDH, PLK2, TSPO, PTGS1, FBXO32, NID2, ATP5D, EXOSC4, NOL9, UBB4B, VPS18, ORMDL2, FSCN1, TMEM33, TUBA4A, EMP3, TMEM201, CRIP2, BRAT1, SERPINH1, CD9, DPYSL2, CDK9, TFRC. PSMB3, FASN, PSMB6, PRSS56, KPNA6, SFT2D2, PARD6B, LPP, SPARC, SCAND1, VASN, SLC26A1, LCLAT1, FBXL18, SLC35F6, RAB3D, MAP1B, VMA21, CYBA , SEZ6L2, PCOLCE, VTN, ALDH16A1, RAVER1, KPNA6, SERINC5, JUP, CPN2, CRIP2, EPT1, PNPO, SSSCA1, POLR2L, LIN7C, UQCR10, PYCRL, AMN, MAP1S, N DUFS7, PHGDH, TSPO, ATP5D, EXOSC4, TUBB4B, TUBA4A, EMP3, CRIP2, BRAT1, CD9, CDK9, PSMB3, PSMB6, PRSS56, SCAND1, AMN, CYBA, PCOLCE, MAP1S, VTN, ALDH16A1 (all preferred for humans) and Dpysl2, Ccnd1, Acox2, Cbx6, Ubc, Ldlr, Nutt22, Pcyox1l, Ankrd1, Tmem37, Tspyl4, Slc7a3, Cst6, Aacs, Nosip, Itga 7. Ccnd2, Ebp, Sf3b5, Fasn, Hmgcs1, Osr1, Lmnb1, Vma21, Kif20a, Cdca8, Slc7a1, Ubqln2, Prps2, Shmt2, Aurkb, Fignl1, Cad, Anln, Slfn9, Ncap h, Pole, Uhrf1, Gja1, Fam64a, Kif2c, Tspan10, Scand1, Gpr84, Fads3, Cers6, Cxcr4, Gprc5c, Fen1, Cspg4, Mrpl34, Comtd1, Armc6, Emr4, Atp5d,1110001J03Rik, Csf2ra, Aarsd1, Kif22, Cth, Tpgs1, Ccl17, Alkbh7, Ms4a8a, Acox2, Ubc, Slpi, Pcyox1l, Igf2bp1, Tmem37, Slc7a3, Cst6, Ebp, Sf3b5, Plk1, Cdca8, Kif22, Cad, Cth, Pole, Kif2c, Scand1, Gpr84, Tpgs1, Ccl17, Alkbh7, Ms4a8a, Mrpl34, Comtd1, Armc6, Atp5d, 1110001J03Rik, Nudt22, Aarsd1 (all preferably mouse), or variants thereof, provided that the 3'-UTR element and / or 5'-UTR element are sequences selected from the group consisting of sequences of SEQ ID NO:1-204.

[0854] In a preferred embodiment, the ORF does not encode human or plant ribosomal proteins, particularly Arabidopsis thaliana ribosomal proteins, especially not human ribosomal protein S6 (RPS6), human ribosomal protein L36a-like (RPL36AL), or Arabidopsis thaliana ribosomal protein S16 (RPS16). In a further preferred embodiment, the read frame (ORF) does not encode ribosomal protein S6 (RPS6), ribosomal protein L36a-like (RPL36AL), or ribosomal protein S16 (RPS16) from any source.

[0855] In one embodiment, the present invention provides an artificial DNA molecule comprising a read frame, preferably a read frame derived from a gene different from the gene from which the 3'-UTR element and / or 5'-UTR element originate; a 3'-UTR element comprising or consisting of a sequence having at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%; even more preferably at least 99%; even more preferably 100% sequence identity with a DNA sequence selected from the group consisting of the following sequences: SEQ ID NO:152, SEQ ID NO:153, SEQ ID NO:154, SEQ ID NO:155, SEQ ID NO:156, SEQ ID NO:157, SEQ ID NO:158, SEQ ID NO:159, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:16 ...69, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO: NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173; SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:184, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:191, SEQ ID NO:204 (or SEQ ID NO:192), SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:201, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:177, and / or a 5'-UTR element, said 5'-UTR element comprising or consisting of a sequence having at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%; even more preferably at least 99%;Even more preferred is 100% sequence identity: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:66, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 72, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 144 (or SEQ ID NO: 42), SEQ ID NO: 58, SEQ ID NO: 62;SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147 (or SEQ ID NO:109), SEQ ID NO:148 (or SEQ ID NO:110), SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149 (or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:151 (or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150 (based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125; and a polyadenylation signal and / or polyadenylic acid sequence.;

[0856] Furthermore, the present invention provides an artificial RNA molecule, preferably an artificial mRNA molecule or an artificial viral RNA molecule, comprising a read frame, preferably derived from a gene different from the gene from which the 3'-UTR element and / or 5'-UTR element originate; a 3'-UTR element comprising or consisting of a sequence having at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%; even more preferably at least 99%; even more preferably 100% sequence identity with an RNA sequence corresponding to a DNA sequence selected from the group consisting of the following sequences: SEQ ID NO:152, SEQ ID NO:153, SEQ ID NO:154, SEQ ID NO:155, SEQ ID NO:156, SEQ ID NO:157, SEQ ID NO:158, SEQ ID NO:159, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:164, SEQ ID NO:165, SEQ ID NO:167, SEQ ID NO:16 ...69, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO:16 NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173; SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 179, SEQ ID NO: 180, SEQ ID NO: 181, SEQ ID NO: 182, SEQ ID NO: 183, SEQ ID NO:184, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:191, SEQ ID NO:204 (or SEQ ID NO:192), SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:201, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:177, and / or 5'-UTR element, comprising or consisting of a sequence having at least about 60%, preferably at least about 70%, more preferably at least about 80%, more preferably at least about 90%, even more preferably at least about 95%; even more preferably at least 99%;Even more preferred is 100% sequence identity: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:137 (or SEQ ID NO:3), SEQ ID NO:138 (or SEQ ID NO:9), SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:32, SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:15, SEQ ID NO:140 (or SEQ ID NO:17), SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:143 (or SEQ ID NO:31), SEQ ID NO:39, SEQ ID NO:41, SEQ ID NO:44, SEQ ID NO:65, SEQ ID NO:146 (or SEQ ID NO:52), SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:57, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69 ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 72, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 144 (or SEQ ID NO: 42), SEQ ID NO: 58, SEQ ID NO: 62;SEQ ID NO:73, SEQ ID NO:75, SEQ ID NO:77, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:147 (or SEQ ID NO:109), SEQ ID NO:148 (or SEQ ID NO:110), SEQ ID SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:117, SEQ ID NO:118, SEQ ID NO:149 (or SEQ ID NO:119), SEQ ID NO:123, SEQ ID NO:126, SEQ ID NO:127, SEQ ID NO:128, SEQ ID NO:129, SEQ ID NO:130, SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:151 (or SEQ ID NO:136), SEQ ID NO:96, SEQ ID NO:111, SEQ ID NO:150 (based on SEQ ID NO:120), SEQ ID NO:121, SEQ ID NO:124, SEQ ID NO:125; and polyadenylation signal and / or polyadenylate sequence.

[0857] This invention provides artificial nucleic acid molecules, preferably artificial mRNA, which can be characterized by high translation efficiency. Without any theoretical limitations, high translation efficiency can result from reduced degradation of the artificial nucleic acid molecules (such as artificial mRNA molecules) described in this invention. Therefore, the 3'-UTR element and / or the 5'-UTR element of this invention can prevent the degradation and breakdown of artificial nucleic acids.

[0858] Preferably, the artificial nucleic acid molecule may additionally include a histone stem-loop. Thus, the artificial nucleic acid molecule of the present invention may, for example, include an ORF, a 3'-UTR element, an optional histone stem-loop sequence, an optional polyadenylated sequence or polyadenylated signal, and an optional polycytidine sequence in the 5'-to-3'-direction; or include a 5'-UTR element, an ORF, an optional histone stem-loop sequence, an optional polyadenylated sequence or polyadenylated signal, and an optional polycytidine sequence in the 5'-to-3'-direction; or include a 5'-UTR element, an ORF, a 3'-UTR element, an optional histone stem-loop sequence, an optional polyadenylated sequence or polyadenylated signal, and an optional polycytidine sequence in the 5'-to-3'-direction. It may also include an ORF, a 3'-UTR element, an optional polyadenylated sequence, an optional polycytidine sequence, and an optional histone stem-loop sequence in the 5'-to-3'-direction, or include a 5'-UTR element, an ORF, an optional polyadenylated sequence, an optional polycytidine sequence, and an optional histone stem-loop sequence in the 5'-to-3'-direction, or include a 5'-UTR element, an ORF, a 3'-UTR element, an optional polyadenylated sequence, an optional polycytidine sequence, and an optional histone stem-loop sequence in the 5'-to-3'-direction.

[0859] In a preferred embodiment, the artificial nucleic acid molecule of the present invention further comprises at least one histone stem-loop sequence.

[0860] The histone stem-loop sequence is preferably selected from histone stem-loop sequences disclosed in WO 2012 / 019780, the disclosure of which is incorporated herein by reference.

[0861] The histone stem-loop sequence suitable for use in this invention is preferably selected from at least one of the following formulas (I) or (II):

[0862] Equation (I) (Stem-loop sequence without stem boundary elements):

[0863]

[0864] Equation (II) (Stem-loop sequence with stem boundary elements):

[0865]

[0866] in:

[0867] Stem 1 or stem 2 boundary element N 1-6 It is a continuous sequence of 1-6, preferably 2-6, more preferably 2-5, even more preferably 3-5, most preferably 4-5 or 5 N, wherein each N is independently selected from the following nucleotides: the nucleotides are selected from A, U, T, G and C, or nucleotide analogs thereof;

[0868] Stem 1[N 0-2 GN 3-5 It is either anticomplementary or partially anticomplementary to element stem 2, and is a continuous sequence of 5-7 nucleotides;

[0869] Where N 0-2 It is a continuous sequence of 0-2, preferably 0-1, more preferably 1 N, wherein each N is independently selected from the following nucleotides: the nucleotides are selected from A, U, T, G and C or their nucleotide analogs;

[0870] Where N 3-5 It is a continuous sequence of 3-5, preferably 4-5, more preferably 4 N, wherein each N is independently selected from the following nucleotides: said nucleotides are selected from A, U, T, G and C or their nucleotide analogs, and

[0871] Wherein G is guanosine or an analogue thereof, and may optionally be replaced by cytidine or an analogue thereof, provided that its complementary nucleotide cytidine in stem 2 is replaced by guanosine.

[0872] Circular sequence [N] 0-4 (U / T)N 0-4 Located between stem 1 and stem 2, and is a continuous sequence of 3-5 nucleotides, more preferably 4 nucleotides;

[0873] Where each N 0-4 The sequence consists independently of 0-4, preferably 1-3, more preferably 1-2 N, consecutive sequences, wherein each N is independently selected from the following nucleotides: said nucleotides are selected from A, U, T, G and C or their nucleotide analogues; and

[0874] Where U / T represents uridine or optional thymidine;

[0875] stem2[N 3-5 CN 0-2 It is either anticomplementary or partially anticomplementary to element stem 1, and is a continuous sequence of 5-7 nucleotides;

[0876] Where N 3-5 It is a continuous sequence of 3-5, preferably 4-5, more preferably 4 N, wherein each N is independently selected from the following nucleotides: said nucleotides are selected from A, U, T, G and C or nucleotide analogs thereof;

[0877] Where N 0-2 It is a continuous sequence of 0-2, preferably 0-1, more preferably 1 N, wherein each N is independently selected from the following nucleotides: said nucleotides are selected from A, U, T, G or C, or nucleotide analogs thereof; and

[0878] Wherein C is cytidine or an analogue thereof, and may optionally be replaced by guanosine or an analogue thereof, provided that its complementary nucleoside guanosine in stem 1 is replaced by cytidine.

[0879] in

[0880] Stem 1 and stem 2 are capable of base pairing with each other to form an inverse complementary sequence. Base pairing can occur between stem 1 and stem 2, for example, through Watson-Crick base pairing of nucleotides A with U / T or G with C, or through non-Watson-Crick base pairing, such as wobbling base pairing, inverse Watson-Crick base pairing, Hoogsteen base pairing, inverse Hoogsteen base pairing, or they can be capable of base pairing with each other to form a partially inverse complementary sequence. Incomplete base pairing can occur between stem 1 and stem 2 based on the fact that one or more bases in one stem do not have complementary bases in the inverse complementary sequence of the other stem.

[0881] According to a further preferred embodiment, a histone stem-loop sequence can be selected according to at least one of the following specific formulas (Ia) or (IIa):

[0882] Equation (Ia) (Stem-loop sequence without stem boundary elements):

[0883]

[0884] Equation (IIa) (Stem-loop sequence with stem boundary elements):

[0885]

[0886] in:

[0887] N, C, G, T, and U are defined as above.

[0888] According to another, more particularly preferred embodiment of the first aspect, the artificial nucleic acid molecule sequence may comprise at least one histone stem-loop sequence as described in at least one of the following specific formulas (Ib) or (IIb):

[0889] Equation (Ib) (Stem-loop sequence without stem boundary elements):

[0890]

[0891] Equation (IIb) (Stem-loop sequence with stem boundary elements):

[0892]

[0893] in:

[0894] N, C, G, T, and U are defined as above.

[0895] A specific preferred histone stem-loop sequence is SEQ ID NO:209:CAAAGGCTCTTTTCAGAGCCACCA, or more preferably, the corresponding RNA sequence of the nucleic acid sequence of SEQ ID NO:209.

[0896] As an example, individual elements can exist in an artificial nucleic acid molecule in the following order:

[0897] 5'-cap – 5'-UTR (element) – ORF – 3'-UTR (element) – histone stem-loop – polyadenylate / polycytidine sequence;

[0898] 5'-cap – 5'-UTR (element) – ORF – 3'-UTR (element) – polyadenylate / polycytidine sequence – histone stem-loop;

[0899] 5'-cap-5'-UTR (element)–ORF–IRES–ORF–3'-UTR (element)-histone stem-loop-polyadenylate / polycytidine sequence;

[0900] 5'-cap-5'-UTR (element)–ORF–IRES–ORF–3'-UTR (element)-histone stem-loop-polyadenylate / polycytidine sequence-polyadenylate / polycytidine sequence;

[0901] 5'-cap – 5'-UTR (element) – ORF – IRES – ORF – 3'-UTR (element) – polyadenylate / polycytidine sequence – histone stem-loop;

[0902] 5'-cap – 5'-UTR (element) – ORF – IRES – ORF – 3'-UTR (element) – polyadenylate / polycytidine sequence – polyadenylate / polycytidine sequence – histone stem-loop;

[0903] 5'-cap – 5'-UTR (element) – ORF – 3'-UTR (element) – polyadenylated acid / polycytidine sequence – polyadenylated acid / polycytidine sequence;

[0904] 5'-cap – 5'-UTR (element) – ORF – 3'-UTR (element) – polyadenylated / polycytosine sequence – polyadenylated / polycytosine sequence – histone stem-loop; etc.

[0905] In some embodiments, the artificial nucleic acid molecule further comprises elements such as a 5'-cap, a polycytidine sequence, and / or an IRES-motif. The 5'-cap can be added to the 5' end of the RNA during or after transcription. Furthermore, especially if the nucleic acid is in the form of mRNA or encodes mRNA, the artificial nucleic acid molecule of the present invention can be modified with a sequence of at least 10 cytidines, preferably at least 20 cytidines, more preferably at least 30 cytidines (the so-called "polycytidine sequence"). In particular, especially if the nucleic acid is in the form of (m)RNA or encodes mRNA, the artificial nucleic acid molecule of the present invention can contain a polycytidine sequence typically of about 10 to 200 cytidine nucleotides, preferably about 10 to 100 cytidine nucleotides, more preferably about 10 to 70 cytidine nucleotides, or even more preferably about 20 to 50 or even 20 to 30 cytidine nucleotides. Most preferably, the nucleic acid of the present invention comprises a polycytidine sequence of 30 cytosines. Therefore, the artificial nucleic acid molecule of the present invention preferably includes, in the 5'-to-3' direction, at least one 5'-UTR element, ORF, at least one 3'-UTR element, polyadenylated sequence or polyadenylated signal and polycytidine sequence as described above, or in the 5'-to-3' direction, optionally additional 5'-UTR, ORF, at least one 3'-UTR element, polyadenylated sequence or polyadenylated signal and polycytidine sequence as described above, or in the 5'-to-3' direction, at least one 5'-UTR element, ORF, optionally additional 3'-UTR, polyadenylated sequence or polyadenylated signal and polycytidine sequence as described above.

[0906] For example, if the artificial nucleic acid molecule encodes more than two peptides or proteins, the internal ribosome entry site (IRES) sequence or IRES-motif can separate some reading frames. The IRES-sequence can be particularly useful if the artificial nucleic acid molecule is a bicistronic or polycistronic nucleic acid molecule.

[0907] Furthermore, the artificial nucleic acid molecule may contain additional 5'-elements, preferably a promoter, or a promoter-containing sequence. The promoter can drive or regulate the transcription of the artificial nucleic acid molecule described in this invention, such as the transcription of the artificial DNA molecule described in this invention.

[0908] Preferably, the artificial nucleic acid molecule of the present invention, particularly the read frame, is at least partially G / C modified. Therefore, the artificial nucleic acid molecule of the present invention can be thermodynamically stabilized by modifying the G (guanosine) / C (cytidine) content of the molecule. Compared to the G / C content of the read frame of the corresponding wild-type sequence, it is preferable to increase the G / C content of the read frame of the artificial nucleic acid molecule of the present invention by using genetic code degeneracy. Therefore, compared to the coding amino acid sequence of a specific wild-type sequence, it is preferable not to change the coding amino acid sequence of the artificial nucleic acid molecule by G / C modification. Therefore, the codons of the coding sequence or the entire artificial nucleic acid molecule (e.g., mRNA) can be different from the wild-type coding sequence, thereby including an increased amount of G / C nucleotides while maintaining the translated amino acid sequence. Due to the fact that some codons encode one and the same amino acid (so-called genetic code degeneracy), it is easy to change codons without changing the encoded peptide / protein sequence (so-called alternative codon use). Therefore, it may be more advantageous to specifically introduce certain codons (replacing their respective wild-type codons that encode the same amino acid) for RNA stability and / or codon usage in subjects (so-called codon optimization).

[0909] Compared to its wild-type coding region, depending on the amino acid encoded by the coding region of the artificial nucleic acid molecule of the present invention as defined herein, there are various possibilities for modifications to the nucleic acid sequence, such as the read frame. In the case of amino acids encoded by codons containing only G or C nucleotides, no codon modification is required. Therefore, since A or U / T are absent, the codons for Pro (CCC or CCG), Arg (CGC or CGG), Ala (GCC or GCG), and Gly (GGC or GGG) do not require modification.

[0910] Conversely, codons containing A and / or U / T nucleotides can be modified by substitutions of other codons encoding the same amino acid but not containing A and / or U / T. For example

[0911] The codons for Pro can be modified from CC(U / T) or CCA to CCC or CCG;

[0912] Arg codons can be modified from CG(U / T) or CGA or AGA or AGG to CGC or CGG;

[0913] Ala's codons can be modified from GC(U / T) or GCA to GCC or GCG;

[0914] Gly's codons can be modified from GG(U / T) or GGA to GGC or GGG.

[0915] In other cases, although it is not possible to eliminate the A or (U / T) nucleotide from a codon, it is possible to reduce the A and (U / T) content by using a codon containing a lower content of A and / or (U / T) nucleotides. Examples of these are:

[0916] The codon of Phe can be modified from (U / T)(U / T)(U / T) to (U / T)(U / T)C;

[0917] Leu's codons can be modified from (U / T)(U / T)A, (U / T)(U / T)G, C(U / T)(U / T)A to C(U / T)C or C(U / T)G;

[0918] The codons of Ser can be modified from (U / T)C(U / T) or (U / T)CA or AG(U / T) to (U / T)CC, (U / T)CG or AGC;

[0919] Tyr's codon can be modified from (U / T)A(U / T) to (U / T)AC;

[0920] The codon of Cys can be modified from (U / T)G(U / T) to (U / T)GC;

[0921] His codons can be modified from CA(U / T) to CAC;

[0922] Gln codons can be modified from CAA to CAG;

[0923] The codons of Ile can be modified from A(U / T)(U / T) or A(U / T)A to A(U / T)C;

[0924] The codons of Thr can be modified from AC(U / T) or ACA to ACC or ACG.

[0925] The codon of Asn can be modified from AA(U / T) to AAC;

[0926] Lys codons can be modified from AAA to AAG;

[0927] The codons of Val can be modified from G(U / T)(U / T) or G(U / T)A to G(U / T)C or G(U / T)G;

[0928] The codon of Asp can be modified from GA(U / T) to GAC;

[0929] The codon for Glu can be modified from GAA to GAG;

[0930] The stop codon (U / T)AA can be modified to (U / T)AG or (U / T)GA.

[0931] On the other hand, in the case of codons Met(A(U / T)G) and Trp((U / T)GG), there is no possibility of sequence modification without changing the encoded amino acid sequence.

[0932] The substitutions listed above can be used alone or in all possible combinations to increase the G / C content of the read frame of an artificial nucleic acid molecule of the present invention as defined herein, compared to its specific wild-type read frame (i.e., the original sequence). Thus, for example, all codons of Thr present in the wild-type sequence can be modified to ACC (or ACG).

[0933] Preferably, the G / C content of the reading frame of the artificial nucleic acid molecule of the present invention, as defined herein, is increased by at least 7%, more preferably at least 15%, and particularly preferably at least 20%, compared to the G / C content of the wild-type coding region, without altering the encoded amino acid sequence, i.e., utilizing the degeneracy of the genetic code. According to a particular embodiment, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, more preferably at least 70%, even more preferably at least 80%, and most preferably at least 90%, 95%, or even 100% of the interchangeable codons of the reading frame, fragments thereof, variants, or derivatives of the artificial nucleic acid molecule of the present invention are interchanged, thereby increasing the G / C content of the reading frame.

[0934] In this case, it is particularly preferable to maximize the G / C content of the read frame of the artificial nucleic acid molecule of the present invention as defined herein (i.e., 100% codon substitution) compared to the wild-type read frame, without changing the encoded amino acid sequence.

[0935] Furthermore, the read frame is preferably at least partially codon-optimized. Codon optimization is based on the discovery that translation efficiency can be determined by the varying frequencies of occurrence of transfer RNA (tRNA) in the cell. Thus, if so-called “rare codons” are present to an increased degree in the coding region of an artificial nucleic acid molecule as defined herein, the translation of the corresponding modified nucleic acid sequence is less efficient than in the presence of codons encoding relatively “common” tRNAs.

[0936] Therefore, the read frame of the artificial nucleic acid molecule of the present invention is preferably modified compared to the corresponding wild-type coding region, such that at least one codon encoding a wild-type sequence of a relatively rare tRNA in the cell is exchanged for a codon encoding a tRNA that is more frequently present in the cell and carries the same amino acid as the relatively rare tRNA. With this modification, the read frame of the artificial nucleic acid molecule of the present invention, as defined herein, is modified so that the codons available for the frequently occurring tRNA can replace the codons corresponding to the rare tRNA. In other words, according to the present invention, with this modification, all codons of the wild-type read frame encoding the rare tRNA can be exchanged for codons encoding a tRNA that is more frequently present in the cell and carries the same amino acid as the rare tRNA. Which tRNAs are relatively frequently present in the cell and conversely, which are relatively rarely present are known to those skilled in the art; see, for example, Akashi, Curr. Opin. Genet. Dev. 2001, 11(6):660-666. Therefore, preferably, preferably with regard to the system in which the artificial nucleic acid molecule of the present invention will be expressed, preferably with regard to the system in which the artificial nucleic acid molecule of the present invention will be translated, the read frame is codon-optimized. Preferably, the codon usage of the readable frame is codon-optimized based on mammalian codon usage, and more preferably based on human codon usage. Preferably, the readable frame is codon-optimized and modified with G / C content.

[0937] To further improve degradation resistance, such as resistance to in vivo (or in vitro as defined above) degradation caused by exonucleases or endonucleases, and / or to further improve the stability of protein expression of the artificial nucleic acid molecule described in this invention, the artificial nucleic acid molecule may further include modifications, such as backbone modifications, sugar modifications, and / or base modifications, for example, lipid modifications. Preferably, the transcription and / or translation of the artificial nucleic acid molecule described in this invention are not significantly impaired by the modifications.

[0938] Typically, the artificial nucleic acid molecules of the present invention can contain any natural (i.e., naturally occurring) nucleotides, such as guanosine, uracil, adenosine, and / or cytosine or analogues thereof. In this respect, nucleotide analogues are defined as natural or non-natural variants of naturally occurring nucleotides adenosine, cytosine, thymidine, guanosine, and uridine. Thus, analogues are, for example, nucleotides chemically derived from nucleotides having non-naturally occurring functional groups (preferably added to naturally occurring nucleotides or naturally occurring functional groups removed from or substituted by naturally occurring nucleotides). Therefore, the various components of naturally occurring nucleotides can be modified, i.e., the base component, the sugar (ribose) component, and / or the phosphate component forming the backbone of the RNA sequence (see above).Analogs of guanosine, uridine, adenosine, thymidine, and cytosine include (but are not intended to limit) any naturally occurring or non-natural guanosine, uridine, adenosine, thymidine, or cytosine, for example, modified chemically, such as by acetylation, methylation, hydroxylation, etc., including 1-methyl-adenosine, 1-methyl-guanosine, 1-methyl-inosine, 2,2-dimethyl-guanosine, 2,6-diaminopurine, 2'-amino-2'-deoxyadenosine, 2'-amino-2'-deoxycytidine, 2'-amino-2'-deoxyguanosine, 2'-amino-2'-deoxyuridine, 2-amino-6-chloropurine riboside, 2-aminopurine-riboside, 2'-amino-6-chloropurine riboside, Glycoadenosine, 2'-cytarabine, 2'-cytarabine, 2'-azido-2'-deoxyadenosine, 2'-azido-2'-deoxyadenosine, 2'-azido-2'-deoxyguanosine, 2'-azido-2'-deoxyuridine, 2-chloroadenosine, 2'-fluoro-2'-deoxyadenosine, 2'-fluoro-2'-deoxyadenosine, 2'-fluoro-2'-deoxyguanosine, 2'-fluoro-2'-deoxyuridine, 2'-fluorothymidine, 2-methyl-adenosine, 2-methyl-guanosine, 2-methyl-thio-N6-isoopenenyl-adenosine, 2'-O-methyl-2-aminoadenosine, 2'-O-methyl-2'-deoxy Adenosine, 2'-O-methyl-2'-deoxycytidine, 2'-O-methyl-2'-deoxyguanosine, 2'-O-methyl-2'-deoxyuridine, 2'-O-methyl-5-methyluridine, 2'-O-methylinosine, 2'-O-methylpseuuridine, 2-thiocytidine, 2-thio-cytosine, 3-methyl-cytosine, 4-acetyl-cytosine, 4-thiouridine, 5-(hydroxymethyl)-uracil, 5,6-dihydrouridine, 5-aminoallylcytidine, 5-aminoallyl-deoxy-uridine, 5-bromouridine, 5-carboxymethylaminomethyl-2-thio-uracil, 5-carboxymethylaminomethyl-uracil, 5-chloro -Agar-cytosine, 5-fluorouridine, 5-iodouridine, 5-methoxycarbonylmethyluridine, 5-methoxyuridine, 5-methyl-2-thiouridine, 6-azacytidine, 6-azauridine, 6-chloro-7-deazaguanidine, 6-chloropurine riboside, 6-mercaptoguanidine, 6-methyl-mercaptopurine riboside, 7-deaza-2'-deoxyguanidine, 7-deazaadenosine, 7-methylguanidine, 8-azaadenosine, 8-bromo-adenosine, 8-bromo-guanidine, 8-mercaptoguanidine, 8-oxoguanidine, benzimidazole-riboside, β-D-mannosyl-queosine, dihydrouracil, inosine, N. 1 -Methyladenosine, N 6 -([6-aminohexyl]carbamoylmethyl)-adenosine, N 6 -Isopropenyl-adenosine, N 61,7-methyl-adenosine, N7-methyl-flavonoid, N-uracil-5-oxyacetic acid methyl ester, puromycin, queosine, uracil-5-oxyacetic acid, uracil-5-oxyacetic acid methyl ester, wybutoxosine, flavonoid, and xylose-adenosine. The preparation of such analogues is known to those skilled in the art, for example, from U.S. Patents 4,373,071, 4,401,796, 4,415,732, 4,458,066, 4,500,707, 4,668,777, 4,973,679, 5,047,524, 5,132,418, 5,153,319, 5,262,530, and 5,700,642. In the case of analogues as described above, particular preferences may be given in certain embodiments of the present invention for those analogues that increase protein expression of encoded peptides or proteins or increase the immunogenicity of the artificial nucleic acid molecules of the present invention and / or do not interfere with further modifications of the artificial nucleic acid molecules already introduced.

[0939] According to a specific implementation, the artificial nucleic acid molecule of the present invention may include lipid modification.

[0940] In a preferred embodiment, the artificial nucleic acid molecule preferably includes the following elements from the 5' to 3' direction:

[0941] A 5'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule (preferably a nucleic acid sequence of any one of SEQ ID NO: 1-151); or another 5'-UTR, preferably a 5'-TOP UTR;

[0942] At least one open reading frame (ORF), wherein the ORF preferably contains at least one modification relative to the wild-type sequence;

[0943] A 3'-UTR element that provides high translation efficiency for the artificial nucleic acid molecule (preferably a nucleic acid sequence of any one of SEQ ID NO: 152-204); or another 3'-UTR, preferably albumin 73'-UTR;

[0944] A polyadenylation sequence, preferably comprising 64 adenosine nucleotides;

[0945] The polycytosine sequence preferably contains 30 adenosine nucleotides;

[0946] Histone stem-loop sequence.

[0947] In a particularly preferred embodiment, the artificial nucleic acid molecule of the present invention may further comprise one or more of the modifications described below:

[0948] Chemical modification:

[0949] As used in this article, the term "modification" in relation to artificial nucleic acid molecules can refer to chemical modifications, including skeletal modifications as well as sugar or base modifications.

[0950] In this case, the artificial nucleic acid molecule, preferably an RNA molecule as defined herein, may contain nucleotide analogs / modifications, such as backbone modifications, sugar modifications, or base modifications. Backbone modifications relevant to this invention are modifications in which the phosphate ester of the backbone of the nucleotide contained in a nucleic acid molecule as defined herein is chemically modified. Sugar modifications relevant to this invention are chemical modifications of the sugars in the nucleotides of a nucleic acid molecule as defined herein. Furthermore, base modifications relevant to this invention are chemical modifications of the base portions of the nucleotides in a nucleic acid molecule. In this case, the nucleotide analog or modification is preferably selected from nucleotide analogs that can be used for transcription and / or translation.

[0951] Sugar modification:

[0952] As described herein, modified nucleosides and nucleotides that can be incorporated into artificial nucleic acid molecules (preferably RNA) can be modified at the sugar moiety. Examples of "oxy"-2'-hydroxyl modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g., R=H, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar); polyethylene glycol (PEG), -O(CH2CH2O)nCH2CH2OR; wherein the 2' hydroxyl group is, for example, a "locked" nucleic acid (LNA) linked to the 4' carbon of the same ribose via a methylene bridge; and amino (-O-amino, wherein the amino group, e.g., NRR, can be alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, polyamino) or aminoalkoxy.

[0953] "Deoxygenation" modification includes hydrogen, amino (e.g., NH2; alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid); or amino can be linked to sugar by a linker, wherein the linker contains one or more of atoms C, N, and O.

[0954] The sugar group may also contain more than one carbon with an opposite stereochemical configuration to the corresponding carbon in ribose. Therefore, modified nucleic acid molecules may include nucleotides containing, for example, arabinose as sugars.

[0955] Skeletal modification:

[0956] The phosphate backbone can be further modified in nucleosides and nucleotides that can be incorporated into artificial nucleic acid molecules (preferably RNA) as described herein. The phosphate groups of the backbone can be modified by replacing one or more oxygen atoms with different substituents. Furthermore, the modified nucleosides and nucleotides can include complete replacement of the unmodified phosphate ester moiety with a modified phosphate ester as described herein. Examples of modified phosphate groups include, but are not limited to, thiophosphate, phosphoroselenates, boranophosphates, boranophosphates, hydrogen phosphates, aminophosphates, alkyl or aryl phosphonates, and phosphate triesters. In dithiophosphates, the two unlinked oxygen atoms are replaced by sulfur. The phosphate linker can also be modified by replacing the linked oxygen atoms with nitrogen (bridged aminophosphate), sulfur (bridged thiophosphate), and carbon (bridged methylene-phosphonate).

[0957] Base modification:

[0958] The nucleosides and nucleotides described herein that can be incorporated into artificial nucleic acid molecules (preferably RNA molecules) can also be modified at their nucleobase sites. Examples of nucleobases found in RNA include, but are not limited to, adenine, guanine, cytosine, and uracil. For example, the nucleosides and nucleotides described herein can be chemically modified at the main groove surface. In some embodiments, the main groove chemical modification may include amino, thiol, alkyl, or halogenated groups.

[0959] In a particularly preferred embodiment of the invention, the nucleotide analog / modification is selected from base modifications, preferably from 2-amino-6-chloropurine riboside-5'-triphosphate, 2-aminopurine-riboside-5'-triphosphate; 2-aminoadenosine-5'-triphosphate, 2'-amino-2'-deoxycytidine-triphosphate, 2-thiocytidine-5'-triphosphate, 2-thiouridine-5'-triphosphate, 2'-fluorothymidine-5'-triphosphate, 2'-O-methyl Inosine-5'-triphosphate, 4-thiouridine-5'-triphosphate, 5-aminoallylcytidine-5'-triphosphate, 5-aminoallylcytidine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, 5-bromouridine-5'-triphosphate, 5-bromo-2'-deoxycytidine-5'-triphosphate, 5-bromo-2'-deoxyuridine-5'-triphosphate, 5-iodocytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5- Iodouridine-5'-triphosphate, 5-iodino-2'-deoxyuridine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 5-methyluridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-propynyl-2'-deoxyuridine-5'-triphosphate, 6-azacytidine-5'-triphosphate, 6-azauridine-5'-triphosphate, 6-chloropurine riboside-5'-triphosphate, 7-deazaadenosine-5'-triphosphate Phosphate, 7-deazaguanosine-5'-triphosphate, 8-azaadenosine-5'-triphosphate, 8-azidoadenosine-5'-triphosphate, benzimidazole-riboside-5'-triphosphate, N1-methyladenosine-5'-triphosphate, N1-methylguanosine-5'-triphosphate, N6-methyladenosine-5'-triphosphate, O6-methylguanosine-5'-triphosphate, pseudouridine-5'-triphosphate, or puromycin-5'-triphosphate, xanthoside-5'-triphosphate. A particular preference for base modifications is given for the group of nucleotides selected from the base-modified nucleotides consisting of 5-methylcytidine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, and pseudouridine-5'-triphosphate.

[0960] In some embodiments, the modified nucleosides include pyridine-4-ketoribonucleotide, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-pseudouridine, 5-propynyluridine, 1-propynyl-pseudouridine, 5-tauronic acid methyluridine, 1-tauronic acid methyl-2-thio-uridine, 1-taurine Methyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.

[0961] In some embodiments, the modified nucleosides include 5-aza-cytidine, pseudocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudocytidine, pyrrolo-cytidine, pyrrolo-pseudocytidine, 2-thio-cytidine, 2-thio-5-methylcytidine, 4-thio-pseudocytidine, 4-thio-1-methyl-pseudocytidine, 4-thio-cytidine 1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.

[0962] In other embodiments, the modified nucleosides include 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, and N6-methyladenosine. N6-Isopentenyl adenosine, N6-(cis-hydroxyisopentenyl) adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl) adenosine, N6-glycylcarbamoyl adenosine, N6-threonylcarbamoyl adenosine, 2-methylthio-N6-threonylcarbamoyl adenosine, N6,N6-dimethyl adenosine, 7-methyl adenosine, 2-methylthio-adenosine, and 2-methoxy-adenosine.

[0963] In other embodiments, the modified nucleosides include inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methyl-guanosine, N2-methyl-guanosine, N2,N2-dimethyl-guanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0964] In some embodiments, the nucleotide may be modified on the main groove surface and may include replacing the hydrogen at C-5 of uracil with a methyl or halogenated group. In specific embodiments, the modified nucleoside is 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine, 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-pseuuridine.

[0965] In a further specific embodiment, the artificial nucleic acid molecule, preferably an RNA molecule, may contain nucleoside modifications selected from the following: 6-aza-cytidine, 2-thio-cytidine, α-thio-cytidine, pseudo-iso-cytidine, 5-aminoallyl-uridine, 5-iodo-uridine, N1-methyl-pseudo-uridine, 5,6-dihydrouridine, α-thio-uridine, 4-thio-uridine, 6-aza-uridine, 5-hydroxy-uridine, deoxy- - Thymidine, 5-methyl-uridine, pyrrolo-cytidine, inosine, α-thio-guanosine, 6-methyl-guanosine, 5-methyl-cytidine, 8-oxo-guanosine, 7-deaza-guanosine, N1-methyl-adenosine, 2-amino-6-chloro-purine, N6-methyl-2-amino-purine, pseudo-iso-cytidine, 6-chloro-purine, N6-methyl-adenosine, α-thio-adenosine, 8-azido-adenosine, 7-deaza-adenosine.

[0966] Lipid modification:

[0967] According to a further embodiment, the artificial nucleic acid molecule, preferably RNA, as defined herein, may contain lipid modifications. The lipid-modified RNA typically comprises RNA as defined herein. The lipid-modified RNA molecule, as defined herein, typically also comprises at least one adapter covalently linked to the RNA molecule, and at least one lipid covalently linked to the respective adapter. Alternatively, the lipid-modified RNA molecule comprises at least one RNA molecule as defined herein and at least one (bifunctional) lipid covalently linked (without adapter) to the RNA molecule. According to a third alternative, the lipid-modified RNA molecule comprises an artificial nucleic acid molecule, preferably RNA, as defined herein, at least one adapter covalently linked to the RNA molecule, and at least one lipid covalently linked to the respective adapter, and further at least one (bifunctional) lipid covalently linked (without adapter) to the RNA molecule. In this case, it is particularly preferred that the lipid modification is present at the end of the linear RNA sequence.

[0968] Modification of the 5' end of the modified RNA:

[0969] According to another preferred embodiment of the invention, an artificial nucleic acid molecule, preferably an RNA molecule, as defined herein, can be modified by adding a so-called "5' cap" structure.

[0970] A 5'-cap is an entity, typically a modified nucleotide entity, that is usually "capped" at the 5' end of mature mRNA. The 5'-cap can typically be formed from modified nucleotides, particularly derivatives of guanine nucleotides. Preferably, the 5'-cap is linked to the 5'-terminus via a 5'-5'-triphosphate linker. The 5'-cap can be methylated, for example, m7GpppN, where N is the terminal 5' nucleotide of the nucleic acid carrying the 5'-cap, typically the 5'-terminus of RNA. m7GpppN is a 5'-cap structure naturally present in mRNA transcribed by polymerase II and is therefore not considered a modification included in the modified RNA described herein. This means that the artificial nucleic acid molecule, preferably an RNA molecule, described herein may contain m7GpppN as a 5'-cap, but furthermore, the artificial nucleic acid molecule, preferably an RNA molecule, contains at least one further modification as defined herein.

[0971] Further examples of 5' cap structures include glycerol groups, reverse deoxygenated non-basic residues (partially), 4',5' methylene nucleotides, 1-(β-D-erythrofuranosyl) nucleotides, 4'-thionucleotides, carbocyclic nucleotides, 1,5-dehydrated hexitol nucleotides, L-nucleotides, α-nucleotides, modified base nucleotides, threo-pentafuranosyl nucleotides, acyclic 3',4'-open nucleotides, acyclic 3,4-dihydroxybutyl nucleotides, acyclic 3,5-dihydroxypentyl nucleotides, 3'-3'-reverse nucleotide moiety, 3'-3'-reverse non-basic moiety, 3'-2'-reverse nucleotide moiety, 3'-2'-reverse non-basic moiety, 1,4-butanediol phosphate, 3'-aminophosphate, hexyl phosphate, aminohexyl phosphate, 3'-phosphate, 3'-thiophosphate, dithiophosphate, or bridged or unbridged methyl phosphonate moiety. These modified 5'-cap structures are considered to be at least one modification contained in the artificial nucleic acid molecule (preferably RNA molecule) described in this invention.

[0972] Particularly preferred modified 5'-cap structures are CAP1 (ribose methylation of the nucleotide adjacent to m7G), CAP2 (ribose methylation of the second nucleotide downstream of m7G), CAP3 (ribose methylation of the third nucleotide downstream of m7G), CAP4 (ribose methylation of the fourth nucleotide downstream of m7G), ARCA (anti-reactive CAP analogs, modified ARCA (e.g., phosphate thioester modified ARCA), inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.

[0973] In a preferred embodiment, at least one read frame encodes a therapeutic protein or peptide. In another embodiment, the antigen is encoded by at least one read frame, such as a pathogenic antigen, tumor antigen, sensitizing antigen, or autoimmune antigen. The administration of the artificial nucleic acid molecule encoding the antigen is used in a gene vaccination method targeting a disease involving said antigen.

[0974] In an alternative embodiment, the antibody or antigen-specific T-cell receptor or a fragment thereof is encoded by at least one read frame of the artificial nucleic acid molecule described in this invention.

[0975] antigen:

[0976] Pathogenic antigen:

[0977] The artificial nucleic acid molecules described in this invention can encode proteins or peptides containing pathogenic antigens or fragments, variants, or derivatives thereof. The pathogenic antigens are derived from pathogenic organisms, particularly bacteria, viruses, or protozoan (multicellular) pathogenic organisms, which elicit an immunological response in subjects, particularly mammalian subjects, and more particularly humans. More specifically, the pathogenic antigens are preferably surface antigens, for example, proteins (or fragments of proteins, e.g., the outer portion of a surface antigen) located on the surface of a viral, bacterial, or protozoan organism.

[0978] The pathogenic antigen is preferably a peptide or protein antigen derived from pathogens associated with the infectious disease, and is preferably selected from antigens derived from the following pathogens: *Acinetobacter baumannii*, *Anaplasma* genus, *Anaplasma phagocytophilum*, *Ancylostoma braziliense*, *Ancylostoma duodenale*, *Arcanobacterium haemolyticum*, *Ascaris lumbricoides*, *Aspergillus* genus, *Astroviridae*, *Babesia* genus, *Bacillus anthracis*, *Bacillus cereus*, *Bartonella henselae*, BK virus, and *Blastocystis*. *Hominis*, *Blastomyces dermatitidis*, *Bordetella pertussis*, *Borrelia burgdorferi*, *Borreliagenus*, *Borrelia spp*, *Brucella genus*, *Brugia malayi*, Bunyaviridae family, *Burkholderia cepacia* and other *Burkholderia* species, *Burkholderia mallei*, *Burkholderia pseudomallei*, Caliciviridae family, *Campylobactergenus*, *Candida albicans*, *Candida* species. Chlamydia trachomatis (spp.), Chlamydia pneumoniae (Chlamydia pneumoniae), Chlamydia psittaci (Chlamydia psittaci), CJD prions, Clonorchis sinensis (Clonorchis sinensis),Clostridium botulinum, Clostridium difficile, Clostridium perfringens, Clostridium spp, Clostridium tetani, Coccidioides spp, coronaviruses, Corynebacterium diphtheriae, Coxiella burnetii, Crimean-Congo hemorrhagic fever virus, Cryptococcus neoformans, Cryptosporidium genus, Cytomegalovirus (CMV), Dengue virus Viruses (DEN-1, DEN-2, DEN-3, and DEN-4), *Dientamoeba fragilis*, Ebola virus (EBOV), *Echinococcus* genus, *Ehrlichia chaffeensis*, *Ehrlichia ewingii*, *Ehrlichia* genus, *Entamoebahistolytica*, *Enterococcus* genus, *Enterovirus* genus, Enteroviruses, mainly Coxsackie A virus and Enterovirus 71 (EV71), *Epidermophyton* species, Ebola virus (Epstein-Barr). Virus (EBV), Escherichia coli O157:H7, O111 and O104:H4, Fasciola hepatica and Fasciola gigantica, FFI prions, Filarioidea superfamily, Flaviviruses, Francisella tularensis,Fusobacterium genus, Geotrichum candidum, Giardia intestinalis, Gnathostoma spp., GSS prions, Guanarito virus, Haemophilus ducreyi, Haemophilus influenzae, Helicobacter pylori, Henipavirus (Hendra virus, Nipah virus), Hepatitis A virus, Hepatitis B virus (HBV), Hepatitis C virus (HCV), Hepatitis D virus, Hepatitis E virus, Herpes simplex virus HSV-1 and HSV-2, Histoplasma capsulatum, HIV (Human Immunodeficiency Virus), Hortaea werneckii, Human bocavirus (HBoV), Human herpesvirus 6 (HHV-6) and Human herpesvirus 7 (HHV-7), Human metapneumovirus (hMPV), Human papillomavirus (HPV), Human parainfluenzaviruses (HPIV), Japanese encephalitis virus, JC virus, Junin virus, Kingella kingae, Klebsiella granulomatis, Kuru prion, Lassa virus (virus), Legionella pneumophila, Leishmania genus, Leptospira genus, Listeria monocytogenes,Lymphocytic choriomeningitis virus (LCMV), Machupo virus, *Malassezia spp* species, Marburg virus, Measles virus, *Metagonimus yokagawai*, Microsporidia phylum, Molluscum contagiosum virus (MCV), Mumps virus, *Mycobacterium leprae* and *Mycobacterium lepromatosis*, *Mycobacterium tuberculosis*, *Mycobacterium ulcerans*, *Mycoplasma pneumoniae*, *Naegleria fowleri* fowleri), Necatoramericanus, Neisseria gonorrhoeae, Neisseria meningitidis, Nocardia asteroides, Nocardia spp, Onchocerca volvulus, Orientia tsutsugamushi, Orthomyxoviridae family (Influenza), Paracoccidioides brasiliensis, Paragonimus spp, Paragonimus westermani, Parvovirus B19, Pasteurella genus, Plasmodium genus, Pneumocystis jirovecii jirovecii), poliovirus, rabies virus, respiratory syncytial virus (RSV), rhinovirus, rhinoviruses, Rickettsia akari,Rickettsia genus, Rickettsia prowazekii, Rickettsia rickettsii, Rickettsia typhi, Rift Valley fever virus, Rotavirus, Rubella virus, Sabia virus, Salmonella genus, Sarcoptes scabiei, SARS coronavirus, Schistosoma genus, Shigella genus, Sin Nombre virus, Hantavirus, Sporothrix schenckii), Staphylococcus genus, Staphylococcus genus, Streptococcus agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, Strongyloides stercoralis, Taenia genus, Taenia solium, Tick-borne encephalitis virus (TBEV), Toxocara canis or Toxocara cati, Toxoplasma gondii, Treponema pallidum, Trichinella spiralis, Trichomonas vaginalis, Trichophyton species spp.), Trichuris trichiura, Trypanosoma brucei, Trypanosoma cruzi, Ureaplasma urealyticum, Varicella zoster virus (VZV), Varicella zoster virus (VZV), Variola major or Variola minor,vCJD prions, Venezuelan equine encephalitis virus, Vibrio cholerae, West Nile virus, Western equine encephalitis virus, Wuchereria bancrofti, Yellow fever virus, Yersinia enterocolitica, Yersinia pestis, and Yersinia pseudotuberculosis.

[0979] In this case, antigens derived from pathogens selected from the following are particularly preferred: influenza virus, respiratory syncytial virus (RSV), herpes simplex virus (HSV), human papillomavirus (HPV), human immunodeficiency virus (HIV), Plasmodium, Staphylococcus aureus, dengue virus, Chlamydia trachomatis, cytomegalovirus (CMV), hepatitis B virus (HBV), Mycobacterium tuberculosis, rabies virus, and yellow fever virus.

[0980] Tumor antigen:

[0981] In a further embodiment, the artificial nucleic acid molecule of the present invention can encode a protein or peptide, comprising a peptide or protein containing a tumor antigen, a fragment of the tumor antigen, a variant of the tumor antigen, or a derivative thereof. Preferably, the tumor antigen is a melanocyte-specific antigen, a testicular cancer antigen, or a tumor-specific antigen, preferably a CT-X antigen, a non-X CT-antigen, a binding partner for the CT-X antigen, or a binding partner for a non-X CT-antigen or a tumor-specific antigen, more preferably a CT-X antigen, a binding partner for a non-X CT-antigen or a tumor-specific antigen, or a fragment of the tumor antigen, a variant of the tumor antigen, or a derivative thereof; and wherein each nucleic acid sequence encodes a different peptide or protein; and wherein at least one nucleic acid sequence encodes 5T4, 707-AP, 9D7, AFP, or AlbZIP. HPG1, α-5-β-1-integrin, α-5-β-6-integrin, α-actin-4 / m, α-methylacyl-CoA racemicase, ART-4, ARTC1 / m, B7H4, BAGE-1, BCL-2, bcr / abl, β-binin / m, BING-4, BRCA1 / m, BRCA2 / m, CA 15-3 / CA 27-29, CA 19-9, CA72-4, CA125, calreticulin, CAMEL, CASP-8 / m, cathepsin B, cathepsin L, CD19, CD20, CD22, CD25, CDE30, CD33, CD4, CD52, CD55, CD56, CD80, CDC27 / m, CDK4 / m, CDKN2A / m, CEA, CLCA2, CML28, CML66, COA-1 / m, coactosin-like protein, collagen XXIII, COX-2, CT-9 / BRD6, Cten, cyclin B1, cyclin D1, cyp-B, CYPB1, DAM-10, DAM-6, DEK-CAN, EFTUD2 / m, EGFR, ELF2 / m, EMMPRIN, EpC am, EphA2, EphA3, ErbB3, ETV6-AML1, EZH2, FGF-5, FN, Frau-1, G250, GAGE-1, GAGE-2, GAGE-3, GAGE-4, GAGE-5, GAGE-6, GAGE7b, GAGE-8, GDEP, GnT-V, gp100, GPC3, GPNMB / m, H AGE, HAST-2, hepsin, Her2 / neu, HERV-K-MEL, HLA-A*0201-R17I, HLA-A11 / m, HLA-A2 / m, HNE, homeobox NKX3.1, HOM-TES-14 / SCP-1, HOM-TES-85, HPV-E6, HPV-E7, HSP70-2M, HST-2,hTERT, iCE, IGF-1R, IL-13Ra2, IL-2R, IL-5, immature laminin receptor, kallikrein-2, kallikrein-4, Ki67, KIAA0205, KIAA0205 / m, KK-LC-1, K-Ras / m, LAGE-A1, LDLR-FUT, MAGE-A1, MAGE-A2, MAGE-A3, MAGE-A4, MAGE-A6, MAGE-A9, MAGE-A10, MAGE-A12, MAGE-B1, MAGE-B2, MAGE-B3, MAGE-B4, MAGE -B5, MAGE-B6, MAGE-B10, MAGE-B16, MAGE-B17, MAGE-C1, MAGE-C2, MAGE-C3, MAGE-D1, MAGE-D2, MAGE-D4, MAGE-E1, MAGE-E2, MAGE-F1, MAGE-H1, MAGEL2, mammary globulin A, MART-1 / melan-A, MART-2, MART-2 / m, matrix protein 22, MC1R, M-CSF, ME1 / m, mesothelin, MG50 / PXDN, MMP11, MN / CA IX-antigen, MRP-3, MUC-1, MUC-2, MUM-1 / m, MUM-2 / m, MUM-3 / m, type I myosin I / m, NA88-A, N-acetylglucosamine transferase-V, Neo-PAP, Neo-PAP / m, NFYC / m, NGEP, NMP22, NPM / ALK, N-Ras / m, NSE, NY-ESO-1, NY-ESO-B, OA1, OFA-iLRP, OGT, OGT / m, OS-9, OS-9 / m, osteocalcin, osteopontin, p15, p190 small bcr-abl, p53, p53 / m, PAGE-4, PAI-1, PAI-2, PAP, PART-1, PATE, PDEF, Pim-1-kinase, Pin-1, Pml / PARα, POTE, PRAME, PRDX5 / m, prostein, protease-3, PSA, PSCA, PSGR, PSM, PSMA, PTPRK / m, RAGE-1, RBAF600 / m, RHAMM / CD168, RU1, RU2, S-100, SAGE, SART-1, SART-2, SART-3, SCC, SIRT2 / m, Sp17, SSX-1, SSX-2 / HOM-MEL-40, SSX-4, STAMP-1, STEAP-1, survivin, survivin-2B, SYT-SSX-1, SYT-SSX-2, TA-90, TAG-72, TARP, TEL-AML1, TGFβTGFβRII, TGM-4, TPI / m, TRAG-3, TRG, TRP-1, TRP-2 / 6b, TRP / INT2, TRP-p8, tyrosinase, UPA, VEGFR1, VEGFR-2 / FLK-1, WT1, and lymphocyte immunoglobulin idiotypes or lymphocyte T cell receptor idiotypes, or fragments, variants, or derivatives of the tumor antigen; preferably survivin or its homologues, MAGE-family antigens or their binding partners, or fragments, variants, or derivatives of the tumor antigen. Particularly preferred in this case are tumor antigens NY-ESO-1, 5T4, MAGE-C1, MAGE-C2, survivin, Muc-1, PSA, PSMA, PSCA, STEAP, and PAP.

[0982] In a preferred embodiment, the artificial nucleic acid molecule encodes a protein or peptide, which comprises a therapeutic protein or a fragment, variant, or derivative thereof.

[0983] Therapeutic proteins, as defined herein, are peptides or proteins that are beneficial for the treatment of any genetic or acquired disease or for improving an individual's condition. Specifically, among other functions, therapeutic proteins play a significant role in generating therapeutic agents that can modify and repair genetic errors, destroy cancer cells or pathogen-infected cells, treat immune system disorders, and treat metabolic or endocrine disorders. For example, erythropoietin (EPO) (a protein hormone) can be used to treat patients with erythrocyte deficiency, a common cause of kidney complications. Furthermore, therapeutic proteins encompass adjuvant proteins, therapeutic antibodies, and hormone replacement therapy, for example, for the treatment of postmenopausal women. In recent approaches, a patient's somatic cells are used to reprogram them into pluripotent stem cells, which replace controversial stem cell therapies. Furthermore, these proteins used for reprogramming somatic cells or for differentiating stem cells are defined herein as therapeutic proteins. In addition, therapeutic proteins can be used for other purposes, such as wound healing, tissue regeneration, angiogenesis, etc. Furthermore, antigen-specific B-cell receptors and their fragments and variants are defined herein as therapeutic proteins.

[0984] Therefore, therapeutic proteins can be used for a variety of purposes, including the treatment of a variety of diseases such as, for example, infectious diseases, tumors (e.g., cancer or neoplasm), diseases of the blood and blood-forming organs, endocrine, nutritional and metabolic diseases, diseases of the nervous system, diseases of the circulatory system, diseases of the respiratory system, diseases of the digestive system, diseases of the skin and subcutaneous tissue, diseases of the musculoskeletal system and connective tissue, and diseases of the reproductive and urinary systems, whether they are genetic or acquired.

[0985] In this context, particularly preferred therapeutic proteins that can be used specifically for the treatment of metabolic or endocrine disorders are selected from the following (the specific disease in parentheses is the one in which the therapeutic protein is used in treatment): acid sphingomyelinase (Niemann-Pick disease), adipotide (obesity), agalsidase-β (human galactosidase A) (Fabry disease; prevents fat accumulation that may lead to renal and cardiovascular complications), alglucosidase (Pompe disease (glycogen storage disease type II)), α-galactosidase A (α-GAL A, agalsidase α) (Fabry disease), α-glucosidase (glycogen storage disease (GSD), Morbus Pompe), α-L-iduronidase (mucopolysaccharidoses (MPS), Hurler syndrome, Scheie syndrome), α-N-acetylglucosidase (Sanfilippo syndrome), bimodalin (cancer, metabolic disorder), angiopoietin (Ang1, Ang2, Ang3, Ang4, ANGPTL2, ANGPTL3, ANGPTL4, ANGPTL5, ANGPTL6, ANGPTL7) (angiogenesis, vascular stabilization), β-animal cellulose (metabolic disorder). Disorders), β-glucuronidase (Slysyndrome), bone morphogenetic proteins BMP (BMP1, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8a, BMP8b, BMP10, BMP15) (regenerative role, skeletal-related conditions, chronic kidney disease (CKD)), CLN6 protein (CLN6 disease – atypical late infancy, late-onset variants, early adolescence, neuronal ceroid lipofuscinoses (NCL)), epidermal growth factor (EGF) (wound healing, regulation of cell growth, proliferation and differentiation), Epigen (metabolic disorders), epidermal regulatory factors (metabolic disorders), fibroblast growth factors (FGF, FGF-1, FGF-2, FGF-3, FGF-4)FGF-5, FGF-6, FGF-7, FGF-8, FGF-9, FGF-10, FGF-11, FGF-12, FGF-13, FGF-14, FGF-16, FGF-17, FGF-17, FGF-18, FGF-19, FGF-20, FGF-21, FGF-22, FGF-23 (wound healing, angiogenesis, endocrine disorders, tissue regeneration), Galsulphase (Mucopolysaccharidosis VI), Ghrelin (irritable bowel syndrome (IBS), obesity, Prader-Willi syndrome, type II diabetes mellitus), glucocerebrosidase (Gaucher's disease) Disease), GM-CSF (regeneration, production of white blood cells, cancer), heparin-bound EGF-like growth factor (HB-EGF) (wound healing, cardiac hypertrophy and cardiac development and function), hepatocyte growth factor HGF (regeneration, wound healing), hepcidin (iron metabolism disorders, beta-thalassemia), human albumin (decreased albumin production (hypoproteinemia), increased albumin loss (nephrotic syndrome), hypovolemia, hyperbilirubinemia), idurose-2-sulfatase (mucopolysaccharidosis II, Hunter syndrome). Integrin αVβ3, αVβ5, and α5β1 (binding matrix macromolecules and proteases, angiogenesis), iduronate sulfatase (Hunter syndrome), laronidase (Hurler and Hurler-Scheie forms of mucopolysaccharidosis I), N-acetylgalactosamine-4-sulfatase (rhASB; galsulfase, arylsulfatase A (ARSA), arylsulfatase B (ARSB)) (arylsulfatase B deficiency),Maroteaux–Lamy syndrome, mucopolysaccharidosis VI, N-acetylglucosamine-6-sulfatase (Sanfilippo syndrome), nerve growth factor (NGF, brain-derived neurotrophic factor (BDNF), neurotrophic factor-3 (NT-3), and neurotrophic factor-4 / 5 (NT-4 / 5)) (regenerative effects, cardiovascular diseases, coronary atherosclerosis, obesity, type 2 diabetes, metabolic syndrome, acute coronary syndromes, dementia, depression, schizophrenia, autism, Rett syndrome, anorexia nervosa, bulimia nervosa, wound healing, skin ulcers, corneal ulcers) Ulcers), Alzheimer's disease, neuromodulatory proteins (NRG1, NRG2, NRG3, NRG4) (metabolic disorders, schizophrenia), neurofeltin (NRP-1, NRP-2) (angiogenesis, axonal guidance, cell survival, migration), obesstatin (irritable bowel syndrome, IBS, obesity, Prader-Willi syndrome, type II diabetes), platelet-derived growth factor (PDGF (PDFF-A, PDGF-B, PDGF-C, PDGF-D)) (regeneration, wound healing, disorders in angiogenesis, arteriosclerosis, fibrosis, cancer), TGFβ receptors (endothelial factor, TGF-β1 receptor, TGF-β2 receptor, TGF-β3 receptor) (renal fibrosis, kidney disease). Diseases, diabetes, end-stage renal disease (ESRD), angiogenesis.Thrombopoietin (THPO) (megakaryocyte growth and development factor (MGDF)) (platelet disorders, platelet donation, recovery of platelet count after myelosuppressive chemotherapy), Transforming growth factor (TGF (TGF-α, TGF-β (TGFβ1, TGFβ2 and TGFβ3)) (regeneration, wound healing, immunity, cancer, heart disease, diabetes, Marfan syndrome, Loeys–Dietz syndrome), VEGF (VEGF-A, VEGF-B, VEGF-C, VEGF-D, VEGF-E, VEGF-F and PIGF) (regeneration, angiogenesis, wound healing, cancer, permeability), Nesiritide (acute decompensated congestive heart failure), Trypsin (decubbitus ulcers) Varicose ulcer, eschar debridement, dehiscent wound, sunburn, meconium ileus, adrenocorticotrophic hormone (ACTH) (Addison's disease, small cell carcinoma, adrenoleukodystrophy, congenital adrenal hyperplasia, Cushing's syndrome, Nelson's syndrome, infantile spasms), atrial-natriuretic peptide (ANP) (endocrine disorders) Disorders), cholecystokinin (different), gastrin (hypogastrinemia), leptin (diabetes, hypertriglyceridemia, obesity),Oxytocin (to stimulate breastfeeding, non-progression of parturition), somatostatin (for carcinoid syndrome, acute variceal bleeding and acromegaly, polycystic diseases of the liver and kidney, acromegaly and symptomatic treatment of symptoms caused by neuroendocrine tumors), vasopressin (antidiuretic hormone) (for diabetes insipidus), calcitonin (for postmenopausal osteoporosis, hypercalcemia, Paget's disease, bone metastases, phantom limb pain, spinal stenosis). Stenosis), Exenatide (for type 2 diabetes resistant to metformin and sulfonylurea), growth hormone (GH), somatotropin (for growth failure due to GH deficiency, chronic renal insufficiency, Prader-Willi syndrome, Turner syndrome, AIDS wasting, or cachexia from antiviral therapy), insulin (for diabetes, diabetic ketoacidosis, and hyperkalemia), insulin-like growth factor 1 (IGF-1) (for growth retardation in children with GH gene deletion or severe primary IGF-1 deficiency, neurodegenerative diseases, cardiovascular diseases, and heart failure). failure), Mecaserminrinfabate.IGF-1 analogs (for growth retardation, neurodegenerative diseases, cardiovascular diseases, and heart failure in children with GH gene deletion or severe primary IGF-1 deficiency), Mecasermin, IGF-1 analogs (for growth retardation, neurodegenerative diseases, cardiovascular diseases, and heart failure in children with GH gene deletion or severe primary IGF-1 deficiency), Pegvisomant (acromegaly), Pramlintide (diabetes, in combination with insulin), Teriparatide (human parathyroid hormone residues 1–34) (severe osteoporosis), Becaplermin (adjunctive debridement for diabetic ulcers), Dibotermin-α (bone morphogenetic protein 2) (spinal fusion surgery, bone injury repair), Histrelin acetate. Acetate (gonadotropin-releasing hormone (GnRH)) for precocious puberty; octreotide (symptom relief for acromegaly, VIP-secreting adenoma, and metastatic carcinoid tumors); and palifermin (keratinocyte growth factor (KGF)) for severe oral mucositis and wound healing in patients undergoing chemotherapy.

[0986] These and other proteins should be understood as therapeutic because they are designed to treat subjects by replacing the deficient endogenous production of their functional proteins in adequate amounts. Therefore, these therapeutic proteins are typically mammalian proteins, particularly human proteins.

[0987] The following therapeutic proteins may be used to treat hematologic disorders, circulatory system diseases, respiratory system diseases, cancer or tumor diseases, infectious diseases, or immunodeficiency: Alteplase (tissue plasminogen activator; tPA) (for pulmonary embolism, myocardial infarction, acute ischemic stroke, and occlusion of central venous access devices); Anistreplase (for thrombolysis); Antithrombin III (AT-III) (for hereditary AT-III deficiency and thromboembolism); and Bivalirudin (for coronary angioplasty and heparin-induced thrombocytopenia). Reduced risk of blood clotting in thrombocytopaenia), darbepoetin-α (for anemia in patients with chronic renal insufficiency and chronic renal failure (+ / - dialysis)), drotrecogin-α (activated protein C) (for severe sepsis with high mortality risk), erythropoietin, epidotin-α, erythropoietin (for anemia in chronic diseases, myleodysplasia, anemia due to renal failure or chemotherapy, preoperative preparation), factor IX (Haemophilia B), factor VIIa (for bleeding in patients with hemophilia A or B and inhibitors of factor VIII or factor IX), facto...

Claims

1. An artificial nucleic acid molecule comprising a. At least one readable bounding box (ORF); and b. At least one 5'-untranslated region element (5'-UTR element), said at least one 5'-UTR element consisting of the nucleic acid sequence of SEQ ID NO:87 or its corresponding RNA sequence; The read frame is derived from a gene that is different from the gene from which the at least one 5'-UTR element originates; The reference nucleic acid molecule includes at least one open reading frame (ORF), which is identical to at least one ORF of the artificial nucleic acid molecule; and the reference nucleic acid molecule does not include at least one 5'-untranslated region element (5'-UTR element) of the artificial nucleic acid molecule. The 5'-UTR element of the reference nucleic acid molecule is a 5'-UTR element composed of the nucleic acid sequence shown in SEQ ID NO: 208 or its corresponding RNA sequence; The translation efficiency of the artificial nucleic acid molecule and the translation efficiency of the reference nucleic acid molecule are measured by a method comprising the steps (i) to (iii) below: (i) Transfect mammalian cells with the artificial nucleic acid molecule and measure the expression level of the protein encoded by the ORF of the artificial nucleic acid molecule at specific time points after transfection. (ii) Transfect mammalian cells with the reference nucleic acid molecule and measure the expression level of the protein encoded by the ORF of the reference nucleic acid molecule at the same time point after transfection. (iii) Calculate the ratio of the amount of protein expressed from the artificial nucleic acid molecule to the amount of protein expressed from the reference nucleic acid molecule. Wherein the ratio calculated in (iii) is > 1, the ratio being the expression level of the protein encoded by the ORF of the artificial nucleic acid molecule / the expression level of the protein encoded by the ORF of the reference nucleic acid molecule.

2. The artificial nucleic acid molecule according to claim 1, wherein the reference nucleic acid molecule includes a 3'-UTR element that is not present in the artificial nucleic acid molecule.

3. The artificial nucleic acid molecule according to claim 1, wherein it comprises at least one 3'-UTR element.

4. The artificial nucleic acid molecule of claim 3, wherein each of the at least one read frame, the at least one 3'-UTR element, and the at least one 5'-UTR element is heterologous to each other.

5. The artificial nucleic acid molecule according to claim 1, wherein, Compared to protein production from the reference nucleic acid molecule, the at least one 5'-UTR element increases the translation efficiency of the artificial nucleic acid molecule by at least 1.2 times.

6. The artificial nucleic acid molecule according to claim 1, wherein, Compared to protein production from the reference nucleic acid molecule, the at least one 5'-UTR element provides the artificial nucleic acid molecule with at least 1.5 times higher translation efficiency.

7. The artificial nucleic acid molecule according to claim 3, wherein the at least one 3'-UTR element is composed of a nucleic acid sequence or its corresponding RNA sequence as shown in any one of SEQ ID NO: 152-161, 164-165, 167-177, 179-184, 187-189 and 191-204.

8. The artificial nucleic acid molecule according to claim 3, further comprising: c. Polyadenylation sequence and / or polyadenylation signal.

9. The artificial nucleic acid molecule according to claim 8, wherein the polyadenylation sequence or the polyadenylation signal is located at the 3' of the 3'-UTR element.

10. The artificial nucleic acid molecule according to claim 8 or 9, wherein the polyadenylation signal comprises a common sequence NN(U / T)ANA, where N = A or U.

11. The artificial nucleic acid molecule according to claim 8 or 9, wherein the polyadenylation signal is located downstream of the 3'-end of the 3'-UTR element by fewer than 50 nucleic acids.

12. The artificial nucleic acid molecule according to claim 8 or 9, wherein the polyadenylate sequence has a length of 20-300 adenine nucleotides.

13. The artificial nucleic acid molecule according to claim 1, wherein the artificial nucleic acid molecule is at least partially G / C modified.

14. The artificial nucleic acid molecule of claim 13, wherein the readable frame is at least partially G / C modified.

15. The artificial nucleic acid molecule of claim 14, wherein the G / C content of the read frame is increased compared to the wild-type read frame.

16. The artificial nucleic acid molecule of claim 1, wherein the readable frame includes a codon-optimized region.

17. The artificial nucleic acid molecule of claim 16, wherein the read frame is codon-optimized.

18. The artificial nucleic acid molecule according to claim 1, wherein it is RNA.

19. The artificial nucleic acid molecule according to claim 1, further comprising a 5'-cap structure, a polycytidine sequence, a histone stem-loop, and / or an IRES motif.

20. The artificial nucleic acid molecule according to claim 18, wherein it is an mRNA molecule.

21. A vector comprising the artificial nucleic acid molecule according to any one of claims 1-20.

22. A cell comprising the artificial nucleic acid molecule of any one of claims 1-20 or the vector of claim 21.

23. A method for increasing the translation efficiency of an artificial nucleic acid molecule, the method comprising the step of concatenating a read frame with a 5'-UTR element, wherein the 5'-UTR element is composed of a nucleic acid sequence of SEQ ID NO: 87 or its corresponding RNA sequence, thereby obtaining the artificial nucleic acid molecule of any one of claims 1-20, or the vector of claim 21.

Citation Information

Patent Citations

  • Complexes of RNA and cationic peptides for transfection and for immunostimulation

    WO2009030481A1

  • Composition comprising a complexed (m)RNA and a naked mRNA for providing or enhancing an immunostimulatory response in a mammal and uses thereof

    WO2010037539A1

  • Disulfide-linked polyethyleneglycol / peptide conjugates for the transfection of nucleic acids

    WO2011026641A1

  • Complexation of nucleic acids with disulfide-crosslinked cationic components for transfection and immunostimulation

    WO2012013326A1

  • Nucleic acid comprising or coding for a histone stem-loop and a poly(a) sequence or a polyadenylation signal for increasing the expression of an encoded protein

    WO2012019780A1