Artificial nucleic acid molecule

By designing artificial nucleic acid molecules containing highly expressed gene UTR variants, the problems of low stability and translation efficiency in mRNA therapy were solved, and more efficient protein expression and reasonable dosing regimen were achieved.

CN120464655APending Publication Date: 2025-08-12SHENZHEN HONGSHENG BIOTECHNOLOGIES CO LTD

Patent Information

Application Number
CN202410178112.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-08
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In existing mRNA therapies, the short half-life and low translation efficiency of mRNA lead to unreasonable dosage and dosage intervals, which affects the therapeutic effect.

Method used

An artificial nucleic acid molecule is designed to include the target 5’ untranslated region (UTR) and 3’ untranslated region (UTR), optimized to be a UTR variant of highly expressed genes, combining the 5’ end cap structure and PolyA tail to improve the stability and translation efficiency of mRNA.

Benefits of technology

By optimizing the UTR sequence, the stability and translation efficiency of mRNA in vivo are significantly improved, the expression of target protein is enhanced, and the dosing regimen is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120464655A_ABST
    Figure CN120464655A_ABST
Patent Text Reader

Abstract

The invention provides an artificial nucleic acid molecule which is used for improving the expression quantity of target amino acid, polypeptide or protein. The artificial nucleic acid molecule at least comprises a target 5'untranslated region (UTR), a target coding region (CDS) and a target 3 'untranslated region (UTR). Wherein the sequence of the target 5 'UTR is one of the following sequences: 5' UTR of a high-expression gene and a 5 'UTR variant of the high-expression gene. The sequence of the target 3 'UTR is one of the following sequences: 3' UTR of a high-expression gene and a 3 'UTR variant of the high-expression gene. Optionally, the artificial nucleic acid molecule may further comprise, for example, a 5 '-end cap structure (Cap), a PolyA tail. The 5 'UTR and the 3' UTR have regulating effects on translation and stability of nucleic acid molecules, so that the 5 'UTR, the 3' UTR and variants thereof are selected from high-expression genes, the nucleic acid molecules can be further stabilized and are not easy to degrade, and the amount of protein or polypeptide obtained by translation of the nucleic acid molecules can be increased. The invention also provides methods for making, delivering, and using such artificial nucleic acid molecules, as well as the use of the artificial nucleic acid molecules for the treatment and / or prevention of related diseases or disorders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioinformatics, and in particular to an artificial nucleic acid molecule for improving protein or polypeptide expression. Background Art

[0002] In vitro transcription of mRNA uses DNA as a template and uses RNA polymerase to transcribe mRNA in vitro in a cell-free system (IVT-mRNA) to simulate the process of mRNA synthesis in vivo. After the in vitro transcribed mRNA is successfully introduced into the cells of an organism, it can produce specific proteins, such as antigens or functional proteins, thereby achieving the purpose of disease prevention or treatment. In addition to the new coronavirus mRNA vaccine, mRNA technology also has great application potential in tumor immunotherapy, gene therapy, infectious disease prevention, gene editing, genetic disease treatment and other fields.

[0003] mRNA technology uses a cell-free manufacturing method, which enables rapid, economical and efficient production, and has a unique advantage in quickly responding to large-scale outbreaks of sudden infectious diseases. Since IVT-mRNA does not undergo RNA splicing, nuclear export and other processes after entering the cell, it can efficiently express the target protein after entering the target cell. At the same time, mRNA is almost never integrated into the genome, which is highly safe and avoids the possibility of insertion mutations. In addition, a single mRNA molecule can encode multiple antigen molecules, enhance the immune response against adaptive pathogens, and can target multiple microorganisms or viral variants with a single formula.

[0004] mRNA therapy involves administering specific mRNA to a recipient to produce the protein encoded by the mRNA in the patient. However, clinical application of mRNA therapy is still limited by issues such as the short half-life of mRNA and low expression efficiency. Improving the in vivo translation efficiency of IVT-mRNA is a key technical issue in the field of mRNA drug or vaccine therapy. In this regard, mRNA stability, particularly mRNA translation efficiency, determines the dosage and dosing interval of mRNA drugs or vaccines.

[0005] The use of chemically modified nucleotides to increase the stability of mRNA and reduce the immunogenic response triggered by mRNA in cells or organisms is a proven strategy. Kormann et al. have shown that replacing only 25% of the uridine and cytidine residues with 2-thiouridine and 5-methylcytidine is sufficient to increase the stability of mRNA and reduce the activation of innate immunity triggered by exogenously administered mRNA in vitro (WO2012 / 0195936A1; WO2007024708A2).

[0006] In addition, the untranslated regions (UTRs) in mRNA play a key role in regulating the stability and translation efficiency of mRNA. UTRs affect translation initiation, elongation and termination, as well as the localization and stability of mRNA in the cell through interactions with RNA-binding proteins (G. Pesole, G. Grillo, A. Larizza and S. Liuni, Briefings in Bioinformatics, 2000, 1, 236-249.). Depending on the specific motivation within the UTR, it can enhance or reduce the translation efficiency of mRNA.

[0007] Although the prior art describes means and methods for increasing mRNA stability, reducing immunogenic responses triggered by mRNA administered to cells or organisms, and increasing translation efficiency, there remains a need for improvements, particularly with respect to additional or alternative means of increasing mRNA translation efficiency and stability. Summary of the Invention

[0008] The present application provides an artificial nucleic acid molecule for increasing the expression level of a target protein or peptide, characterized in that the artificial nucleic acid molecule comprises at least: a target 5' untranslated region (UTR); a target coding region (CDS) for being translated to produce the target amino acid; and a target 3' untranslated region (UTR). The sequence of the target 5'UTR is one of the following sequences: a 5'UTR of a highly expressed gene and a 5'UTR variant of the highly expressed gene. The sequence of the target 3'UTR is one of the following sequences: a 3'UTR of a highly expressed gene and a 3'UTR variant of the highly expressed gene. Optionally, the artificial nucleic acid molecule may further comprise a 5' end cap structure (Cap) and a PolyA tail. The highly expressed gene is a gene with the 1st to 25th highest number of transcripts per million mapped reads (nTPM) per kilobase of transcription in all types of cells of all tissues of the human body.

[0009] In some embodiments, the artificial nucleic acid molecule comprises at least one selected from the group consisting of messenger RNA (mRNA), self-amplifying RNA (saRNA), circular RNA (circRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA) and microRNA (miRNA), primary-miRNA, antisense oligonucleotide (ASO), transfer RNA (tRNA), plasmid DNA (pDNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), deoxyribozyme (DNAzyme), ribozyme (RNAzyme), nucleic acid aptamer (aptamer), clustered regularly interspaced short palindromic repeats (CRISPR)-related nucleic acid, single guide RNA (sgRNA), CRISPR-RNA (crRNA), trans-activating crRNA (tracrRNA), guide RNA (gRNA), long non-coding RNA (LncRNA), single-stranded RNA (ssRNA) and double-stranded RNA (dsRNA).

[0010] In some embodiments of the artificial nucleic acid molecule of the present invention, the target protein or peptide is a protein associated with a reporter gene, preferably the protein comprises at least one selected from the group consisting of green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), orange fluorescent protein (OFP), firefly luciferase (Fluc), Renilla luciferase (Rluc), spiny shrimp luciferase (Nluc), chloramphenicol acetyltransferase (CAT), β-galactosidase (β-gal), secretory human placental alkaline phosphatase (SEAP).

[0011] In some embodiments of the artificial nucleic acid molecule of the present invention, the target protein or peptide is an immune-related antigenic protein, and preferably the protein comprises at least one selected from the following: an antigenic protein associated with the new coronavirus (such as RBD, spike, etc.), an antigenic protein associated with Klebsiella pneumoniae, an antigenic protein associated with Pseudomonas aeruginosa, and an antigenic protein associated with Helicobacter pylori.

[0012] In some embodiments of the artificial nucleic acid molecules of the present invention, the target protein or peptide is a protein involved in translation or transcription. In some embodiments, the target protein is associated with the CRISPR process. In some embodiments, the target protein is a CRISPR-related protein. In some embodiments of the artificial nucleic acid molecules of the present invention, the artificial nucleic acid molecule can be further translated into a target protein or peptide and targeted for delivery to a target organ, tissue, or cell in a subject. The targeting moiety is preferably an encoded polypeptide and / or protein.

[0013] In some embodiments of the artificial nucleic acid molecule of the present invention, the target protein or peptide is a protein related to gene therapy. In some embodiments, the target protein is related to the treatment of pulmonary cystic fibrosis, such as the cystic fibrosis transmembrane conductance regulator (CFTR). In some embodiments, the target protein is a protein related to primary ciliary dyskinesia (PCD). In some embodiments of the artificial nucleic acid molecule of the present invention, the artificial nucleic acid molecule, i.e., the target protein or peptide, is further packaged with a targeting ligand for targeted delivery to a target organ, tissue, or cell in a subject. Preferably, the targeting moiety is a polypeptide and / or protein encoding an antibody.

[0014] In some embodiments of the artificial nucleic acid molecules of the present invention, the target protein or peptide is an adjuvant-associated protein, preferably comprising at least one selected from the group consisting of flagellin or an immunomodulatory protein such as IL-12, IL-12, GM-CSF, and TSLP. In some embodiments of the artificial nucleic acid molecules of the present invention, the target protein or peptide comprises a potentiator that enhances transfection of the artificial nucleic acid molecule, preferably comprising at least one selected from the group consisting of a pulmonary surfactant protein, a cell-penetrating peptide, an amphiphilic polypeptide, and a mucolytic enzyme.

[0015] The present application also provides a use of the above-mentioned artificial nucleic acid molecule in increasing the expression level of a target amino acid.

[0016] The present application also provides a use of the above-mentioned artificial nucleic acid molecule in the preparation of medical supplies, wherein the medical supplies include mRNA drugs and vaccines.

[0017] The present application also provides a method for delivering the artificial nucleic acid molecule to a cell, comprising: contacting the cell with the artificial nucleic acid molecule when the artificial nucleic acid molecule is capable of being taken up into the cell, wherein the cell is a mammalian cell.

[0018] The present application also provides a use of the above-mentioned artificial nucleic acid molecule in the preparation of a drug or a pharmaceutical composition, wherein the drug is used to prevent and / or treat a subject, and wherein the artificial nucleic acid molecule in the drug or the pharmaceutical composition is targeted at a disease or condition.

[0019] The present application also provides a use of the above-mentioned artificial nucleic acid molecule in the treatment and / or prevention of diseases or conditions, characterized in that the diseases or conditions include: immune system diseases, metabolic diseases, genetic diseases, cancer, blood diseases, bacterial infections or viral infections.

[0020] Other objects, features and advantages of the present disclosure will become apparent from the following detailed description. However, it should be understood that although the detailed description and specific examples indicate certain embodiments of the present disclosure, they are given by way of illustration only, as those skilled in the art will appreciate various changes and modifications within the spirit and scope of the present disclosure from this detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] This application will illustrate the embodiments with reference to the accompanying drawings. The drawings in this application are only used to describe the embodiments for illustrative purposes. Without departing from the principles of this application, those skilled in the art can easily make other embodiments according to the steps described below by following the description.

[0022] Figure 1 A schematic structural diagram of an artificial nucleic acid molecule provided in an embodiment of the present application is shown. This schematic diagram is a schematic diagram of the artificial nucleic acid molecule in subsequent embodiments and is a schematic diagram of one of the artificial nucleic acid molecules.

[0023] Figure 2 The artificial nucleic acid molecules constructed with the CDS of the Fluc gene as the target CDS and the genes TMSB10, SFTPD, AGER, LRRK2, MSR1, SCGB3A2, and SFTPA2 highly expressed in the lung as the target UTR sources were transfected into 16HBE (a), 293T (b), A549 (c), and DC (d) for 24 hours. The expression levels of the target proteins in each cell line were measured. The UTRs in α-globin, mRNA-1273 vaccine, and BNT162b2 vaccine were selected as controls.

[0024] Figure 3 The expression distribution of the target protein in mice (a) and lung tissue (b) 6 hours after the artificial nucleic acid molecule constructed with the CDS of the Fluc gene as the target CDS and the genes SCGB3A2 and AGER, which are highly expressed in the lungs, were administered to mice by intratracheal spray. The UTRs in α-globin, mRNA-1273 vaccine, and BNT162b2 vaccine were selected as controls.

[0025] Figure 4The expression distribution of the target protein in mice (a) and lung tissue (b) 6 hours after the artificial nucleic acid molecule constructed with the CDS of the Fluc gene as the target CDS and the genes TMSB10, AGER, SCGB3A2, SFTPA1, SFTPA2, and MSR1, which are highly expressed in the lungs, were administered to mice by intratracheal spray. The UTRs in α-globin, mRNA-1273 vaccine, and BNT162b2 vaccine were selected as controls.

[0026] Figure 5 The figure shows an artificial nucleic acid molecule constructed with the CDS of the Fluc gene as the target CDS, and the A7 and R3U elements added to the 5'UTR and 3'UTR of the lung highly expressed gene TMSB10 as the target UTR source, respectively, wherein the microRNA binding site is removed from the 3'UTR of TMSB10. 6 hours after the constructed artificial nucleic acid molecule was administered to mice by intratracheal spray, the expression level distribution of the target protein in the mice (a) and lung tissue (b), and the expression level of luciferase in 16HBE cells 24 hours after the constructed artificial nucleic acid was transfected into the cells (c).

[0027] Figure 6 To demonstrate the effect of the insertion position of a functional element on 5'UTR function, artificial nucleic acid molecules were constructed using the CDS of the Fluc gene as the target CDS and the A7 element as the target UTR source at different positions in the 5'UTR of the highly expressed lung gene AGER. Six hours after the constructed artificial nucleic acid molecules were administered to mice by intratracheal spray, the expression distribution of the target protein in mice (a) and lung tissue (b) was observed.

[0028] Figure 7 The following diagram shows the expression distribution of the target protein in mice (a) and lung tissue (b) 6 hours after the constructed artificial nucleic acid molecule was administered to mice by intratracheal spray, using the CDS of the Fluc gene as the target CDS and the A7 element added to the 5'UTR of the B3A2B6, A2B6, and R1B6 groups as the target UTR source.

[0029] Figure 8The CDS of the constructed Fluc gene is used as the target CDS, and the UTRs in the B6, B6A7, B3A2-1, B3A2B6, A2B6, A2A7B6, R1B6, and R1A7B6 groups are used as the target UTR-derived artificial nucleic acid molecules. The constructed artificial nucleic acid molecules were administered to mice by intratracheal spray for 6 hours. The expression level distribution of the target protein in the mouse liver (Liver) (a) and spleen (Spleen) (b) tissues was observed.

[0030] Figure 9 This figure shows the optimization of the 5'UTR sequence using computer-aided design. The CDS of the Fluc gene was used as the target CDS, and the 5'UTR sequence of the R1A7B6 group was selected for further sequence optimization using LinearDesign. This was used to construct a new artificial nucleic acid molecule. Six hours after the constructed artificial nucleic acid molecule was administered to mice via intratracheal spray, the expression distribution of the target protein in mice (a) and lung tissue (b) was observed.

[0031] Figure 10 The figure shows the expression distribution of the target protein in mice (a) and lung tissue (b) 6 hours after the constructed artificial nucleic acid molecules were administered to mice by intratracheal spray. The CDS of the Fluc gene was used as the target CDS, and the IRESvI and PABPv3 elements were added to the 5'UTR of the R1B6 and R1A7B6 groups, respectively, as the target UTR sources.

[0032] Figure 11 The figure shows an artificial nucleic acid molecule constructed with the CDS of the gene of the novel coronavirus SARS-CoV-2 spike protein receptor domain (RBD) as the target CDS and the UTR of group B6 as the source of the target UTR. The constructed artificial nucleic acid molecule was immunized with mice by intratracheal spray, and the IgG titer in the mouse serum was 14 days (a) and 21 days (b) after immunization.

[0033] Figure 12 The figure shows an artificial nucleic acid molecule constructed with the CDS encoding the Klebsiella pneumoniae vaccine antigen (Ag2) gene as the target CDS and the UTRs of the B6 and R1A7B6 groups as the target UTR sources. The constructed artificial nucleic acid molecule was immunized with mice by intratracheal spray. The IgG levels in the mouse serum were measured 14 days (a), 21 days (b) and 7 days after the second immunization (c), and the protection rate of the immunized mice against virus infection was also measured (d). DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0035] For clarity and readability, the following definitions are provided. Any technical features mentioned in these definitions can be read with respect to each embodiment of the present invention. In the case of these embodiments, additional definitions and explanations can be specifically provided.

[0036] When used in conjunction with the term "comprising" in the claims and / or the specification, the use of the word "a" or "an" can mean "one", but it is also consistent with the meaning of "one or more", "at least one", and "one or more than one".

[0037] The terms "comprise," "have," and "include" are open-ended linking verbs. Any form or tense of one or more of these verbs, such as "comprise," "have," and "include," is also open-ended. For example, any method that "comprises," "has," or "includes" one or more steps is not limited to having only those one or more steps and also encompasses other, unlisted steps.

[0038] The terms "first," "second," and the like in this application are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0039] As used in this specification and / or claims, the term "effective" means sufficient to achieve a desired, expected, or intended result. When used in the context of treating a patient or subject with an artificial nucleic acid molecule, "effective amount," "therapeutically effective amount," or "pharmaceutically effective amount" means an amount of the artificial nucleic acid molecule that, when administered to a subject or patient for treatment of a disease, is sufficient to achieve such treatment of the disease.

[0040] As used herein, the terms "improve," "increase," or "decrease," or grammatical equivalents, refer to values relative to a baseline measurement, such as the measurement of the same individual before the start of a treatment described herein, or the measurement of a control sample or subject (or multiple control samples or subjects) in the absence of a treatment described herein. A "control sample" is a sample that has been subjected to the same conditions as the test sample, except for the test article. A "control subject" is a subject having the same form of disease as the subject being treated and who is about the same age as the subject being treated.

[0041] As used herein, the term "in vitro" refers to events that occur in an artificial environment, such as in a test tube or reaction vessel, in cell culture, etc., rather than in a multicellular organism.

[0042] As used herein, the term "in vivo" refers to events that occur within multicellular organisms such as humans and non-human animals. In the context of cell-based systems, the term can be used to refer to events that occur within living cells (as opposed to, for example, in vitro systems).

[0043] As used herein, the term "patient" or "subject" refers to a living mammalian organism, such as a human, monkey, cow, sheep, goat, dog, cat, mouse, rat, guinea pig, or a transgenic species thereof. In certain embodiments, the patient or subject is a primate. Non-limiting examples of human subjects are adults, adolescents, infants, and fetuses.

[0044] Artificial nucleic acid molecules

[0045] Artificial nucleic acid molecules can typically be understood as nucleic acid molecules that do not exist in nature, such as DNA or RNA. Artificial nucleic acid molecules can be non-natural due to their own sequence and or other nucleotide modifications. Artificial nucleic acid molecules can be DNA molecules, RNA molecules and / or hybrid molecules comprising DNA and RNA parts. Typically, artificial nucleic acid molecules can be designed and / or produced by means of genetic modification, in which case the artificial nucleic acid molecules are generally not naturally occurring, i.e., there is a difference of at least one nucleotide with the naturally occurring sequence. The artificial nucleic acid molecule according to the present invention can preferably comprise ribonucleic acid (RNA), such as single-stranded RNA, more preferably messenger RNA (mRNA), such as in vitro transcribed mRNA.

[0046] In some embodiments, the artificial nucleic acid molecule comprises at least one selected from the group consisting of messenger RNA (mRNA), self-amplifying RNA (saRNA), circular RNA (circRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA) and microRNA (miRNA), primary-miRNA, antisense oligonucleotide (ASO), transfer RNA (tRNA), plasmid DNA (pDNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), deoxyribozyme (DNAzyme), ribozyme (RNAzyme), nucleic acid aptamer (aptamer), clustered regularly interspaced short palindromic repeats (CRISPR)-related nucleic acid, single guide RNA (sgRNA), CRISPR-RNA (crRNA), trans-activating crRNA (tracrRNA), guide RNA, long non-coding RNA (LncRNA), single-stranded RNA (ssRNA) and double-stranded RNA (dsRNA). In some embodiments, the nucleic acid is a therapeutic nucleic acid. More preferably, the nucleic acid comprises mRNA.

[0047] In principle, any type of RNA can be used in the context of the present invention. In a preferred embodiment, the RNA is single-stranded RNA. The term "single-stranded RNA" means a single continuous ribonucleotide chain, which is distinguished from RNA that is a double-stranded molecule formed by hybridization of two or more separate chains. The term "single-stranded RNA" does not exclude that the single-stranded molecule itself forms a double-stranded structure, such as a secondary (e.g., loop and stem-loop) structure or a tertiary structure.

[0048] The term "RNA" encompasses RNA that encodes an amino acid sequence. RNA can be prepared by synthetic chemistry and enzymatic methods known to those of ordinary skill in the art, or by using recombinant technology, or can be isolated from natural sources, or by a combination thereof.

[0049] Messenger RNA (mRNA) is a copolymer composed of nucleoside phosphate building blocks, primarily adenosine, cytidine, uridine, and guanosine. It acts as an intermediate to carry genetic information from DNA in the cell nucleus into the cytoplasm, where it is translated into proteins. Therefore, mRNA is a suitable surrogate for gene expression.

[0050] In the context of the present invention, mRNA should be understood to mean any polyribonucleotide molecule, if it enters the cell, is then suitable for protein or its fragmentary expression, or can be translated into protein or its fragment.Term " protein " contains any kind of amino acid sequence in this article, i.e. two or more amino acid whose chains connected by peptide bonds separately, and also comprises peptide and fusion protein.

[0051] MRNA contains a ribonucleotide sequence that encodes a protein or a fragment thereof of a function needed or useful in or near a cell. MRNA can contain the sequence of a complete protein or its functional variant. Further, a ribonucleotide sequence can encode a protein or its functional fragment that acts as a factor, an inducer, a regulator, a stimulant or an enzyme, wherein such protein is a protein necessary for its function to make up for an obstacle (particularly a metabolic disorder) or to start a process in the body (such as the formation of new blood vessels, tissues, etc.) or to induce the immune system to produce an adaptive immune response. Herein, functional variant means following fragment: it can assume the function of a protein in a cell, the function of the protein is needed in a cell, or the form of the absence or defect of the protein is pathogenic.

[0052] Typically, mRNA synthesis includes adding a "cap" to the 5' end and a "tail" to the 3' end. The presence of the cap is important for providing resistance to nucleases present in most eukaryotic cells. The presence of the "tail" is used to protect the mRNA from exonuclease degradation. Therefore, in some embodiments, the mRNA includes a 5' cap structure. In some embodiments, the mRNA includes a 5' and / or 3' untranslated region. In some embodiments, the 5' untranslated region includes one or more elements that affect the stability or translation of the mRNA. In some embodiments, the 3' untranslated region includes one or more polyadenylation signals, protein binding sites that affect the stability of the mRNA location in the cell, or one or more miRNA binding sites. As described above, unless otherwise noted in the specific context, the term mRNA used herein encompasses modified mRNA, i.e., the mRNA can be a modified mRNA.

[0053] In some embodiments, the modification of mRNA may include modification of the nucleotides of RNA. Modified mRNA according to the present invention may include, for example, backbone modifications, sugar modifications, phosphate modifications, or base modifications. In some embodiments, mRNA may be synthesized from naturally occurring nucleotides and / or nucleotide analogs (modified nucleotides), including but not limited to purines (adenine (A), guanine (G)) or pyrimidines (thymine (T), cytosine (C), uracil (U)) and modified nucleotide analogs or derivatives of purines and pyrimidines, the preparation of such analogs being known to those skilled in the art, for example, from U.S. Patent No. 4,373,071, U.S. Patent No. 4,373,071. No. 4,401,796, U.S. Patent No. 4,415,732, U.S. Patent No. 4,458,066, U.S. Patent No. 4,500,707, U.S. Patent No. 4,668,777, U.S. Patent No. 4,973,679, U.S. Patent No. 5,047,524, U.S. Patent No. 5,132,418, U.S. Patent No. 5,153,319, U.S. Patent Nos. 5,262,530 and 5,700,642, the entire disclosures of which are incorporated herein by reference.

[0054] For RNA, preferably mRNA, according to the present invention, all uridine nucleotides and cytidine nucleotides can be modified in the same form, or a mixture of modified nucleotides can be used for each. The modified nucleotides can have natural or non-naturally occurring modifications. Mixtures of various modified nucleotides can be used.

[0055] In one embodiment of the invention, one type of nucleotide uses at least two different modifications, wherein the modified nucleotide of one type has a functional group through which other groups can be attached. Nucleotides with different functional groups can also be used to provide binding sites for attaching different groups.

[0056] In a preferred embodiment, the RNA, preferably mRNA, according to the invention is characterized in that the modified uridine is selected from the group consisting of 2-thiouridine, 5-methyluridine, pseudouridine (ψ), 5-methyluridine 5'-triphosphate (m5U), N-1-methyl-pseudouridine (N1mΨ), N-1-methyl-pseudouridine-triphosphate, 2-thiouridine 5'-triphosphate (S2U), 5-iodouridine 5'-triphosphate (I5U), 4-thiouridine 5'-triphosphate (S4U), 5-bromouridine 5'-triphosphate (Br5U), 2'-methyl-2'-deoxyuridine 5'-triphosphate (U2'm), 2'-amino-2'-deoxyuridine 5'-triphosphate (U2'NH2), 2'-azido-2'-deoxyuridine 5'-triphosphate (U2'N3) and 2'-fluoro-2'-deoxyuridine 5'-triphosphate (U2'F).

[0057] In another preferred embodiment, the RNA, preferably the mRNA, according to the invention is characterized in that the modified cytidine is selected from the group consisting of 5-methylcytidine, 5-hydroxymethylcytidine, 5-methoxycytidine, 3-methylcytidine, 2-thio-cytidine, 2'-methyl-2'-deoxycytidine 5'-triphosphate (C2'm), 2'-amino-2'-deoxycytidine 5'-triphosphate (C2'NH2), 2'-fluoro-2'-deoxycytidine 5'-triphosphate (C2'F), 5-iodocytidine 5'-triphosphate (I5C), 5-bromocytidine 5'-triphosphate (Br5C), 5-methylcytidine 5'-triphosphate (m5C), 2-thiocytidine 5'-triphosphate (S2C) and 2'-azido-2'-deoxycytidine 5'-triphosphate (C2'N3).

[0058] In another preferred embodiment, the RNA, preferably the mRNA, according to the invention is characterized in that the modified adenosine is selected from the group consisting of N6-methyladenosine 5'-triphosphate (m6A), N1-methyladenosine 5'-triphosphate (m1A), 2'-O-methyladenosine 5'-triphosphate (A2'm), 2'-amino-2'-deoxyadenosine 5'-triphosphate (A2'NH2), 2'-azido-2'-deoxyadenosine 5'-triphosphate (A2'N3) and 2'-fluoro-2'-deoxyadenosine 5'-triphosphate (A2'F).

[0059] In another preferred embodiment, the RNA, preferably mRNA, according to the invention is characterized in that the modified guanosine is selected from N1-methylguanosine 5-triphosphate (m1G), 2'-O-methylguanosine 5'-triphosphate (G2'm), 2-amino-2-deoxyguanosine 5'-triphosphate (G2'NH2), 2'-azido-2'-deoxyguanosine 5'-triphosphate (G2'N3), and 2'-fluoro-2'-deoxyguanosine 5'-triphosphate (G2'F).

[0060] Representative U.S. patents that teach the preparation of some of the above-mentioned modified nucleobases, as well as other modified nucleobases, include, but are not limited to, U.S. Patents 3,687,808; 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540; 5,587,469; 5,594,121; 5,596,091; 5,614,617; 5,645,985; 5,681,941; 5,750,692; 5,763,588; 5,830,653 and 6,005,096, each of which is incorporated herein by reference in its entirety.

[0061] In some embodiments, the invention provides oligonucleotides comprising connected nucleosides. In such embodiments, nucleosides can be linked together using any internucleoside bond. The two main categories of internucleoside linking groups are defined by the presence or absence of a phosphorus atom. Compared to natural phosphodiester bonds, modified bonds can be used to change (usually increase) the nuclease resistance of oligonucleotides. In some embodiments, internucleoside bonds with chiral atoms can be prepared as racemic mixtures or separate enantiomers. Representative chiral bonds include, but are not limited to, alkyl phosphonates and thiophosphates. The preparation methods of phosphorus-containing and non-phosphorus-containing internucleoside bonds are well known to those skilled in the art.

[0062] Representative U.S. patents that teach the preparation of such oligonucleotide conjugates include, but are not limited to, U.S. Patents 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118, 802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958, 013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,416, 5,599,923; 5,599,928 and 5,688,941, each of which is incorporated herein by reference.

[0063] In a preferred embodiment, the mRNA is an mRNA containing a combination of modified and unmodified nucleotides. Preferably, it is an mRNA containing a combination of modified and unmodified nucleotides as described in WO2011 / 012316. The mRNA described therein shows improved stability and reduced immunogenicity.

[0064] mRNA can be synthesized according to any of a variety of known methods. For example, mRNA according to the present invention can be synthesized via in vitro transcription (IVT).

[0065] Furthermore, modified RNA, preferably mRNA molecules can be chemically synthesized, for example, by conventional chemical synthesis on an automated nucleotide sequence synthesizer using a solid support and standard techniques, or by chemically synthesizing the corresponding DNA sequence and subsequently transcribing it in vitro or in vivo.

[0066] The artificial nucleic acid molecules of the present invention can be mRNAs of various lengths. In some embodiments, the artificial nucleic acid molecule is an in vitro synthesized mRNA having a length equal to or greater than about 0.1 kb, 0.5 kb, 1 kb, 1.5 kb, 2 kb, 2.5 kb, 3 kb, 3.5 kb, 4 kb, 4.5 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 11 kb, 12 kb, 13 kb, 14 kb, 15 kb, or 20 kb. In some embodiments, the artificial nucleic acid molecule of the present invention is an in vitro synthesized mRNA having a length ranging from about 0.1 to 2 kb, about 1 to 20 kb, about 1 to 15 kb, about 1 to 10 kb, about 5 to 20 kb, about 5 to 15 kb, about 5 to 12 kb, about 5 to 10 kb, about 8 to 20 kb, or about 8 to 15 kb.

[0067] The cross expression of the exogenous artificial nucleic acid molecule (preferred mRNA) introduced can be intended to compensate or complement endogenous gene expression, particularly in the case of endogenous gene defect or silence, cause gene expression product to not exist, not enough or defective or malfunction, multiple metabolic and hereditary diseases, such as cystic fibrosis, hemophilia or muscular dystrophy etc. are often such situations. The cross expression of the exogenous artificial nucleic acid molecule (preferred mRNA) introduced can also be intended to make expression product interact with any endogenous cell process or interfere with any endogenous cell process, such as the regulation and control of gene expression, signal transduction and other cellular processes. The cross expression of the exogenous artificial nucleic acid molecule (preferred mRNA) introduced can also be intended to cause immune response under the background of the organism that the cell of transfection or transduction resides or makes it reside. Example is the genetic modification of antigen presenting cells (such as dendritic cells) so that it presents antigens for vaccination purpose. Other examples are the cross expression of cytokines in tumors, so as to cause tumor-specific immune response. Furthermore, overexpression of introduced exogenous artificial nucleic acid molecules (preferably mRNA) can also be aimed at generating transiently genetically modified cells for cell therapy in vivo or ex vivo, such as modified T cells or precursors or stem cells or other cells for regenerative medicine.

[0068] In addition, a variety of genetic disorders caused by mutations in a single gene are known and are candidates for RNA, preferably mRNA treatment methods. Regarding the possibility that a certain trait will occur in offspring, the disorder caused by a single gene mutation, such as cystic fibrosis, hemophilia and a variety of other diseases can be dominant or recessive. In contrast, polygenic disorders are caused by two or more genes, and the manifestation of the corresponding disease is often variable and related to environmental factors. Examples of polygenic disorders are high blood pressure, elevated cholesterol levels, cancer, neurodegenerative diseases, mental illnesses, etc. In these cases as well, therapeutic RNA, preferably mRNA, representing one or more of these genes can be beneficial to those patients. In addition, genetic disorders are not necessarily transmitted from parental genes, but may also be caused by new mutations. In these cases as well, therapeutic RNA, preferably mRNA, representing the correct gene sequence can be beneficial to patients.

[0069] An online catalogue with 22,993 entries for human genes and genetic disorders with descriptions of their corresponding genes and their phenotypes is available at the ONIM (Online Mendelian Inheritance in Man) website (http: / / onim.org); each sequence is available from the Uniprot database (http: / / www.uniprot.org).

[0070] In some embodiments, the coding sequence of the artificial nucleic acid molecule of the present invention, preferably mRNA, can be transcribed and translated into a partial or full-length protein comprising a level of cellular activity equal to or greater than that of the native protein.

[0071] In some embodiments, genetic diseases may be involved, for example, those affecting the lungs, such as SPB (surfactant protein B) deficiency, ABCA3 deficiency, cystic fibrosis (CF), primary ciliary dyskinesia (PCD), and dyskinesia), asthma, chronic obstructive pulmonary disease (COPD) and alpha 1-antitrypsin deficiency, or which affect plasma proteins (e.g. congenital hemochromatosis (hepcidin deficiency), thrombotic thrombocytopenic purpura (TPP, ADAMTS13 deficiency) and cause coagulation defects (e.g. hemophilia a and b) and complement deficiencies (e.g. protein C deficiency), immunodeficiencies such as, for example, SCID (caused by mutations in various genes such as RAG1, RAG2, JAK3, IL7R, CD45, CD3δ, CD3ε) or deficiency, due to deficiency of adenosine deaminase, for example (ADA-SCID), septic granulomatosis (caused by mutations in the gp-91-phox gene, the p47-phox gene, the p67-phox gene or the p33-phox gene) and storage diseases such as Gaucher's disease, Fabry's disease, Krabbe's disease, MPS I, MPS II (Hunter syndrome), MPS Type VI and II glycogen storage diseases or muccopolysacchaidoses.

[0072] Other diseases for which the artificial nucleic acid molecules of the present invention, preferably mRNA, can exert therapeutic effects include, for example, SMN1-related spinal muscular atrophy (SMA); amyotrophic lateral sclerosis (ALS); GALT-related galactosemia; SLC3A1-related disorders, including cystinuria; COL4A5-related disorders, including Alport syndrome; galactocerebrosidase deficiency; X-linked leukoreflexia and adrenoneuropathy; Friedreich's ataxia; Perelman-Merzheimer disease; TSC1 and TSC2-related tuberous sclerosis; Sanfilip B syndrome (MPS) IIIB); CTNS-related cystinosis; FMR1-related disorders, including fragile X syndrome, fragile X-linked tremor / ataxia syndrome, and fragile X premature ovarian failure syndrome; Prader-Willi syndrome; hereditary hemorrhagic telangiectasia (AT); Nieto-Philosophyll disease type C1; neuronal ceroid lipofuscinosis-related disorders, including juvenile neuronal ceroid lipofuscinosis (JNCL), juvenile Batten disease, Santavuori-Haltia disease, Jansky-Bielschowsky disease, and PTT-1 and TPP1 deficiency; EIF2B1-, EIF2B2-, EIF2B3-, EIF2B4-, and EIF2B5-related childhood ataxia with hypomyelination / vanishing white matter in the central nervous system; CACNA1A and CAC NB4-related paroxysmal ataxia type 2; MECP2-related disorders, including classic Rett syndrome, MECP2-related severe neonatal encephalopathy, and PPM-X syndrome; CDKL5-related atypical Rett syndrome; Kennedy disease (SBMA); Notch-3-related cerebral autosomal dominant arteriopathy with subcortical infarcts and leukoencephalopathy (CADASIL); SCN1A and SCN1B-related seizure disorders; polymerase G-related disorders, including Alpers-Huttenlocher syndrome, POLG-related sensory ataxia neuropathy, dysphonia, and ophthalmoparesis, and autosomal dominant and recessive progressive external ophthalmoplegia with mitochondrial DNA deletions; X-linked adrenal hypoplasia; X-linked agammaglobulinemia; Fabry disease; and Wilson disease.

[0073] In all of these diseases, a protein, such as an enzyme, is defective and can be treated by treatment with an artificial nucleic acid molecule, preferably mRNA, encoding any of the above proteins of the present invention, which makes the protein encoded by the defective gene or a functional fragment thereof available. Transcript replacement therapy / enzyme replacement therapy does not affect the underlying genetic defect, but increases the concentration of the enzyme that the patient lacks. As an example, in Pompe disease, transcript replacement therapy / enzyme replacement therapy replaces the deficient lysosomal enzyme acid α-glucosidase (GAA).

[0074] Thus, non-limiting examples of proteins that can be encoded by the mRNA of the present invention are erythropoietin (EPO), growth hormone (somatotropin, hGH), cystic fibrosis transmembrane conductance regulator (CFTR), growth factors such as GM-SCF, G-CSF, MPS, protein C, hepcidin, ABCA3, and surfactant protein B. Other examples of diseases that can be treated with the artificial nucleic acid molecules according to the present invention are hemophilia A / B, Fabry disease, CGD, ADAMTS13, Hurler's disease, X-linked A-gammaglobulinemia, adenosine deaminase-related immunodeficiency, and respiratory distress syndrome in newborns, which is associated with SP-B. Particularly preferably, the artificial nucleic acid molecule according to the present invention, preferably the mRNA, contains a coding sequence for cystic fibrosis transmembrane conductance regulator (CFTR), α1-antitrypsin, dynein axonemal intermediate chain 1 (DNAI1), surfactant protein B (SP-B), or erythropoietin. Further examples of proteins that can be encoded by the inventive RNA, preferably mRNA, according to the invention are growth factors, such as the human growth hormone hGH, BMP-2 or angiogenic factors.

[0075] Alternatively, the artificial nucleic acid molecule, preferably mRNA, may contain a ribonucleotide sequence encoding a full-length antibody or Nanobody (e.g., both heavy and light chains), which may be used in therapeutic settings, such as to confer immunity on a subject. Corresponding antibodies and their therapeutic application(s) are known in the art.

[0076] In another embodiment, the artificial nucleic acid molecule, preferably mRNA, can encode a functional monoclonal or polyclonal antibody that can be used to target and / or inactivate a biological target (e.g., a stimulatory cytokine such as tumor necrosis factor). Similarly, the RNA, preferably mRNA sequence, can encode a functional anti-nephrotic factor antibody, for example, for the treatment of type II membranoproliferative glomerulonephritis or acute hemolytic uremic syndrome, or alternatively, can encode an anti-vascular endothelial growth factor (VEGF) antibody for the treatment of VEGF-mediated diseases such as cancer.

[0077] In another embodiment, the artificial nucleic acid molecule, preferably mRNA, may contain a ribonucleotide sequence encoding a polypeptide or protein that can be used for genome editing technology. A variety of genome editing systems utilizing different polypeptides or proteins are known in the art, i.e., for example, CRISPR-Cas systems, large-range nucleases (homing nucleases, meganucleases), zinc finger nucleases (ZFNs), and nucleases based on transcription activator-like effectors (TALENs). Trends in Biotechnology, 2013, 31 (7), 397-405 summarizes methods for genome engineering.

[0078] Therefore, in a preferred embodiment, the artificial nucleic acid molecule, preferably mRNA, may contain a ribonucleotide sequence encoding a polypeptide or protein of the Cas (CRISPR-associated protein) protein family, preferably Cas9 (CRISPR-associated protein 9). Proteins of the Cas protein family, preferably Cas9, can be used in CRISPR / Cas9-based methods and / or CRISPR / Cas9 genome editing technologies. CRISPR-Cas systems for genome editing, regulation, and targeting are reviewed in Nat. Biotechnol., 2014, 32(4): 347-355.

[0079] In another preferred embodiment, the artificial nucleic acid molecule, preferably mRNA, may contain a ribonucleotide sequence encoding a transcription activator-like effector nuclease (TALEN). In another preferred embodiment, the artificial nucleic acid molecule, preferably mRNA, may contain a ribonucleotide sequence encoding a zinc finger nuclease (ZFN). In another preferred embodiment, the RNA, preferably mRNA, may contain a ribonucleotide sequence encoding a meganuclease.

[0080] In certain embodiments, the present invention provides a method for preparing a therapeutic composition comprising a polymer-lipid composition as described herein for delivering a specific antigen and / or a nucleic acid encoding a specific antigen, wherein the antigen can be an antigen from bacteria, virus, fungus, or cancer cells.

[0081] The term "antigen" refers to a peptide or nucleotide-based biological material (natural, recombinant, or synthetic) that stimulates a protective immune response in an animal. Antigens suitable for the present invention can be amino acid sequences such as peptides or proteins, or nucleic acid sequences such as genomic DNA, cDNA, mRNA, saRNA, circRNA, tRNA, rRNA, small interfering RNA (iRNA) hybridization sequences, or modified or unmodified synthetic or semisynthetic oligonucleotide sequences.

[0082] Antigens suitable for the present invention may be obtained from an organism selected from the group consisting of bacteria, viruses, parasites, rickettsiae, protozoa and cancer cells.

[0083] The artificial nucleic acid molecules of the present invention can be used to treat or prevent a number of diseases and disorders, for example:

[0084] Diseases and disorders involving the following viruses: Retroviridae (e.g., human immunodeficiency virus, including HIV-1); Flaviviridae (e.g., dengue virus, encephalitis virus, yellow fever virus); Coronaviridae (e.g., coronavirus); Rhabdoviridae (e.g., vesicular stomatitis virus, rabies virus); Filoviridae (e.g., Ebola virus); Paramyxoviridae (e.g., parainfluenza virus, mumps virus, measles virus, respiratory syncytial virus); Orthomyxoviridae (e.g., influenza virus); Reoviridae (e.g., reovirus, orbivirus, and rotavirus); Binaviridae; Hepadnaviridae (hepatitis B virus); Parvoviridae (parvovirus); Herpesviridae (herpes simplex virus (HSV) 1 and 2, varicella-zoster virus, cytomegalovirus (CMV), herpes simplex virus); Poxyiridae (monkeypox virus, smallpox virus, vaccinia virus, Poxvirus); and Iridoviridae (e.g., African swine fever virus); and unclassified viruses (e.g., causative agents of spongiform encephalopathies, agents of hepatitis delta (thought to be defective satellites of hepatitis B virus), HCV virus (causing non-A, non-B hepatitis); Norwalk and related viruses and astroviruses). HIV, hepatitis A, hepatitis B, hepatitis C, coronavirus, rabies virus, poliovirus, influenza virus, meningitis virus, measles virus, mumps virus, rubella, pertussis, encephalitis virus, papillomavirus, yellow fever virus, respiratory syncytial virus, parvovirus, chikungunya virus, hemorrhagic fever virus and herpes virus (particularly varicella), cytomegalovirus and Epstein-Barr virus are particularly preferred. In the above embodiments, the antigens selected for use in the polymer-lipid composition are derived from those antigens present in naturally occurring viruses (or expressed / induced during infection) (or designed with reference to those antigens).

[0085] Diseases and disorders involving the following Gram-negative and Gram-positive bacteria: Helicobacter pylori, Legionella pneumophilia, Mycobacterium (e.g., M. tuberculosis), Staphylococcus aureus, Neisseria gonorrhoeae, Neisseria meningitidis, Listeria monocytogenes, Streptococcus pyogenes (Group A Streptococcus), Streptococcus agalactiae (Group B Streptococcus), Streptococcus viridans, Streptococcus pneumoniae, Klebsiella spp.) (including Klebsiella pneumoniae (K. pneumoniae)), Rickettsia spp., and Actinomyces spp. (including A. israelii). In the above embodiments, the antigens selected for use in the polymer-lipid composition are derived from those antigens present in naturally occurring bacteria (or expressed / induced during infection) (or artificially designed with reference to those antigens).

[0086] Diseases and disorders involving the following fungi: Cryptococcus neoformans, Histoplasma capsulatum, Coccidioides immitis, Blastomyces dermatitidis, Chlamydia trachomatis, and Candida albicans, in which embodiments the antigens selected for use in the vaccine are derived from (or designed with reference to) those antigens present in naturally occurring fungi (or expressed / induced during infection).

[0087] Diseases and disorders involving the following protozoa: Plasmodium spp. (including Plasmodium falciparum, Plasmodium malariae, Plasmodium ovale, and Plasmodium vivax), Toxoplasma spp. (including T. gondii and T. cruzii), and Leishmania spp.

[0088] Diseases and disorders involving cancer cells of the blood and lymphatic systems (including Hodgkin's disease, leukemias, lymphomas, multiple myeloma, and Diseases), melanoma (including melanoma of the eye), adenoma, sarcoma, cancer of solid tissue, melanoma, lung cancer, thyroid cancer, salivary gland cancer, leg cancer, tongue cancer, lip cancer, bile duct cancer, pelvic cancer, mediastinal cancer, urethral cancer, Kaposi's sarcoma (e.g., when associated with AIDS); skin cancer (including malignant melanoma), digestive tract cancer (including head and neck cancer, esophageal cancer, stomach cancer, pancreatic cancer, liver cancer, colon and rectal cancer, anal cancer), reproductive and urinary tract cancer (including kidney cancer, bladder cancer, testicular cancer, prostate cancer), female cancer (including breast cancer, cervical cancer, ovarian cancer, gynecological cancer and choriocarcinoma) and brain cancer, bone carcinoid tumor, nasopharyngeal cancer, retroperitoneal tumor, thyroid cancer, soft tissue tumor and unknown primary site cancer. In the above embodiment, the antigen selected for the vaccine is a homologous neoantigen or tumor-associated antigen present in malignant cells and / or tissues.

[0089] Diseases and disorders involving multicellular parasites such as helminths (eg, Schistosoma spp.).

[0090] In a preferred embodiment, the antigen useful in the present invention is mRNA.

[0091] The antigen is naturally associated with the polymer-lipid composition of the present invention in an immunogenic effective amount. An "immunogenic effective amount" means that when a polymer-lipid composition of the present invention and an antigen (e.g., mRNA) are administered to an individual, the antigen contains a protective component at a concentration sufficient to protect the animal from the target disease. As an example of an immunogenic effective amount of an antigen, an amount of 0.01-100 μg can be mentioned.

[0092] In certain embodiments, the present invention provides a method for preparing a composition as described herein comprising delivering an immunomodulator and / or mRNA encoding an immunomodulator. The immunomodulator includes, but is not limited to, interleukin 2 (IL-2), interleukin 12 (IL-12), granulocyte-macrophage colony-stimulating factor (GM-CSF), interleukin 23 (IL-23), CC domain chemokine ligand 28 (CCL28), interleukin 36γ (IL-36γ), constitutively active variants of stimulator of interferon genes (STING) protein, etc.

[0093] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0094] like Figure 1 As shown, Figure 1 The following is a schematic diagram of the structure of an artificial nucleic acid molecule provided in an embodiment of the present application. The artificial nucleic acid molecule comprises: a 5' cap structure (Cap), a target 5' untranslated region (UTR), a target coding region (CDS), a target 3' untranslated region (UTR), and a PolyA tail.

[0095] The target CDS is used to be translated to produce a target protein or polypeptide. The target protein or polypeptide refers to a protein or polypeptide that is desired to be expressed in large quantities in cells. In practical applications, the experimenter can determine which protein or polypeptide expression level they wish to increase in the cell based on the experimental needs, thereby determining the corresponding CDS as the target CDS. This application does not limit the target CDS and the target protein or polypeptide.

[0096] The target 5'UTR sequence is one of the following sequences: a 5'UTR of a highly expressed gene and a 5'UTR variant of the highly expressed gene. The target 3'UTR sequence is one of the following sequences: a 3'UTR of a highly expressed gene and a 3'UTR variant of the highly expressed gene.

[0097] Because the 5'UTR and 3'UTR contain specific regulatory sequence elements, they have a regulatory effect on mRNA translation and stability. Therefore, by combining the CDS sequence that can translate the target protein or polypeptide with the sequences of the 5'UTR and 3'UTR, it is possible to affect the translation efficiency of the CDS sequence, thereby affecting the expression level of the target protein or polypeptide. Based on this principle, it can be understood that if the number of mRNAs detected in a certain tissue cell is greater, the more stable the mRNA is in the cell. It can be further concluded that the UTR of the mRNA is the optimal UTR that enhances mRNA expression and stabilizes the structure in the tissue cell.

[0098] In this embodiment, the native 5'UTR and native 3'UTR of highly expressed genes can be used as target 5'UTR and target 3'UTR. Specifically, the genetic information contained in various human organs and tissue cells can be obtained by sequencing. For example, single-cell sequencing technology can be used, and mitochondrial genes and ribosomal genes can be removed. This application does not limit the specific sequencing technology or method. The genetic information of each type of cell in each organ obtained can then be sorted from high to low according to the mRNA content. For example, the genetic information of each type of cell in each organ obtained can be sorted according to the number of transcripts per kilobase of exon model per million mapped reads (TPM). This application does not limit the TPM database. The higher the TPM ranking, the higher the expression level of the gene in the cell, that is, the higher the mRNA content of the gene in the cell, and the 5'UTR and 3'UTR of the gene can then be used as the target 5'UTR and target 3'UTR in the embodiments of this application. Among them, the TBtools tool can be used to determine the UTR sequence of the highly expressed gene. In this embodiment, genes ranked 1 to 25 by TPM can be considered highly expressed genes. In actual applications, experimenters can also define the TPM rankings that can be considered highly expressed genes based on experimental needs. For example, genes ranked 1 to 20 by TPM can be considered highly expressed genes; or genes ranked 1 to 30 by TPM can be considered highly expressed genes, etc.

[0099] In some embodiments, the tissues from which genetic information is collected include at least: lungs, bronchi, trachea, liver, spleen, myocardial tissue, small intestine, stomach, colon, salivary glands, lymph nodes, brain, kidneys, blood vessels, reproductive tract, and esophagus, etc.

[0100] In some embodiments, the type of cells from which genetic information is collected includes at least one of the following: epithelial cells, alveolar cells, bronchial epithelial cells, ciliated airway epithelial cells, goblet cells, ionocytes, exocrine cells, microfold cells, dendritic cells, stromal cells, macrophages, T cells, endothelial cells, granulocytes, B cells, bile duct cells, red blood cells, fibroblasts, hepatocytes, plasma cells, hepatic macrophages, basal cells, ion transport cells, smooth muscle cells, mucous gland cells, salivary duct cells, serous gland cells, distal intestinal epithelial cells, enteroendocrine cells, intestinal goblet cells, Paneth cells, stem cells, gastric mucus-secreting cells, collecting duct cells, glial cells, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocyte precursor cells, cardiomyocytes, spleen cells, kidney cells, muscle cells, etc.

[0101] For example, in this embodiment, according to the above method, the highly expressed genes obtained include but are not limited to the sequences mentioned in Table 1:

[0102] Table 1 Examples of highly expressed gene names and transcript IDs

[0103]

[0104]

[0105] Note: Gene names and transcript IDs are from Ensemble (https: / / www.ensembl.org / index.html)

[0106] It is understood that in some embodiments, the ranking of mRNA levels from high to low can be performed separately for each cell type. For example, basal cells have genes ranked 1 to 25 by TPM, and epithelial cells also have genes ranked 1 to 25 by TPM. The same gene may have different expression rankings in different cell types.

[0107] For example, in the present embodiment, RNA single cell type cluster data (RNA single cell type tissue cluster data) are obtained from the public database human proteome atlas. These data include the standardized expression matrix of human genes in each cell type in each tissue of human body. Subsequently, R language dplyr package filter function is used to filter out specific tissue, after obtaining the expression matrix of each tissue, the expression matrix ribosomal genes, i.e., the genes and mitochondrial genes MD at the beginning of RPS and RPL, are removed, Excel screening and sorting function is used, and the expression amount nTPM of the various genes included in the various cells of each tissue is arranged from high to low. The present embodiment has selected 15 tissues in total, and subsequently selected the genes (as shown in the following table 2-table 3) whose expression amount is arranged in the 1st to the 25th in various cells, i.e., highly expressed genes described above. In order to simplify description, the application does not enumerate all tissue cells, and other unselected tissues can also screen and draw their respective highly expressed genes in the same way.

[0108] Table 2 Highly expressed genes in lung cells

[0109]

[0110]

[0111]

[0112] Table 3 Highly expressed genes in bronchus

[0113]

[0114]

[0115] In actual applications, the experimenter can decide which gene's UTR to use as the target UTR based on the experimental requirements. For example, if a certain cell type needs to be used for the experiment, genes with higher TPM in cells of this type will be given priority. In other cases, when deciding which gene's UTR to choose as the target UTR, the experimenter can also comprehensively consider the ranking of a certain gene in all or multiple tissue cell types. For example, an experiment based on hepatocytes is being conducted, but a certain gene is a gene with a higher TPM ranking in multiple cells such as bile duct cells, hepatocytes, plasma cells, liver macrophages, basal cells, distal intestinal epithelial cells, and enteroendocrine cells. Even if the gene is not the gene with the highest TPM ranking in hepatocytes, the UTR of the gene can be used as the target UTR. This application does not limit the selection method or selection conditions for cell types or tissue types. According to the above screening conditions, the highly expressed genes include but are not limited to the following genes:

[0116] HTN3、STATH、EEF1A1、TPT1、TMSB4X、B2M、IGKC、BTG1、CD74、HTN1、HLA-DRA、 PTMA, PRB3, TXNIP, FAU, IGHM, RACK1, UBA52, ACTB, ZG16B, EEF1B2, CXCR4, NA CA、EIF1、FTH1、NTS、CHGA、REG4、S100A6、SST、PHGR1、PCSK1N、TFF3、PYY、TM SB10, FTL, LCN15, SPINK1, SCT, CHGB, TPH1, KRT8, H3-3B, GAPDH, DPP10, SLC1 A2, PCDH9, NRXN1, GPC5, GPM6A, CTNNA2, NPAS3, LSAMP, NRG3, CTNND2, ADGRV 1、NTM、RORA、LRP1B、ERBB4、ZBTB20、CADM2、ADGRB3、MAGI2、DTNA、SLC1A3、PP P2R2B, SOX5, ARHGAP24, KRT15, S100A2, FABP5, FOS, HSPB1, KRT5, ACTG1, AN XA1, PPIA, ZFP36L1, JUN, SFN, KRT19, HBB, SAA1, IGLC2, CD52, SAA2, HSP90AA 1、CCL4、IGHG4、AREG、ZFP36L2、CD69、NFKBIA、IL7R、TIMP1、MT2A、MGP、DCN、 CFD、JUNB、CCN1、VIM、SAT1、SFTPC、A2M、TAGLN、ACTA2、IGFBP7、TPM2、C11orf 96、MYL6、MYL9、JUND、CALD1、IGHA1、IGLC3、HLA-B、HLA-DPA1、IGHG3、PRB4、 SPP1, PRH2, PRH1, SMR3B, CCL3, S100A9, NAMPT, GUCA2B, GUCA2A, MT1G, MT1E MT1H, MT1X, CA7, PDPPF, LGALS3, FABP6, DOCK4, PLXDC2, FRMD4A, LRMDA, RBM S3、ST6GALNAC3、ITPR2、PTPRG、MEF2C、QKI、MEF2A、LAMA2、SFMBT2、ST6GAL1、 AUTS2, MAML2, SLC8A1, ABCB1, MBNL1, HS3ST4, NAV3, ELMO1, CST3, APOD, PTGD S、GSN、SPARCL1、PLA2G2A、IGFBP6、MFAP4、LUM、ADH1B、HBA1、HBA2、CA1、HBD、<h2 style=";text-align:left;direction:ltr">AHSP、HBM、PRDX2、H4C3、BLVRB、TUBA1B、CA2、SLC25A37、HMGB2、UBB、YBX1、A TP5F1E、IGFBP5、TIMP3、APOE、GPX3、CXCL14、CCL5、GNLY、GZMA、TPSB2、RGS1 、CD7、IL32、ZFP36、SRGN、KCNIP4、ROBO2、CNTN5、CNTNAP2、RBFOX1、NRXN3、C SMD1、GRIK1、SGCZ、ZNF385D、FGF14、DLGAP1、SYT1、ADARB2、GRIK2、IL1RAPL1 、GALNTL6、GRIP1、CCL21、IFITM3、ID3、IFITM1、ACKR1、S100A10、IFITM2、AL B、SERPINA1、CLU、AMBP、TFF2、ANXA4、TM4SF4、SCGB3A1、DEFB1、SLPI、GC、COX 7C、MUC7、WFDC2、FDCSP、KRT7、LCN2、AZGP1、UBC、LTF、DEFA5、DEFA6、REG3A、 PRSS2、LYZ、REG1A、ITLN2、OLFM4、PLP1、DLG2、PDE4B、ST18、CTNNA3、MBP、PTP RD、SLC24A2、NKAIN2、PEX5L、FMNL2、SLC44A1、NCKAP5、RNF220、PLCL1、UNC5 C、EDIL3、TMTC2、MAN2A1、JCHAIN、IGLV1-40、IGKV3-20、IGLV2-14、IGKV4-1 IGKV1-5, IGLV1-51, IGHV4-39, IGHV3-74, IGLV3-19, IGHV3-30, IGLV3-1, IGHA2, IGHV3-23, IGKV1D-13, IGHV2-70, IGKV3-11, IGHV4-59, IGHV5-51, IG LV2-23、IGHV1-3、IGKV1-12、IGLV3-25、APOC3、ORM1、HP、APOA2、TTR、APOC1 、APOA1、FGB、FGA、FGG、RBP4、FGL1、APOH、ORM2、CRP、KRT17、FOSB、S100A11、E GR1、CEBPD、DUSP1、COL1A2、COL1A1、CCL2、COL3A1、THBS4、HSPA1A、LGALS1、 S100A4、ADIRF、GADD45B、C1QB、C1QA、CTSB、GPX1、TYROBP、HLA-DPB1、CD163、<h2 style=";text-align:left;direction:ltr">PSAP、PRB1、CA6、PRB2、S100A8、SH3BGRL3、DLK1、GNAS、CD63、NNMT、PFN1、HS PA8、ID1、HLA-E、DNASE1L3、FABP1、PIGR、LGALS4、CKB、C15orf48、FXYD3、SE RF2、TSC22D3、KLF6、FABP4、CXCL2、RGS5、IFI27、TM4SF1、SCGB1A1、CCL20、C XCL1 、 CSTB 、 SERPINB3 、 BPIFB1 、 CXCL8 、 GDF15 、 SFTPB 、 EMP2 、 KRT18 、 DSTN 、 hop X 、 Ager 、 GPRC5A 、 anxa2 、 NAPSA 、 GSTP1 、 Rab11fip1 、 zg16 、 FCGBP 、 muc2 、 Agr2 、TFF1、KLK1、FXYD4、AQP2、ITM2B、ATP1B1、CRYAB、NPPA、TTN、MYL2、COX7A1、M YH7、MYL7、ANKRD1、ATP5ME、TPM1、TNNI3、MYH6、ATP5PF、MYL12A、COX6A2、CM YA5、SLC25A4、MYL3、COX5B、UQCRB、UQCRQ、ATP5MD、ATP5IF1、DYNLL1、CALM2、 PRDX5、MSMB、SCGB3A2、FCER1G、AIF1、CYBA、C7、HLA-DRB1、CTSD、SLC26A3、T SPAN8、CLDN3、HLA-C、HLA-A、NKG7、CD37、IER2、IGKV2-30、IGKV2-24、IGLV5 -45、IGLV1-44、IGLV7-46、IGLV1-47、IGHV7-4-1、IGKV1-16、IGKV3-15、IGH G1、IGKV1D-16、IGHV3-7、IGLV2-11、FCN3、EPAS1、ITLN1、FN1、IGKV1-6、IGLV 2-8、IGHV1-18、IGKV1-17、IGHV3-48、CREM、KLRB1、S100B、PMP22、GPM6B、AH NAK、CD9、ELF3、HES1、MYH11、DES、SPINK4、CLCA1、LRRTM4、ASIC2、FAM155A、R ALYL、PDE4D、SNTG1、LRRC4C、HSP90AB1、OAZ1、OIT3、STAB1、ALDOB、APOA4、P RAP1、RBP2、SELENOP、LHFPL3、DSCAM、NLGN1、OPCML、PCDH15、PTPRZ1、MMP16、GRID2, TNR, KCNMB2, MDGA2, RNASE1, CRHBP, FCN2, CLEC1B, SSR4, MZB1, HSP90B1, CAV1, SPARC, TUBA1A , BPIFA2, MUC5AC, GKN1, PGC, PSCA, PGA3, GKN2, CYSTM1, S100P, KRT13, KRT4, CSTA, S100A14, PERP, SP INK5, IGHG2, IGHV4-34, IGLV6-57, IGHV5-10-1, IGHV1-46, TNNC2, ACTA1, MYLPF, TNNI2, CKM, NEB, TN NT3, MYL1, MB, YBX3, TCAP, CA3, TNNC1, PI3, MMP10, CMC1, INSL5, GCG, H3-3A, LTB, EEF1D, SFTPA2, SFTP A1, CTSH, NPC2, CYB5A, HMGB1, HNRNPA1, FXYD2, NDUFA4, COX6B1, COX7A2, ATP5MG, UQCR10, UQCR11, CO X4I1, LDHB, SRP14, CD36, CALM1, VWF, MTRNR2L1, GNG11, H2AZ1, ATP5F1B, PRDX1, PRSS1, CLPS, PNLIP, C TRB1, CPA1, CPB1, PLA2G1B, CTRC, GP2, CTRB2, SYCN, CEL, CELA3A, CELA2A, CPA2, CELA2B, PRSS3, CELA 3B, MT1F, MIOX, DCXR, PDZK1IP1, PEBP1, IL1B, TXN, MMRN1, TNFAIP3, DUSP2, TPSAB1, CPA3, CA4, F13A1. ,

[0117] Furthermore, the target 5'UTR and target 3'UTR can be derived from different genes. For example, to achieve increased expression of a target protein in a cell, the 5'UTR of the TMSB10 gene can be used as the target 5'UTR for the CDS of the target protein, and the 3'UTR of the SFTPA1 gene can be used as the target 3'UTR for the CDS of the target protein. In other words, in the artificial nucleic acid molecule prepared according to this embodiment, the target 5'UTR, target CDS, and target 3'UTR can be derived from three or more different genes.

[0118] In this embodiment, after obtaining the highly expressed genes by the above method, the TBtools tool was used to obtain the 5'UTR sequence of the highly expressed genes, and a Kozak sequence was added to the 3' end of the 5'UTR, specifically, the 5'UTR sequences derived from TMSB10, SFTPA1, SFTPA2, SFTPC, SCGB3A2, SFTPB, SFTPD, AGER, MS4A15, NAPSA, RTKN2, SCGB1A1, SLC34A2, LAMP3, HTR3C, CCL18, LRRK2 and MSR1 genes as shown in SEQ ID NOs. 9-26; or the sequences optimized by LinearDesign, such as the 5'UTR derived from the MSR1 gene as shown in SEQ ID NOs. 27-31 after adding the A7 element and then optimized by LinearDesign; similarly, in this embodiment, the 3'UTR sequences of the highly expressed genes were obtained by using the TBtools tool, specifically, the 5'UTR sequences derived from SEQ ID NOs. 3'UTR sequences of TMSB10, SFTPD, SFTPC, SCGB3A2, AGER, SCGB1A1 and HTR3C genes shown in NO.32-38; microRNA binding sites are deleted in the 3'UTR, such as sequences obtained by deleting part or all of the microRNA binding sites from the 3'UTR of TMSB10, AGER, and SCGB3A2 genes shown in SEQ ID NO.39-43.

[0119] In the above examples, the native 5'UTR and 3'UTR of the highly expressed gene were used as target UTRs for increasing the expression of the target CDS. In other examples, after obtaining the 5'UTR and 3'UTR of the highly expressed gene, they can be further adjusted to obtain variants of the 5'UTR and 3'UTR of the highly expressed gene, and these variants can be used as the target 5'UTR and target 3'UTR.

[0120] In some embodiments, the 5'UTR variant may be a sequence that is completely complementary to the microRNA seed region with gene silencing function or an element that inhibits gene expression deleted from the native 5'UTR of a highly expressed gene, and the 3'UTR variant may be a sequence that is completely complementary to the microRNA seed region with gene silencing function or an element that inhibits gene expression deleted from the native 3'UTR of a highly expressed gene. MicroRNA seed sequences with gene silencing function can be obtained in a variety of ways, for example, using prediction tools such as miRBase, RNA22, Traget Scan 8.0, miRcode, miRDB, microRNA.org, PicTar, PITA, etc., to predict possible microRNA seed sequences. For example, sequences with a prediction score of >70% are used as microRNA seed sequences.

[0121] The above-mentioned microRNA seed sequence can be obtained through multiple predictions. In one embodiment, after obtaining the natural 5'UTR sequence of a gene, a prediction can be performed based on the natural 5'UTR sequence to obtain a first microRNA seed region. Based on the first microRNA seed region, the sequence complementary to the first microRNA seed region is deleted from the natural 5'UTR sequence to obtain a first 5'UTR deleted sequence. Thereafter, a second prediction can be performed based on the first 5'UTR deleted sequence to obtain a second microRNA seed region. Based on the second microRNA seed region, the region complementary to the second microRNA seed region is deleted from the first 5'UTR deleted sequence to obtain a second 5'UTR deleted sequence. It is understood that the first and second 5'UTR deleted sequences in this embodiment are the above-mentioned 5'UTR variants. It is understood that this prediction method can be performed multiple times, that is, after obtaining the second 5'UTR deleted sequence, prediction can still be performed again based on the second 5'UTR deleted sequence. This article does not list the number of times exhaustively.

[0122] Among them, the natural 5'UTR, the first 5'UTR deleted sequence, the second 5'UTR deleted sequence, and other 5'UTR deleted sequences obtained after more predictions and deletions can all be used as target 5'UTRs. In actual applications, the experimenter can select according to needs.

[0123] Similarly, for 3'UTR, the above operations can also be performed to obtain a natural 3'UTR and at least one 3'UTR deletion sequence (ie, at least one 3'UTR variant), and any one of them can be used as the target 3'UTR.

[0124] Therefore, by combining the above prediction tools and using the natural 5'UTR and natural 3'UTR of the above genes as templates, the sequence of any microRNA seed region predicted, as well as any 5'UTR deletion sequence and 3'UTR deletion sequence obtained after deletion, are all within the scope of protection of this application.

[0125] For the 5'UTR, elements that inhibit gene expression include at least one of the following: the TOP sequence CCUCUCUU, the AUG contained in the 5'UTR and the upstream open reading frame, the RNAG-quadruplex (RG4), and an A-rich element (poly A). Specifically, any one or more of these sequences can be deleted from the native 5'UTR of a highly expressed gene to generate a 5'UTR variant, which can be used as a target 5'UTR for increasing the expression of the target CDS. For example, the final target 5'UTR may not contain the start codon AUG or the open reading frame.

[0126] It is understandable that when the 5'UTR and 3'UTR variants are used as the target 5'UTR and target 3'UTR for increasing the expression level of the target CDS, only the sequence that is completely complementary to the microRNA seed region having a gene silencing function can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene; or only the element that inhibits gene expression can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene; or both the sequence that is completely complementary to the microRNA seed region having a gene silencing function and the element that inhibits gene expression can be deleted.

[0127] In some embodiments, bases or sequences for increasing efficient gene expression may be added to the natural 5'UTR and 3'UTR of highly expressed genes. The added bases or sequences should ensure that the 5'UTR and 3'UTR have stable secondary structures.

[0128] For example, any one or any combination of the following sequences can be added to the native 5'UTR of a highly expressed gene: sequence elements ACTC ACT ATT TGT TTT CGC GCC CAG TTGCAA AAAGTG TCG for aptamer recruitment of proteins that promote translation or capping, PABPs (PolyA-binding proteins) recognition motifs, PCBPs (PolyC-binding proteins) recognition motifs, YTHDFs recognition motifs, IRES sequences, conserved sequences in the 5'UTR of homologous genes in eukaryotes, such as the Kozak sequence (GCCACCAUGG), conserved motifs in the 5'UTR sequences of highly expressed genes across all mammals, microRNA miR-24-1 that plays a positive regulatory role in gene expression, and variants of the above elements, etc. The conserved motifs in the 5'UTR sequences of highly expressed genes across all mammals can be obtained using sequence alignment tools, including DNAman, MEGA, sequencelogo, the online tool WebLogo3, and the MEME Suite (Motif-based sequence analysis tools). Among them, variants of the elements include naturally occurring DNA sequences and homologues, variants, fragments and corresponding RNA sequences thereof, for example, DNA sequences or fragments or variants thereof having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity compared to the native DNA sequence.

[0129] In the examples of this application, when inserting an element, the insertion site at the 5' end of the UTR sequence should have a low secondary structure free energy. The optimal position includes, but is not limited to, the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th position upstream of the start codon AUG.

[0130] The free energy is predicted based on RNA secondary structure prediction tools, including LinearDesign (https: / / rna.baidu.com / ), RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi), MC-Fold MC-Sym (https: / / major.iric.ca / MC-Pipeline / ), DNAman, RNAComposer (http: / / rnacomposer.ibch.poznan.pl / ), and ARES (http: / / drorlab.stanford.edu / ares.html).

[0131] To the natural 3'UTR of a highly expressed gene, any one or any combination of the following sequences is added, and the optimal position for insertion into the 3'UTR includes but is not limited to after the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th or 10th nucleotide downstream of the 5' end or before the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th or 10th nucleotide position upstream of the Poly A tail:

[0132] Sequence element that recruits the RNA-binding protein HUR, which enhances RNA stability: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA;

[0133] The cis-acting element R3U (Repeated sequence element 3 and U-rich element, AAAACUCAAUGUAUUUCUGAGGAAGCGUGGUGCAUAAUGCCACGCAGCGUCUGCAUAACUUUUAUUUCUUUUUAUUAAUCAACAAA) present in the 3'UTR;

[0134] QRE1(QKI response element 1,GCCGUAACCACGUCUACUAACGCCG);

[0135] QRE2(QKI response element 2,AACUCCAGGACUGUAUUUGUGACUAAUUGUAUAACAGGUU);

[0136] ARE (AU-rich element, UAUGUCUGUUUUUGUAUCUUUAUGCUGUAUUUUAACACUUUGUAUUACUUAGGUUAUU);

[0137] The ribosome binding site (AUAAGAUACACCUGCAAAGGCGGCACAACCCCAGUGCCACGUUGUGAGUUGGAUAGUUGUGGAAAGAGUCAAAUGGCUCUCCUCAAGCGUAUUCAACAAGGGGCUGAAGGAUGCCCAGAAGGUACCCCAUUGUAUGGGAUCUGAUCUGGGGCCUCGGUGCACAUGCUUUACAUGUGUUUAGUCGAGGUUAAAAAACGUCUAGGCCCCCCGAACCACGGGGACGUGGUUUUUCCUUUGAAAAACACGAUGAUAAU), etc.

[0138] In some embodiments, variants of the natural 5'UTR and 3'UTR of highly expressed genes may also comprise modifications to at least one base in the 5'UTR and 3'UTR, wherein the modifications include one or more of: N6-methyladenosine (m6A), N1-methyladenosine (m1A), 5-methylcytosine (m5C), 5-hydroxymethylcytosine (hm5C), pseudouridine (Ψ), inosine (I), uridine (U), and ribose methylation (2'-O-Me).

[0139] Among them, m6A modification refers to the modification of the nitrogen at position 6 of adenylic acid (A) by methylation. m1A refers to the modification of the nitrogen at position 1 of adenylic acid (A) by methylation. m5C refers to the methylation of the carbon at position 5 of cytosine (C). Pseudouridine (Ψ) is an isomer of uridine, which leads to the formation of different base structures. Inosine (I) is formed by deamination of adenylic acid, resulting in the removal of the amino group. During RNA translation, inosine can pair with cytosine, uracil or adenine. Ribose methylation (2'-O-Me) refers to the addition of a methyl group to the 2' oxygen of ribose in RNA.

[0140] It is understood that when variants of the natural 5'UTR and 3'UTR of highly expressed genes are used as target 5'UTR and 3'UTR, the above variant methods can be performed alone or in combination with each other. For example, only the sequence that is completely complementary to the microRNA seed region with gene silencing function can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene; only the element that inhibits gene expression can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene; bases or sequences for increasing efficient gene expression can be added to the 5'UTR and 3'UTR; only the natural 5'UTR and 3'UTR of the highly expressed gene can be base modified; the sequence that is completely complementary to the microRNA seed region with gene silencing function and the element that inhibits gene expression can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene at the same time; the above-mentioned sequences can be deleted from the natural 5'UTR and 3'UTR of the highly expressed gene, and bases or sequences for increasing efficient gene expression can be added; bases or sequences for increasing efficient gene expression can be added to the 5'UTR and 3'UTR, and then base modification can be performed; and so on. Various combinations are provided, which are not exhaustively listed in this application. In addition, only one of the 5'UTR and 3'UTR may be mutated, while the other one still uses its native sequence as the target UTR for increasing the expression level of the target protein or polypeptide.

[0141] When designing variants of the natural 5'UTR and 3'UTR of highly expressed genes, screening for conserved sequences that can be added or deleted can be performed by performing multiple sequence analysis on the 5'UTR or 3'UTR of multiple highly expressed genes according to the present invention using software such as Clustalw, Clustalx, and MEGA, and conserved site analysis can be performed using sequencelogo, the online tool WebLogo3, and The MEMESuite (Motif-based sequence analysis tools). The obtained conserved elements are added or deleted in the natural 5'UTR and 3'UTR, such as the previously mentioned Kozak sequence (GCCACCAUGG).

[0142] At the same time, the highly expressed UTR sequences obtained by the present invention can be further used as training sets to obtain de novo designed UTR sequences through artificial intelligence learning algorithms including but not limited to genetic algorithms, deep convolutional neural network algorithms, transfer learning algorithms, regression algorithms, etc.

[0143] It can be understood from the description of the above embodiments that each UTR (3' and 5') in the present application can be a naturally occurring DNA sequence and its homologs, variants, fragments and corresponding RNA sequences. For example, the target UTR can be a DNA sequence or a fragment or variant thereof having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% identity compared to the DNA sequence of a certain highly expressed gene; or an RNA sequence or a fragment or variant thereof having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% identity compared to the corresponding RNA sequence of a certain highly expressed gene.

[0144] In addition to optimizing the 5'UTR and 3'UTR as described above, in the present embodiments, the target CDS corresponding to the protein / polypeptide to be amplified and expressed can also be optimized. The target CDS can be the original CDS sequence, or the original CDS sequence can be partially or fully codon-optimized. For example, compared to the original CDS sequence, the optimized sequence has an increased G / C content.

[0145] To summarize the above technical solutions, in the embodiments of the present application, the target 5'UTR is selected from one of the following sequences:

[0146] Any one of SEQ ID. NO: 9-31;

[0147] Inserting SEQ ID. NO: 1 into position 2 near the 5' end of SEQ ID. NO: 9, and / or inserting any one, two, three, four or five sequences of SEQ ID. NO: 1-5 into positions 1-15 near the 5' end and / or positions 1-15 near the 3' end of SEQ ID. NO: 9;

[0148] A sequence formed by inserting any one, two, three, four, or five of SEQ ID NOs: 1-5 into positions 1-15 of the 5' end or positions 1-15 of the 3' end of SEQ ID NO: 10;

[0149] A sequence consisting of: adding SEQ ID.NO: 1 to the 9th base near the 3' end of SEQ ID.NO: 11, and / or inserting SEQ ID.NO: 1 at positions 1-15 of the 5' end of SEQ ID.NO: 11 or at any one of positions 1, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, and 15 near its 3' end, and / or inserting any one, two, three, four, or five of SEQ ID.NO: 1-5 at positions 1-15 of the 5' start end of SEQ ID.NO: 11 or at positions 1-15 near its 3' end;

[0150] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 12 or near positions 1-15 of the 3' end thereof;

[0151] A sequence consisting of: adding SEQ ID.NO: 1 to the sixth base from the 5' end of SEQ ID.NO: 13, and / or inserting SEQ ID.NO: 1 at any one of positions 1, 2, 3, 4, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 near the 5' end of SEQ ID.NO: 13 or positions 1-15 near its 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO: 1-5 at positions 1-15 from the 5' end of SEQ ID.NO: 13 or positions 1-15 near its 3' end;

[0152] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 14 or positions 1-15 near its 3' end;

[0153] A sequence composed of any one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 15 or positions 1-15 near the 3' end thereof;

[0154] A sequence consisting of inserting SEQ ID.NO:1 into position 2 or 15 near the 5' end of SEQ ID.NO:16, and / or inserting SEQ ID.NO:1 into position 7 near the 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO:1-5 into positions 1-15 of the 5' end or positions 1-15 near the 3' end of SEQ ID.NO:16;

[0155] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 17 or positions 1-15 near the 3' end thereof;

[0156] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 18 or near positions 1-15 of the 3' end thereof;

[0157] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 19 or near positions 1-15 of the 3' end thereof;

[0158] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 20 or near positions 1-15 of the 3' end thereof;

[0159] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 21 or near positions 1-15 of the 3' end thereof;

[0160] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 22 or near positions 1-15 of the 3' end thereof;

[0161] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 23 or near positions 1-15 of the 3' end thereof;

[0162] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 24 or positions 1-15 near its 3' end;

[0163] A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 25 or near positions 1-15 of the 3' end thereof;

[0164] A sequence consisting of inserting SEQ ID.NO:1 at the 27th base from the 3' end of SEQ ID.NO:26, and / or inserting SEQ ID.NO:5 at the 19th base near the 3' end, and / or adding SEQ ID.NO:2 at the 27th base near the 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO:1-5 at positions 1-15 from the 5' end of SEQ ID.NO:26 or at positions 1-15 near its 3' end;

[0165] The target 3'UTR is selected from one of the following sequences:

[0166] Any one of SEQ ID. NO: 32-43;

[0167] 50%, 60%, 70%, 80%, 90% and 100% of all microRNA binding sites with prediction scores > 70% are removed from SEQ ID. NO: 32-38, and any one, any two, or three of the sequences SEQ ID. NO: 6-8 are added to the 3' end to form the sequence; or the sequence element of RNA binding protein HUR is added to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, any one, any two, three, four, five or six of the following, followed by a sequence;

[0168] The sequence is formed by adding any one, any two, or three of the sequences SEQ ID. NO: 6-8 to the 3' end of SEQ ID. NO: 39; or adding the sequence element of the RNA binding protein HUR to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, any one, any two, three, four, five or six of which are followed by a sequence consisting of:

[0169] A sequence composed of any one, any two, or three of SEQ ID. NOs: 6-8 is added to the 3' end of SEQ ID. NO: 40; or a sequence element of RNA binding protein HUR is added to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, any one, any two, three, four, five or six of the following, followed by a sequence;

[0170] The sequence is formed by adding any one, any two, or three of the sequences SEQ ID. NO: 6-8 to the 3' end of SEQ ID. NO: 41; or adding the sequence element of the RNA binding protein HUR to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, any one, any two, three, four, five or six of the following, followed by a sequence;

[0171] The sequence is formed by adding any one, any two, or three of the sequences SEQ ID. NO: 6-8 to the 3' end of SEQ ID. NO: 42; or adding the sequence element of the RNA binding protein HUR to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, any one, any two, three, four, five or six of the following, followed by a sequence;

[0172] The sequence is formed by adding any one, any two, or three of the sequences SEQ ID. NO: 6-8 to the 3' end of SEQ ID. NO: 43. Alternatively, the sequence element of the RNA binding protein HUR is added to the 3' end: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Iron responsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and variants thereof, followed by any one, any two, three, four, five or six of the following.

[0173] The source genes and functional element names of the above sequences are summarized in Tables 4 and 5 below:

[0174] Table 4 Serial numbers corresponding to functional elements

[0175]

[0176]

[0177] Table 5 Sequence numbers corresponding to target UTRs

[0178]

[0179] The following application will illustrate the technical effects of the artificial nucleic acid molecules of the present application in increasing protein or polypeptide expression through multiple experimental examples. In the following examples, firefly luciferase (Fluc) or the novel coronavirus SARS-CoV-2 spike protein receptor domain (RBD) is used as the target protein (protein whose expression needs to be increased) for description. That is, the target CDS sequence is the gene sequence encoding firefly luciferase or the novel coronavirus SARS-CoV-2 spike protein receptor domain.

[0180] First, a DNA template for in vitro transcription was prepared. A vector for in vitro transcription was constructed containing a T7 promoter, a gene sequence encoding firefly luciferase (FLuc), and an interrupted A110 polyadenylation sequence (derived from BNT162b2, BioNTech). Restriction sites were added to the front of the T7 promoter and the end of the polyadenylation sequence to linearize the vector prior to in vitro transcription.

[0181] Then perform in vitro transcription. The DNA template according to the embodiment is linearized using restriction endonucleases, and then in vitro transcription is performed using T7-polymerase, and uridine is completely replaced by N1-methyl pseudouracil (APExBIO, B7972-.1). The DNA template is then digested by DNA enzyme treatment. After transcription, 7.5mM lithium chloride precipitate is used for purification. 30μL lithium chloride and 30μL RNA-free water precipitate (Thermo fisher, AM9480) are added to each 20μL reaction system and mixed well, and then placed in -20°C for 30min; then it is placed in a 4°C pre-cooled centrifuge and centrifuged at 12000rpm for 15min; after the end, the supernatant is discarded, 70% ethanol pre-cooled at 4°C is added, and centrifuged at 12000rpm for 2min after gentle pipetting, and this step is repeated three times, the supernatant is discarded, and 60-100μL RNase-free water is added. After transcription, a Cap1 structure was added to the 5' end using a Novoprotein kit (Novoprotein, M082-01B). After capping, the mRNA was purified using a 7.5 mM lithium chloride precipitation solution, following the same steps as above, followed by mRNA electrophoresis.

[0182] Subsequently, luciferase expression was expressed via mRNA lipid delivery / transfection. Female Balb / c mice weighing 18-20 g and approximately 6-8 weeks old (Wei Tong Li Hua Laboratory Animal Co., Ltd.) were selected and cultured in an SPF-grade animal room for approximately one week. Each mouse was administered 50 μL of the mRNA-LNP formulation (approximately 3 μg of mRNA for all routes except nebulization) via intratracheal spray (it), aerosol inhalation (nebulization), intramuscular injection (im), or intranasal injection (in).

[0183] Finally, a luciferase assay was performed. The IVIS in vivo imaging device was used to detect the expression of the Luciferase reporter gene in mice. The substrate D-luciferin solution (3 mg D-luciferin dissolved in 100 μL PBS) was injected into the mouse via intraperitoneal injection (ip) at a dose of 150 mg / kg. After 10 minutes, the mice were anesthetized in an isoflurane chamber, and then bioluminescence imaging was performed using the imaging system in IVIS (Perkin Elmer, USA). After the in vivo imaging of the mice was completed, the mice were killed by cervical dislocation, and organs such as the lungs, liver or spleen were quickly removed, and bioluminescence imaging was performed again on the obtained isolated organs. The software used was Living Image (PerkinElmer, USA), and the images were analyzed by measuring the radiance in the region of interest (ROI), and the values were displayed as the total radiance (Total Flux [p / s]). Alternatively, Luciferase activity was measured as relative light units (RLU) in a Navigator microplate luminometer. 90 μL of lysate and 100 μL of Firefly Luciferase Assay Reagent were used to measure FLuc activity with a measurement time of 10 s.

[0184] Alternatively, human bronchial epithelial cells (16HBE) or lung cancer human alveolar basal epithelial cells (A549) or 293T cells derived from human embryonic kidney cells were seeded in 96-well plates at a density of 2.5×104 cells / well. After overnight culture, the cells were washed in Opti-MEM and then transfected with 0.4 μg / well of Lipofectamine 3000-complexed FireLuc encoding mRNA dissolved in Opti-MEM. (Prepare mRNA encoding FireLuc, add the mRNA to Opti-MEM containing Lipofectamine 3000, and add 0.4 μg mRNA / well to the 96-well plate seeded with cells) 6 hours after transfection, Opti-MEM was replaced with culture medium. After 24 hours, the culture medium was aspirated and the cells were lysed in 100 μL of lysis buffer. The cleavage product was detected using a Firefly Luciferase Reptor Gene Assay Kit (Beyotime).

[0185] Experimental Example 1

[0186] Preparation of DNA template for in vitro transcription.

[0187] Construct a vector for in vitro transcription. The vector backbone can be a commonly used AT cloning vector such as PUC57, PBS-T or PMD19. In this example, we used the PUC57 cloning vector. The vector contains the T7 promoter 5'-TAATACGACTCACTATA-3, 5'UTR (designed in this patent), an open reading frame encoding the target gene such as the gene sequence of firefly luciferase (FLuc) or the new coronavirus SARS-CoV-2 spike protein receptor domain, 3'UTR (designed in this patent), and 110 nucleotides of polyadenylic acid residues A (the tail consists of 30 adenosine residues, followed by a 10-nucleotide linker sequence GCATATGACT and another 70 adenosine residues). Restriction sites such as HindIII and XbaI are added to the front end of the T7 promoter and the end of the polyadenylic acid sequence to linearize the vector before in vitro transcription. Figure 1 The designed gene sequence was submitted to Shanghai GenScript for sequence synthesis to obtain the plasmid template and recombinant strain.

[0188] Experimental Example 2

[0189] The DNA template according to Example 1 was linearized using restriction endonucleases and subsequently transcribed in vitro using T7 polymerase, completely replacing uridine with N1-methylpseudouracil (APExBIO, B7972-.1). Transcription was performed at 37°C for 3 h. The DNA template was then digested by DNase treatment. After transcription, the DNA template was purified using 7.5 mM lithium chloride precipitation solution. For every 20 μL reaction system, 30 μL lithium chloride and 30 μL RNA-free water precipitation solution (Thermo Fisher, AM9480) were added, mixed, and placed at -20°C for 30 min. Next, the reaction was placed in a 4°C pre-cooled centrifuge and centrifuged at 12,000 rpm for 15 min. After completion, the supernatant was discarded, 4°C pre-cooled 70% ethanol was added, and the mixture was gently pipetted. The reaction was centrifuged at 12,000 rpm for 2 min. This step was repeated three times, the supernatant discarded, and 60-100 μL RNA-free water was added. After transcription, a Cap1 structure was added to the 5' end using a nearshore kit (Novoprotein, M082-01B). After capping, the mRNA was purified using 7.5 mM lithium chloride precipitation solution, following the same steps as above. mRNA electrophoresis was then performed, and nucleic acid quality and concentration were determined using a NanoDrop Microvolume Spectrophotometer (ThermoFisher, USA).

[0190] Experimental Example 3

[0191] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS, and the genes TMSB10, SFTPD, AGER, LRRK2, MSR1, SCGB3A2, and SFTPA2, which are highly expressed in the lungs, were used as the source of target UTRs. The following eleven sets of cell-based experiments were set up:

[0192] TM group: the target 5'UTR of the target CDS is the native 5'UTR of the TMSB10 gene (ie, sequence SEQ ID NO: 9); the target 3'UTR of the target CDS is the native 3'UTR of the TMSB10 gene (ie, sequence SEQ ID NO: 32).

[0193] B3A2 group: the target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene (ie, sequence SEQ ID NO: 13); the target 3'UTR of the target CDS is the native 3'UTR of the SCGB3A2 gene (ie, sequence SEQ ID NO: 35).

[0194] TPD group: the target 5'UTR of the target CDS is the native 5'UTR of the SFTPD gene (ie, sequence SEQ ID NO: 15); the target 3'UTR of the target CDS is the native 3'UTR of the SFTPD gene (ie, sequence SEQ ID NO: 33).

[0195] AGER group: The target 5'UTR of the target CDS is the native 5'UTR of the AGER gene (ie, sequence SEQ ID NO: 16); the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene (ie, sequence SEQ ID NO: 36).

[0196] Group B6: The target 5'UTR of the target CDS is the native 5'UTR of the TMSB10 gene (i.e., sequence SEQ ID NO: 9); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0197] Group A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SFTPA2 gene (i.e., sequence SEQ ID NO: 11); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0198] K2B6 group: The target 5'UTR of the target CDS is the native 5'UTR of the LRRK2 gene (i.e., sequence SEQ ID NO: 25); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, in which the sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0199] R1B6 group: The target 5'UTR of the target CDS is the native 5'UTR of the MSR1 gene (i.e., sequence SEQ ID NO: 26); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which the sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0200] BF group: The target 5'UTR and 3'UTR of the target CDS are the 5'UTR and 3'UTR of the mRNA vaccine (BNT162b2 of BioNTech, BF) that has been marketed. The target 5'UTR of the target CDS is: the 5'UTR variant of the α-globin gene (GAGAATAAACTAGTATTCTTCTGGTCCCCACAGACTCAGAGAGAACCCGCCACC) obtained by modifying the end of the natural 5'UTR of the α-globin gene to the Kozak sequence (GCCACC); the target 3'UTR of the target CDS is the AES (Amino-terminal enhancer of split) gene sequence and mtrnr1 (mitochondrial 12S rRNA) combined (CTCGAGCTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTACCCCGAGTCTCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACCTGCCCCACTCACCACCTCTGCTAGTTCCAGACACCTCC CAAGCACGCAGCAATGCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACAGCAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATACTAACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACCCTGGAGCTAGC

[0201] ). This group was the control group.

[0202] MF group: The target 5'UTR and 3'UTR of this target CDS were the 5'UTR and 3'UTR of a marketed mRNA vaccine (Moderna mRNA-1273, MF). The target 5'UTR of this target CDS was a gp160 gene 5'UTR variant (GGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGACCCCGGCGCCGCCACC) obtained by modifying the end of the native 5'UTR of the gp160 gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS was the native 3'UTR of the HBA1 gene (GCTGGAGCCTCGGTGGCCTAGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGTGGGCGGCA). This group served as the control group.

[0203] Group α: The target 5'UTR of this target CDS was modified by adding the Kozak sequence (GCCACC) to the native 5'UTR of the α-globin gene (GACTCTTCTGGTCCCCACAGACTCAGAGAGAACGCCACC). The target 3'UTR of this target CDS was the native 3'UTR of the α-globin gene (GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGTGGGCGGCA). This group served as the control group.

[0204] The mRNA prepared by in vitro transcription of TM, B3A2, B6, A2B6, MF, BF and α groups in the above groups was transfected into human bronchial epithelial 16HBE cells using Lipofectamine 3000 reagent (Lipofectamine 3000 Transfection Reagent, ThermoFisher, USA). The expression of luciferase was measured after 24 hours. Figure 2 As shown in a. It can be seen that the expression level of the target protein Fluc in the B3A2 group was significantly higher than that in the three control groups; the expression level of the target protein Fluc in the TM group was not significantly different from that in the α group, but was higher than that in the control groups MF and BF; the expression level of the target protein Fluc in the modified B6 group was significantly higher than that in the three control groups; similarly, the expression level of the target protein Fluc in the A2B6 group was also significantly higher than that in the three control groups, indicating that the natural 5'UTR and 3'UTR or variant UTR can all enhance the expression of the target protein in cells.

[0205] The mRNAs of TPD, AGER, K2B6, R1B6 and α synthesized in vitro in the above groups were transfected into mouse dendritic cells (DC2.4) by Lip3000, and the expression of luciferase was measured after 24 hours. Figure 2 As shown in b, it can be seen that the expression levels of the target protein Fluc in the TPD, AGER, K2B6, and R1B6 groups were significantly higher than that in the control group α, approximately 5 to 8 times higher than that in the α group.

[0206] The mRNAs synthesized in vitro by TM, B6, BF and MF in the above groups were transfected into human embryonic kidney 293T cells and human alveolar adenocarcinoma basal epithelial cells A549 cells respectively by Lip3000, and the expression of luciferase was measured after 24 hours. Figure 2As shown in Figures 2c and 2d, the expression levels of the target protein Fluc in the B1 and B6 groups were higher than those in the control groups BF and MF. The expression level of the target protein Fluc in the modified B6 group of the TM group was the highest in both cells, approximately 2 times that of the control group BF and 6 times that of the control group MF in 293T cells, and approximately 4 times that of the control group BF and 6 times that of the control group MF in A549 cells.

[0207] It can be seen that the combination of the native 5'UTR and 3'UTR, or the native 5'UTR and 3'UTR variants, can enhance the expression of the target protein in cells. This demonstrates that the combination of the native UTR and its variants of highly expressed lung genes can enhance the expression of exogenous genes in cells.

[0208] Experimental Example 4

[0209] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS, and the highly expressed genes SCGB3A2 and AGER from the lung were selected as the source of the target UTR. The following five groups of animal experiments were set up:

[0210] Group α: The target 5'UTR of this target CDS was added to the native 5'UTR of the α-globin gene, forming a 5'UTR variant with a Kozak sequence (GCCACC); the target 3'UTR of this target CDS was the native 3'UTR of the α-globin gene. This group served as the control group.

[0211] The BF group: The target 5'UTR and 3'UTR of this target CDS are the 5'UTR and 3'UTR of the marketed mRNA vaccine (BioNTech's BNT162b2, BF). The target 5'UTR of this target CDS is a 5'UTR variant of the α-globin gene obtained by modifying the end of the native 5'UTR of the α-globin gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS is a combination of the AES (Amino-terminal enhancer of split) gene sequence and mtrnr1 (mitochondrial 12S rRNA). This group served as the control group.

[0212] MF Group: The target 5'UTR and 3'UTR of this target CDS were the 5'UTR and 3'UTR of a marketed mRNA vaccine (Moderna mRNA-1273, MF). The target 5'UTR of this target CDS was a gp160 5'UTR variant obtained by modifying the end of the native 5'UTR of the gp160 gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS was the native 3'UTR of the HBA1 gene. This group served as the control group.

[0213] Group B3A2: The target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is the native 3'UTR of the SCGB3A2 gene.

[0214] AGER group: the target 5'UTR of the target CDS is the native 5'UTR of the AGER gene; the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene.

[0215] The five groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 h after delivery to the lungs of mice. Figure 3 As shown in Figures 3a and 3b. The expression of the target protein Fluc in mice was significantly higher in both the B3A2 and AGER groups compared to the three control groups: approximately 20-fold higher in the B3A2 group than in the control α group, 11-fold higher in the MF group, and 10-fold higher in the BF group; and approximately 8-fold higher in the AGER group than in the control α group, 5-fold higher in the MF group, and 4-fold higher in the BF group. In lung tissue, the expression levels of the target protein Fluc in both the B3A2 and AGER groups were similar, all significantly higher than in the three control groups: approximately 10-fold higher in the control α group, 7-fold higher in the MF group, and 3-fold higher in the BF group.

[0216] This example verifies that UTRs derived from highly expressed genes in the lungs can increase the expression of exogenous proteins or polypeptides in mice, especially in lung tissue.

[0217] Experimental Example 5

[0218] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS, and the highly expressed genes TMSB10, AGER, SCGB3A2, SFTPA1, SFTPA2, and MSR1 selected from the lungs were used as the source of target UTRs. The following thirteen animal experiments were set up:

[0219] Group α: The target 5'UTR of this target CDS was added to the native 5'UTR of the α-globin gene, forming a 5'UTR variant with a Kozak sequence (GCCACC); the target 3'UTR of this target CDS was the native 3'UTR of the α-globin gene. This group served as the control group.

[0220] The BF group: The target 5'UTR and 3'UTR of this target CDS are the 5'UTR and 3'UTR of the marketed mRNA vaccine (BioNTech's BNT162b2, BF). The target 5'UTR of this target CDS is a 5'UTR variant of the α-globin gene obtained by modifying the end of the native 5'UTR of the α-globin gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS is a combination of the AES (Amino-terminal enhancer of split) gene sequence and mtrnr1 (mitochondrial 12S rRNA). This group served as the control group.

[0221] MF Group: The target 5'UTR and 3'UTR of this target CDS were the 5'UTR and 3'UTR of a marketed mRNA vaccine (Moderna mRNA-1273, MF). The target 5'UTR of this target CDS was a gp160 5'UTR variant obtained by modifying the end of the native 5'UTR of the gp160 gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS was the native 3'UTR of the HBA1 gene. This group served as the control group.

[0222] TM group: the target 5'UTR of the target CDS is the native 5'UTR of the TMSB10 gene; the target 3'UTR of the target CDS is the native 3'UTR of the TMSB10 gene.

[0223] Group B6: The target 5'UTR of the target CDS is the native 5'UTR of the TMSB10 gene; the target 3'UTR of the target CDS is a 3'UTR variant obtained by removing a sequence completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0224] AGER group: the target 5'UTR of the target CDS is the native 5'UTR of the AGER gene; the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene.

[0225] AGER-4 group: The target 5'UTR of the target CDS is the native 5'UTR of the AGER gene (i.e., sequence SEQ ID NO: 16); the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene, with the 3'UTR variant having the sequence completely complementary to the seed sequence of the microRNA removed (i.e., sequence SEQ ID NO: 40).

[0226] Group B3A2: The target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is the native 3'UTR of the SCGB3A2 gene.

[0227] Group B3A2-1: The target 5'UTR of the target CDS is the natural 5'UTR of the SCGB3A2 gene (i.e., sequence SEQ ID NO: 13); the target 3'UTR of the target CDS is the natural 3'UTR of the SCGB3A2 gene, with the 3'UTR variant (i.e., sequence SEQ ID NO: 42) having the sequence completely complementary to the seed sequence of the microRNA removed.

[0228] Group B3A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0229] Group A1B6: The target 5'UTR of the target CDS is the native 5'UTR of the SFTPA1 gene (i.e., sequence SEQ ID NO: 10); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0230] Group A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SFTPA2 gene (i.e., sequence SEQ ID NO: 11); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0231] R1B6 group: The target 5'UTR of the target CDS is the native 5'UTR of the MSR1 gene (i.e., sequence SEQ ID NO: 26); the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which the sequence completely complementary to the seed sequence of the microRNA is removed (i.e., sequence SEQ ID NO: 39).

[0232] The mRNAs of the above groups were encapsulated with novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 h after delivery to the lungs of mice. Figure 4 As shown in Figures 4a and 4b, the expression of the target protein Fluc in mice in the TM group was lower than in the three control groups. The native 3'UTR of the TM group gene was then modified (i.e., the sequence completely complementary to the microRNA seed sequence was removed from the native 3'UTR of the TMSB10 gene) to generate a 3'UTR variant, thereby generating the target 3'UTR of the B6 group. The expression level of the target protein Fluc in the B6 group increased significantly, by more than 16 times, and was significantly higher than that in the control group. After the 3'UTR of the AGER group was modified, the expression level of Fluc also increased, and the Fluc expression level in the AGER-4 group was 1.5 times that of the AGER group. Similarly, after the 3'UTR of the B3A2 group was modified, the expression level of Fluc also increased, and the Fluc expression level in the B3A2-1 group was 1.5 times that of the B3A2 group. Especially in the lung tissue, the Fluc expression level of highly expressed genes in the lung was significantly increased after the 3'UTR of the lung was modified, among which the Fluc expression level in the B6 group was 3 times that of the TM group, the AGER-4 group was 1.5 times that of the AGER group, and the B3A2-1 group was 3.5 times that of the B3A2 group. The combination of the modified 3'UTR and the 5'UTR of highly expressed genes from the lungs can also increase the expression of target proteins in mice. The Fluc expression levels of the B3A2B6 group, A1B6 group, A2B6 group, and R1B6 group were all higher than those of the three control groups, approximately 8 to 75 times that of the α group, 4 to 37 times that of the MF group, and 5 to 44 times that of the BF group; the differences in lung tissue were also more significant, approximately 5 to 34 times that of the α group, 1.5 to 9 times that of the MF group, and 3.5 to 20 times that of the BF group.

[0233] This example verifies that claim 17 of the application mentions deletion of a completely complementary sequence of the microRNA seed region with gene silencing function from the 3'UTR and that claim 21 mentions that the highly expressed gene corresponding to the target 5'UTR and the target 3'UTR are derived from different gene and variant combinations.

[0234] Experimental Example 6

[0235] To investigate the effects of adding different elements to the 5'UTR and 3'UTR of highly expressed lung genes on mRNA protein expression during mRNA synthesis, the following four sets of experiments were conducted at the cellular and animal levels.

[0236] Group B6: The target 5'UTR of the target CDS was the native 5'UTR of the TMSB10 gene; the target 3'UTR of the target CDS was a 3'UTR variant of the native 3'UTR of the TMSB10 gene, in which the sequence completely complementary to the microRNA seed sequence was removed. This group served as the control group.

[0237] B6A7 group: The target 5'UTR of this target CDS is a 5'UTR variant obtained by adding the A7 element to the second base near the 5' end of the natural 5'UTR of the TMSB10 gene (GACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGGTTTCTTGCTGCAGCA ACGCGAGTGGGAGCACCAGGATCTCGGGCTCGGAACGAGACTGCACGGATTGTTTTAAGAAAGCCACC); the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0238] B6R3U group: The target 5'UTR of this target CDS is the native 5'UTR of the TMSB10 gene; the target 3'UTR of this target CDS is: a sequence completely complementary to the seed sequence of the microRNA is removed from the native 3'UTR of the TMSB10 gene, and then the R3U element (GATCCTGGAGGATTTCCTCCTCTTCGAGTCGCCGGTCGGTTCTCCGTAAATCGTGGC AAAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA) is added to its 3' end to obtain a 3'UTR variant.

[0239] B6A7R3U group: The target 5'UTR of this target CDS is a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the TMSB10 gene; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene and then adding the R3U element.

[0240] The above four groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 hours after delivery to the lungs of mice. Figure 5 The results showed that after the 3'UTR and 5'UTR were modified separately, the expression levels of the target protein Fluc in mice in the B6A7 and B6R3U groups were higher than those in the B6 group, 3 times and 1.2 times, respectively. The expression level of the target protein Fluc in the B6A7R3U group was comparable to that in the B6 group, with no significant difference. However, in lung tissue, the expression pattern of Fluc was the opposite. The expression levels of the target protein Fluc in mice in the B6A7 and B6R3U groups were lower than those in the B6 group, while the expression level of the target protein Fluc in the B6A7R3U group was higher than that in the B6 group, approximately 1.5 times that of the B6 group.

[0241] The above four groups of mRNA were transfected into 16HBE cells via Lip3000, and the expression of luciferase was measured after 24 h. Figure 5 c. The results show that adding a functional element to just one of the native 5'UTR and 3'UTR can increase target protein expression in cells. Adding functional elements to both the native 5'UTR and 3'UTR can potentially increase target protein expression even more significantly.

[0242] This example verifies that modifying the natural UTR of a highly expressed gene in the lung can increase the expression of an exogenous gene in tissues and cells in vivo. It also explains that the at least one 5'UTR or 5'UTR variant and the at least one 3'UTR or 3'UTR variant mentioned in claim 21 increase the protein production from the artificial nucleic acid molecule through addition or synergy.

[0243] Experimental Example 7

[0244] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS, and the highly expressed gene AGER from the lung was selected as the source of the target UTR. The following four sets of experiments were set up:

[0245] AGER group: the target 5'UTR of the target CDS is the native 5'UTR of the AGER gene; the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene.

[0246] AGER-1 group: The target 5'UTR of the target CDS is the native 5'UTR of the AGER gene, and the functional element A7 (GACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGAGACAGAGCCAGGACCCTGGAAGGAAGCAGGGCCACC) is added to its 5' end; the target 3'UTR of the target CDS is the native 3'UTR of the AGER gene.

[0247] AGER-2 group: The target 5'UTR of this target CDS is the natural 5'UTR of the AGER gene, and the functional element A7 (GAGACAGAGCCAGGACCCTGGAAGGAAGCAGGACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGGCCACC) is added to its 3' end, and has a lower secondary structure free energy; the target 3'UTR of this target CDS is the natural 3'UTR of the AGER gene.

[0248] AGER-3 group: The target 5'UTR of this target CDS is the natural 5'UTR of the AGER gene, and the functional element A7 (GAGACAGAGCCAGGACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGAC CCTGGAAGGAAGCAGGGCCACC) is added at the 14th base near the 5' end, which is predicted to have the lowest secondary structure free energy; the target 3'UTR of this target CDS is the natural 3'UTR of the AGER gene.

[0249] The above four groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 hours after delivery to the lungs of mice. Figure 6As shown in Figures 6a and 6b. In vivo, compared with the AGER group, the expression level of the target protein Fluc in the AGER-2 group was higher than that in the AGER group, while the expression levels of the target protein Fluc in the AGER-1 and AGER-3 groups were not significantly different from those in the AGER group. In lung tissue, the expression level of the target protein Fluc in the AGER-2 group was higher than that in the AGER group, approximately twice that of the AGER group, while the expression levels of the target protein Fluc in the AGER-1 and AGER-3 groups were not significantly different from those in the AGER group. Among the three groups, the 5'UTR of the AGER-1 group had the highest secondary structure free energy (-20.30 kcal / mol), while the 5'UTRs of the AGER-2 and AGER-3 groups had lower secondary structure free energies (-14.90 kcal / mol and -14.70 kcal / mol, respectively). This indicates that the position of the inserted element at the 5' end has a significant effect on protein expression. Inserting an element with the function of recruiting ribosomes near the start codon and maintaining the lowest secondary structure free energy into the 5'UTR can significantly improve protein expression.

[0250] Experimental Example 8

[0251] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS. The UTRs in the B3A2B6, A2B6, and R1B6 groups in Example 5 were selected, and the A7 element mentioned in Example 6 was added to the 5' end. The following seven groups of experiments were set up according to the addition strategy obtained in Example 5:

[0252] Group α: The target 5'UTR of this target CDS was added to the native 5'UTR of the α-globin gene, forming a 5'UTR variant with a Kozak sequence (GCCACC); the target 3'UTR of this target CDS was the native 3'UTR of the α-globin gene. This group served as the control group.

[0253] Group B3A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0254] B3A2A7B6 group: The target 5'UTR of this target CDS is a 5'UTR variant obtained by adding the A7 element to the 5th base from the 5' end of the natural 5'UTR of the SCGB3A2 gene (GACACACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGTTTGTGCAAGTG GAACCACTGGCTTGGTGGATTTTGCTAGATTTTTCTGATTTTTAAACTCCTGAAAAATATCCCAGATAACTGTCGCCACC), and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence that is completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0255] Group A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SFTPA2 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0256] A2A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the second base near the 3' end of the natural 5'UTR of the SFTPA2 gene (ACTTGGAGGCAGAGACCCAAGCAGCTGGAGGCTCTGTGTGTGGGTCGCTGATTTCT TGGAGCCTGAAAAGAAGGAGCAGCGACTGGACCCAACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCGGAGCCACC), and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0257] R1B6 group: the target 5'UTR of the target CDS is the native 5'UTR of the MSR1 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0258] R1A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element (i.e., sequence SEQ ID NO: 1) to the natural 5'UTR of the MSR1 gene near the 21st base of the 3' end, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the natural 3'UTR of the TMSB10 gene.

[0259] The six groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 h after delivery to the lungs of mice. Figure 7 As shown in Figures 7a and 7b. In vivo, after the addition of the A7 element to the 5'UTR, the expression levels of the target protein Fluc in the A2A7B6 and R1A7B6 groups were higher than those in the A2B6 and R1B6 groups, increasing by 0.75 and 2.4 times, respectively. Similarly, in lung tissue, the expression levels of the target protein Fluc in the A2A7B6 and R1A7B6 groups were higher than those in the A2B6 and R1B6 groups, increasing by 0.8 and 3.3 times, respectively. Although the Fluc expression level in the B3A2A7B6 group was lower than that in the B3A2B6 group in vivo, the Fluc expression level in the lung tissue of the B3A2A7B6 group was higher than that in the B3A2B6 group, approximately 1.3 times that in the B3A2B6 group. This example serves as a supplement to Examples 4 and 5, and illustrates that inserting a functional element at the 5' end, and inserting the functional element near the start codon while maintaining the lowest secondary structure free energy, can significantly improve protein expression.

[0260] Experimental Example 9

[0261] When synthesizing mRNA, in order to study the effects of the above strategies on the expression of target proteins in the liver and spleen, the following nine groups of animal level experiments were set up.

[0262] Group α: The target 5'UTR of this target CDS was added to the native 5'UTR of the α-globin gene, forming a 5'UTR variant with a Kozak sequence (GCCACC); the target 3'UTR of this target CDS was the native 3'UTR of the α-globin gene. This group served as the control group.

[0263] Group B6: The target 5'UTR of the target CDS was the native 5'UTR of the TMSB10 gene; the target 3'UTR of the target CDS was a 3'UTR variant of the native 3'UTR of the TMSB10 gene, in which the sequence completely complementary to the microRNA seed sequence was removed. This group served as the control group.

[0264] B6A7 group: The target 5'UTR of this target CDS is a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the TMSB10 gene; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0265] Group B3A2-1: The target 5'UTR of the target CDS is the natural 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is the natural 3'UTR of the SCGB3A2 gene, with the 3'UTR variant having the sequence completely complementary to the seed sequence of the microRNA removed.

[0266] Group B3A2-B6: The target 5'UTR of the target CDS is the native 5'UTR of the SCGB3A2 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0267] Group A2B6: The target 5'UTR of the target CDS is the native 5'UTR of the SFTPA2 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0268] A2A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the SFTPA2 gene, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0269] R1B6 group: the target 5'UTR of the target CDS is the native 5'UTR of the MSR1 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0270] R1A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0271] The above nine groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 h after delivery to the lungs of mice. Figure 8 As shown in a and 8b. In the liver and spleen of mice, after the microRNA binding site was removed by 3'UTR, its expression in the liver and spleen was significantly increased. For example, after B6, B3A2-1, B3A2-B6, A2B6 and R1B6 were transformed, their expression in the liver was 11, 29, 85, 28 and 5 times that of the control group α, respectively, and their expression in the spleen was 14, 22, 107, 30 and 17 times that of the control group α, respectively; on this basis , and after adding the functional element A7 to the 5'UTR, its expression level was further improved. For example, in liver tissue, B6A7 was 4.6 times that of the B6 group, A2A7B6 was 1.5 times that of the A2B6 group, and R1A7B6 was 15.7 times that of the R1B6 group. In spleen tissue, B6A7 was 6.5 times that of the B6 group, A2A7B6 was 4.8 times that of the A2B6 group, and R1A7B6 was 10.2 times that of the R1B6 group. This example shows that after 5'UTR and 3'UTR modification, the expression of the target protein can be increased not only in lung tissue, but this principle can also be applied to enhance protein expression in the liver and spleen.

[0272] Experimental Example 10

[0273] This example uses computer-aided design to optimize existing 5'UTR sequences. When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS. The 5'UTR sequence of the R1A7B6 group was selected and further sequence optimized using LinearDesign. Compared to the pre-optimized sequence, the predicted MRL (Mean Ribosome Load) was higher. The following six sets of experiments were set up:

[0274] R1A7B6 group: R1A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element (i.e., sequence SEQ ID NO: 1) to the natural 5'UTR of the MSR1 gene near the 21st base of the 3' end, and has the lowest secondary structure free energy, with a predicted MRL of 6.561705;; The target 3'UTR of this target CDS is a 3'UTR variant in which a sequence that is completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0275] R1A7B6-1 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy. The 5'UTR sequence obtained after further sequence optimization using LinearDesign (i.e., sequence SEQ ID NO: 27) has the highest predicted MRL of 6.994167; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0276] R1A7B6-2 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy. The 5'UTR sequence obtained after further sequence optimization using LinearDesign (i.e., sequence SEQ ID NO: 28) has a higher predicted MRL of 6.676221; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0277] R1A7B6-3 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy. The 5'UTR sequence obtained after further sequence optimization using LinearDesign (i.e., sequence SEQ ID NO: 29) has a higher predicted MRL of 6.803836; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0278] R1A7B6-4 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene (i.e., sequence SEQ ID NO: 30), and has the lowest secondary structure free energy. The 5'UTR sequence obtained after further sequence optimization using LinearDesign has a high predicted MRL of 6.894272; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0279] R1A7B6-5 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy. The 5'UTR sequence obtained after further sequence optimization using LinearDesign (i.e., sequence SEQ ID NO: 31) has a high predicted MRL of 6.86496; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0280] The six groups of mRNA were encapsulated by novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 h after delivery to the lungs of mice. Figure 9 a and 9b. In vivo, after the 5'UTR was optimized by LinearDesign, the expression levels of the target protein Fluc in the R1A7B6-2 and R1A7B6-4 groups were significantly increased, 2.8 and 1.4 times that of the R1A7B6 group, respectively. Except for the decreased expression level of the target protein Fluc in the R1A7B6-1 group, the expression levels of the target protein Fluc in the other two groups were not significantly different from those in the R1A7B6 group. In lung tissue, the expression levels of the target protein Fluc in the R1A7B6-2 and R1A7B6-5 groups were increased to a certain extent, 1.1 and 1.4 times that of the R1A7B6 group, respectively. Similarly, except for the decreased expression level of the target protein Fluc in the R1A7B6-1 group, the expression levels of the target protein Fluc in the lungs of the other two groups were not significantly different from those in the R1A7B6 group. This example shows that optimizing the 5'UTR sequence using the principles of AI-assisted design may help increase the expression of the target protein.

[0281] Experimental Example 11

[0282] When synthesizing mRNA, the CDS of the Fluc gene was used as the target CDS. UTRs in the R1B6 and R1A7B6 groups were selected, and IRESvI (i.e., sequence SEQ ID NO: 5) and PABPv3 elements (i.e., sequence SEQ ID NO: 2) were added to the 5' end. The following five experiments were performed based on the addition strategy obtained in Example 5:

[0283] R1B6 group: the target 5'UTR of the target CDS is the native 5'UTR of the MSR1 gene; the target 3'UTR of the target CDS is a 3'UTR variant of the native 3'UTR of the TMSB10 gene, from which a sequence completely complementary to the seed sequence of the microRNA is removed.

[0284] R1-IRESv1-B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant (GAATTGTAAAGAGAGAGAAGTGGATAAATCAGTGCTGCTTTCTTTAGTTTCATTTTAT TCCTATACTGGCTGCTTATGGTGACAATTGAGAGATCGTTACCATATAGCGACGAAAGA AGTGCCACC) obtained by inserting the IRESv1 element into the natural 5'UTR of the MSR1 gene near the 13th base from the 3' end, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence that is completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0285] R1-PABPv3-B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the PABPv3 element to the 21st base near the 3' end of the natural 5'UTR of the MSR1 gene (GAATTGTAAAGAGAGAGAAGTGGATAAATCAGTGCTGCTAAAAAAAAAAAACCAA AAAAAAAAAACAAAAAAAAAAAATTCTTTAGGACGAAAGAAGTGCCACC), and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence that is completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0286] R1A7B6 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the native 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the native 3'UTR of the TMSB10 gene.

[0287] R1A7-IRESv1-B6 group: The target 5'UTR of this target CDS is: the IRESv1 element is added to the 12th base at the 3' end of the 5'UTR of the R1A7B6 group, resulting in a 5'UTR variant (GAATTGTAAAGAGAGAGAAGTGGATAAATCAGTGCTGCTTTCTTTAGTTTCATTTTAT TCCTATACTGGCTGCTTATGGTGACAATTGAGAGATCGTTACCATATAGCGACGAAAGA AGTGCCACC); the target 3'UTR of this target CDS is a 3'UTR variant in which a sequence completely complementary to the seed sequence of the microRNA is removed from the native 3'UTR of the TMSB10 gene.

[0288] The above five groups of mRNA were encapsulated with novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, and the luciferase levels were measured 6 hours after delivery to the lungs of mice. Figure 10 As shown in Figures 10a and 10b. In vivo, the addition of the PABPv3 element to the 5'UTR of the R1B6 group increased the expression of the target protein Fluc to a certain extent, while the addition of the IRESv1 element decreased the expression of the target protein Fluc. The addition of the IRESv1 element to the 5'UTR of the R1A7B6 group significantly decreased the expression of the target protein Fluc. However, in mouse lung tissue, the addition of both the PABPv3 and IRESv1 elements to the 5'UTR of the R1B6 group significantly increased the expression of the target protein Fluc, both by approximately 2.3 times that of the R1B6 group. Similarly, the addition of the IRESv1 element to the 5'UTR of the R1A7B6 group also significantly increased the expression of the target protein Fluc, by approximately 1.8 times that of the R1A7B6 group. This example serves as a supplement to Examples 6 and 8, illustrating that inserting a functional element at the 5' end, and inserting the functional element near the start codon while maintaining the lowest secondary structure free energy, can significantly improve protein expression; at the same time, this example can explain that the 5'UTR variant involved in claim 16 is a combination of at least one and / or several variants and / or 5'UTR elements, and that these elements act in combination to improve the expression of the target protein.

[0289] Experimental Example 12

[0290] When synthesizing mRNA, the CDS of the gene encoding the receptor domain (RBD) of the novel coronavirus SARS-CoV-2 spike protein was used as the target CDS, and the gene TMSB10, which is highly expressed in the lungs, was used as the source of the target UTR. The following two sets of experiments were set up:

[0291] B6RBD group: The target 5'UTR of the target CDS is the natural 5'UTR of the TMSB10 gene; the target 3'UTR of the target CDS is a 3'UTR variant in which a sequence completely complementary to the seed sequence of the microRNA is removed from the natural 3'UTR of the TMSB10 gene.

[0292] BRBD group: The target 5'UTR of this target CDS is: the 5'UTR variant of the α-globin gene obtained by modifying the end of the natural 5'UTR of the α-globin gene to the Kozak sequence (GCCACC); the target 3'UTR of this target CDS is a combination of the AES (Amino-terminal enhancer of split) gene sequence and mtrnr1 (mitochondrial 12S rRNA).

[0293] The two groups of mRNA were encapsulated with novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs, which were then administered intramuscularly (IM) to mice for immunization. The immunization dose was 3 μg mRNA. The IgG titer in the mouse serum was measured 14 and 21 days after immunization. Figure 11 As shown in the figure, it can be seen that the antibodies produced in mice by the mRNA vaccine preparation of the B6RBD group were significantly higher than those of the BRBD group. This may be due to the higher expression level of the target protein RBD in the B6RBD group than that of the BRBD group, thus verifying the feasibility of increasing the expression level of the artificial nucleic acid molecule according to the patented strategy.

[0294] Experimental Example 13

[0295] When synthesizing mRNA, the CDS encoding the Klebsiella pneumoniae vaccine antigen (Ag2) gene was used as the target CDS, and the B6 and R1A7B6 groups were used as the source of the target UTRs. The following two sets of experiments were set up:

[0296] Group B-Ag2: The target 5'UTR of this target CDS is a variant of the α-globin gene 5'UTR obtained by modifying the terminus of the native 5'UTR of the α-globin gene to a Kozak sequence (GCCACC). The target 3'UTR of this target CDS is a combination of the AES (Amino-terminal enhancer of split) gene sequence and mtrnr1 (mitochondrial 12S rRNA). This group served as the control group.

[0297] R1A7B6-Ag2 group: The target 5'UTR of this target CDS is: a 5'UTR variant obtained by adding the A7 element to the natural 5'UTR of the MSR1 gene, and has the lowest secondary structure free energy; the target 3'UTR of this target CDS is a 3'UTR variant obtained by removing the sequence that is completely complementary to the seed sequence of the microRNA from the natural 3'UTR of the TMSB10 gene.

[0298] The two groups of mRNA were encapsulated with novel lipid nanoparticles (LNPs) to prepare mRNA-LNPs. The mice were administered intrapulmonary (it) for primary immunization at a dose of 3 μg mRNA. The IgG concentration in the mouse serum was measured 21 days after immunization. A secondary immunization was performed simultaneously. The immunization method and dosage were the same as those for the primary immunization. The IgG concentration in the mouse serum was measured seven days after the secondary immunization. A challenge experiment was also conducted. PBS treatment was used as a control experiment. The results are shown in the figure. Figure 12 As shown. It can be seen that 21 days after the first immunization, the antibody level of the B-Ag2 group decreased, and there was no significant difference with the R1A7B6-Ag2 group ( Figure 12 a); However, 7 days after the second immunization, the antibody level of mice in the R1A7B6-Ag2 group increased significantly and was significantly higher than that in the B-Ag2 group ( Figure 12 b). At the same time, mice immunized with R1A7B6-Ag2 group showed a 50% protection rate, which was significantly better than mice immunized with B-Ag2 group (12.5% protection rate) ( Figure 12 c) This is likely due to the stable, sustained expression of the target protein Ag2 in mice in the R1A7B6-Ag2 group, which induced the mice to produce higher antibody levels. This also verifies the feasibility of increasing the expression of the artificial nucleic acid molecule based on the patented strategy.

[0299] The present application also provides a transfectant, comprising: the artificial nucleic acid molecule and a vector as described above, wherein the vector is a DNA vector, a plasmid vector, or a viral vector.

[0300] The present application also provides an artificial cell, comprising: the transfection molecule as described above and a host cell of the vector, wherein the vector is used to carry the artificial nucleic acid molecule to the host cell, wherein the host cell is a mammalian cell.

[0301] The present application also provides a use of the artificial nucleic acid molecule and / or vector and / or artificial cell as described above in the preparation of medical supplies, wherein the medical supplies include RNA drugs and vaccines.

[0302] The present application also provides a method for delivering the artificial nucleic acid molecule as described above to a cell, wherein the artificial nucleic acid molecule is carried in the vector as described above; the method comprises: contacting the cell with the artificial nucleic acid molecule and / or vector when the artificial nucleic acid molecule and / or vector are capable of being taken up into the cell; wherein the cell is a mammalian cell.

[0303] The present application also provides a use of the artificial nucleic acid molecule as described above, characterized in that the artificial nucleic acid molecule is administered to a subject in a pharmaceutically effective amount for preventing and / or treating a disease or condition in a mammal.

[0304] The present application also provides a use of the artificial nucleic acid molecule and / or vector and / or cell as described above in the preparation of a medicament and / or pharmaceutical composition, wherein the medicament is used to prevent and / or treat a subject, wherein the artificial nucleic acid molecule and / or vector in the medicament and / or pharmaceutical composition is for preventing or treating a disease or condition.

[0305] The present application also provides a use of the artificial nucleic acid molecule and / or vector and / or cell as described above in the treatment and / or prevention of diseases or conditions, wherein the diseases or conditions include but are not limited to: immune system diseases, metabolic diseases, genetic diseases, cancer, blood diseases, bacterial infections or viral infections.

[0306] The present application also provides a kit comprising the artificial nucleic acid molecule and / or vector and / or cells described above. The kit also includes: cells for transfection, an adjuvant, a device for administering the pharmaceutical composition, a pharmaceutical carrier, and / or a pharmaceutical device for dissolving or diluting the artificial nucleic acid molecule, the vector, or the pharmaceutical composition.

[0307] Through the above experimental examples, it can be understood that the artificial nucleic acid molecules provided by the present application are used to increase the expression level of the target protein or peptide. The artificial nucleic acid molecule includes: a 5' end cap structure (Cap), a target 5' untranslated region (UTR), a target coding region (CDS), a target 3' untranslated region (UTR) and a PolyA tail. Wherein, the target CDS is used to be translated to produce the target protein or peptide. The sequence of the target 5'UTR is one of the following sequences: the 5'UTR of a highly expressed gene and the 5'UTR variant of the highly expressed gene. The sequence of the target 3'UTR is one of the following sequences: the 3'UTR of a highly expressed gene and the 3'UTR variant of the highly expressed gene. Wherein, the highly expressed gene is a gene with the 1st to 25th highest number of transcripts per million mapped reads (nTPM) per kilobase of transcription in all types of cells of all tissues of the human body. Since 5'UTR and 3'UTR have a regulatory effect on the translation and stability of mRNA, by selecting 5'UTR and 3'UTR and their variants from highly expressed genes, the mRNA molecule can be further stabilized, making it less susceptible to degradation and increasing the amount of protein or polypeptide translated from the mRNA molecule.

[0308] The above is only a preferred embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An artificial nucleic acid molecule for increasing the expression of a target amino acid, polypeptide or protein, characterized in that: The artificial nucleic acid molecule comprises: (a) at least one target 5' untranslated region (UTR), wherein the target 5'UTR sequence is a 5'UTR of a highly expressed gene in mammalian tissues or cells or a 5'UTR variant of the highly expressed gene; (b) at least one target coding region (CDS) for being translated to produce the target amino acid, polypeptide or protein; (c) at least one target 3' untranslated region (UTR), wherein the sequence of the target 3'UTR is a 3'UTR of a highly expressed gene in mammalian tissues or cells and a 3'UTR variant of the highly expressed gene; The highly expressed genes are genes with the 1st to 25th highest number of transcripts per million mapped reads (nTPM) in all types of cells in all tissues of mammals; the UTRs include naturally occurring DNA sequences and their homologs, variants, fragments and corresponding RNA sequences; Furthermore, the at least one 5'UTR and / or the at least one 3'UTR are used to increase the production of amino acids, polypeptides or proteins of the artificial nucleic acid molecule.

2. The artificial nucleic acid molecule according to claim 1, wherein The tissues where the highly expressed genes derived from the 5'UTR and 3'UTR are located include but are not limited to the lung, bronchi, trachea, liver, spleen, myocardial tissue, small intestine, stomach, colon, salivary gland, lymph node, brain, kidney, blood vessel, reproductive tract and esophagus.

3. The artificial nucleic acid molecule according to claim 1, wherein The cells where the highly expressed genes derived from the 5'UTR and 3'UTR are located include, but are not limited to, epithelial cells, alveolar cells, bronchial epithelial cells, ciliated airway epithelial cells, goblet cells, ionocytes, exocrine cells, microfold cells, dendritic cells, stromal cells, macrophages, T cells, endothelial cells, granulocytes, B cells, bile duct cells, erythrocytes, fibroblasts, hepatocytes, plasma cells, hepatic macrophages, basal cells, ion transport cells, smooth muscle cells, mucous gland cells, salivary duct cells, serous gland cells, distal intestinal epithelial cells, enteroendocrine cells, intestinal goblet cells, Paneth cells, stem cells, gastric mucus-secreting cells, collecting duct cells, glial cells, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocyte precursor cells, cardiomyocytes, spleen cells, kidney cells, and muscle cells.

4. The artificial nucleic acid molecule according to claim 1, wherein the highly expressed genes include but are not limited to the following genes: HTN3、STATH、EEF1A1、TPT1、TMSB4X、B2M、IGKC、BTG1、CD74、HTN1、HLA-DRA、 PTMA, PRB3, TXNIP, FAU, IGHM, RACK1, UBA52, ACTB, ZG16B, EEF1B2, CXCR4, NA CA、EIF1、FTH1、NTS、CHGA、REG4、S100A6、SST、PHGR1、PCSK1N、TFF3、PYY、TM SB10, FTL, LCN15, SPINK1, SCT, CHGB, TPH1, KRT8, H3-3B, GAPDH, DPP10, SLC1 A2, PCDH9, NRXN1, GPC5, GPM6A, CTNNA2, NPAS3, LSAMP, NRG3, CTNND2, ADGRV 1、NTM、RORA、LRP1B、ERBB4、ZBTB20、CADM2、ADGRB3、MAGI2、DTNA、SLC1A3、PP P2R2B, SOX5, ARHGAP24, KRT15, S100A2, FABP5, FOS, HSPB1, KRT5, ACTG1, AN XA1, PPIA, ZFP36L1, JUN, SFN, KRT19, HBB, SAA1, IGLC2, CD52, SAA2, HSP90AA 1、CCL4、IGHG4、AREG、ZFP36L2、CD69、NFKBIA、IL7R、TIMP1、MT2A、MGP、DCN、 CFD、JUNB、CCN1、VIM、SAT1、SFTPC、A2M、TAGLN、ACTA2、IGFBP7、TPM2、C11orf 96、MYL6、MYL9、JUND、CALD1、IGHA1、IGLC3、HLA-B、HLA-DPA1、IGHG3、PRB4、 SPP1, PRH2, PRH1, SMR3B, CCL3, S100A9, NAMPT, GUCA2B, GUCA2A, MT1G, MT1E MT1H, MT1X, CA7, PDPPF, LGALS3, FABP6, DOCK4, PLXDC2, FRMD4A, LRMDA, RBM S3、ST6GALNAC3、ITPR2、PTPRG、MEF2C、QKI、MEF2A、LAMA2、SFMBT2、ST6GAL1、 AUTS2, MAML2, SLC8A1, ABCB1, MBNL1, HS3ST4, NAV3, ELMO1, CST3, APOD, PTGD S、GSN、SPARCL1、PLA2G2A、IGFBP6、MFAP4、LUM、ADH1B、HBA1、HBA2、CA1、HBD、<h2 style=";text-align:left;direction:ltr">AHSP、HBM、PRDX2、H4C3、BLVRB、TUBA1B、CA2、SLC25A37、HMGB2、UBB、YBX1、A TP5F1E、IGFBP5、TIMP3、APOE、GPX3、CXCL14、CCL5、GNLY、GZMA、TPSB2、RGS1 、CD7、IL32、ZFP36、SRGN、KCNIP4、ROBO2、CNTN5、CNTNAP2、RBFOX1、NRXN3、C SMD1、GRIK1、SGCZ、ZNF385D、FGF14、DLGAP1、SYT1、ADARB2、GRIK2、IL1RAPL1 、GALNTL6、GRIP1、CCL21、IFITM3、ID3、IFITM1、ACKR1、S100A10、IFITM2、AL B、SERPINA1、CLU、AMBP、TFF2、ANXA4、TM4SF4、SCGB3A1、DEFB1、SLPI、GC、COX 7C、MUC7、WFDC2、FDCSP、KRT7、LCN2、AZGP1、UBC、LTF、DEFA5、DEFA6、REG3A、 PRSS2、LYZ、REG1A、ITLN2、OLFM4、PLP1、DLG2、PDE4B、ST18、CTNNA3、MBP、PTP RD、SLC24A2、NKAIN2、PEX5L、FMNL2、SLC44A1、NCKAP5、RNF220、PLCL1、UNC5 C、EDIL3、TMTC2、MAN2A1、JCHAIN、IGLV1-40、IGKV3-20、IGLV2-14、IGKV4-1 IGKV1-5, IGLV1-51, IGHV4-39, IGHV3-74, IGLV3-19, IGHV3-30, IGLV3-1, IGHA2, IGHV3-23, IGKV1D-13, IGHV2-70, IGKV3-11, IGHV4-59, IGHV5-51, IG LV2-23、IGHV1-3、IGKV1-12、IGLV3-25、APOC3、ORM1、HP、APOA2、TTR、APOC1 、APOA1、FGB、FGA、FGG、RBP4、FGL1、APOH、ORM2、CRP、KRT17、FOSB、S100A11、E GR1、CEBPD、DUSP1、COL1A2、COL1A1、CCL2、COL3A1、THBS4、HSPA1A、LGALS1、 S100A4、ADIRF、GADD45B、C1QB、C1QA、CTSB、GPX1、TYROBP、HLA-DPB1、CD163、<h2 style=";text-align:left;direction:ltr">PSAP、PRB1、CA6、PRB2、S100A8、SH3BGRL3、DLK1、GNAS、CD63、NNMT、PFN1、HS PA8、ID1、HLA-E、DNASE1L3、FABP1、PIGR、LGALS4、CKB、C15orf48、FXYD3、SE RF2、TSC22D3、KLF6、FABP4、CXCL2、RGS5、IFI27、TM4SF1、SCGB1A1、CCL20、C XCL1 、 CSTB 、 SERPINB3 、 BPIFB1 、 CXCL8 、 GDF15 、 SFTPB 、 EMP2 、 KRT18 、 DSTN 、 hop X 、 Ager 、 GPRC5A 、 anxa2 、 NAPSA 、 GSTP1 、 Rab11fip1 、 zg16 、 FCGBP 、 muc2 、 Agr2 、TFF1、KLK1、FXYD4、AQP2、ITM2B、ATP1B1、CRYAB、NPPA、TTN、MYL2、COX7A1、M YH7、MYL7、ANKRD1、ATP5ME、TPM1、TNNI3、MYH6、ATP5PF、MYL12A、COX6A2、CM YA5、SLC25A4、MYL3、COX5B、UQCRB、UQCRQ、ATP5MD、ATP5IF1、DYNLL1、CALM2、 PRDX5、MSMB、SCGB3A2、FCER1G、AIF1、CYBA、C7、HLA-DRB1、CTSD、SLC26A3、T SPAN8、CLDN3、HLA-C、HLA-A、NKG7、CD37、IER2、IGKV2-30、IGKV2-24、IGLV5 -45、IGLV1-44、IGLV7-46、IGLV1-47、IGHV7-4-1、IGKV1-16、IGKV3-15、IGH G1、IGKV1D-16、IGHV3-7、IGLV2-11、FCN3、EPAS1、ITLN1、FN1、IGKV1-6、IGLV 2-8、IGHV1-18、IGKV1-17、IGHV3-48、CREM、KLRB1、S100B、PMP22、GPM6B、AH NAK、CD9、ELF3、HES1、MYH11、DES、SPINK4、CLCA1、LRRTM4、ASIC2、FAM155A、R ALYL、PDE4D、SNTG1、LRRC4C、HSP90AB1、OAZ1、OIT3、STAB1、ALDOB、APOA4、P RAP1、RBP2、SELENOP、LHFPL3、DSCAM、NLGN1、OPCML、PCDH15、PTPRZ1、MMP16、<h2 style=";text-align:left;direction:ltr">GRID2、TNR、KCNMB2、MDGA2、RNASE1、CRHBP、FCN2、CLEC1B、SSR4、MZB1、HSP90B1、CAV1、SPARC、TUBA1A、BPIFA2 、MUC5AC、GKN1、PGC、PSCA、PGA3、GKN2、CYSTM1、S100P、KRT13、KRT4、CSTA、S100A14、PERP、SPINK5、IGHG2、IGH V4-34、IGLV6-57、IGHV5-10-1、IGHV1-46、TNNC2、ACTA1、MYLPF、TNNI2、CKM、NEB、TNNT3、MYL1、MB、YBX3、TCAP 、CA3、TNNC1、PI3、MMP10、CMC1、INSL5、GCG、H3-3A、LTB、EEF1D、SFTPA2、SFTPA1、CTSH、NPC2、CYB5A、HMGB1、HNR NPA1、FXYD2、NDUFA4、COX6B1、COX7A2、ATP5MG、UQCR10、UQCR11、COX4I1、LDHB、SRP14、CD36、CALM1、VWF、MTRN R2L1、GNG11、H2AZ1、ATP5F1B、PRDX1、PRSS1、CLPS、PNLIP、CTRB1、CPA1、CPB1、PLA2G1B、CTRC、GP2、CTRB2、SYC N, CEL, CELA3A, CELA2A, CPA2, CELA2B, PRSS3, CELA3B, MT1F, MIOX, DCXR, PDZK1IP1, PEBP1, IL1B, TXN, MMRN1, TNFAIP3, DUSP2, TPSAB1, CPA3, CA4, F13A1, MSR1, SFTPD, MS4A15, RTKN, SL34A2, LAMP3, HTR3C, CCL18, LRRK2, etc.

5. The artificial nucleic acid molecule according to any one of claims 1 to 4, wherein the UTR comprises at least one 5'UTR element derived from the 5'UTR of a highly expressed gene in the right 4 or its corresponding RNA sequence, homologue, fragment or variant, and / or at least one 3'UTR element derived from the 3'UTR of a highly expressed gene in the right 4 or its corresponding RNA sequence, homologue, fragment or variant.

6. The artificial nucleic acid molecule according to claim 5 , wherein the 5'UTR element derived from the highly expressed gene comprises or consists of a DNA sequence according to the highly expressed gene, or a DNA sequence having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the nucleic acid sequence thereof in ascending order of priority, or a fragment or variant thereof; or an RNA sequence according to the highly expressed gene, or a RNA sequence having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the nucleic acid sequence thereof in ascending order of priority, or a fragment or variant thereof.

7. The artificial nucleic acid molecule according to any one of claims 1 to 6, characterized in that The target 5'UTR is selected from one of the following sequences: Any one of SEQ ID. NO: 9-31; Inserting SEQ ID. NO: 1 into position 2 near the 5' end of SEQ ID. NO: 9, and / or inserting any one, two, three, four or five sequences of SEQ ID. NO: 1-5 into positions 1-15 near the 5' end and / or positions 1-15 near the 3' end of SEQ ID. NO: 9; A sequence formed by inserting any one, two, three, four, or five of SEQ ID NOs: 1-5 into positions 1-15 of the 5' end or positions 1-15 of the 3' end of SEQ ID NO: 10; A sequence consisting of: adding SEQ ID.NO: 1 to the 9th base near the 3' end of SEQ ID.NO: 11, and / or inserting SEQ ID.NO: 1 at positions 1-15 of the 5' end of SEQ ID.NO: 11 or at any one of positions 1, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, and 15 near its 3' end, and / or inserting any one, two, three, four, or five of SEQ ID.NO: 1-5 at positions 1-15 of the 5' start end of SEQ ID.NO: 11 or at positions 1-15 near its 3' end; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 12 or near positions 1-15 of the 3' end thereof; A sequence consisting of: adding SEQ ID.NO: 1 to the sixth base from the 5' end of SEQ ID.NO: 13, and / or inserting SEQ ID.NO: 1 at any one of positions 1, 2, 3, 4, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 near the 5' end of SEQ ID.NO: 13 or positions 1-15 near its 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO: 1-5 at positions 1-15 from the 5' end of SEQ ID.NO: 13 or positions 1-15 near its 3' end; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 14 or positions 1-15 near its 3' end; A sequence composed of any one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 15 or positions 1-15 near the 3' end thereof; A sequence consisting of inserting SEQ ID.NO:1 into position 2 or 15 near the 5' end of SEQ ID.NO:16, and / or inserting SEQ ID.NO:1 into position 7 near the 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO:1-5 into positions 1-15 of the 5' end or positions 1-15 near the 3' end of SEQ ID.NO:16; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 17 or positions 1-15 near the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 18 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 19 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 20 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 21 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 22 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 23 or near positions 1-15 of the 3' end thereof; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 24 or positions 1-15 near its 3' end; A sequence consisting of one, two, three, four, or five of SEQ ID NOs: 1-5 inserted into positions 1-15 of the 5' end of SEQ ID NO: 25 or near positions 1-15 of the 3' end thereof; A sequence consisting of inserting SEQ ID.NO:1 at the 27th base from the 3' end of SEQ ID.NO:26, and / or inserting SEQ ID.NO:5 at the 19th base near the 3' end, and / or adding SEQ ID.NO:2 at the 27th base near the 3' end, and / or inserting any one, two, three, four or five of SEQ ID.NO:1-5 at positions 1-15 from the 5' end of SEQ ID.NO:26 or at positions 1-15 near its 3' end; The target 3'UTR is selected from one of the following sequences: Any one of SEQ ID. NO: 32-43; A sequence obtained by removing 50%, 60%, 70%, 80%, 90% and 100% of all microRNA binding sites with a predicted score of >70% from SEQ ID. NOs: 32-38, and adding any one, any two, or three of SEQ ID. NOs: 6-8 to the 3' end; A sequence formed by adding any one, any two, or three of the sequences SEQ ID. NO: 6-8 to the 3' end of SEQ ID. NO: 39; A sequence formed by adding any one, any two, or three of the sequences SEQ ID. NOs: 6-8 to the 3' end of SEQ ID. NO: 40; A sequence formed by adding any one, any two, or three of the sequences SEQ ID. NOs: 6-8 to the 3' end of SEQ ID. NO: 41; A sequence formed by adding any one, any two, or three of the sequences SEQ ID. NOs: 6-8 to the 3' end of SEQ ID. NO: 42; The sequence is formed by adding any one, any two, or three of the sequences of SEQ ID. NOs: 6-8 to the 3' end of SEQ ID. NO:

43.

8. The artificial nucleic acid molecule according to any one of claims 1 to 7, wherein The at least one 5'UTR and the at least one 3'UTR increase the protein production from the artificial nucleic acid molecule through additive or synergistic effects, and the 5'UTR and the 3'UTR may be derived from the same highly expressed gene or from different highly expressed genes.

9. The artificial nucleic acid molecule according to any one of claims 1 to 8, wherein the 5'UTR does not comprise the start codon AUG or an open reading frame (ORF).

10. The artificial nucleic acid molecule according to claim 1, wherein the 5'UTR variant is a first 5'UTR deleted sequence obtained by deleting a sequence that is completely complementary to a microRNA seed region having a gene silencing function from the 5'UTR of the highly expressed gene or a sequence derived therefrom; and another 5'UTR deleted sequence obtained by further deleting a sequence that is completely complementary to a microRNA seed region having a gene silencing function from the first 5'UTR deleted sequence obtained. in, The microRNA seed region is predicted by at least one microRNA seed region prediction tool, including miRBase, RNA22, Traget Scan 8.0, miRcode, miRDB, microRNA.org, PicTar, and PITA, with a prediction score of >70% or all.

11. The artificial nucleic acid molecule according to claim 1, wherein the 5'UTR variant is obtained by deleting an element that inhibits downstream gene expression from the 5'UTR of the highly expressed gene or a derivative sequence thereof; in, The elements include, but are not limited to, a 5' terminal oligopyrimidine tract derived from the 5' UTR of the TOP gene, or a 5' UTR comprising at least one of AUG and an upstream open reading frame (uORF).

12. The artificial nucleic acid molecule according to claim 1, wherein the 5'UTR variant comprises at least one additional 5'UTR element to increase or / and prolong the production of the artificial nucleic acid molecule protein. in, The 5'UTR elements include, but are not limited to, sequence element A7 for adding adaptor recruitment to promote translation or capping proteins: ACTC ACT ATT TGT TTT CGC GCC CAG TTG CAA AAAGTG TCG, PABPs (PolyA-binding proteins) recognition motif, PCBPs (PolyC-binding proteins) recognition motif, YTHDFs recognition motif, IRES sequence, Kozak sequence: GCCACCAUGG, and / or conserved motifs of 5'UTR sequences of highly expressed genes in all mammals, and variants of the above elements; The conserved motifs of the 5'UTR sequences of the highly expressed genes in all mammals can be obtained by sequence alignment tools, including DNAman, MEGA, sequencelogo, online tool WebLogo3, and The MEMESuite (Motif-based sequence analysis tools); Variants of the elements include naturally occurring DNA sequences and homologs, variants, fragments and corresponding RNA sequences, RNA sequences having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity, in ascending order of priority, to the nucleic acid sequence, or fragments or variants thereof.

13. The artificial nucleic acid molecule according to claim 12, wherein the 5'UTR element has a low secondary structure free energy at its 5' end insertion site; the optimal position for the 5'UTR element to be inserted into the 5'UTR includes but is not limited to positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 upstream of the start codon AUG or at nucleotide positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 from the 5' end of the 5'UTR; in, The free energy is predicted based on RNA secondary structure prediction tools, including LinearDesign (https: / / rna.baidu.com / ), RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi), MC-Fold MC-Sym (https: / / major.iric.ca / MC-Pipeline / ), DNAman, RNAComposer (http: / / rnacomposer.ibch.poznan.pl / ), and ARES (http: / / drorlab.stanford.edu / ares.html).

14. The artificial nucleic acid molecule according to claim 1 and claims 10-13, wherein the 5'UTR variant is the 5'UTR of the highly expressed gene and its variants, which are sequence optimized using RNA sequence optimization software / program (such as LinearDesign (https: / / rna.baidu.com / )); in, The optimized sequence is a 5'UTR with a higher mean ribosome load (MRL) that can increase the expression of amino acids; The optimized 5'UTR has a sequence similarity of no less than 80% with the original 5'UTR.

15. The artificial nucleic acid molecule according to claims 10-14, wherein the 5'UTR variant is a combination of at least one and / or several variants and / or 5'UTR elements.

16. The artificial nucleic acid molecule according to claim 1, wherein the 3'UTR variant is a first 3'UTR deleted sequence obtained by deleting a sequence completely complementary to a microRNA seed region having a gene silencing function from the 3'UTR of the highly expressed gene; and another 3'UTR deleted sequence obtained by further deleting a sequence completely complementary to a microRNA seed region having a gene silencing function from the first 3'UTR deleted sequence obtained; in, The microRNA seed region is predicted by at least one microRNA seed region prediction tool, including miRBase, RNA22, Traget Scan 8.0, miRcode, miRDB, microRNA.org, PicTar, and PITA, with a prediction score of >70% or all.

17. The artificial nucleic acid molecule according to claim 1, wherein the 3'UTR variant is obtained by adding one or more 3'UTR elements to increase and / or prolong the production of the artificial nucleic acid molecule protein; in, The 3'UTR elements include but are not limited to the sequence element that recruits the RNA-binding protein HUR that enhances RNA stability: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; the cis-acting element R3U (Repeated sequence element 3 and U-rich element present in the 3'UTR: AAAACTCAATGTATTTCTGAGGAAGCGTGGTGCATAATGCCACGCAGCGTCTGCATAACTTTTATTATTTCTTTTATTAATCAACAAA; QRE1 (QKI response element 1): GCCGTAACCACGTCTACTAACGCCG; QRE2 (QKI response element 2): AACTCCAGGACTGTATTTGTGACTAATTGTATAACAGGTT; ARE (AU-rich ribosome binding site in the 3'UTR of the EMCV genome: ATAAGATACACCTGCAAAGGCGGCACAACCCCAGTGCCACGTTGTGAGTTGGATAGTTGTGGAAAGAGTCAAATGGCTCTCCTCAAGCGTATTCAACAAGGGGCTGAAGGATGCCCAGAAGGTACCCCATTGTATGGGATCTGATCTGGGGCCTCGGTGCACATGCTTTACATGTGTTTAGTCGAGGTTAAAAAACGTCTAGGCCCCCCGAACCACGGGGACGTGGTTTTCCTTTGAAAAACACGATGATAAT; human α-globin 1 (HBA1): GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA); Ironresponsive element (IRE): GCCATCAGAGCCAGTGTGTTTCTATGGT and its variants; Among them, the variants include: naturally occurring DNA sequences and their homologues, variants, fragments and corresponding RNA sequences, RNA sequences or their fragments or variants having at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity according to the nucleic acid sequence in ascending order of priority.

18. The artificial nucleic acid molecule according to claim 17, wherein the 3'UTR element is inserted into the 3'UTR at an optimal position including but not limited to after the nucleotide position 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 downstream of the 5' end or before the nucleotide position 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 upstream of the Poly A tail.

19. The artificial nucleic acid molecule according to claims 16-18, wherein the 3'UTR variant is at least one and / or a combination of several variants.

20. The artificial nucleic acid molecule according to claim 1, wherein the UTR variants are a database constructed based on the 5'UTR and 3'UTR of the highly expressed gene and their variants, and can be trained by artificial intelligence (AI) to obtain UTRs with higher mean ribosome load (MRL) and / or half-life, thereby increasing amino acid expression; in, The artificial intelligence learning algorithms include: genetic algorithm (Genetic Algorithm), deep convolutional neural network (Deep Convolutional Neural Networks) algorithm, transfer learning algorithm (TransferLearning), and regression algorithm (Regression).

21. The artificial nucleic acid molecule according to claim 20, wherein the at least one database consists of the following nucleic acid sequences: The group consisting of the nucleic acid sequences SEQ ID NO.9-43, wherein, when present, the symbols G, A, T, U, C, R, Y, M, K, S, W, H, B, V, D, N and * in the nucleic acid sequence are defined as follows: guanine; adenine; thymine; uracil; cytosine; guanine or adenine; thymine / uracil or cytosine; adenine or cytosine; guanine or thymine / uracil; guanine or cytosine; adenine or thymine / uracil; adenine or cytosine or thymine / uracil; guanine or thymine / uracil or cytosine; guanine or cytosine or adenine; guanine or adenine or thymine / uracil; guanine or cytosine or thymine / uracil or adenine; present or absent.

22. The artificial nucleic acid molecule according to claim 1, 15, or 19-21, wherein the at least one 5'UTR or 5'UTR variant and the at least one 3'UTR or 3'UTR variant increase protein production from the artificial nucleic acid molecule by additive or synergistic action.

23. The artificial nucleic acid molecule according to any one of claims 1 to 22, wherein the 5'UTR or 5'UTR variant thereof can be supplemented with one or more 3'UTR elements according to claim 19 to increase and / or prolong the production of the artificial nucleic acid molecule protein.

24. The artificial nucleic acid molecule according to claims 1-22, wherein the 3'UTR or 3'UTR variant thereof can be supplemented with one or more 5'UTR elements according to claim 12 to increase and / or prolong the production of the artificial nucleic acid molecule protein.

25. The artificial nucleic acid molecule according to claim 1, wherein the target coding region can be the original coding region sequence, or the original coding region sequence can be partially or completely codon-optimized; in, Compared with the original coding region sequence, the G / C content of the optimized sequence increased.

26. The artificial nucleic acid molecule according to claim 1, wherein the artificial nucleic acid molecule may be an unmodified original sequence or may be modified with at least one of the following nucleotides, including but not limited to N6-methyladenosine (m6A), N1-methyladenosine (m1A), 5-methylcytosine (m5C), 5-hydroxymethylcytosine (hm5C), pseudouridine (Ψ), inosine (I), uridine (U), and ribose methylation (2'-O-Me).

27. The artificial nucleic acid molecule according to claim 1, wherein the nucleic acid molecule further comprises: (d) target 5' cap structure (5'Cap); wherein, The 5'Cap is Cap0 or Cap1 or Cap2 or a cap structure analogue; (e) A target Poly A tail, wherein the Poly A tail consists of 20 to 200 poly (A) sequences, and the poly (A) sequences may contain a consensus sequence of GCATATGACT.

28. The artificial nucleic acid molecule according to claims 1-27, wherein the nucleic acid molecule comprises at least one selected from the group consisting of messenger RNA (mRNA), self-amplifying RNA (saRNA), circular RNA (circRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA) and microRNA (miRNA), primary miRNA, antisense oligonucleotide (ASO), transfer RNA (tRNA), plasmid DNA (pDNA), single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), deoxyribozymes (DNAzymes), ribozymes (RNAzymes), aptamers, clustered regularly interspaced short palindromic repeats (CRISPR)-related nucleic acids, single guide RNA (sgRNA), CRISPR RNA (crRNA), trans-activating crRNA (tracrRNA), guide RNA, long non-coding RNA (LncRNA), single-stranded RNA (ssRNA) and double-stranded RNA (dsRNA).

29. A transfectant comprising: The artificial nucleic acid molecule and vector according to any one of claims 1 to 28, wherein the vector is a DNA vector, a plasmid vector or a viral vector.

30. An artificial cell, comprising: The transfection molecule according to claim 29 and the host cell of the vector, wherein the vector is used to carry the artificial nucleic acid molecule to the host cell, wherein the host cell is a mammalian cell.

31. Use of the artificial nucleic acid molecule according to any one of claims 1 to 28 and / or the vector according to claim 29 and / or the artificial cell according to claim 30 in the preparation of a medical product, wherein: The medical supplies include RNA drugs and vaccines.

32. A method for delivering an artificial nucleic acid molecule to a cell, wherein: The artificial nucleic acid molecule is the artificial nucleic acid molecule according to any one of claims 1 to 28, and the artificial nucleic acid molecule is carried in the vector according to claim 29; the method comprises: contacting the cell with the artificial nucleic acid molecule and / or the vector in a condition where the artificial nucleic acid molecule and / or the vector are capable of being taken up into the cell; Wherein, the cells are mammalian cells.

33. Use of an artificial nucleic acid molecule according to any one of claims 1 to 29, characterized in that: The artificial nucleic acid molecule is administered to a subject in a pharmaceutically effective amount for preventing and / or treating a disease or disorder in a mammal.

34. Use of the artificial nucleic acid molecule according to any one of claims 1 to 28 and / or the vector according to claim 29 and / or the cell according to claim 30 in preparing a medicament and / or a pharmaceutical composition, wherein: The medicament is used for preventing and / or treating a subject, wherein the artificial nucleic acid molecule and / or the vector in the medicament and / or the pharmaceutical composition is for preventing or treating a disease or disorder.

35. Use of the artificial nucleic acid molecule according to any one of claims 1 to 28 and / or the vector according to claim 29 and / or the cell according to claim 30 in treating and / or preventing a disease or disorder, wherein: The diseases or conditions include, but are not limited to, immune system diseases, metabolic diseases, genetic diseases, cancer, blood diseases, bacterial infections or viral infections.

36. A kit comprising the artificial nucleic acid molecule according to any one of claims 1 to 28 and / or the vector according to claim 29 and / or the cell according to claim 30.

37. The kit according to claim 36, further comprising: Cells for transfection, adjuvants, devices for administering the pharmaceutical composition, pharmaceutical carriers and / or pharmaceutical devices for dissolving or diluting the artificial nucleic acid molecule, the vector or the pharmaceutical composition.

Citation Information

Patent Citations

  • Synthetic polynucleotides

    US3687808A

  • Solid-phase synthesis of polynucleotides

    US4373071A

  • Solid-phase synthesis of polynucleotides

    US4401796A

  • Phosphoramidite compounds and processes

    US4415732A

  • Process for preparing polynucleotides

    US4458066A

Cited By

  • MRNA translation enhancement tool and application thereof

    CN121203042A

  • UTR element NS1-G and construction method and application thereof

    CN121718550A

  • Animal model for verifying efficacy of small nucleic acid drug and construction method of animal model

    CN122214375A

  • Trimethylchitosan-hyaluronic acid-FMNL2 siRNA nano-drug and application thereof

    CN122272529A

  • A trimethyl chitosan-hyaluronic acid-FMNL2 siRNA nanomedicine and its application

    CN122272529B