Engineered DNA molecules to encode RNA
Engineered DNA molecules with defined poly(A) tail coding sequences address poly(A) tail instability in E. coli replication, ensuring stable mRNA production and regulated expression in eukaryotic cells.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2026-03-04
AI Technical Summary
Existing methods for preparing mRNA drugs in vitro face challenges with poly(A) tail instability during replication in E. coli, leading to deletion mutations and affecting the in vivo stability and biological activity of mRNA.
Engineered DNA molecules with specific poly(A) tail coding sequences, comprising elements a, b, c, and d, are designed to stabilize the poly(A) tail during replication and regulate RNA expression levels in eukaryotic cells.
The engineered DNA molecules ensure stable poly(A) tail preservation and enable controlled RNA expression, enhancing the production and stability of mRNA drugs.
Smart Images

Figure 2026507667000009 
Figure 2026507667000010 
Figure 2026507667000011
Abstract
Description
[Technical Field]
[0001] This application relates to the field of biotechnology, and in particular to RNAs containing poly(A) tails, which can be used to stabilize replication of DNA encoding the RNA in prokaryotic systems and to regulate the expression level of the RNA in eukaryotic cells. [Background technology]
[0002] The primary structure of a translatable mRNA drug molecule consists of a 5' cap structure, a 5' non-coding region (5' UTR), a coding region, a 3' non-coding region, and a polyadenosine tail (poly(A) tail). Known functions of the poly(A) tail include maintaining the in vivo stability of the mRNA molecule and participating in the initiation of protein translation. Protein translation initiation is achieved through the interaction of poly(A) tail-binding protein (PABP) with the translation initiation complex. In eukaryotic cells, the poly(A) tail is synthesized by post-transcriptional modification, typically through the action of poly(A) polymerase.
[0003] The first step in preparing mRNA drugs in vitro is to synthesize the mRNA drug by in vitro transcription (IVT) using a linearized plasmid containing the designed product sequence as a template. A poly(A) tail is typically added downstream of the 3′ UTR by co-transcriptional methods. To achieve this co-transcriptional addition of poly(A), the corresponding poly(dA:dT) sequence must be included in the template plasmid. However, poly(dA:dT) repeat sequences in the plasmid are unstable during replication in E. coli, and deletion mutations frequently occur in such sequences, resulting in shortened poly(dA:dT) sequences. This is unhelpful for the preparation of in vitro transcription template plasmids by large-scale fermentation, and poly(A) cleavage significantly affects the in vivo stability and biological activity of the mRNA. Summary of the Invention [Problem to be solved by the invention]
[0004] The present application provides a novel poly(A) tail for improving the preservation of poly(A) tails during in vitro preparation processes, and also provides a method for regulating the expression level of RNA in eukaryotic cells based on the poly(A) tail. [Means for solving the problem]
[0005] Specifically, a first aspect of the present application provides an engineered DNA molecule capable of replicating in a cell, comprising a polyadenosine tail (polyA tail) coding sequence, wherein the poly(A) tail coding sequence is a single element a, at least one element b, and at least one element c; a single element a, at least one element b, and at least one element d; or comprising a single element a, at least one element b, at least one element c, and at least one element d; In the poly(A) tail coding sequence, element a consists of multiple consecutive adenine (A) nucleotides, and the length of element a is in the range of 20 nt or more; Element b consists of multiple consecutive A nucleotides, and the length of element b is in the range of 3 nt or more and less than 20 nt; element c consists of one non-A nucleotide, the nucleotide being selected from T, C and G nucleotides; element d consists of any two or more consecutive nucleotides, the nucleotides are selected from A, T, C, and G nucleotides, the nucleotides at the 5' and 3' ends of element d are not A nucleotides, element d does not contain three or more consecutive A nucleotides, and the length of element d is in the range of 2 nt to 20 nt; Element a and element b are not adjacent, element c and element d are not adjacent, The poly(A) tail coding sequence does not include any two elements b that are adjacent to each other, does not include any two elements c that are adjacent to each other, and does not include any two elements d that are adjacent to each other.
[0006] In some embodiments, the poly(A) tail coding sequence further comprises a single element e consisting of one or two consecutive As located at the 3' end of the poly(A) tail coding sequence and adjacent to element d or element c.
[0007] In some embodiments, the poly(A) tail coding sequence does not include any elements other than element a, element b, element c, and element d.
[0008] In some embodiments, the poly(A) tail coding sequence does not include any elements other than element a, element b, element c, element d, and element e.
[0009] In some embodiments, the poly(A) tail coding sequence comprises at least two elements d. In some embodiments, the poly(A) tail coding sequence comprises at least two elements c. In some embodiments, the poly(A) tail coding sequence comprises at least one element d and one element c.
[0010] In some embodiments, the number of elements b is between 2 and 10, for example, 3, 4, 5, 6, 7, 8, or 9.
[0011] In some embodiments, the number of elements c is 0-10, for example 1, 2, 3, 4, 5, 6, 7, 8, or 9.
[0012] In some embodiments, the number of elements d is 0-5, for example 1, 2, 3, or 4.
[0013] In some embodiments, when elements c and d are present together, the total number of elements c and d is 2 to 15, for example 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14.
[0014] In some embodiments, the number of elements a is 1, the number of elements b is 3, the number of elements c is 2, and the number of elements d is 1.
[0015] In some embodiments, the number of elements a is 1, the number of elements b is 4, the number of elements c is 4, and the number of elements d is 1.
[0016] In some embodiments, the number of elements a is 1, the number of elements b is 5, the number of elements c is 4, and the number of elements d is 1.
[0017] In some embodiments, the number of elements a is 1, the number of elements b is 3, the number of elements c is 3, and the number of elements d is 1.
[0018] In some embodiments, the number of elements a is 1, the number of elements b is 3, the number of elements c is 2, and the number of elements d is 1.
[0019] In some embodiments, the length range of element a is 80 nt or less. In some embodiments, element a is 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, 50nt, 51nt, 52nt, 53nt, 54nt, 55nt, 56nt, 57nt, 58nt, 59nt, 60nt, 61nt, 62nt, 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, or 79nt.
[0020] In some embodiments, element b is 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, or 19 nt in length.
[0021] In some embodiments, the length of element d is in the range of 2 nt to 20 nt, 3 to 18 nt, 5 to 16 nt, 4 to 10 nt, or 6 to 12 nt, for example, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt, preferably 6 nt.
[0022] In some embodiments, the length of the poly(A) tail coding sequence is greater than 40 nt, e.g., 41 nt, 42 nt, 43 nt, 44 nt, 45 nt, 46 nt, 47 nt, 48 nt, 49 nt, 50 nt, 51 nt, 52 nt, 53 nt, 54 nt, 55 nt, 56 nt, 57 nt, 58 nt, 59 nt, 60 nt, 61 nt, 62 nt, 63 nt, 64 nt, 65 nt, 66 nt, 67 nt, 68 nt, 69 nt, 70 nt, 71 nt, 72 nt, 73 nt, 74 nt, 75 nt, 76 nt, 77 nt, 78 nt, 79 nt, 80 nt, 81 nt, 82 nt, 83 nt, 84 nt, 85 nt, 86 nt, 87 nt, 88 nt, 89 nt, 90 nt, 91 nt, 92 nt, 93 nt, 94 nt, 95 nt, 96 nt, 97 nt, 98 nt, 99 nt, 100 nt, 101 nt, 102 nt, 103 nt, 104 nt, 105 nt, 106 nt, 107 nt, 108 nt, 109 nt, 110 nt, 111 nt, 112 nt, 113 nt, 114 nt, 115 nt, 116 nt, 117 nt, 1 nt、83nt、84nt、85nt、86nt、87nt、88nt、89nt、90nt、91nt、92nt、93nt、94nt、95nt、96nt、97nt、98nt、99nt、100nt、101nt、102nt、103nt、104nt、105nt、 106nt, 107nt, 108nt, 109nt, 110nt, 111nt, 112nt, 113nt, 114nt, 115nt, 116nt, 117nt, 118nt, 119nt, 120nt, 121nt, 122nt, 123nt, 124nt, 125nt, 126nt , 127nt, 128nt, 129nt, 130nt, 131nt, 132nt, 133nt, 134nt, 135nt, 136nt, 137nt, 138nt, 139nt, 140nt, 141nt, 142nt, 143nt, 144nt, 145nt, 146nt, 147 nt, 148nt, 149nt, 150nt, 151nt, 152nt, 153nt, 154nt, 155nt, 156nt, 157nt, 158nt, 159nt, 160nt, 161nt, 162nt, 163nt, 164nt, 165nt, 166nt, 167nt, 16 8nt, 169nt, 170nt, 171nt, 172nt, 173nt, 174nt, 175nt, 176nt, 177nt, 178nt, 179nt, 180nt, 181nt, 182nt, 183nt, 184nt, 185nt, 186nt, 187nt, 188nt, 1 89nt、190nt、191nt、192nt、193nt、194nt、195nt、196nt、197nt、198nt、199nt、200nt、201nt、202nt、203nt、204nt、205nt、206nt、207nt、208nt、209nt、210nt, 211nt, 212nt, 213nt, 214nt, 215nt, 216nt, 217nt, 218nt, 219nt, 220nt, 221nt, 222nt, 223nt, 224nt, 225nt, 226nt, 227nt, 228nt, 229nt, 230n t、231nt、232nt、233nt、234nt、235nt、236nt、237nt、238nt、239nt、240nt、 241nt、242nt、243nt、244nt、245nt、246nt、247nt、248nt、249nt、250nt、251nt nt、252nt、253nt、254nt、255nt、256nt、257nt、258nt、259nt、260nt、261nt 、262nt、263nt、264nt、265nt、266nt、267nt、268nt、269nt、270nt、271nt、2 72nt, 273nt, 274nt, 275nt, 276nt, 277nt, 278nt, 279nt, 280nt, 281nt, 282nt, 283nt, 284nt, 285nt, 286nt, 287nt, 288nt, 289nt, 290nt, 291nt, 292nt 293nt, 294nt, 295nt, 296nt, 297nt, 298nt, 299nt, 300nt, 301nt, 302nt, 303nt, 304nt, 305nt, 306nt, 307nt, 308nt, 309nt, 310nt, 311nt, 312nt, 313nt t、314nt、315nt、316nt、317nt、318nt、319nt、320nt、321nt、322nt、323nt、 324nt、325nt、326nt、327nt、328nt、329nt、330nt、331nt、332nt、333nt、334nt nt、335nt、336nt、337nt、338nt、339nt、340nt、341nt、342nt、343nt、344nt 、345nt、346nt、347nt、348nt、349nt、350nt、351nt、352nt、353nt、354nt、3 55nt, 356nt, 357nt, 358nt, 359nt, 360nt, 361nt, 362nt, 363nt, 364nt, 365nt, 366nt, 367nt, 368nt, 369nt, 370nt, 371nt, 372nt, 373nt, 374nt, 375nt376nt, 377nt, 378nt, 379nt, 380nt, 381nt, 382nt, 383nt, 384nt, 385nt, 386nt, 387nt, 388nt, 389nt, 390nt, 391nt, 392nt, 393nt, 394nt, 395nt, 396nt, 397nt, 398nt, 399nt, or 400nt, etc.
[0023] In some embodiments, 50% or more of the polynucleotides of element a are located within the 5' or 3' portion of the poly(A) tail coding sequence. In some embodiments, 50% or more of the polynucleotides of element a are located within the 5' portion of the poly(A) tail coding sequence. In some embodiments, 50% or more of the polynucleotides of element a are located within the 3' portion of the poly(A) tail coding sequence. In some embodiments, the number of nucleotides of element a located in the 3' portion of the poly(A) tail coding sequence is equal to the number of nucleotides located in the 5' portion of the poly(A) tail coding sequence.
[0024] In some embodiments, element c is G, C, or T.
[0025] In some embodiments, element d comprises a palindromic sequence. Element d is a palindromic sequence. In some embodiments, element d comprises a sequence selected from the group consisting of GATATC (SEQ ID NO: 15), GTATAC (SEQ ID NO: 16), GAATCT (SEQ ID NO: 17), GCATATGACT (SEQ ID NO: 18), and GATATCGTATAC (SEQ ID NO: 19). In some embodiments, element d is a sequence selected from the group consisting of GATATC (SEQ ID NO: 15), GTATAC (SEQ ID NO: 16), GAATCT (SEQ ID NO: 17), GCATATGACT (SEQ ID NO: 18), and GATATCGTATAC (SEQ ID NO: 19). In some embodiments, element d comprises a polynucleotide sequence represented by SEQ ID NO: 15. In some embodiments, the polynucleotide sequence of element d is represented by SEQ ID NO: 15.
[0026] In some embodiments, the nucleotide at the 3' end of the poly(A) tail coding sequence is A. In some embodiments, the nucleotide at the 3' end of the poly(A) tail coding sequence is G. In some embodiments, the nucleotide at the 3' end of the poly(A) tail coding sequence is C. In some embodiments, the nucleotide at the 3' end of the poly(A) tail coding sequence is T.
[0027] In some embodiments, the 3' portion of the poly(A) tail coding sequence comprises one or more non-A nucleotides. In some embodiments, one-half of the poly(A) tail coding sequence near the 3' end comprises one or more non-A nucleotides. In some embodiments, one-third of the poly(A) tail coding sequence near the 3' end comprises one or more non-A nucleotides. In some embodiments, one-quarter of the poly(A) tail coding sequence near the 3' end comprises one or more non-A nucleotides.
[0028] In some embodiments, the structure of the poly(A) tail coding sequence is: element a - element c - element b - element c - element b - element c - element b - element c - element b, element b - element c - element b - element c - element a - element d - element b - element c - element b - element c - element b, element b - element c - element b - element c - element b - element d - element a - element c, element a-element d-element b-element c-element b-element c-element b, or Element b-element c-element b-element c-element b-element d-element a.
[0029] In some embodiments, the structure of the poly(A) tail coding sequence is: The sequence is element a-element c-element b-element c-element b-element c-element b-element c-element b, and a specific example is 60A-G-19A-G-19A-G-19A-G-3A.
[0030] In some embodiments, the structure of the poly(A) tail coding sequence is: 7A-C-18A-G-60A-GG-7A-C-18A-G-14A, 19A-G-19A-G-19A-element d-60A-G, 60A - element d-19A-G-19A-G-17A, 19A-G-19A-G-19A-Element d-60A, 19A-G-19A-G-19A-Element d-60A, 19A-C-19A-C-19A-Element d-60A, 19A-T-19A-T-19A-Element d-60A, 19A-G-19A-G-19A-Element d-60A, or 19A-G-19A-G-19A-element d-60A, Element d consists of 6 or 12 nucleotides.
[0031] In some embodiments, the poly(A) tail coding sequence is represented by any one of SEQ ID NOs: 1-10.
[0032] In some embodiments, the poly(A) tail coding sequence is represented by SEQ ID NO:3 or SEQ ID NO:4.
[0033] In some embodiments, the engineered DNA molecule is further linked to a gene fragment of interest at the 5' end of its poly(A) tail coding sequence, such that the gene fragment of interest and the poly(A) tail coding sequence jointly encode an RNA. In some embodiments, the engineered DNA molecule is further linked to a gene fragment of interest at the 5' end of its poly(A) tail coding sequence, such that the gene fragment of interest and the poly(A) tail coding sequence jointly encode an mRNA. In some embodiments, the gene fragment of interest comprises a protein coding sequence or a non-protein coding sequence, such as a functional RNA coding sequence. In some embodiments, the gene fragment of interest further comprises a 5' UTR coding sequence at the 5' end of the protein coding sequence or functional RNA coding sequence. In some embodiments, the gene fragment of interest further comprises a 3' UTR coding sequence at the 3' end of the protein coding sequence or functional RNA coding sequence. In some embodiments, the gene fragment of interest further comprises a 3'UTR coding sequence 3'-end of the protein-coding sequence or functional RNA-coding sequence and a 5'UTR coding sequence 5'-end of the protein-coding sequence or functional RNA-coding sequence. In some embodiments, the engineered DNA molecule further comprises a replicon, e.g., an origin of replication such as an ORI. In some embodiments, the engineered DNA molecule further comprises a marker gene to facilitate screening of cells containing the engineered DNA molecule, e.g., an antibiotic resistance gene, a fluorescent protein, or the like. In some embodiments, the DNA molecule further comprises a promoter that initiates transcription of the RNA jointly encoded by the gene fragment of interest and the poly(A) tail coding sequence. In some embodiments, the promoter is a prokaryotic promoter. In some embodiments, the promoter is a eukaryotic promoter. In some embodiments, the DNA molecule further comprises a replicon, e.g., an origin of replication, a promoter, a 5'UTR coding sequence, a protein-coding sequence, and a 3'UTR coding sequence.In some embodiments, the DNA molecule further comprises a replicon, e.g., an origin of replication, a resistance gene, a 5' UTR coding sequence, a protein coding sequence, and a 3' UTR coding sequence. In some embodiments, the DNA molecule further comprises a replicon, e.g., an origin of replication, a resistance gene, a promoter, a 5' UTR coding sequence, a protein coding sequence, and a 3' UTR coding sequence. In some embodiments, the protein coding sequence encodes an HPV viral antigen protein. In some embodiments, the HPV protein is derived from HPV types 16 and / or 18. In some embodiments, the protein coding sequence encodes an HPV E2, E6, or E7 protein. In some embodiments, the protein coding sequence encodes a fusion protein of HPV E6 and E7 proteins. In some embodiments, the protein coding sequence encodes a fusion protein of HPV E2, E6, and E7 proteins. In some embodiments, the protein coding sequence encodes a fusion protein of HPV E2, E6, and E7 proteins. In some embodiments, the polypeptide fragment of the fusion protein is derived from HPV types 16 and / or 18. In some embodiments, the polypeptide fragments of the fusion protein are derived from the E2, E6, and E7 proteins of HPV types 16 and / or 18. In some embodiments, the protein coding sequence encodes a polypeptide set forth in SEQ ID NO: 26 or a conservatively substituted variant thereof.
[0034] In some embodiments, the engineered DNA molecule comprises a polynucleotide sequence set forth in any one of SEQ ID NOs: 22-25 or a synonymous variant thereof, or a polynucleotide sequence having 85% or greater sequence identity to a polynucleotide sequence set forth in any one of SEQ ID NOs: 22-25 or a synonymous variant thereof.
[0035] In some embodiments, the DNA molecule is a DNA plasmid. In some embodiments, the DNA molecule is a linear or circular plasmid. In some embodiments, the DNA molecule is single-stranded or double-stranded. In some embodiments, the plasmid is a pUC-, pTZ-, pMB1-, or pCoIE1-based plasmid. In some embodiments, the plasmid is a pUC57 vector-based plasmid.
[0036] In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a recA bacterium. In some embodiments, the cell is Escherichia coli. In some embodiments, the E. coli is selected from the group consisting of K-12 strains and their derivatives, and B strains and their derivatives. In some embodiments, the E. coli is selected from the group consisting of MG1655, DH5 or DH5α, DH10B, BL21, DB3.1, HB101, JM109, JM110, MC1061, MG1655, Pir1, Stbl2, Stbl3, Top10, XL1 Blue, XL10 Gold, BLR, HMS174, Tuner, Rostetta2, Lemo21, T7Express, and Origami2.
[0037] A second aspect of the present application discloses a cell comprising the DNA molecule of the first aspect. In some embodiments, the DNA molecule of the first aspect can be replicated and / or transcribed in the cell. In some embodiments, the cell is a prokaryotic cell, and the DNA molecule of the first aspect can be replicated in the prokaryotic cell. In some embodiments, the cell is a recA bacterium. In some embodiments, the cell is E. coli. In some embodiments, the prokaryotic cell is a competent cell. In some embodiments, the prokaryotic cell is an engineered cell. In some embodiments, the prokaryotic cell is an engineered prokaryotic cell. In some embodiments, the cell is E. coli, and the E. coli is selected from the group consisting of K-12 strains and derivatives thereof, and B strains and derivatives thereof. In some embodiments, the E. coli is selected from the group consisting of MG1655, DH5 or DH5α, DH10B, BL21, DB3.1, HB101, JM109, JM110, MC1061, MG1655, Pir1, Stbl2, Stbl3, Top10, XL1 Blue, XL10 Gold, BLR, HMS174, Tuner, Rostetta2, Lemo21, T7 Express, and Origami2. In some embodiments, the cell is a eukaryotic cell, and the DNA molecule of the first aspect described above can be transcribed in the eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the eukaryotic cell is selected from the group consisting of yeast or mold.
[0038] In a third aspect of the present application, there is provided a poly(A) tail, the poly(A) tail comprising: (1) obtained by transcription of an engineered DNA molecule according to the first aspect; (2) has a polynucleotide sequence identical to the poly(A) tail obtained by transcription of the engineered DNA molecule of the first aspect and is obtained by chemical synthesis; or (3) Obtained by further modifying the poly(A) tail in item (1) or (2) above.
[0039] In some embodiments, the further modification comprises substituting one or more ribonucleotides in the poly(A) tail obtained by item (1) or (2) above with one or more deoxyribonucleotides. In some embodiments, one or more ribonucleotides are substituted with corresponding deoxyribonucleotides, for example, one or more ribonucleotides A in the poly(A) tail are substituted with deoxyribonucleotides A, one or more ribonucleotides U are substituted with deoxyribonucleotides T, one or more ribonucleotides C are substituted with deoxyribonucleotides C, one or more ribonucleotides G are substituted with deoxyribonucleotides G, or one or more ribonucleotides G are substituted with ribonucleotide I (inosine) or deoxyribonucleotide I. In some embodiments, the modification is a chemical modification. In some embodiments, the modification is base editing. In some embodiments, the modification is deamination of one or more ribonucleotides in the poly(A) tail obtained by item (1) or (2) above.
[0040] In some embodiments, the poly(A) tail comprises a polynucleotide sequence selected from any one of SEQ ID NOs: 1-10. In some embodiments, the polynucleotide sequence of the poly(A) tail is represented by any one of SEQ ID NOs: 1-10.
[0041] The present application also provides uses of the aforementioned poly(A) tail. In some embodiments, the use of a poly(A) tail to increase the stability of an RNA molecule is provided, where the poly(A) tail is located at the 3' end of the RNA, and "increasing stability" refers to increasing the stability outside a cell, inside a cell, or in an animal compared to an RNA molecule containing another poly(A) tail. In some embodiments, the use of a poly(A) tail to decrease the stability of an RNA molecule is provided, where the poly(A) tail is located at the 3' end of the RNA, and "decreasing stability" refers to decreasing the stability outside a cell, inside a cell, or in an animal compared to an RNA molecule containing another poly(A) tail. In some embodiments, the use of a poly(A) tail to increase the expression level of an RNA molecule within the same period of time is provided, where "increasing expression level" refers to increasing the expression level outside a cell, inside a cell, or in an animal compared to an RNA molecule containing another poly(A) tail. In some embodiments, the use of a poly(A) tail to decrease the expression level of an RNA molecule within the same period of time is provided, where "decreasing expression level" refers to decreasing the expression level outside a cell, inside a cell, or in an animal compared to an RNA molecule containing another poly(A) tail. In some embodiments, there is provided the use of a poly(A) tail to extend the expression time of an RNA molecule, where extending expression time refers to extending the expression time extracellularly, intracellularly, or in an animal compared to an RNA molecule comprising a different poly(A) tail. In some embodiments, there is provided the use of a poly(A) tail to shorten the expression time of an RNA molecule, where shortening expression time refers to shortening the expression time extracellularly, intracellularly, or in an animal compared to an RNA molecule comprising a different poly(A) tail. In some embodiments, there is provided the use of a poly(A) tail to extend the half-life of an RNA molecule, where extending half-life refers to extending the half-life extracellularly, intracellularly, or in an animal compared to an RNA molecule comprising a different poly(A) tail. In some embodiments, there is provided the use of a poly(A) tail to shorten the half-life of an RNA molecule, where shortening half-life refers to shortening the half-life extracellularly, intracellularly, or in an animal compared to an RNA molecule comprising a different poly(A) tail.
[0042] In some embodiments, the RNA molecule is an mRNA molecule. In some embodiments, the poly(A) tail and the additional poly(A) tail are two different poly(A) tails belonging to the poly(A) tails according to the third aspect of the present application. In some embodiments, the poly(A) tail is a poly(A) tail according to the third aspect of the present application, and the additional poly(A) tail is a poly(A) tail other than a poly(A) tail according to the third aspect of the present application. In some embodiments, "inside a cell" refers to a host cell, and the host cell is a eukaryotic cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell.
[0043] The present application also provides use of a DNA fragment or DNA-RNA hybrid molecule fragment encoding a poly(A) tail according to the third aspect of the present application, and a DNA fragment or DNA-RNA hybrid molecule fragment encoding an RNA, for more conservative replication in a host cell. In the use, the DNA fragment or DNA-RNA hybrid molecule fragment encoding a poly(A) tail according to the third aspect of the present application is located 3' to the RNA-coding sequence in the DNA molecule or DNA-RNA hybrid molecule. In some embodiments, the host cell is a prokaryotic cell. In some embodiments, the host cell is a recA bacterium. In some embodiments, the host cell is Escherichia coli. In some embodiments, the E. coli is selected from the group consisting of K-12 strains and their derivatives, and B strains and their derivatives. In some embodiments, the E. coli is selected from the group consisting of MG1655, DH5 or DH5α, DH10B, BL21, DB3.1, HB101, JM109, JM110, MC1061, MG1655, Pir1, Stbl2, Stbl3, Top10, XL1 Blue, XL10 Gold, BLR, HMS174, Tuner, Rostetta2, Lemo21, T7 Express, and Origami2.
[0044] A fourth aspect of the present application also provides an RNA molecule comprising a poly(A) tail according to the third aspect. In some embodiments, the RNA molecule is an mRNA molecule. In some embodiments, the RNA molecule is (1) obtained by transcription of an engineered DNA molecule according to the first aspect; (2) has the same polynucleotide sequence as the RNA molecule of item (1) above and is obtained by chemical synthesis; or (3) It is obtained by further modifying the RNA molecule in the above item (1) or (2).
[0045] In some embodiments, the further modification comprises substituting one or more ribonucleotides in the RNA molecule obtained by (1) or (2) above with one or more deoxyribonucleotides. In some embodiments, one or more ribonucleotides are substituted with corresponding deoxyribonucleotides, for example, one or more ribonucleotides A in the RNA molecule are substituted with deoxyribonucleotides A, one or more ribonucleotides U are substituted with deoxyribonucleotides T, one or more ribonucleotides C are substituted with deoxyribonucleotides C, one or more ribonucleotides G are substituted with deoxyribonucleotides G, or one or more ribonucleotides G are substituted with ribonucleotide I (inosine) or deoxyribonucleotide I. In some embodiments, the modification is a chemical modification. In some embodiments, the modification is base editing. In some embodiments, the modification is deamination of one or more ribonucleotides in the RNA molecule obtained by (1) or (2) above. In some embodiments, the further modification is a post-transcriptional modification. In some embodiments, the further modification comprises capping. In some embodiments, the further modification comprises splicing. In some embodiments, further modifications include splicing and capping processes.
[0046] In some embodiments, the RNA molecule comprises coding RNA or non-coding RNA (ncRNA). In some embodiments, the RNA molecule is a pre-mRNA. In some embodiments, the RNA is a mature mRNA. In some embodiments, the RNA molecule is a long non-coding RNA (lncRNA). In some embodiments, the RNA molecule further comprises a 5' cap structure. In some embodiments, the polynucleotide sequence of the RNA molecule is represented by any one of SEQ ID NOs: 22-25. In some embodiments, the RNA molecule comprises a polynucleotide sequence represented by any one of SEQ ID NOs: 22-25.
[0047] The present application also provides a DNA and RNA hybrid molecule that carries the same genetic information as the engineered DNA molecule of the first aspect, the same genetic information as the poly(A) tail of the third aspect, or the same genetic information as the RNA molecule of the fourth aspect.
[0048] The present application also provides a nucleic acid molecule library, in some embodiments, the nucleic acid molecule library comprises engineered DNA molecules according to the first aspect, poly(A) tails according to the third aspect, DNA fragments or DNA-RNA hybrid molecule fragments encoding the poly(A) tails according to the third aspect of the present application, or RNA molecules according to the fourth aspect.
[0049] The present application also provides a method for modulating protein expression, comprising introducing into a cell of interest a plurality of nucleic acid molecules in the aforementioned library of nucleic acid molecules at different times and / or in different amounts. In some embodiments, the nucleic acid molecules are engineered DNA molecules according to the first aspect above. In some embodiments, the nucleic acid molecules are RNA molecules according to the fourth aspect above.
[0050] It should be understood that the aspects and embodiments of the present application described herein include aspects and embodiments that "comprise," "consist," or "consist essentially of" them. Although preferred embodiments of the present application have been described in detail above, the present application is not limited thereto. Within the technical spirit of the present application, various simple modifications can be made to the technical solutions of the present application, including combining various technical features in any other suitable manner. These simple modifications and combinations should also be considered as the contents disclosed in the present application and fall within the protection scope of the present application. [Brief explanation of the drawings]
[0051] [Figure 1] 1 shows the replication stability of the 10 poly(A) fragments in this application in E. coli DH5α in Example 2. [Figure 2] 1 shows the deletion statistics of 10 poly(A) bases in the present application in E. coli DH5α in Example 2. [Figure 3] 1 shows the replication stability of poly(A) P1, P2, P3, P4 and poly(A) controls C1 and C2 in Escherichia coli DH5α under the plasmid system containing the HPV antigen sequence in Example 3. [Figure 4] 1 shows base deletion statistics for poly(A) P3, P4 and poly(A) controls C1 and C2 in Escherichia coli DH5α under the plasmid system containing the HPV antigen sequence in Example 3. [Figure 5] 1 shows a comparison of the replication stability of poly(A) P1, P2, P3 and poly(A) control C1 in Escherichia coli DH5α under two temperature conditions of 30°C and 37°C in the plasmid system containing the HPV antigen sequence in Example 3. [Figure 6] 1 shows an example of a universal vector plasmid DNA profile in an example. [Figure 7] 1 shows the animal imaging results of the expression levels of luciferase with poly(A) P3, P4, P5, P8, P9 and control C2 in mice in Example 4. [Figure 8]This shows the quantitative results of fluorescence intensity after animal imaging of the expression levels of luciferase with poly(A) P3, P4, P5, P8, P9 and control C2 in mice in Example 4 (ns: no significant difference, ★★: significant difference, p<0.01). DETAILED DESCRIPTION OF THE INVENTION
[0052] This application first provides a method for stably amplifying poly(A)-tailed transcription template DNA in vitro, reducing the mutation frequency of the poly(A)-tailed transcription template sequence during large-scale DNA replication in cells. This allows for the production of large amounts of RNA based on DNA, including a poly(A) tail with a defined sequence. Based on this, RNA with a poly(A) tail engineered to have a specific function, such as mRNA, can be produced on a large scale by in vitro fermentation.
[0053] The present application also provides DNA containing a poly(A)-tailed transcription template that can be stably amplified in vitro, and RNA transcribed from the DNA. Furthermore, the present application further provides a group of poly(A) tails that have different regulatory effects on RNA stability and / or expression efficiency, RNA containing a poly(A) tail, DNA containing a poly(A) tail-encoding sequence, and a library consisting of poly(A) tails, RNA, or DNA, provided that stable amplification in vitro is achieved.
[0054] Additionally, the present application provides uses of the aforementioned poly(A) tails, RNA, DNA, and libraries.
[0055] term As used herein, "element a," "element b," "element c," "element d," and "element e" refer to types of elements contained in poly(A). Element a consists of multiple consecutive adenine (A) nucleotides, and the length of element a ranges from 20 nt or more. Element b consists of multiple consecutive A nucleotides, and the length of element b ranges from 3 to less than 20. Element c consists of one non-A nucleotide, and the nucleotide is selected from T, C, and G nucleotides. Element d consists of any two or more consecutive nucleotides, and the nucleotide is selected from A, T, C, and G nucleotides. The nucleotides at the 5' and 3' ends of element d are not A nucleotides, and element d does not contain three or more consecutive A nucleotides, and the length of element d ranges from 2 nt to 20 nt. Element e consists of one or two consecutive A nucleotides, and is located at the 3' end of the poly(A) tail coding sequence and is adjacent to element d or element c, if present. When poly(A) contains two or more "elements b," "elements c," and "elements d," the sequences of each two elements b may be the same or different, the sequences of each two elements c may be the same or different, and the sequences of each two elements d may be the same or different, so long as they satisfy the definitions of elements a, b, c, and d above. In this application, "element a," "element b," "element c," "element d," "element e," etc. in the poly(A) tail may be referred to by the term "element."
[0056] As used herein, when the positional relationship between two or more elements is described as "non-adjacent," this means that the two or more elements are not adjacent to each other. In other words, the two or more elements contain at least one or more nucleotides or bases other than the nucleotides of the two elements between each two elements.
[0057] As used herein, "encode" refers to i) a DNA sequence containing genetic information that can be transcribed into an RNA molecule, and / or ii) an RNA molecule containing genetic information that can be translated into an amino acid sequence. Therefore, as used herein, a "coding sequence" refers to a ribonucleotide (RNA) sequence or a fragment thereof in a pre-mRNA or mature mRNA that can be translated into a protein, and also refers to a complementary sequence of a deoxyribonucleotide (DNA) sequence or a fragment thereof that serves as a template for transcribing the pre-mRNA or mature mRNA. Furthermore, the "coding sequence" of the present application may further include polynucleotide sequences that encode proteins, functional nucleic acids, or fragments thereof, such as miRNA, shRNA, dsRNA, guide RNA, poly(A) tail, 5' UTR, and 3' UTR. Among these, a DNA molecule containing genetic information that can be transcribed into an RNA molecule is referred to as the "encoding nucleic acid" of the RNA molecule, and an RNA molecule containing genetic information that can be translated into an amino acid sequence is referred to as the "encoding nucleic acid" of the amino acid sequence.
[0058] In this application, nucleotides in all polynucleotide sequences are numbered from the 5' end to the 3' end, i.e., the nucleotide at the 5' end is the first nucleotide and the nucleotide at the 3' end is the last nucleotide. Unless otherwise specified, the terms "5' end" and "5' end" can be used interchangeably, and "3' end" and "3' end" can also be used interchangeably. The terms "5' end" and "3' end" focus on describing the relative positional relationship between nucleotides, nucleotide sequence segments, or nucleotides and nucleotide sequence segments within the same nucleic acid sequence. "5' end" and "3' end" are used to describe the positions of the first and last nucleotides or segments of a nucleic acid sequence, respectively. "5' end" is used to describe the relative positional relationship between two non-overlapping sequences within the same polynucleotide sequence. When a sequence is described as being located at the 5' end of another sequence, this means that the sequence is closer to the "5' end" of the polynucleotide sequence than the other sequence. Similarly, when a sequence is described as being located at the 3' end of another sequence, this means that the sequence is closer to the "3' end" of the polynucleotide sequence than the other sequence and does not include any overlapping portions of the other sequence. Specifically, for example, "the DNA coding sequence for the poly(A) tail is located at the 3' end of the RNA coding sequence" means that the DNA coding sequence for the poly(A) tail, which is a constituent element of the RNA coding sequence, includes nucleotides at the 3' end of the RNA coding sequence. Furthermore, as used herein, the "5' portion" refers to the portion of the polynucleotide sequence that is approximately halfway to the 5' end, with the "center position" of the polynucleotide sequence as the boundary. The "3' portion" refers to the portion of the polynucleotide sequence that is approximately halfway to the 3' end, with the "center position" of the polynucleotide sequence as the boundary. The number of nucleotides from the "center position" to the 5' end, as described herein, is equal to the number of nucleotides from the "center position" to the 3' end, as described herein.
[0059] As used herein, when referring to the replication of nucleic acid molecules, the term "conservative" means a low probability of mutation during the replication process. In this context, "conservative" is a relative concept. For example, when it is stated that "a DNA coding sequence in a poly(A) tail is used to make the replication of a DNA molecule encoding an RNA more conservative in a host cell," this refers to the probability that a parent DNA molecule encoding an RNA will be replicated into a child DNA molecule. When an RNA molecule contains a poly(A) tail, the child DNA molecule has a higher probability of showing 100% sequence identity with the parent DNA molecule compared to the coding DNA of an RNA molecule that does not contain a poly(A) tail (e.g., an RNA molecule that contains several other poly(A) tails). This increases the average sequence identity between the parent DNA molecule and multiple child DNA molecules obtained by replication of the parent DNA molecule.
[0060] In the present application, when referring to "modulating" the expression of an RNA molecule, "modulating" means increasing or decreasing the total amount of protein or functional RNA expressed by the RNA molecule within the same period of time, or allowing the RNA to express a protein or functional RNA within a longer or shorter period of time, where the increase or decrease, or longer or shorter period of time, is compared to another RNA molecule expressing the same protein or functional RNA. When referring to "modulating" protein expression, this means modulating the expression of an RNA molecule containing a protein-coding sequence. The modulating effect described herein can be achieved by linking the poly(A) tail of the present application to the 3' end of an RNA molecule that does not contain a poly(A) tail, or by replacing the RNA's native poly(A) tail with the poly(A) tail of the present application.
[0061] As used herein, the percentage of "identity" (e.g., 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, or 99.5% identity) refers to the degree of similarity between amino acid sequences or nucleotide sequences as determined by sequence alignment, and is 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, or 99.5%. For example, by introducing gaps, the two sequences are made to have identical residues at as many positions as possible, and then the ratio of the number of positions occupied by identical bases or amino acid residues to the total number of positions is determined. The percentage of "identity" can be determined using software programs known in the art. Preferably, alignment is performed using default parameters. A preferred alignment program is BLAST. Preferred programs are BLASTN and BLASTP. Details of these programs are available at ncbi.nlm.nih.gov / cgi-bin / BLAST.
[0062] As used herein, "complementarity" of a nucleic acid refers to the ability of one nucleic acid to form hydrogen bonds with another nucleic acid through conventional Watson-Crick base pairing. The percentage of complementarity refers to the proportion of residues in another nucleic acid molecule that can form hydrogen bonds (i.e., Watson-Crick base pairing) with one nucleic acid molecule (e.g., about 5, 6, 7, 8, 9, and 10 out of 10 correspond to about 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Fully complementary" means that all contiguous residues of a nucleic acid sequence form hydrogen bonds with the same number of contiguous residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to two nucleic acids that hybridize under stringent conditions with at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% complementarity over a region of about 40, 50, 60, 70, 80, 100, 150, 200, 250, or more nucleotides. According to the Watson-Crick base-pairing principle, a single base or nucleotide is said to be complementary or matching if A pairs with T or U and C pairs with G or I, and vice versa. Any other base pairing is said to be non-complementary. In this application, a "complementary polynucleotide sequence" of a given polynucleotide sequence refers to a polynucleotide sequence that is completely complementary to that polynucleotide sequence.
[0063] As used herein, a "conservative substitution variant" of a protein, polypeptide, or amino acid sequence refers to one or more amino acid substitutions made without altering the overall form and function of the protein or enzyme. This includes, but is not limited to, the substitution of amino acids in the amino acid sequence of a parent protein, as described above under "conservative substitution." Therefore, the degree of similarity between two proteins or amino acid sequences with similar functions may vary. For example, the similarity (identity) may be between 70 and 99% based on the MEGALIGN algorithm. "Conservative substitution variants" also include polypeptides or enzymes with amino acid identity of 60% or more, preferably 75% or more, more preferably 85% or more, and most preferably 90% or more, as determined by the BLAST or FASTA algorithm, and that have the same or substantially similar properties or functions as the native or parent protein or enzyme.
[0064] In the context of this application, the terms "DNA" and "RNA" refer to single-stranded or double-stranded DNA or RNA molecules. Unless otherwise indicated, the terms "DNA" and "DNA molecule" refer to double-stranded DNA molecules composed of A, C, G, and / or T nucleotides, and the terms "RNA" and "RNA molecule" refer to single-stranded RNA molecules composed of A, C, G, and / or U nucleotides. As used herein, A, C, G, T, and U nucleotides refer to nucleotides containing adenine, guanine, cytosine, thymine, and uracil as distinct nitrogenous bases.
[0065] RNA molecules include coding RNA or non-coding RNA (ncRNA) such as pre-mRNA, mature mRNA, or long non-coding RNA (lncRNA).
[0066] As used herein, a "DNA and RNA hybrid molecule" is a molecule that contains a polynucleotide sequence composed of deoxyribonucleotides and ribonucleotides. Substituting one or more deoxyribonucleotides in the DNA with ribonucleotides, or Substituting one or more ribonucleotides in the RNA with deoxyribonucleotides, or They can be obtained by a method of newly synthesizing them by biological or chemical synthesis using deoxyribonucleotides and ribonucleotides as raw materials. Note that the method for obtaining a DNA-RNA hybrid molecule is not limited to the above methods, and a DNA-RNA hybrid molecule obtained by any method also falls within the category of "DNA-RNA hybrid molecule" as defined in this application.
[0067] As used herein, when two nucleic acid molecules are described as having the same genetic information, this means that the two nucleic acid molecules are complementary, contain the exact same base sequence, or that a nucleic acid molecule having the exact same base sequence as another nucleic acid molecule can be obtained by converting one or more thymines in the base sequence of one nucleic acid molecule to uracil. Thus, any two DNA, RNA, and DNA-RNA hybrid molecules can also have the same genetic information. As used herein, the term "base sequence" refers to the order of the arrangement of bases in a polynucleotide molecule. Unless otherwise specified, those skilled in the art should understand that when a base sequence or polynucleotide sequence described in this application is used to describe a DNA sequence, thymine may be represented by "T," while when a base sequence or polynucleotide sequence is used to describe RNA (such as mRNA), "T" is replaced by "U" (uracil). Thus, whenever a DNA is disclosed herein by a particular SEQ ID NO, the complementary or corresponding RNA (e.g., mRNA or poly(A) tail) sequence to that DNA is also disclosed, in which each "T" in that DNA sequence is replaced with a "U."
[0068] Poly(A) tails and uses thereof As used herein, the term "polyA tail" or "poly(A) sequence" refers to an uninterrupted or uninterrupted sequence of adenylate residues, typically located at the 3' end of an RNA molecule. If a 3'-UTR is present in the RNA, the polyA sequence is linked to the 3' end of the 3'-UTR. An uninterrupted polyA tail is characterized by consecutive adenylate residues. The polyA tail can be any length. In some embodiments, the polyA tail comprises or consists of at least 20, at least 30, at least 40, at least 80, or at least 100, and up to 500, up to 400, up to 300, up to 200, or up to 150 adenylate nucleotides (A), particularly about 120 A. Typically, the majority of the nucleotides in the poly-A tail are adenosines, where majority refers to at least 75%, at least 80%, at least 85%, at least 90%, etc. of the nucleotides, although the remaining nucleotides may be nucleotides other than A (non-A nucleotides), such as U (uridylic acid), guanine G (nylic acid), or C (cytidylic acid).
[0069] In some embodiments, the in vitro process for preparing RNA is a prokaryotic fermentation process, i.e., introducing a coding nucleic acid of an RNA molecule containing a poly(A) tail into a prokaryotic cell, amplifying the prokaryotic cell to achieve the goal of amplifying the coding nucleic acid, and then transcribing the amplified coding nucleic acid into RNA. In some embodiments, the in vitro process for preparing RNA is ligating an RNA fragment containing a protein-coding sequence to a poly(A) tail by homologous recombination, enzymatic digestion and ligation, or other non-homologous recombination methods, and the poly(A) tail is prepared by a prokaryotic fermentation process. During the prokaryotic fermentation process, the coding nucleic acid containing a poly(A) tail is introduced into a prokaryotic cell, and the prokaryotic cell is amplifying to achieve the goal of amplifying the coding nucleic acid. The amplified coding nucleic acid is then transcribed into RNA containing a poly(A) tail. In some embodiments, the coding nucleic acid is linear. In some embodiments, the coding nucleic acid is circular. In some embodiments, the coding nucleic acid is a plasmid. In some embodiments, the coding nucleic acid is single-stranded or double-stranded. In some embodiments, the encoding nucleic acid is chemically modified before being introduced into the prokaryotic cell. In some embodiments, the encoding nucleic acid is chemically synthesized before being introduced into the prokaryotic cell. In some embodiments, the encoding nucleic acid is inserted into the nucleoid / nucleoid genomic DNA of the prokaryotic cell. In some embodiments, the encoding nucleic acid is released into the cytoplasm of the prokaryotic cell or into the nucleoid / nucleoid in vitro. In some embodiments, the prokaryotic cell is E. coli.
[0070] Based on this, the present application provides a series of poly(A) tails that are highly conserved during in vitro preparation of RNA, the poly(A) tails containing one or more non-A nucleotides at one or more positions.
[0071] In some embodiments, the poly(A) tail coding sequence is contains a single element a, at least one element b, and at least one element c, or a single element a, at least one element b, and at least one element d, or a single element a, at least one element b, at least one element c, and at least one element d; Element a and element b are not adjacent, element c and element d are not adjacent, The poly(A) tail coding sequence does not include any two elements b that are adjacent to each other, does not include any two elements c that are adjacent to each other, and does not include any two elements d that are adjacent to each other.
[0072] In some embodiments, the poly(A) tail coding sequence further comprises a single element e consisting of one or two consecutive As located at the 3' end of the poly(A) tail coding sequence and adjacent to element d or element c.
[0073] The poly(A) tail herein can be a segment of RNA or a hybrid molecule of DNA and RNA.
[0074] The present application also provides poly(A) tails that can regulate protein expression levels. The present application also provides poly(A) tails that can regulate protein expression levels while maintaining a high degree of preservation during an in vitro preparation process. In some embodiments, the poly(A) tails that can regulate protein expression levels, or the poly(A) tails that can regulate protein expression levels while maintaining a high degree of preservation during an in vitro preparation process, are: element a - element c - element b - element c - element b - element c - element b - element c - element b, element b - element c - element b - element c - element a - element d - element b - element c - element b - element c - element b, element b - element c - element b - element c - element b - element d - element a - element c, Element a - element d - element b - element c - element b - element c - element b, Element b-element c-element b-element c-element b-element d-element a.
[0075] In some embodiments, the poly(A) tail capable of modulating protein expression levels, or capable of modulating protein expression levels while maintaining a high degree of conservation during the in vitro preparation process, is 60A-G-19A-G-19A-G-19A-G-3A and 7A-C-18A-G-60A-element d-7A-C-18A-G-14A and 60A-element d-19A-G-19A-G-17A, 19A-G-19A-G-19A-Element d-60A and 19A-G-19A-G-19A-Element d-60A-G and 19A-G-19A-G-19A-Element d-60A and 19A-C-19A-C-19A-Element d-60A and 19A-T-19A-T-19A-element d-60A.
[0076] In some embodiments, the poly(A) tail capable of modulating protein expression levels, or capable of modulating protein expression levels while maintaining a high degree of conservation during the in vitro preparation process, is 60A-element d-19A-G-19A-G-17A or 19A-G-19A-G-19A-element d-60A.
[0077] In particular, the two elements linked by "-" are directly linked, with no nucleotides between the two elements.
[0078] In the above poly(A) tail structure, "yA" represents the number of consecutive As in element a or element b, where y is a natural number. For example, 19A means that 19 consecutive As are contained, and 60A means that 60 consecutive As are contained.
[0079] In some embodiments, the poly(A) tail capable of modulating protein expression levels, or the poly(A) tail capable of modulating protein expression levels while maintaining a high degree of conservation during the in vitro preparation process, is selected from any one of the polynucleotide sequences represented by SEQ ID NOs: 1-10.
[0080] The present application also provides a use of the poly(A) tail described above for regulating protein expression. In the use, the poly(A) tail is located at the 3' end of an mRNA, for example, at the 3' end of the 3' UTR. In some embodiments, the protein expression is regulated using the methods for regulating protein expression described below.
[0081] Engineered DNA molecules and libraries The present application also provides an engineered DNA molecule capable of replicating within a cell. This engineered DNA molecule comprises a coding sequence for the aforementioned polyA tail or its complementary sequence. Those skilled in the art should understand that, in addition to the coding sequence for the polyA tail, the engineered DNA molecule should also comprise structural elements necessary for the DNA molecule to replicate or efficiently replicate within a cell. Structural elements necessary for the DNA molecule to replicate or efficiently replicate within a cell are known in the art and include, for example, an origin of replication (ORI). In some embodiments, the engineered DNA molecule further comprises a marker gene or fragment thereof, and / or a reporter gene or fragment thereof, and a unique restriction endonuclease site that allows insertion of a DNA element, preferably a restriction endonuclease site that functions as a multiple cloning site (MCS). The marker gene facilitates identification of cells containing a plasmid containing a marker gene, which may be selected from, for example, an antibiotic resistance gene. Each restriction endonuclease site within the MCS can be specifically recognized by a different restriction endonuclease.
[0082] In some embodiments, the DNA molecule is a DNA plasmid. As used herein, the term "DNA plasmid" refers to a plasmid consisting of a double-stranded DNA molecule. In some embodiments, a "plasmid" is a circular DNA molecule. In some embodiments, a "plasmid" may also encompass a linear DNA molecule. Specifically, the term "plasmid" encompasses molecules obtained by linearizing a circular plasmid, for example, by cleaving the circular plasmid with a restriction endonuclease to convert the circular plasmid molecule into a linear molecule, as well as linear molecules that are replicable in prokaryotes. Plasmids can replicate, i.e., amplify, intracellularly independently of the genomic genetic information stored in the nuclear material or nucleoid of a prokaryotic cell, and can be used for cloning, i.e., amplifying genetic information in bacterial cells. Preferably, the DNA plasmids of the present application are medium-copy or high-copy plasmids, more preferably high-copy plasmids. Examples of such high-copy plasmids include vectors based on pUC, pTZ plasmids, or any other plasmids (e.g., pMB1, pCoIE1, etc.) containing an ORI that supports high copy numbers of the plasmid.
[0083] In some embodiments, the DNA molecule is a DNA molecule or fragment thereof that constitutes a prokaryotic nucleoid or nucleoid, i.e., the coding sequence or its complement, including the aforementioned poly(A) tail, can be replicated along with the prokaryotic genome.
[0084] In some embodiments, the DNA molecule is further linked to a gene fragment of interest at the 5' end of the poly(A) tail coding sequence, and the gene fragment of interest and the poly(A) tail coding sequence together encode an RNA. In some embodiments, the gene fragment of interest and the poly(A) tail coding sequence together encode an mRNA. The gene fragment of interest comprises a coding sequence for a protein, polypeptide, or fragment thereof. In some embodiments, the gene fragment of interest also comprises a coding sequence for elements that can be used to initiate or regulate expression of the protein, polypeptide, or fragment thereof after transcription, including, but not limited to, a 5' UTR, a 3' UTR, and the like. In some embodiments, the gene fragment of interest comprises a coding sequence for at least one untranslated region (UTR). In some embodiments, the gene fragment of interest comprises a coding sequence for at least a 5' UTR and a coding sequence for a protein, polypeptide, or fragment thereof. In some embodiments, a gene fragment of interest comprises, from 5' to 3', at least a 5' UTR coding sequence, a protein, polypeptide, or fragment thereof coding sequence, and a 3' UTR coding sequence. The coding sequence for a protein, polypeptide, or fragment thereof can ultimately be translated into one or more proteins or polypeptides, such as short peptides, oligopeptides, polypeptides, fusion proteins, proteins, and fragments thereof (e.g., portions of known proteins, such as functional portions). A functional portion can be, for example, a biologically active portion of a protein or an antigenic portion capable of effectively generating antibodies, such as an antigenic epitope. The coding sequence for a protein, polypeptide, or fragment thereof includes an initiation codon (5' end) and a termination codon (3' end), which are the first three nucleotides and last three nucleotides of a translatable mRNA molecule, respectively. The 5' UTR typically contains at least one ribosome binding site (RBS), such as a Shine-Dalgarno sequence in prokaryotes, or at least one translation initiation site, such as a Kozak sequence in eukaryotes. The RBS promotes efficient and accurate translation of mRNA molecules by recruiting ribosomes during translation initiation.The activity of a RBS can be optimized by varying the length and sequence of a given RBS or translation initiation site, as well as the distance from the given RBS or translation initiation site to the start codon. Alternatively or alternatively, the 5'UTR contains an internal ribosome entry site (IRES). The 3'UTR may contain one or more regulatory sequences, such as a binding site for an amino acid sequence that enhances the stability of the mRNA molecule, a binding site for a regulatory RNA molecule (such as an miRNA molecule), and / or a signal sequence involved in the intracellular transport of the mRNA molecule.
[0085] Based on the above embodiments, in some embodiments, the gene fragment of interest further comprises one or more additional regulatory sequences, such as a binding site for an amino acid sequence that enhances the stability of the mRNA molecule, a binding site for an amino acid sequence that enhances the translation of the mRNA molecule, a binding site for a regulatory element (e.g., a riboswitch), a binding site for a regulatory RNA molecule (e.g., an miRNA molecule), and / or a nucleotide sequence that positively influences translation initiation. Furthermore, the 5'UTR preferably does not contain a functional upstream open reading frame, an out-of-frame upstream translation initiation site, an out-of-frame upstream start codon, and / or a nucleotide sequence that produces a secondary structure that reduces or prevents translation. The presence of such nucleotide sequences in the 5'UTR may adversely affect translation.
[0086] The coding sequence of a protein, polypeptide, or fragment thereof includes codons that can be translated into an amino acid sequence. All codons included in the coding sequence may be naturally occurring codons that encode amino acids, or may be partially or entirely composed of artificially synthesized codons. In some embodiments, some or all of the codons are codon-optimized. In some embodiments, some or all of the codons encode unnatural amino acids.
[0087] In some embodiments, the DNA molecule further comprises structural elements necessary for initiating or regulating transcription of RNA located 5' of the gene fragment of interest, and these structural elements are known in the art. In some embodiments, the structural elements include at least a promoter. Promoters and their sequences are known in the art and include weak promoters, medium-strength promoters, strong promoters, minipromoters, or core promoters. In some specific embodiments, the promoter is a strong promoter. In some embodiments, the promoter is capable of initiating transcription of the gene fragment of interest and / or poly(A) tail in prokaryotic cells. In some embodiments, the promoter is capable of initiating transcription of the gene fragment of interest and / or poly(A) tail in eukaryotic cells. A "promoter" comprises at least one transcription recognition site followed by a transcription factor binding site. The recognition site and binding site may interact with an amino acid sequence that mediates or regulates transcription. Compared to the recognition site, the binding site is closer to the gene fragment of interest. The binding site may be, for example, a Pribnow box in prokaryotes or a TATA box in eukaryotes. For example, in some embodiments, when a Pribnow box is used, the transcription recognition site may be located approximately 35 bp upstream of the transcription start site, while the transcription factor binding site may be located approximately 10 bp upstream of the transcription start site. In some embodiments, the promoter contains at least one additional regulatory element, such as an AT-rich upstream element located approximately 40 and / or 60 nucleotides before the transcription start site, and / or an additional regulatory element located between the recognition and binding sites to enhance promoter activity. In some embodiments, the promoter is a strong promoter, i.e., the promoter contains a sequence for promoting transcription of the aforementioned RNA coding sequence. Strong promoters are known to those skilled in the art, such as the OXB18, OXB19, and OXB20 promoters derived from the E. coli RecA promoter, or can be identified or synthesized by routine laboratory procedures. In some embodiments, the promoter is a T7 promoter.In some embodiments, the promoter further comprises additional regulatory elements, such as an enhancer, contained in a DNA plasmid that can promote transcription of said RNA coding sequence.
[0088] The present application also provides a library comprising the aforementioned engineered DNA molecules. In some embodiments, the library comprises at least two DNA molecules with different poly(A) tail coding sequences.
[0089] The present application also provides uses of the above-described engineered DNA molecules for stably amplifying poly(A)-tailed coding sequences or coding sequences of poly(A)-tailed RNA. In some embodiments, the method for amplifying poly(A)-tailed coding sequences or coding sequences of poly(A)-tailed RNA is as described below in the method for stably amplifying poly(A)-tailed transcription template DNA in vitro.
[0090] Engineered RNA and libraries The present application provides an RNA comprising the aforementioned poly(A) tail and a gene fragment of interest located 5' to the poly(A) tail coding sequence. In some embodiments, the RNA further comprises a 5' cap structure. In some embodiments, the RNA is mRNA.
[0091] As used herein, "mRNA" (messenger RNA) is any naturally occurring, non-naturally occurring, or modified RNA that encodes at least one protein, polypeptide, or fragment thereof and can be translated to produce the encoded protein, polypeptide, or fragment thereof in vivo, in vitro, in situ, or ex vivo. Thus, mRNA may be mature or immature, and elements or structures that mRNA necessarily or alternatively contains are known in the art. In some embodiments, mRNA contains coding sequences for multiple functional elements necessary to express, regulate, or enhance the expression level of a protein, polypeptide, or fragment thereof. Functional elements include, but are not limited to, a 5' cap, a 5' UTR, a 3' UTR, etc. Both the 5' UTR and the 3' UTR are usually transcribed from genomic DNA, elements present in immature mRNA. In the case of mature mRNA, The term "5' cap" refers to a methylated guanylate attached to the 5' end of an mRNA via pyrophosphate, forming a 5',5'-triphosphate bond with the adjacent nucleotide. There are three types of 5' cap structures (m7G5'ppp5'Np, m7G5'ppp5'NmpNp, and m7G5'ppp5'NmpNmpNp), commonly referred to as type O, type I, and type II. Type O refers to an unmethylated ribose at the terminal nucleotide; type I refers to a methylated ribose at one terminal nucleotide; and type II refers to a methylated ribose at both terminal nucleotides. In some embodiments, for the 5′ cap, a 5′-guanosine cap structure can be produced by completing the 5′ capping of a polynucleotide during an in vitro transcription reaction using the following chemical RNA cap analogs: 3′-O-Me-m7G(5′)ppp(5′)G [ARCA cap], G(5′)ppp(5′)A, G(5′)ppp(5′)G, m7G(5′)ppp(5′)A, m7G(5′)ppp(5′)G (New England BioLabs, Ipswich, MA), or m7G(5′)ppp(5′)(2′-OMeA)pG (CleanCapAG), according to the manufacturer's protocol. For example, in some embodiments, 5′-capping of modified RNAs can be achieved by post-transcriptionally generating a type O cap structure (m7G(5′)ppp(5′)G (New England BioLabs, Ipswich, MA) using vaccinia virus capping enzyme. Type I cap structures can be generated by using both vaccinia virus capping enzyme and a 2′-O-methyltransferase to generate m7G(5′)ppp(5′)(2′-OMeA)pG. Type II cap structures can be generated from type I cap structures by 2′-O-methylating the third nucleotide from the 5′ end using a 2′-O-methyltransferase.Type III cap structures can be generated from type I cap structures by 2'-O-methylating the fourth nucleotide from the 5' end using a 2'-O-methyltransferase.
[0092] In some embodiments, some or all of the uridines in the mRNA are chemically modified uridines.
[0093] In some embodiments, some or all of the uridines in the mRNA are pseudouridine or 1-methyl-pseudouridine.
[0094] In some embodiments, some or all of the uracil nucleotides in the mRNA are substituted with pseudouridine (ψ) or N1-methyl-pseudouridine (m1ψ) nucleotides.
[0095] In some embodiments, the mRNA further comprises a stabilizing element. The stabilizing element may include, for example, a histone stem loop. In some embodiments, the mRNA comprises a coding region, at least one histone stem loop, and optionally a poly(A) sequence or polyadenylation signal. The poly(A) sequence or polyadenylation signal generally enhances the expression level of the encoded protein. In some embodiments, the mRNA comprises a combination of a poly(A) sequence or polyadenylation signal and at least one histone stem loop. Although both have alternative mechanisms in nature, they act synergistically to increase protein expression to levels exceeding those observed with either element alone. The synergistic effect of the combination of poly(A) and at least one histone stem loop is independent of the order of the elements or the length of the poly(A) sequence. In some embodiments, the histone stem loop is generally derived from a histone gene and comprises an intramolecular base-paired loop formed by two adjacent partial or complete reverse-complementary sequences separated by a spacer region (composed of a short sequence). The unpaired loop region typically cannot base-pair with any of the stem-loop elements. The stability of the stem-loop structure generally depends on the length of the paired region, the number of mismatches or bulges, and the base composition. In some embodiments, wobble base-pairing (non-Watson-Crick base-pairing) can be produced. In some embodiments, the at least one histone stem-loop sequence has a length of 15 to 45 nucleotides.
[0096] In some embodiments, one or more AU-rich sequences of mRNA may be removed. Such sequences may be referred to as AURES, which are destabilizing sequences found in the 3'UTR. AURES may be removed from mRNA. Alternatively, AURES may be retained in mRNA.
[0097] In some embodiments, the mRNA is formulated in lipid nanoparticles (LNPs). In some embodiments, lipids are mixed with the mRNA to form lipid nanoparticles. In some embodiments, the RNA is formulated in lipid nanoparticles. In some embodiments, the lipid nanoparticles are first formed as empty lipid nanoparticles, and then mixed or encapsulated with the vaccine mRNA immediately before administration (e.g., within minutes to an hour).
[0098] Lipid nanoparticles generally contain ionizable lipids, non-cationic lipids, sterols, PEG-lipid components, and target nucleic acids such as mRNA. Lipid nanoparticles according to the present disclosure can be produced using components, compositions, and methods generally known in the art. For example, see PCT / US2016 / 052352, PCT / US2016 / 068300, PCT / US2017 / 037551, PCT / US2015 / 027400, PCT / US2016 / 047406, PCT / US2016 / 000129, PCT / US2016 / 014280, PCT / US2016 / 014280, PCT / US2017 / 0384 26, PCT / US2014 / 027077, PCT / US2014 / 055394, PCT / US2016 / 52117, PCT / US2012 / 069610, PCT / US2017 / 027492, PCT / US2016 / 059575, and PCT / US2016 / 069491, all of which are incorporated herein by reference in their entireties.
[0099] The present application also provides a library comprising the aforementioned mRNA molecules, wherein the library comprises at least two mRNA molecules having different poly(A) tails.
[0100] The present application also provides the use of mRNA and mRNA library. At least two or more mRNA molecules having poly(A) tails with different degrees of influence on the mRNA expression level can be used to regulate the expression level of the coding sequence of the aforementioned protein, polypeptide, or fragment thereof, for example, by adjusting the ratio of different mRNA molecules in a library containing the two or more mRNA molecules, or by introducing one or more of the two or more mRNA molecules at different times with the same or different content.
[0101] cell The present application also provides cells containing the aforementioned engineered DNA molecules, where the DNA molecules can be stored and / or amplified within the storage cells. In some embodiments, the cells are prokaryotic cells capable of replicating the DNA molecules. In some embodiments, the cells are prokaryotic cells capable of replicating and / or transcribing the DNA molecules. In some embodiments, the DNA molecules are eukaryotic cells capable of replicating the DNA molecules. In some embodiments, the DNA molecules can be transcribed and / or replicated within the DNA-containing cells.
[0102] In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a bacterium, actinomycete, cyanobacterium, mycoplasma, rickettsia, or chlamydia. In some embodiments, the cell is selected from the group consisting of Bacillus subtilis, Lactobacillus, Acetobacterium, Corynebacterium, Brevibacterium, Arthrobacter, Pseudomonas, and Pediococcus. In some embodiments, the cell is a recA bacterium. In some embodiments, the cell is Escherichia coli. In some embodiments, the cell is Escherichia coli selected from the group consisting of K-12 strains and their derivatives, and B strains and their derivatives. In some embodiments, the E. coli is selected from the group consisting of MG1655, DH5 or DH5α, DH10B, BL21, DB3.1, HB101, JM109, JM110, MC1061, MG1655, Pir1, Stbl2, Stbl3, Top10, XL1 Blue, XL10 Gold, BLR, HMS174, Tuner, Rostetta2, Lemo21, T7 Express, and Origami2, etc. In some embodiments, the cell is selected from the group consisting of Streptomyces, Micromonospora, and Nocardia. In some embodiments, the cell is a fungus. In some embodiments, the cell is selected from a yeast or a mold.
[0103] method The present application provides a method for stably amplifying poly(A)-tailed transcription template DNA in vitro to reduce the mutation frequency of the poly(A)-tailed transcription template sequence during large-scale DNA replication in cells, which method comprises growing cells containing engineered DNA molecules.
[0104] In some embodiments, the method further comprises introducing the engineered DNA molecule into the cells prior to expanding the cells. In some embodiments, the introduction may comprise chemical transformation or electrical transformation. In some embodiments, the introduction is a natural plasma membrane invagination process of the engineered DNA molecule carried out by the cells.
[0105] In some embodiments, the method further comprises expanding the cells, extracting cellular DNA, and synthesizing RNA by in vitro transcription. In some embodiments, the method further comprises expanding the cells, inducing transcription of RNA in the cells, and then extracting and isolating the RNA in the cells. In some embodiments, the method further comprises extracting cellular DNA and transducing it into a second cell capable of transcribing the RNA. In some embodiments, the transduction comprises administration to a human, wherein the administration is selected from the group consisting of intravenous, intraperitoneal, subcutaneous, intracranial, intrathecal, intra-arterial (e.g., via the carotid artery), intramuscular, and intratumoral injection or perfusion.
[0106] The present application also provides a method for modulating protein expression, the method comprising: transducing two or more of the aforementioned engineered DNA molecules into a cell of interest at different times and / or in different ratios; or transducing two or more of the aforementioned engineered RNA molecules at different times and / or in different ratios into a cell of interest; Two or more of the engineered DNA molecules and two or more of the RNA molecules have different poly(A) tails, which have different degrees of influence on the expression level of the RNA.
[0107] In some embodiments, the present application also provides a method for modulating protein expression, the method comprising introducing the engineered DNA molecule or the RNA molecule into a cell of interest. In some embodiments, the poly(A) tail encoded by the DNA and the coding sequence for the poly(A) tail contained in the RNA are element a - element c - element b - element c - element b - element c - element b - element c - element b, element b - element c - element b - element c - element a - element d - element b - element c - element b - element c - element b, element b - element c - element b - element c - element b - element d - element a - element c, element a-element d-element b-element c-element b-element c-element b, or It comprises a structure selected from the group consisting of element b-element c-element b-element c-element b-element d-element a.
[0108] In some embodiments, the coding sequence for the poly(A) tail encoded by the DNA and contained in the RNA is: 60A-G-19A-G-19A-G-19A-G-3A and 7A-C-18A-G-60A-element d-7A-C-18A-G-14A and 60A-element d-19A-G-19A-G-17A, 19A-G-19A-G-19A-Element d-60A and 19A-G-19A-G-19A-Element d-60A-G and 19A-G-19A-G-19A-Element d-60A and 19A-C-19A-C-19A-Element d-60A and 19A-T-19A-T-19A-element d-60A.
[0109] In some embodiments, the coding sequence for the poly(A) tail encoded by the DNA and contained in the RNA comprises the structure 60A-element d-19A-G-19A-G-17A or 19A-G-19A-G-19A-element d-60A.
[0110] In particular, the two elements linked by "-" are directly linked, with no nucleotides between the two elements.
[0111] In the above poly(A) tail structure, "yA" represents the number of consecutive As in element a or element b, where y is a natural number. For example, 19A means that 19 consecutive As are contained, and 60A means that 60 consecutive As are contained.
[0112] In some embodiments, the coding sequence of the poly(A) tail encoded by the DNA and the poly(A) tail contained in the RNA comprises or consists of any one polynucleotide sequence selected from the polynucleotide sequences represented by SEQ ID NOs: 1 to 10.
[0113] It should be understood that the present application encompasses various aspects, embodiments, and combinations of aspects and / or embodiments described herein. The above description and the following examples are intended to illustrate, rather than limit, the scope of the present application. Other aspects, improvements, and modifications within the scope of the present application will be apparent to those skilled in the art. Therefore, those skilled in the art should understand that such improvements and modifications to these aspects and embodiments are also included within the scope of the present application. [Example]
[0114] Example 1: Construction of a poly(A) tail
[0115] The poly(A) tails and DNA sequences encoding the poly(A) tails shown in Table 1 below were constructed by conventional genetic engineering techniques.
[0116] [Table 1]
[0117] Example 2: Testing poly(A) function using the luciferase coding sequence as an example
[0118] 2.1 Testing the replication stability of DNA molecules encoding mRNA in prokaryotic cells
[0119] Using luciferase as the protein coding region, we investigated the stability of different poly(A) mutants in E. coli and their effect on the expression of luciferase in the cells.
[0120] 1) Construction of a universal vector containing the luciferase protein coding region In this universal vector, the E. coli cloning vector pUC57 was used as the vector backbone, and the T7 promoter sequence (5′-TAATACGACTCACTATAAGG-3′), 5′UTR, luciferase protein, 3′UTR, and polyadenylate string poly(dA:dT) were sequentially placed between the Xba I and EcoR I restriction sites in the multiple cloning site.
[0121] 2) The polyadenylate string poly(dA:dT) in the universal vector was replaced with P1 to P10 of the present application and controls C1 to C4 (A60-10 nt spacer-A60 (control C1), A30-10 nt spacer-A70 (control C2), A60-1 nt spacer-A60 (control C3), or A60-6 nt spacer-A60 (control C4). Of these, C1 and C2 are derived from a reference patent document (U.S. Patent No. 10,717,982 B2).
[0122] All primers required for constructing P1 to P10 and the control C1 to C4 were synthesized, and poly(dA:dT) was removed from the universal vector constructed in step 1 by double digestion with two restriction endonucleases. Then, P1 to P10 and C1 to C4 were ligated with T4 DNA ligase 1 to the vector from which poly(dA:dT) had been removed, thereby completing the replacement of poly(dA:dT) in the universal vector.
[0123] 3) Detection of replication stability of different poly(A) variants in E. coli The vector plasmid constructed in step 2 was confirmed to be correct by sequencing and then transformed into E. coli DH5α. After growing the transformed culture dish at 30°C, the plasmid was extracted and sequenced. After sequencing, the stability and base deletion of different poly(A) variants were analyzed and calculated based on the sequencing results. Replication stability was expressed as the percentage of clones without any base changes; the higher the percentage, the higher the replication stability of the plasmid in E. coli.
[0124] The results are shown in Figures 1 and 2, and the specific experimental results are as follows.
[0125] In control C1, a total of 100 clones were tested, of which 15 clones had base deletions, accounting for 15%, and the number of correct clones was 85, accounting for 85%.
[0126] In control C2, a total of 50 clones were tested, and all clones were correct without base changes or deletions, resulting in a correct clone rate of 100%.
[0127] In control C3, a total of 50 clones were tested, of which 9 clones had base deletions, accounting for 18%, and the number of correct clones was 41, accounting for 82%.
[0128] In control C4, a total of 50 clones were tested, of which 14 clones had base deletions, accounting for 28%, and the number of correct clones was 36, accounting for 72%.
[0129] In control P1, a total of 100 clones were tested, of which 9 clones had base deletions, accounting for 9%, and the number of correct clones was 91, accounting for 91%.
[0130] In control P2, a total of 100 clones were tested, of which 12 clones had base deletions, accounting for 12%, and the number of correct clones was 88, accounting for 88%.
[0131] In control P3, a total of 62 clones were tested, of which 5 clones had base deletions, accounting for 8%, and the number of correct clones was 57, accounting for 92%.
[0132] In control P4, a total of 50 clones were tested, of which 3 clones had base deletions, accounting for 6%, and the number of correct clones was 47, accounting for 94%.
[0133] In control P5, a total of 50 clones were tested, of which 4 clones had base deletions, accounting for 8%, and the number of correct clones was 46, accounting for 92%.
[0134] In control P6, a total of 50 clones were tested, of which 5 clones had base deletions, accounting for 10%, and the number of correct clones was 45, accounting for 90%.
[0135] In control P7, a total of 50 clones were tested, of which 7 clones had base deletions, accounting for 14%, and the number of correct clones was 43, accounting for 86%.
[0136] In control P8, a total of 50 clones were tested, of which 3 clones had base deletions, accounting for 6%, and the number of correct clones was 47, accounting for 94%.
[0137] In the control P9, a total of 50 clones were tested, of which 4 clones had base deletions, accounting for 8%, and the number of correct clones was 46, accounting for 92%.
[0138] In control P10, a total of 50 clones were tested, of which 5 clones had base deletions, accounting for 10%, and the number of correct clones was 45, accounting for 90%.
[0139] Considering the ratio of correct clones and the number of deleted bases, the poly(A) mutants designed in this application are superior or comparable to conventional techniques in terms of replication stability in E. coli cells. In particular, the replication stability of poly(A) mutants P3, P4, and P8 is the highest. The replication stability of P3, P4, and P8 is comparable to that of C2, with no statistically significant difference (p>0.05, χ 2 Furthermore, the replication stability of P3, P4, and P8 was superior to that of the controls C1, C3, and C4, and the difference was statistically significant (p<0.05, χ 2 Certification).
[0140] Example 3: Testing poly(A) function using HPV antigen protein coding sequences as an example
[0141] 1) Construction of plasmids containing HPV protein coding regions
[0142] As described above, in Example 1, vectors were constructed that combined a luciferase-encoding gene with different poly(A) sequences. Based on these vectors, the luciferase-encoding gene was replaced with an HPV-encoding gene using conventional molecular cloning techniques. The main elements were arranged in the following order: T7 promoter sequence (5'-TAATACGACTCACTATAAGG-3'), 5'UTR, HPV antigen protein coding sequence, 3'UTR, and poly(A) coding sequence.
[0143] 2) Conducting small-scale bacterial cultures to test the stability of four HPV poly(A) variants in E. coli
[0144] The HPV vector plasmids containing P1, P2, P3, and P4 constructed in step 1 were confirmed to be correct by sequencing and then transformed into E. coli DH5α. After growing the transformed culture plates at 30°C, the plasmids were extracted and sequenced. After sequencing, the stability and base deletions of different poly(A) variants were analyzed and calculated based on the sequencing results. Stability was expressed as the percentage of clones without any base changes; a higher percentage indicates higher stability.
[0145] The results (Figures 3-4) showed that when the target gene was replaced with the HPV antigen coding sequence, In control C1, a total of 88 clones were tested, of which 47 clones had base deletions, accounting for 53%, and the number of correct clones was 41, accounting for 47%.
[0146] In control C2, a total of 50 clones were tested, of which one clone had a base deletion, accounting for 2%, and the number of correct clones was 49, accounting for 98%.
[0147] In control P1, a total of 101 clones were tested, of which 14 clones had base deletions, accounting for 14%, and the number of correct clones was 87, accounting for 86%.
[0148] In control P2, a total of 100 clones were tested, of which 12 clones had base deletions, accounting for 12%, and the number of correct clones was 88, accounting for 88%.
[0149] In control P3, a total of 70 clones were tested, of which 4 clones had base deletions, accounting for 6%, and the number of correct clones was 66, accounting for 94%.
[0150] In control P4, a total of 50 clones were tested, of which 9 clones had multiple base deletions, accounting for 18%, and the number of correct clones was 41, accounting for 82%.
[0151] Considering the mutation rate and the average number of base deletions, when the luciferase protein-encoding gene in Example 1 was replaced with an HPV antigen protein, the different poly(A) variants of the present application still maintained high replication stability. In particular, P3 and C2 had comparable stability, with no statistically significant difference (p>0.05, χ 2(test). Furthermore, compared with other groups of novel mutants in this application, the cloning stability is optimal. The probability of large fragment deletion in C2 is 1 / 50 (=2%), while the probability of large fragment deletion in P3 is 1 / 70 (=1.4%). Because large fragment deletion affects the in vivo expression and efficacy of mRNA products, P3 is more suited to product requirements. These results demonstrate that the poly(A) mutants designed in this application are universally applicable to examples involving different protein-coding regions.
[0152] As described above, in Examples 1 and 2, the culture plates containing transformed E. coli were placed in a biochemical incubator and cultured overnight at 30°C, and the resulting clones were sequenced to evaluate replication stability. This is because the culture temperature of E. coli also affects the DNA replication rate, which in turn affects replication stability. For Example 3, the culture plates containing transformed E. coli were placed in a biochemical incubator at 37°C and then cultured. The sequencing results were also compared in this application. The results (Figure 5) show that the base deletion rate of control C1 significantly increased from 53% when cultured at 30°C to 98% when cultured at 37°C. Unlike the control, there was no significant difference in the mutation rates of P1, P2, and P3 under the two temperature conditions. This indicates that the poly(A) mutants designed in this application still maintained high replication stability in the examples despite different temperature culture conditions of E. coli, demonstrating their universal applicability.
[0153] 3) Large-scale fermentation and stability of the three poly(A) variants during different generations Large-scale fermentation is essential for mRNA drug production to prepare sufficient template plasmids. Plasmid stability during fermentation (in this application, plasmid stability specifically refers to the stability of poly(dA):dT) is crucial for the production of mRNA drugs with consistent quality. Meanwhile, to meet the stability requirements of different production batches, strain libraries containing the desired plasmids (including strain libraries of different generations, such as primary and secondary libraries) must be established. Therefore, it is necessary to evaluate the stability of the plasmids in E. coli at different generations. In response to these two challenges, in Example 3, this application investigated the stability of P1 and P3 during different generations of the fermentation process. Based on sequencing results, four correct E. coli clones for each poly(A) variant were selected and subcultured by fermentation. The results showed that the plasmids of the four clones, P1 and P3, were stable without any base changes at the third, fifth, seventh, and ninth passages.
[0154] Example 4: Detection of luciferase mRNA expression levels in mice
[0155] In eukaryotic cells, a certain length of poly(A) tail is essential for protecting the 3' end of mRNA, maintaining mRNA stability, and promoting protein expression. In vitro, under the influence of physiological or environmental factors, poly(A) gradually shortens, accelerating mRNA degradation. In this application, we investigated the effects of different poly(A) variants on protein expression levels. In vivo expression assays were performed in mice to evaluate the effects of different poly(A) variants on luciferase activity. During this application, mice were intramuscularly injected with luciferase mRNA-LNPs containing control C2 and P3, P4, P5, P8, and P9 variants. Six hours after injection, animal imaging was performed, and fluorescence intensity was quantitatively measured to compare the effects of different poly(A) variants on luciferase activity in vivo. The specific experimental procedure is as follows:
[0156] Luciferase mRNA was synthesized by in vitro transcription and linearized DNA was obtained by digestion with the type II restriction endonuclease BspQI. The 3' end of the linearized DNA was expressed with different poly(A) fragments: control C2, P3, P4, P5, P8, or P9. The linearized DNA was used as a template for in vitro transcription. A 100 μl reaction system contained 1X reaction buffer, 5 mM each (final concentrations) of ATP, CTP, N1M-UTP, and GTP, 4 mM (final concentrations) of CleanCap AG, and 5 μl of in vitro transcription enzyme. After thorough mixing of the reaction mixture, the reaction was carried out at 37°C for 3 hours. The in vitro transcribed mRNA was recovered by LiCl precipitation and finally dissolved in enzyme-free water.
[0157] Animal experiments: In vitro synthesized luciferase mRNA was encapsulated in LNPs. The resulting mRNA stock solution was dispersed in a 20 mM acetic acid solution (pH 5.0) to obtain an RNA solution with an mRNA concentration of 200 μg / mL. The lipid mixture was prepared by mixing ionized lipid, cholesterol, DSPC, and DMG-PEG2000 in a molar ratio of ionized lipid:cholesterol:DSPC:DMG-PEG2000 = 50:38.5:10:1.5. The mRNA and lipid mixture were mixed by controlling the flow rates of the aqueous and oil phases using a T mixer. The injection pump was started to mix the mRNA solution and lipid mixture to form LNPs. The solution was then diluted 10-fold with diluent and concentrated by centrifugation in an ultrafiltration tube, followed by three cycles of solvent exchange. The pH of the resulting solution was adjusted to 7.0-8.0 by adding a Tris aqueous solution to obtain an LNP-encapsulated mRNA solution. LNP stands for lipid nanoparticle. The concentration and particle size of LNP-encapsulated mRNA were determined using a Ribogreen RNA quantification kit (Invitrogen, R11490) and a Darwin ZetaSizer particle size analyzer, respectively. In the four-component LNPs, the molar ratio of each component was SM102:DSPC:cholesterol:DMG-PEG2000 = 50:10:38.5:1.5. After encapsulation, quality control was performed on the LNPs by measuring particle size, encapsulation efficiency, PDI, and other indicators. According to the quality control results, the prepared LNPs met the standard of a particle size range of 50 nm to 150 nm, a PDI < 0.3, and an encapsulation efficiency > 90%, and were suitable for subsequent experiments. The mRNA content in the LNPs was determined by the Ribogreen method, and the LNPs were then diluted to an mRNA content of 100 ng / µL.
[0158] BALB / c mice were randomly divided by weight and administered the treatment after 2–3 days of adaptive feeding. Five mice were injected with the control C2, P3, P4, P5, P8, and P9, each intramuscularly at 100 μl (10 μg mRNA). Five mice were also injected with PBS as a control group. Six hours after injection, animal imaging was performed, and then fluorescence values were calculated. The results (Figures 7–8) show that compared with the control C2, the expression activity of P3 increased significantly by 1.8-fold, with a statistically significant difference (p<0.01, Student's t-test). The expression levels of P4, P5, P8, and P9 were comparable to those of C2, with no significant difference (p>0.05, Student's t-test).
[0159] The sequences used in the above examples of the present application are shown in the sequence listing below. It should be understood that the sequences below are merely exemplary sequences for the embodiments of the present application and do not impose any limitations on the embodiments of the present application. The nucleic acid sequences in the sequence listing below may represent DNA sequences or RNA sequences, and when representing an RNA sequence, "T" represents uridine.
[0160] [Table 2]
[0161] [Table 3]
[0162] [Table 4]
[0163] [Table 5]
[0164] [Table 6]
[0165] Table 7
[0166] Table 8
Claims
1. An engineered DNA molecule capable of replicating in a cell, comprising a polyadenosine tail (polyA tail) coding sequence, wherein the poly(A) tail coding sequence comprises a single element a, at least one element b, and at least one element c and / or at least one element d; the element a consists of a plurality of consecutive adenine (A) nucleotides, and the length of the element a is in the range of 20 nt or more; the element b consists of a plurality of consecutive A nucleotides, and the length of the element b is in the range of 3 nt or more and less than 20 nt; said element c consists of one non-A nucleotide, said nucleotide being selected from T, C and G nucleotides; The element d consists of any two or more consecutive nucleotides, the nucleotides being selected from A, T, C and G nucleotides, the nucleotides at the 5' and 3' ends of the element d are not A nucleotides, the element d does not contain three or more consecutive A nucleotides, and the length of the element d is in the range of 2 nt to 20 nt, The element a and the element b are not adjacent to each other, and the element c and the element d are not adjacent to each other, An engineered DNA molecule wherein the poly(A) tail coding sequence does not contain any two elements b adjacent to each other, does not contain any two elements c adjacent to each other, and does not contain any two elements d adjacent to each other.
2. 2. The DNA molecule of claim 1, wherein the length of the poly(A) tail coding sequence is 101 to 200 nt, 101 to 150 nt, 120 to 150 nt, 130 to 140 nt, 120 to 135 nt, or 123 to 125 nt.
3. 3. The DNA molecule of claim 1, wherein the 3' end of the poly(A) tail coding sequence is an A nucleotide or a non-A nucleotide.
4. The DNA molecule according to any one of claims 1 to 3, wherein the length of element a is 80 nt or less.
5. The DNA molecule according to any one of claims 1 to 3, wherein the length of element a is 30 to 70 nt, 35 to 65 nt, 40 to 60 nt or 45 to 55 nt, preferably 60 nt.
6. The DNA molecule of any one of claims 1 to 5, wherein 50% or more of the polynucleotides of element a are located in the 5' or 3' portion of the poly(A) tail coding sequence.
7. The DNA molecule according to any one of claims 1 to 6, wherein the length of element b is 3 to 10 nt, 10 to 19 nt, 12 to 15 nt, 14 to 17 nt, or 16 to 19 nt, preferably 19 nt.
8. The DNA molecule according to any one of claims 1 to 7, wherein the number of elements b is 2 to 10, preferably 2 to 5, and more preferably 3.
9. The DNA molecule according to any one of claims 1 to 8, wherein the element c is G.
10. The DNA molecule according to any one of claims 1 to 9, wherein the number of elements c is 2 to 10, 3 to 8, 4 to 6 or 2 to 5, and preferably 2.
11. The DNA molecule according to any one of claims 1 to 10, wherein the element d comprises a palindromic sequence.
12. The DNA molecule according to any one of claims 1 to 11, wherein the length of the element d is 3 to 18 nt, 5 to 16 nt, 4 to 10 nt or 6 to 12 nt, preferably 6 nt.
13. The DNA molecule according to any one of claims 1 to 12, wherein the element d is any one or more selected from the sequences GATATC (SEQ ID NO: 15), GTATAC (SEQ ID NO: 16), GAATCT (SEQ ID NO: 17), GCATATGACT (SEQ ID NO: 18), and GATATCGTATAC (SEQ ID NO: 19).
14. The DNA molecule according to any one of claims 1 to 13, wherein the element d is any one or more selected from the nucleotide sequences of SEQ ID NO:15, SEQ ID NO:16, and SEQ ID NO:
17.
15. The DNA molecule according to any one of claims 1 to 14, wherein the nucleotide sequence of element d is represented by SEQ ID NO:
15.
16. The DNA molecule according to any one of claims 1 to 15, wherein the number of elements d is 0 to 5, preferably 1 to 3, and more preferably 1.
17. The DNA molecule according to any one of claims 1 to 16, wherein when the element c and the element d are present simultaneously, the total number of the elements c and the elements d is 2 to 15, preferably 3 to 5, and more preferably 3.
18. The DNA molecule of any one of claims 1 to 17, wherein the 3' portion of the poly(A) tail coding sequence, preferably a half portion near the 3' end of the poly(A) tail coding sequence, comprises one or more non-A nucleotides.
19. The structure of the poly(A) tail coding sequence is element a - element c - element b - element c - element b - element c - element b - element c - element b, element b - element c - element b - element c - element a - element d - element b - element c - element b - element c - element b, element b - element c - element b - element c - element b - element d - element a - element c, element a - element d - element b - element c - element b - element c - element b, or 19. The DNA molecule according to any one of claims 1 to 18, which is element b - element c - element b - element c - element b - element d - element a.
20. The structure of the poly(A) tail coding sequence is The DNA molecule according to any one of claims 1 to 19, comprising: element a - element d - element b - element c - element b - element c - element b, wherein the length of element a is 60 nt, the length of element b is 16 to 19 nt, and the length of element d is 6 nt.
21. The structure of the poly(A) tail coding sequence is The DNA molecule according to any one of claims 1 to 20, comprising: element b - element c - element b - element c - element b - element d - element a, wherein the length of element a is 60 nt, the length of element b is 16 to 19 nt, and the length of element d is 6 nt.
22. The DNA molecule according to any one of claims 1 to 21, wherein the poly(A) tail coding sequence is represented by any one of SEQ ID NOs: 1 to 10.
23. The DNA molecule according to any one of claims 1 to 22, wherein the poly(A) tail coding sequence is represented by SEQ ID NO: 3 or SEQ ID NO:
4.
24. The DNA molecule of any one of claims 1 to 23, wherein the DNA molecule is further linked to a gene fragment of interest at the 5' end of the poly(A) tail coding sequence, and the gene fragment of interest and the poly(A) tail coding sequence jointly encode an RNA.
25. The DNA molecule of any one of claims 1 to 24, further comprising a replicon.
26. The DNA molecule of any one of claims 1 to 25, further comprising a resistance gene.
27. The DNA molecule of any one of claims 1 to 26, further comprising a promoter for initiating transcription of the RNA.
28. The DNA molecule of any one of claims 1 to 27, wherein the gene fragment of interest comprises a 5'UTR coding sequence.
29. The DNA molecule of any one of claims 1 to 28, wherein the gene fragment of interest comprises a protein-coding sequence or a non-protein-coding sequence.
30. The DNA molecule of any one of claims 1 to 29, wherein the gene fragment of interest comprises a 3'UTR coding sequence.
31. 31. The DNA molecule of any one of claims 1 to 30, comprising a replicon, an antibiotic resistance gene, a promoter, a 5'UTR coding sequence, a protein coding sequence, and a 3'UTR coding sequence.
32. 32. The DNA molecule according to any one of claims 1 to 31, wherein said protein coding sequence encodes an HPV (Human Papillomavirus) protein, preferably said HPV protein is from HPV type 16 and / or 18.
33. 33. The DNA molecule of any one of claims 1 to 32, wherein the protein coding sequence encodes an HPV E2, E6 or E7 protein, a fusion protein of E6 and E7 protein polypeptide fragments, or a fusion protein of E2, E6 and E7 protein polypeptide fragments, preferably wherein the HPV protein is from HPV type 16 and / or 18.
34. The DNA molecule according to any one of claims 1 to 33, wherein the protein coding sequence encodes the polypeptide represented by SEQ ID NO:
26.
35. 35. The DNA molecule of any one of claims 1 to 34, comprising a polynucleotide sequence represented by any one of SEQ ID NOs: 22 to 25, or a synonymous variant of said polynucleotide sequence represented by any one of SEQ ID NOs: 22 to 25, or a polynucleotide sequence having more than 85% sequence identity with said polynucleotide sequence represented by any one of SEQ ID NOs: 22 to 25 or a synonymous variant thereof.
36. The DNA molecule of any one of claims 1 to 35, wherein the DNA molecule is a DNA plasmid.
37. A cell comprising a DNA molecule according to any one of claims 1 to 36.
38. 38. The cell of claim 37, wherein the cell is a prokaryotic cell.
39. 38. The cell of claim 37, wherein the cell is Escherichia coli.
40. An RNA molecule encoded by the DNA molecule of any one of claims 1 to 36.
41. The RNA molecule of claim 40, wherein the RNA molecule further comprises a 5' cap structure and / or some or all of the uridines in the RNA are chemically modified uridines, preferably some or all of the uridines in the RNA are pseudouridine or 1-methyl-pseudouridine.
42. A DNA coding sequence for the poly(A) tail of any one of claims 1 to 36.
43. A poly(A) tail sequence encoded by a DNA coding sequence for a poly(A) tail according to any one of claims 1 to 36.
44. 43. Use of a DNA coding sequence of the poly(A) tail of claim 42 to make the replication of a DNA molecule encoding an RNA more conservative in a host cell, wherein the poly(A) tail is located at the 3' end of the RNA.
45. 45. The use according to claim 44, wherein the host cell is a prokaryotic cell, preferably E. coli.
46. 44. Use of the poly(A) tail sequence described in claim 43 for regulating the expression of an RNA molecule in a host cell, wherein the poly(A) tail is located at the 3' end of the RNA.
47. 47. The use according to claim 46, wherein the host cell is a eukaryotic cell, preferably a mammalian cell, more preferably a human cell.
48. A library comprising the DNA molecule of any one of claims 1 to 36.
49. A library comprising RNA molecules encoded by the DNA molecules of any one of claims 1 to 36.
50. 1. A method for modulating protein expression, comprising: Introduction of multiple DNAs in the DNA library according to claim 48 into target cells at different times and / or different amount ratios; or 50. A method comprising introducing multiple RNAs in the RNA library of claim 49 into a target cell at different times and / or in different amount ratios.
51. A DNA and RNA hybrid molecule comprising genetic information similar to the DNA molecule of any one of claims 1 to 36, genetic information similar to the poly(A) tail of claim 42, or genetic information similar to the RNA molecule of claim 40.