Method for designing and screening UTR combination, UTR combination, and use thereof

By designing and screening UTR combinations, and utilizing 49 indicators and multiple cell studies, we solved the problem of systematically evaluating the effects of UTR combinations on protein expression in different cell lines. This enabled us to screen UTR combinations that improve translation efficiency in various cell types, making them applicable to cell therapy and mRNA vaccine fields.

WO2025260375A1PCT designated stage Publication Date: 2025-12-26SHENZHEN BGI HUO-YAN ENGINEERING TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/100785
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies lack systematic evaluation of the effects of UTR combinations on the expression of mRNA proteins in different cell lines, resulting in the lack of universality of the selected sequence combinations in various cell types. Furthermore, independently designing 5'UTR or 3'UTR may lead to suboptimal combined effects.

Method used

We designed and screened UTR combinations, and conducted a comprehensive analysis of UTR combinations using 49 indicators, including primary structure, special sequences or elements, and secondary structure, to assess their impact on cellular translation efficiency. We also conducted studies in various cell types to identify features that enhance translation efficiency.

Benefits of technology

It provides a combination of UTRs with good translation efficiency in multiple cells, simplifies the design and screening process, and is suitable for efficient RNA expression in cell therapy and mRNA vaccine fields, improving the translation efficiency of translated RNA sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024100785-FTAPPB-I100001
    Figure PCTCN2024100785-FTAPPB-I100001
  • Figure PCTCN2024100785-FTAPPB-I100002
    Figure PCTCN2024100785-FTAPPB-I100002
  • Figure PCTCN2024100785-FTAPPB-I100003
    Figure PCTCN2024100785-FTAPPB-I100003
Patent Text Reader

Abstract

Disclosed are a method for designing and screening a UTR combination, a UTR combination, and a use thereof. In the method for designing and screening a UTR combination, the UTR combination comprises a 5'UTR and a 3'UTR, and the method comprises the following steps: dividing groups of the UTR combination, selecting a difference feature, and evaluating the UTR combination. The method provided by the present disclosure can simplify UTR design and screening processes, and obtain, by means of indexes and a formula, a UTR combination widely applicable to improving cell translation efficiency in organisms. A UTR combination obtained by screening according to the present invention has good translation efficiency in multiple cells. The method provided by the present invention can quickly and effectively assist in designing RNA sequences that can be efficiently translated in the field of cell therapy or mRNA vaccines.
Need to check novelty before this filing date? Find Prior Art

Description

A method for designing and screening UTR combination, UTR combination and application thereof TECHNICAL FIELD

[0001] The present application belongs to the field of biotechnology, and particularly relates to a method for designing and screening UTR combination, UTR combination and application thereof. BACKGROUND

[0002] As a storage of genetic information in the post-transcriptional regulation process, multiple regions of mRNA have a regulatory effect on protein translation. Among them, UTR plays an important role in the post-transcriptional regulation of gene expression by regulating mRNA transport, translation efficiency, genetic information stability and subcellular localization.

[0003] At present, the design of mRNA untranslated region (UTR) depends on the widely used sequences, including human endogenous alpha globin, beta globin and the like. In addition, there are studies on the de novo design of 5'UTR or 3'UTR or the combination of endogenous sequences to find the UTR sequence with the highest protein expression level in 293T cells.

[0004] The prior art lacks a method for systematically evaluating the influence of UTR combination on the expression of mRNA in different cell lines to express proteins, and the most suitable UTR sequence needs to be screened from scratch. At the same time, since many studies only screen UTR sequences in one type of cell, this may lead to the sequence combination finally screened only promoting efficient translation of proteins in certain special cell lines, and not having universality in many cells. In addition, many studies only design 5'UTR or 3'UTR separately, and independent design may lead to the effect of combination not being optimal.

[0005] SUMMARY

[0006] The technical problem to be solved by the present application is to overcome the problem that the prior art lacks a method for systematically evaluating the influence of UTR combination on the expression of mRNA in different cell lines to express proteins, and a method for designing and screening UTR combination, UTR combination and application thereof are provided. The method provided by the present application can comprehensively analyze the UTR combination from 49 indicators, including primary structure, special sequence or element and secondary structure, and also includes evaluation of the influence of UTR combination on cell translation efficiency by 19 Motif sequences (Motif 1-19). The method provided by the present application can simplify the UTR design and screening process, and obtain UTR combination widely applicable to improving cell translation efficiency in organisms through indicators and formulas. The UTR combination obtained by screening according to the present application has good translation efficiency in multiple cells. The method provided by the present application can quickly and effectively help to design RNA sequences with high efficiency of translation in the field of cell therapy or mRNA vaccine.

[0007] The present application solves the above technical problems by the following technical solutions.

[0008] The present application designs 36 UTR combinations and 3 control UTR sequences, evaluates the influence of UTR on protein translation efficiency from the perspective of combination, and studies the designed combinations in various common cells including 293T, Hela, A549, Vero, Huh-7 and DC2.4 cells, and realizes the wide applicability of UTR combinations.

[0009] The present application studies the influence of UTR combinations on the protein expression levels of various cells including 293T, Hela, A549, Vero, Huh-7 and DC2.4 cells compared with widely used endogenous beta globin 5'UTR (5U2) and 3'UTR (3U2), and aims to find the UTR combination characteristics for improving cell translation efficiency.

[0010] The first aspect of the present application provides a method for designing and screening UTR combinations, wherein the UTR combinations include 5'UTR and 3'UTR, and the method comprises the following steps:

[0011] (1) Grouping of UTR combinations: obtaining the protein expression amount information of expression elements containing different UTR combinations, and dividing the UTR combinations into excellent group, general group and poor group;

[0012] (2) Selection of differential characteristics: in the excellent group and the poor group as described in step (1), the characteristics with p<0.05 are obtained from the candidate characteristics as the differential characteristics for evaluating the UTR combinations by the t-test method;

[0013] (3) Evaluation of UTR combinations: evaluating the UTR combinations by the differential characteristics described in step (2).

[0014] In some embodiments of the application, in step (2), the candidate features include one or more of: 5' UTR length, 5' UTR GC content, 5' UTR G percentage, 5' UTR C percentage, 5' UTR A percentage, 5' UTR T percentage, 5' UTR thermodynamic free energy, 5' UTR minimum free energy, 5' UTR MFE secondary structure stem number, 5MSS stem total paired bases, 5MSS hairpin structure to Cap distance, 5' UTR centroid method secondary structure stem number, 5CSS stem total paired bases, 5CSS hairpin structure to Cap distance, presence or absence of a classic kozak sequence, Motif 1 similarity, Motif 2 similarity, Motif 3 similarity, Motif 4 similarity, Motif 5 similarity, Motif 6 similarity, Motif 7 similarity, Motif 8 similarity, Motif 9 similarity, 3' UTR length, 3' UTR GC content, 3' UTR G percentage, 3' UTR C percentage, 3' UTR A percentage, 3' UTR T percentage, 3' UTR thermodynamic free energy, 3' UTR MEF, 3MSS stem number, 3MSS stem total paired bases, 3CSS stem number, 3CSS stem total paired bases, ARE element, Motif 10 similarity, Motif 11 similarity, Motif 12 similarity, Motif 13 similarity, Motif 14 similarity, Motif 15 similarity, Motif 16 similarity, Motif 17 similarity, Motif 18 similarity, Motif 19 similarity, 5' UTR to CDS paired bases, and 5' UTR to 3' UTR paired bases;

[0015] In the formula, the nucleotide sequence of the Motif 1 is shown in SEQ ID NO: 14, the nucleotide sequence of the Motif 2 is GRSASM, wherein R is A or G; S is C or G; M is A or C; the nucleotide sequence of the Motif 3 is CAAVMTYA, wherein V is A, C or G; M is A or C; Y is C or T; the nucleotide sequence of the Motif 4 is GKGMWM, wherein K is G or T; M is A or C; W is A or T; the nucleotide sequence of the Motif 5 is AMRCAW, wherein M is A or C; R is A or G; W is A or T; the nucleotide sequence of the Motif 6 is KYMMYM, wherein K is G or T; Y is C or T; M is A or C; the nucleotide sequence of the Motif 7 is NCATMM, wherein N is A, C, G or T; M is A or C; the nucleotide sequence of the Motif 8 is CMWCMN, wherein M is A, C; W is A or T; N is A, C, G or T; the sequence of the Motif 9 is KSWMHC, wherein K is G or T; S is C or G; W is A or T; M is A or C; H is A, C or T; the nucleotide sequence of the Motif 10 is AATAAA; the nucleotide sequence of the Motif 11 is CSTKGWGC, wherein S is C or G; K is G or T; W is A or T; the nucleotide sequence of the Motif 12 is TTCTTK, wherein K is G or T; the nucleotide sequence of the Motif 13 is CTGDGKR, wherein D is A, G or T; K is G or T; R is A or G; the nucleotide sequence of the Motif 14 is MTGSAT, wherein M is A or C; S is C or G; the nucleotide sequence of the Motif 15 is CCAASCMS, wherein S is C or G; M is A or C; the nucleotide sequence of the Motif 16 is MCTCKK, wherein M is A or C; K is G or T; the nucleotide sequence of the Motif 17 is AACMTWWA, wherein M is A or C; W is A or T; the nucleotide sequence of the Motif 18 is TCAWTK, wherein W is A or T; K is G or T; the nucleotide sequence of the Motif 19 is GAAGSA, wherein S is C or G.

[0016] In the present application, the sequence of "Motif X" (X = 1-19) is obtained by meme software prediction analysis according to the UTR sequence (5U1-5U6, 3U1-3U6), wherein Motif 1-9 is the sequence in 5'UTR, and Motif 10-19 is the sequence in 3'UTR.

[0017] In the present application, "Motif X similarity" (X = 1-19) refers to the significance of the UTR sequence in the UTR combination to be evaluated and the Motif sequence as the evaluation, which is based on the p-value obtained by the meme software (see Figure 9), and the smaller the value represents the higher similarity of the actual sequence to the corresponding motif, that is, the negative correlation, and the value of 1 indicates the lack of the corresponding Motif.

[0018] In the present application, T in the UTR sequence means U.

[0019] In the present application, the classical kozak sequence is GCCRCC, wherein R can be G or A, and preferably A. The ARE sequence is AUUUA.

[0020] In the present application, the candidate features are selected from one or more of the following: 5'UTR length of 40-60 nt; 5'UTR GC content of 40-55%; 5'UTR G percentage of 10-20%; 5'UTR C percentage of 30-35%; 5'UTR A percentage of 30-35%; 5'UTR T percentage of 15-25%; 3'UTR length of 80-300 nt; 3'UTR GC content of 25-70%; 3'UTR G percentage of 8-30%; 3'UTR C percentage of 20-50%; 3'UTR A percentage of 25-35%; 3'UTR T percentage of 18-35%;

[0021] 5'UTR or 3'UTR contains at least one or more Motif 1: THWKCTTCTS (SEQ ID NO: 14); wherein, in the mRNA transcription template DNA, A is adenine deoxyribonucleotide; C is cytosine deoxyribonucleotide; G is guanine deoxyribonucleotide; T is thymine deoxyribonucleotide; H is A, C or T; W is A or T; K is G or T; S is C or G;

[0022] 5'UTR or 3'UTR contains at least one or more Motif 2: GRSASM, wherein R is A or G; S is C or G; M is A or C;

[0023] 5'UTR or 3'UTR contains at least one or more Motif 3: CAAVMTYA, wherein V is A, C or G; M is A or C; Y is C or T;

[0024] 5'UTR or 3'UTR contains at least one or more Motif 4: GKGMWM, wherein K is G or T; M is A or C; W is A or T;

[0025] 5'UTR or 3'UTR comprises at least one or more Motif 5: AMRCAW, wherein M is A or C; R is A or G; W is A or T;

[0026] 5'UTR or 3'UTR comprises at least one or more Motif 6: KYMMYM, wherein K is G or T; Y is C or T, M is A or C;

[0027] 5'UTR or 3'UTR comprises at least one or more Motif 7: NCATMM, wherein N is A, C, G or T; M is A or C;

[0028] 5'UTR or 3'UTR comprises at least one or more Motif 8: CMWCMN, wherein M is A, C; W is A or T; N is A, C, G or T;

[0029] 5'UTR or 3'UTR comprises at least one or more Motif 9: KSWMHC, wherein K is G or T; S is C or G; W is A or T; M is A or C; H is A, C or T;

[0030] 5'UTR or 3'UTR comprises at least one or more Motif 10: AATAAA;

[0031] 5'UTR or 3'UTR comprises at least one or more Motif 11: CSTKGWGC, wherein S is C or G; K is G or T; W is A or T;

[0032] 5'UTR or 3'UTR comprises at least one or more Motif 12: TTCTTK, wherein K is G or T;

[0033] 5'UTR or 3'UTR comprises at least one or more Motif 13: CTGDGKR, wherein D is A, G or T; K is G or T; R is A or G;

[0034] 5'UTR or 3'UTR comprises at least one or more Motif 14: MTGSAT, wherein M is A or C; S is C or G;

[0035] 5'UTR or 3'UTR comprises at least one or more Motif 15: CCAASCMS, wherein S is C or G; M is A or C;

[0036] 5'UTR or 3'UTR comprises at least one or more Motif 16: MCTCKK, wherein M is A or C; K is G or T;

[0037] 5'UTR or 3'UTR comprises at least one or more Motif 17: AACMTWWA, wherein M is A or C; W is A or T;

[0038] 5'UTR or 3'UTR comprises at least one or more Motif 18: TCAWTK, wherein W is A or T; K is G or T;

[0039] 5'UTR or 3'UTR comprises at least one or more Motif 19: GAAGSA, wherein S is C or G;

[0040] 5'UTR comprises one classical kozak sequence, which is GCCRCC, wherein R can be G or A, and preferably A;

[0041] 3'UTR comprises at least one or more ARE: AUUUA;

[0042] 5'UTR has a thermodynamic free energy of -9 to -11 kcal / mol;

[0043] 5'UTR has a minimum free energy of -9 to -11 kcal / mol;

[0044] 5'UTR MFE secondary structure has a stem number of 2-3;

[0045] 5'UTR MFE secondary structure has a total paired base number of 10-13;

[0046] 5'UTR MFE secondary structure has a distance of 1-8 between the hairpin and the Cap;

[0047] 5'UTR centroid method secondary structure has a stem number of 1-3;

[0048] 5'UTR centroid method secondary structure has a total paired base number of 7-16;

[0049] 5'UTR centroid method secondary structure has a distance of 4-20 between the hairpin and the Cap;

[0050] 3'UTR has a thermodynamic free energy of -5 to -100 kcal / mol;

[0051] 3'UTR has a MEF of -5 to -100 kcal / mol;

[0052] 3'UTR MFE secondary structure has a stem number of 1-20;

[0053] 3'UTR MFE secondary structure has a total paired base number of 20-100;

[0054] 3'UTR centroid method secondary structure has a stem number of 1-8;

[0055] 3'UTR Centroid method secondary structure stem total number of paired bases is 10-60;

[0056] 5'UTR and CDS number of paired bases is 30-40;

[0057] 5'UTR and 3'UTR number of paired bases is 20-30.

[0058] In some preferred embodiments of the present application, the differential features in step (2) include one or more of the following: 5'UTR length, 3'UTR GC content, 3'UTR C percentage, 3'UTR T percentage, Motif 3 similarity, Motif 4 similarity, Motif 6 similarity, Motif 7 similarity, 5'UTR and CDS number of paired bases, and / or 5'UTR and 3'UTR number of paired bases.

[0059] In some specific embodiments of the present application, the evaluation in step (3) requires that the UTR combination satisfies one or more of the following differential features: 5'UTR length is 40-60 nt; 3'UTR GC content is 25-70%; 3'UTR C percentage is 20-50%; 3'UTR T percentage is 18-35%; the UTR combination contains one or more Motif 3; the UTR combination contains one or more Motif 4, the UTR combination contains one or more Motif 6, the UTR combination contains one or more Motif 7, 5'UTR and CDS number of paired bases is 20-30 nt; and / or 5'UTR and 3'UTR number of paired bases is 20-30 nt.

[0060] And / or, the evaluation in step (3) uses the following formula: total protein level = 4.91759 + 0.913853 x 5'UTR length - 975.05415 x Motif 3 similarity + 966.75783 x Motif 4 similarity - 4.54893 x Motif 6 similarity + 8.33044 x Motif 7 similarity - 0.034045 x 3'UTR GC content + 0.060051 x 3'UTR C percentage + 0.004150 x 3'UTR T percentage - 1.88234 x 5'UTR and CDS number of paired bases + 0.018206 x 5'UTR and 3'UTR number of paired bases.

[0061] The second aspect of the present application provides a method for constructing a model for evaluating a UTR combination, the UTR combination comprising a 5'UTR and a 3'UTR, characterized in that the method comprises the following steps:

[0062] (1) Grouping of UTR combination: obtaining protein expression information of expression elements containing different UTR combinations, and grouping the UTR combinations into an excellent group, a general group and a poor group;

[0063] (2) Selection of differential characteristics: in the excellent group and the poor group as described in step (1), the characteristics with p<0.05 are obtained from the candidate characteristics as the differential characteristics for constructing the evaluation model by the t-test method;

[0064] (3) Construction of evaluation model: taking the differential characteristics described in step (2) as the independent variable and the protein expression information described in step (1) as the dependent variable for response surface analysis, an evaluation model of UTR combination is obtained.

[0065] In some embodiments of the present application, the candidate characteristics and the differential characteristics are as defined in the method of the first aspect; and / or, the evaluation model is a two-factor model or a linear model, preferably a linear fitting model.

[0066] The third aspect of the present application provides a model for evaluating UTR combination, which is constructed by the method of the second aspect.

[0067] In some embodiments of the present application, the model comprises the following formula: total protein level = 4.91759 + 0.913853 x 5'UTR length - 975.05415 x Motif3 similarity + 966.75783 x Motif4 similarity - 4.54893 x Motif6 similarity + 8.33044 x Motif 7 similarity - 0.034045 x 3'UTR GC content + 0.060051 x 3'UTR C percentage + 0.004150 x 3'UTR T percentage - 1.88234 x 5'UTR and CDS pairing base number + 0.018206 x 5'UTR and 3'UTR pairing base number.

[0068] The fourth aspect of the present application provides a system for screening UTR combination, which comprises the following modules:

[0069] An input module for inputting sample data to be evaluated, wherein the sample data to be evaluated comprises the numerical value of evaluation index, and the evaluation index is selected from the candidate characteristics or the differential characteristics as defined in the method of the first aspect;

[0070] An analysis module, which obtains an analysis result through the sample data to be evaluated; wherein the analysis module comprises the model of the third aspect, and when the sample data to be evaluated meets the judgment condition, the analysis result is output as "compliance"; when the sample data to be evaluated does not meet the judgment condition, the analysis result is output as "non-compliance"; and,

[0071] a judging module, configured to judge whether the sample to be evaluated is suitable for being combined as a UTR for improving translation efficiency in cells according to the sample to be evaluated, and output a judging result; when the analysis result is "conform", the judging result is "suitable"; when the analysis result is "not conform", the judging result is "not suitable".

[0072] The fifth aspect of the present application provides a readable medium, which stores a program, and the program is executed by a processor to realize the function of the system according to the fourth aspect.

[0073] The sixth aspect of the present application provides a device for screening a UTR combination, which comprises:

[0074] (1) the readable medium according to the fifth aspect;

[0075] (2) a processor, configured to execute the program to realize the function of the system according to the fourth aspect;

[0076] In some embodiments of the present application, the device further comprises an output device, configured to output the judging result.

[0077] The seventh aspect of the present application provides a UTR combination, which comprises a 5'UTR and a 3'UTR, the 5'UTR comprises a sequence selected from SEQ ID NO: 1-6, and the 3'UTR comprises a sequence selected from SEQ ID NO: 7-12.

[0078] When the 5'UTR comprises the sequence of SEQ ID NO: 2, the 3'UTR does not comprise the sequence of SEQ ID NO: 8;

[0079] When the 5'UTR comprises the sequence of SEQ ID NO: 3, the 3'UTR does not comprise the sequence of SEQ ID NO: 11;

[0080] When the 5'UTR comprises the sequence of SEQ ID NO: 4, the 3'UTR does not comprise the sequence of SEQ ID NO: 7 or SEQ ID NO: 12;

[0081] When the 5'UTR comprises the sequence of SEQ ID NO: 5, the 3'UTR does not comprise the sequence of SEQ ID NO: 11 or SEQ ID NO: 12;

[0082] when the 5' UTR comprises a sequence as set forth in SEQ ID NO: 6, the 3' UTR does not comprise a sequence as set forth in SEQ ID NO: 8 or SEQ ID NO: 10-12.

[0083] In some embodiments of the present application, the 5' UTR is selected from a sequence as set forth in SEQ ID NO: 1-3, and the 3' UTR comprises a sequence selected from a sequence as set forth in SEQ ID NO: 7-10 or SEQ ID NO: 12.

[0084] In some specific embodiments of the present application, the sequence of the 5' UTR is as set forth in SEQ ID NO: 1, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 8; or,

[0085] the sequence of the 5' UTR is as set forth in SEQ ID NO: 1, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 10; or,

[0086] the sequence of the 5' UTR is as set forth in SEQ ID NO: 1, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 12; or,

[0087] the sequence of the 5' UTR is as set forth in SEQ ID NO: 2, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 7; or,

[0088] the sequence of the 5' UTR is as set forth in SEQ ID NO: 2, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 9; or,

[0089] the sequence of the 5' UTR is as set forth in SEQ ID NO: 2, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 10; or,

[0090] the sequence of the 5' UTR is as set forth in SEQ ID NO: 3, and the sequence of the 3' UTR is as set forth in SEQ ID NO: 9.

[0091] In a eighth aspect of the present application, there is provided an expression element comprising a promoter and a UTR combination as described in the seventh aspect.

[0092] In a ninth aspect of the present application, there is provided a gene expression cassette comprising an expression element as described in the eighth aspect.

[0093] In some embodiments of the present application, the gene expression cassette further comprises a gene of interest and a terminator.

[0094] The tenth aspect of the present application provides a recombinant expression vector comprising the UTR combination according to the seventh aspect, the expression element according to the eighth aspect, or the expression cassette according to the ninth aspect.

[0095] The eleventh aspect of the present application provides a transformant comprising the expression cassette according to the ninth aspect, or the recombinant expression vector according to the tenth aspect.

[0096] In some embodiments of the present application, the host cell of the transformant is a mammalian cell, for example, a 293T cell, a Hela cell, an A549 cell, a Vero cell, a HuH-7 cell, or a DC2.4 cell.

[0097] The twelfth aspect of the present application provides use of the UTR combination according to the seventh aspect, the expression element according to the eighth aspect, the gene expression cassette according to the ninth aspect, the recombinant expression vector according to the tenth aspect, or the transformant according to the eleventh aspect in improving protein translation efficiency.

[0098] The positive progress effect of the present application is that the method for screening the UTR combination disclosed in the present application analyzes the UTR combination from 49 indexes, including primary structure, special sequence or element, secondary structure, and 19 motif sequences which affect cell translation efficiency are proposed for the first time. The method provided in the present application can simplify the UTR design and screening process, and obtain the UTR combination which is widely applicable to improving cell translation efficiency in organisms through indexes and formulas. The UTR combination obtained by screening in the present application has good translation efficiency in multiple cells. The UTR combination can quickly and effectively help design RNA sequences with high-efficiency translation in the field of cell therapy or mRNA vaccine. BRIEF DESCRIPTION OF DRAWINGS

[0099] FIG. 1A and FIG. 1B are protein translation levels of luciferase mRNA containing different UTR sequences in 293T cells transfected for 24 h (relative to p02-20, i.e., 5U2_3U2). FIG. 1A and FIG. 1B are classified analysis according to 5'UTR and 3'UTR, respectively.

[0100] FIG. 2A and FIG. 2B are protein translation levels of luciferase mRNA containing different UTR sequences in Hela cells transfected for 24 h (relative to p02-20, i.e., 5U2_3U2). FIG. 2A and FIG. 2B are classified analysis according to 5'UTR and 3'UTR, respectively.

[0101] Figures 3A and 3B are the protein translation levels of luciferase mRNA containing different UTR sequences transfected in A549 cells for 24h (relative to p02-20, i.e. 5U2_3U2). Figures 3A and 3B are analyzed according to 5'UTR and 3'UTR respectively.

[0102] Figures 4A and 4B are the protein translation levels of luciferase mRNA containing different UTR sequences transfected in Vero cells for 24h (relative to p02-20, i.e. 5U2_3U2). Figures 4A and 4B are analyzed according to 5'UTR and 3'UTR respectively.

[0103] Figures 5A and 5B are the protein translation levels of luciferase mRNA containing different UTR sequences transfected in HuH-7 cells for 24h (relative to p02-20, i.e. 5U2_3U2). Figures 5A and 5B are analyzed according to 5'UTR and 3'UTR respectively.

[0104] Figures 6A and 6B are the protein translation levels of luciferase mRNA containing different UTR sequences transfected in DC2.4 cells for 24h (relative to p02-20, i.e. 5U2_3U2). Figures 6A and 6B are analyzed according to 5'UTR and 3'UTR respectively.

[0105] Figures 7A and 7B are the protein translation levels of luciferase mRNA transfected in 293T, Hela, A549, Vero, HuH-7 DC2.4 cells for 24h (relative to p02-20, i.e. 5U2_3U2) respectively. Figure 7A is the experimental group, i.e. containing different UTR sequences; Figure 7B is the negative control group (compared with the control group, p02-49 removes 5U2, p02-50 removes 3U2, p02-51 removes 5U2 and 3U2 at the same time).

[0106] Figure 8A-8C are the parameter evaluation. Figure 8A, 8B and 8C are the relevant parameter evaluation of UTR primary structure, specific sequence or element, and secondary structure, respectively. Among them, Max8 represents the 8 UTR combinations with the highest translation efficiency improvement effect on multiple cell proteins, i.e. p02-14 (5U1_3U2), p02-15 (5U1_3U3), p02-16 (5U1_3U4), p02-18 (5U1_3U6), p02-19 (5U2_3U1), p02-21 (5U2_3U3), p02-22 (5U2_3U4) and p02-27 (5U3_3U3); Min8 represents the 8 UTR combinations with the weakest translation efficiency improvement effect on multiple cell proteins, i.e. p02-29 (5U3_3U5), p02-36 (5U4_3U6), p02-41 (5U5_3U5), p02-42 (5U5_3U6), p02-44 (5U6_3U2), p02-46 (5U6_3U4), p02-47 (5U6_3U5) and p02-48 (5U6_3U6). *, **, ***, and **** represent p≤0.05, p≤0.01, p≤0.001 and p≤0.0001, respectively.

[0107] Figure 9 is a schematic diagram of the p-value value of the UTR sequence in the UTR combination to be evaluated and the Motif X (X = 1-19) sequence obtained by the meme software, which is used as the evaluation value of the "Motif X similarity" (X = 1-19) in the present application. DETAILED DESCRIPTION

[0108] Experimental materials

[0109] 1. Equipment and instruments

[0110] Vertical transparent door refrigerator (Qingdao Aokema Biomedical Equipment Co., Ltd., SC-387NE), low temperature storage box (Qingdao Aokema Biomedical Equipment Co., Ltd., DW-25W525), -86°C upright ultra-low temperature refrigerator (Thermo Scientific, FORMA 900 SERIES), A2 type secondary biological safety cabinet (Esco Airstream, AC2-4S1), carbon dioxide incubator (BIOBASE, QP-160), table centrifuge (Hunan Kaida Scientific Instrument Co., Ltd., TGL20M-II), multifunctional enzyme label instrument (BioTek, SYNERGY H1).

[0111] 2. Experimental reagents

[0112] DMEM medium (Gibco, 11965118), penicillin-streptomycin (Gibco, 15140122), trypsin-EDTA (Gibco, 25200056), fetal bovine serum (Gibco, A3160902), Lipofectamine TM MessengerMAX TM Reagent (Invitrogen, LMRNA015), Fire-Lumi TM Luciferase Assay System (Promega, E1980), Luciferase Assay Reagent (Invitrogen, LMRNA015), Fire-Lumi

[0113] 3. Cell lines

[0114] 293T, Hela, A549, Vero and HuH-7 cells were purchased from the China Academy of Sciences Typical Culture Preservation Committee Cell Library; DC2.4 cells were purchased from the Shanghai Cell Library.

[0115] Experimental description

[0116] 1. Abbreviation list

[0117] Table 1. Abbreviation list

[0118] 2. UTR sequence reference

[0119] Table 2. 5' UTR sequence reference

[0120] Table 3. 3' UTR sequence reference

[0121] 3. Experimental group plasmid UTR sequence type

[0122] Table 4. Experimental group plasmid UTR sequence type list

[0123] 4. Negative control group plasmid UTR sequence type

[0124] Table 5. Negative control group plasmid UTR sequence

[0125] 5. Sequence

[0126] The Luciferase reporter gene sequence in the present application is shown in SEQ ID NO: 13.

[0127] Experimental method

[0128] (1) Each cell was seeded at a certain density (293T: 2.0 x 10 4cells / well; Hela: 1.0 x 10 4 cells / well; A549: 1.5 x 10 4 cells / well; Vero: 1.0 x 10 4 cells / well; HuH-7: 2.0 x 10 4 cells / well; DC2.4: 2.5 x 10 4 cells / well) were seeded in 96-well plates and incubated at 37°C, 5% CO2 overnight.

[0129] (2) Transfection was performed according to the transfection kit (Invitrogen, LMRNA015) by adding 0.15 μL Lipofectamine TM MessengerMAX TM Reagent and 100 ng mRNA (n = 3) per well, and incubated in a carbon dioxide incubator.

[0130] (3) After 24 h, the culture medium was removed, and a luciferin mixed solution (luciferin solution: DMEM = 1:1) was added at a volume of 60 μL / well.

[0131] (4) After the cells were sufficiently lysed (3 min), 50 μL of the mixed solution was transferred to a 96-well white flat-bottom plate, and the chemiluminescence signal intensity was detected using a multifunctional enzyme label meter to evaluate the mRNA sample transfection effect.

[0132] Example 1 Influence of UTR sequences on translation efficiency in different cells

[0133] By comparing with the widely used 5'UTR (5U2) and 3'UTR (3U2) of endogenous beta globin, the influence of UTR combinations on the protein expression levels of various cells, including 293T, Hela, A549, Vero, Huh-7, and DC2.4 cells, was studied to find the UTR combination characteristics for improving cell translation efficiency.

[0134] The UTR sequences in the present application are as follows:

[0135] Table 6 UTR sequences

[0136] 1. Influence of different UTR sequences on protein expression in 293T cells

[0137] 5'UTR and 3'UTR on the expression of luciferase in 293T cells. Among the 36 combinations of 5'UTR and 3'UTR, p02-21 (5U2_3U3) increased the expression of luciferase in 293T cells by 1.46-fold (relative to the control group p02-20 (5U2_3U2)). p02-19 (5U2_3U1) and p02-22 (5U2_3U4) followed, which increased the expression of luciferase mRNA by 1.41 and 1.33-fold, respectively, compared with the control group p02-20 (5U2_3U2).

[0138] 2. Different UTR sequences affect protein expression in Hela cells

[0139] 5'UTR and 3'UTR on the expression of luciferase in 293T cells. Among the 36 combinations of 5'UTR and 3'UTR, p02-21 (5U2_3U3) increased the expression of luciferase in 293T cells by 1.46-fold (relative to the control group p02-20 (5U2_3U2)). p02-19 (5U2_3U1) and p02-22 (5U2_3U4) followed, which increased the expression of luciferase mRNA by 1.41 and 1.33-fold, respectively, compared with the control group p02-20 (5U2_3U2).

[0140] 3. Different UTR sequences affect protein expression in A549 cells

[0141] 5'UTR and 3'UTR on the expression of luciferase in 293T cells. Among the 36 combinations of 5'UTR and 3'UTR, p02-21 (5U2_3U3) increased the expression of luciferase in 293T cells by 1.46-fold (relative to the control group p02-20 (5U2_3U2)). p02-19 (5U2_3U1) and p02-22 (5U2_3U4) followed, which increased the expression of luciferase mRNA by 1.41 and 1.33-fold, respectively, compared with the control group p02-20 (5U2_3U2).

[0142] 4. Different UTR sequences affect protein expression in Vero cells

[0143] 5. UTR sequences affect protein expression in Vero cells

[0144] 5. UTR sequences affect protein expression in HuH-7 cells

[0145] 5. UTR sequences affect protein expression in HuH-7 cells

[0146] 6. UTR sequences affect protein expression in DC2.4 cells

[0147] 6. UTR sequences affect protein expression in DC2.4 cells

[0148] 7. Cell line preference for UTR sequences

[0149] Different UTR sequences affect the expression of luciferase in each cell line to some extent. Among them, the UTR sequence has the greatest impact on the protein translation efficiency of A549 cells (Figure 7A). The presence of UTR has a significant effect on improving cell translation efficiency. Among them, after removing the 5' UTR of 293T, Hela, A549, Vero, HuH-7 and DC2.4 cells, the expression of luciferase decreased by 30.00%, 71.36%, 47.89%, 10.52%, 56.58% and 55.26%, respectively. While removing the 3' UTR, the protein expression level decreased by 38.92%, 38.36%, 18.47%, 25.47%, 43.03% and 23.22%, respectively (Figure 7B). The negative control group plasmid UTR sequence is shown in Table 5.

[0150] In various cell lines, p02-14 (5U1_3U2), p02-15 (5U1_3U3), p02-16 (5U1_3U4), p02-18 (5U1_3U6), p02-19 (5U2_3U1), p02-21 (5U2_3U3), p02-22 (5U2_3U4) and p02-27 (5U3_3U3) all obtained high levels of protein expression (Figure 7A). However, in addition to the negative control group (remove 5'UTR or / and 3'UTR), p02-29 (5U3_3U5), p02-36 (5U4_3U6), p02-41 (5U5_3U5), p02-42 (5U5_3U6), p02-44 (5U6_3U2), p02-46 (5U6_3U4), p02-47 (5U6_3U5) and p02-48 (5U6_3U6) UTR combinations have a negative effect on the improvement of protein translation efficiency in various cell lines such as 293T, Hela, A549, Vero, HuH-7 and DC2.4.

[0151] Example 2 UTR combination characteristics for improving cell translation efficiency

[0152] The present embodiment analyzes the UTR combination through 49 characteristic indexes, including but not limited to 5'UTR length, 5'UTR GC content, 5'UTR G percentage, 5'UTR C percentage, 5'UTR A percentage, 5'UTR T percentage, 5'UTR thermodynamic free energy, 5'UTR minimum free energy (MFE), 5'UTR MFE secondary structure (5'UTR MFE secondary structure, 5MSS) stem number, 5MSS total paired base number of stem, 5MSS hairpin structure and Cap distance, 5'UTR centroid secondary structure (5'UTR Centroid secondary structure, 5CSS) stem number, 5CSS total paired base number of stem, 5CSS hairpin structure and Cap distance, presence or absence of classic kozak sequence, similarity with Motif1, Motif2, Motif3, Motif4, Motif5, Motif6, Motif7, Motif8, Motif9, Motif10, Motif11, Motif12, Motif13, Motif14, Motif15, Motif16, Motif17, Motif18 and Motif19, 3'UTR length, 3'UTR GC content, 3'UTR G percentage, 3'UTR C percentage, 3'UTR A percentage, 3'UTR T percentage, 3'UTR thermodynamic free energy, 3'UTR MEF, 3MSS stem number, 3MSS total paired base number of stem, 3CSS stem number, 3CSS total paired base number of stem, ARE (AU-rich elements) element, 5'UTR and CDS paired base number and 5'UTR and 3'UTR paired base number, etc.

[0153] The present application first proposes the influence of 19 special sequences on cell translation efficiency, which are as follows:

[0154] Motif 1: THWKCTTCTS (SEQ ID NO: 14). Among them, in the mRNA transcription template DNA, A is adenine deoxyribonucleotide; C is cytosine deoxyribonucleotide; G is guanine deoxyribonucleotide; T is thymine deoxyribonucleotide; H can be A, C or T; W can be A or T; K can be G or T; S can be C or G.

[0155] Motif 2: GRSASM. Among them, R can be A or G; S can be C or G; M can be A or C.

[0156] Motif 3: CAAVMTYA. Among them, V can be A, C or G; M can be A or C; Y can be C or T.

[0157] Motif 4: GKGMWM. Wherein, K can be G or T; M can be A or C; W can be A or T.

[0158] Motif 5: AMRCAW. Wherein, M can be A or C; R can be A or G; W can be A or T.

[0159] Motif 6: KYMMYM. Wherein, K can be G or T; Y can be C or T, M can be A or C.

[0160] Motif 7: NCATMM. Wherein, N can be A, C, G or T; M can be A or C.

[0161] Motif 8: CMWCMN. Wherein, M can be A, C; W can be A or T; N can be A, C, G or T.

[0162] Motif 9: KSWMHC. Wherein, K can be G or T; S can be C or G; W can be A or T; M can be A or C; H can be A, C or T.

[0163] Motif 10: AATAAA.

[0164] Motif 11: CSTKGWGC. Wherein, S can be C or G; K can be G or T; W can be A or T.

[0165] Motif 12: TTCTTK. Wherein, K can be G or T.

[0166] Motif 13: CTGDGKR. Wherein, D can be A, G or T; K can be G or T; R can be A or G.

[0167] Motif 14: MTGSAT. Wherein, M can be A or C; S can be C or G.

[0168] Motif 15: CCAASCMS. Wherein, S can be C or G; M can be A or C.

[0169] Motif 16: MCTCKK. Wherein, M can be A or C; K can be G or T.

[0170] Motif 17: AACMTWWA. Wherein, M can be A or C; W can be A or T.

[0171] Motif 18: TCAWTK. Wherein, W can be A or T; K can be G or T.

[0172] Motif 19: GAAGSA. Wherein, S can be C or G.

[0173] In the above Motifs, Motifs 1-9 are sequences in the 5' UTR, and Motifs 10-19 are sequences in the 3' UTR. The Motif sequences are obtained from the meme software prediction analysis according to the UTR sequences (5U1-5U6, 3U1-3U6).

[0174] The eight UTR combinations with the highest enhancement effect on the translation efficiency of various cellular proteins (Max8; including p02-14 (5U1_3U2), p02-15 (5U1_3U3), p02-16 (5U1_3U4), p02-18 (5U1_3U6), p02-19 (5U2_3U1), p02-21 (5U2_3U3), p02-22 (5U2_3U4) and p02-27 (5U3_3U3)) and the eight UTR combinations with the weakest enhancement effect (Min8; including p02-29 (5U3_3U5), p02-36 (5U4_3U6), p02-41 (5U5_3U5), p02-42 (5U5_3U6), p02-44 (5U6_3U2), p02-46 (5U6_3U4), p02-47 (5U6_3U5) and p02-48 (5U6_3U6)) were respectively evaluated for the primary structure (results shown in FIG. 8A), specific sequence or element (results shown in FIG. 8B) and secondary structure (results shown in FIG. 8C).wherein the primary structure related feature indicators include but are not limited to 5'UTR length, 5'UTR GC content, 5'UTR G percentage, 5'UTR C percentage, 5'UTR A percentage, 5'UTR T percentage, 3'UTR length, 3'UTR GC content, 3'UTR G percentage, 3'UTR C percentage, 3'UTR A percentage, 3'UTR T percentage; the specific sequence or element includes but is not limited to Motif 1, Motif 2, Motif 3, Motif 4, Motif 5, Motif 6, Motif 7, Motif 8, Motif 9, Motif 10, Motif 11, Motif 12, Motif 13, Motif 14, Motif 15, Motif 16, Motif 17, Motif 18, Motif 19, classic kozak sequence and ARE (AU-rich elements) element; the secondary structure related feature indicators include but are not limited to 5'UTR thermodynamic free energy, 5'UTR Minimum Free Energy (MFE), 5'UTR MFE secondary structure (5MSS) stem number, 5MSS stem total paired base number, 5MSS hairpin structure distance to Cap, 5'UTR Centroid secondary structure (5CSS) stem number, 5CSS stem total paired base number, 5CSS hairpin structure distance to Cap, 3'UTR thermodynamic free energy, 3'UTR MEF, 3MSS stem number, 3MSS stem total paired base number, 3CSS stem number, 3CSS stem total paired base number, 5'UTR paired base number with CDS, 5'UTR paired base number with 3'UTR. Among them, the Motif 1-19 similarity takes the significance of its actual sequence in the UTR sequence as the evaluation, which is the p-value obtained by the meme software (for example, see FIG. 9), the smaller the value, the higher the similarity of the actual sequence and the corresponding motif, that is, the negative correlation, and the value of 1 indicates the lack of the corresponding Motif. The indicators of whether there is a classic kozak sequence and ARE (AU-rich elements) element are quantified as 1 and 0 in the test process.

[0175] T-test results showed that Max8 and Min8 had significant differences in 10 characteristic indexes, including 5'UTR length, 3'UTR GC content, 3'UTR C percentage, 3'UTR T percentage, similarity to Motif 3, Motif 4, Motif 6 and Motif 7, 5'UTR and CDS pairing base number, and 5'UTR and 3'UTR pairing base number (p<0.05; FIGS. 8A-8C and Table 7). Since the values of Motif 10 in Min8 and Max8 were consistent, t-test could not be performed, so there was no value in Table 7. Since the protein expression levels in the six different cells in Example 1 were not normal, the median was taken as the dependent variable, and the above 10 indexes of Max8 and Min8 that had significant differences were taken as the covariates, and Design Expert was used for response surface analysis. Among them, the linear model p<0.0001, the 2FI model p=0.4849, so the linear fitting was enabled. The linear model (determination coefficient Adjusted R 2 =0.7649) is: total protein level (relative to 5U2_3U2) = 4.91759 + 0.913853x5'UTR length - 975.05415xMotif 3 similarity + 966.75783xMotif 4 similarity - 4.54893xMotif 6 similarity + 8.33044xMotif 7 similarity - 0.034045x3'UTR GC content + 0.060051x3'UTR C percentage + 0.004150x3'UTR T percentage - 1.88234x5'UTR and CDS pairing base number + 0.018206x5'UTR and 3'UTR pairing base number.

[0176] Table 7 T-test results of Max8 and Min8 indexes

[0177] Although the specific embodiments of the present application are described above, those skilled in the art should understand that these are only illustrative, and various changes or modifications can be made to these embodiments without departing from the principles and essence of the present application. Therefore, the protection scope of the present application is defined by the appended claims.

Claims

1. A method for designing and screening UTR combinations, wherein the UTR combination includes a 5' UTR and a 3' UTR, characterized in that, The method includes the following steps: (1) Grouping UTR combinations: Obtain protein expression information of expression elements containing different UTR combinations, and divide the UTR combinations into excellent group, average group and poor group; (2) Selection of differential features: In the excellent group and the poor group as described in step (1), the features with p<0.05 are obtained from the candidate features by t test as differential features for evaluating UTR combinations. (3) Evaluation of UTR combinations: The UTR combinations are evaluated based on the difference characteristics described in step (2).

2. The method as described in claim 1, characterized in that, In step (2), the candidate features include one or more of the following: 5'UTR length, 5'UTR GC content, 5'UTR G percentage, 5'UTR C percentage, 5'UTR A percentage, 5'UTR T percentage, 5'UTR thermodynamic free energy, 5'UTR minimum free energy, 5'UTR RFE secondary structure stem number, 5MSS stem total paired base number, 5MSS hairpin structure to Cap distance, 5'UTR centroid method secondary structure stem number, 5CSS stem total paired base number, 5CSS hairpin structure to Cap distance, presence or absence of classical Kozak sequence, Motif 1 similarity, Motif 2 similarity, Motif 3 similarity, Motif 4 similarity, Motif 5 similarity, Motif 6 similarity, Motif 7 similarity, Motif 8 similarity, Motif 9 similarity, 3'UTR length, 3'UTR GC content, 3'UTR G percentage, 3'UTR C percentage, 3'UTR A percentage, 3'UTR T percentage, 3'UTR thermodynamic free energy, 3'UTR MEF, 3MSS stem number, 3MSS stem total number of paired bases, 3CSS stem number, 3CSS stem total number of paired bases, ARE element, Motif 10 similarity, Motif 11 similarity, Motif 12 similarity, Motif 13 similarity, Motif 14 similarity, Motif 15 similarity, Motif 16 similarity, Motif 17 similarity, Motif 18 similarity, Motif 19 similarity, 5'UTR to CDS paired bases and 5'UTR to 3'UTR paired bases; The nucleotide sequence of Motif1 is shown in SEQ ID NO: 14; The nucleotide sequence of the Motif 2 is GRSASM, where R is A or G; S is C or G; and M is A or C. The nucleotide sequence of the Motif 3 is CAAVMTYA, where V is A, C, or G; M is A or C; and Y is C or T. The nucleotide sequence of the Motif 4 is GKGMWM, where K is G or T; M is A or C; and W is A or T. The nucleotide sequence of Motif 5 is AMRCAW, where M is A or C; R is A or G; W is A or T; and the nucleotide sequence of Motif 6 is KYMMYM, where K is G or T; Y is C or T; M is A or T. It can be A or C; The nucleotide sequence of the Motif7 is NCATMM, where N is A, C, G or T; and M is A or C. The nucleotide sequence of the Motif 8 is CMWCMN, where M is A or C; W is A or T; and N is A, C, G, or T. The nucleotide sequence of Motif9 is KSWMHC, where K is G or T; S is C or G; W is A or T; M is A or C; and H is A, C, or T. The nucleotide sequence of Motif 10 is AATAAA; The nucleotide sequence of Motif 11 is CSTKGWGC, where S is C or G; K is G or T; and W is A or T. The nucleotide sequence of Motif12 is TTCTTK, where K is G or T; The nucleotide sequence of Motif13 is CTGDGKR, where D is A, G, or T; K is G or T; and R is A or G. The nucleotide sequence of Motif 14 is MTGSAT, where M is A or C; S is C or G; The nucleotide sequence of Motif 15 is CCAASCMS, where S is C or G; M is A or C; The nucleotide sequence of Motif 16 is MCTCKK, where M is A or C; K is G or T; The nucleotide sequence of Motif 17 is AACMTWWA, where M is A or C; W is A or T; The nucleotide sequence of Motif18 is TCAWTK, where W is A or T; K is G or T; and / or, The nucleotide sequence of Motif19 is GAAGSA, where S is C or G; Preferably, the differential features described in step (2) include one or more of the following: 5'UTR length, 3'UTR GC content, 3'UTR C percentage, 3'UTR T percentage, Motif 3 similarity, Motif 4 similarity, Motif 6 similarity, Motif 7 similarity, number of base pairs between 5'UTR and CDS, and / or number of base pairs between 5'UTR and 3'UTR; More preferably, the evaluation in step (3) requires the UTR combination to meet one or more of the following differential characteristics: the 5'UTR length is 40-60 nt; the 3'UTR GC content is 25-70%; the 3'UTR C percentage is 20-50%; the 3'UTR T percentage is 18-35%; the UTR combination contains one or more Motif 3; the UTR combination contains one or more Motif 4; the UTR combination contains one or more Motif 6; the UTR combination contains one or more Motif 7; the number of bases paired between the 5'UTR and CDS is 20-30 nt; and / or the number of bases paired between the 5'UTR and 3'UTR is 20-30 nt. And / or, the evaluation described in step (3) adopts the following formula: Total protein level = 4.91759 + 0.913853 × 5'UTR length - 975.05415 × Motif 3 similarity + 966.75783 × Motif 4 similarity - 4.54893 × Motif 6 similarity + 8.33044 × Motif 7 similarity - 0.034045 × 3'UTR GC content + 0.060051 × 3'UTR C percentage + 0.004150 × 3'UTR T percentage - 1.88234 × 5'UTR to CDS base pairings + 0.018206 × 5'UTR to 3'UTR base pairings.

3. A method for constructing a model to evaluate a combination of UTRs, wherein the combination of UTRs includes a 5' UTR and a 3' UTR, characterized in that, The method includes the following steps: (1) Grouping UTR combinations: Obtain protein expression information of expression elements containing different UTR combinations, and divide the UTR combinations into excellent group, average group and poor group; (2) Selection of difference features: In the excellent group and the poor group as described in step (1), the features with p<0.05 are obtained from the candidate features by t test as difference features for constructing the evaluation model. (3) Construct an evaluation model: Using the differential characteristics described in step (2) as independent variables and the protein expression information described in step (1) as dependent variables, a response surface analysis is performed to obtain an evaluation model of the UTR combination. Preferably, the candidate features and the difference features are as defined in the method of claim 2; and / or, the evaluation model is a two-factor model or a linear model, preferably a linear fitting model.

4. A model for evaluating UTR combinations, characterized in that, The model is constructed by the method described in claim 3; Preferably, the model includes the following formula: Total protein level = 4.91759 + 0.913853 × 5'UTR length - 975.05415 × Motif 3 similarity + 966.75783 × Motif 4 similarity - 4.54893 × Motif 6 similarity + 8.33044 × Motif 7 similarity - 0.034045 × 3'UTR GC content + 0.060051 × 3'UTR C percentage + 0.004150 × 3'UTR T percentage - 1.88234 × number of base pairs between 5'UTR and CDS + 0.018206 × number of base pairs between 5'UTR and 3'UTR.

5. A system for screening UTR combinations, characterized in that, The system includes the following modules: An input module is used to input sample data to be evaluated, the sample data to be evaluated including the values ​​of evaluation indicators, the evaluation indicators being selected from candidate features or differential features as defined in the method as described in claim 2. An analysis module, wherein the analysis module obtains analysis results from the sample data to be evaluated; wherein the analysis module includes the model as described in claim 4, wherein when the sample data to be evaluated meets the judgment conditions, the analysis result is output as "compliant"; when the sample data to be evaluated does not meet the judgment conditions, the analysis result is output as "non-compliant"; and, The judgment module determines whether the sample to be evaluated is suitable as a UTR combination to improve translation efficiency in cells based on the sample data to be evaluated, and outputs the judgment result; wherein, when the analysis result is "compliant", the judgment result is "suitable"; when the analysis result is "incompatible", the judgment result is "unsuitable".

6. A readable medium, characterized in that, The readable medium stores a program that, when executed by a processor, enables the functionality of the system as described in claim 5.

7. An apparatus for screening UTR combinations, characterized in that, The device includes: (1) The readable medium as claimed in claim 6; (2) A processor for executing a program to perform the functions of the system as described in claim 5; Preferably, the device further includes an output device for outputting the judgment result.

8. A UTR combination, said UTR combination comprising a 5' UTR and a 3' UTR, characterized in that, The 5'UTR contains sequences selected from those shown in SEQ ID NO: 1-6, and the 3'UTR contains sequences selected from those shown in SEQ ID NO: 7-12; When the 5'UTR contains the sequence shown in SEQ ID NO: 2, the sequence of the 3'UTR does not contain the sequence shown in SEQ ID NO: 8; When the 5'UTR contains the sequence shown in SEQ ID NO: 3, the sequence of the 3'UTR does not contain the sequence shown in SEQ ID NO: 11; When the 5'UTR contains the sequence shown in SEQ ID NO: 4, the sequence of the 3'UTR does not contain the sequence shown in SEQ ID NO: 7 or SEQ ID NO: 12; When the 5'UTR contains the sequence shown in SEQ ID NO: 5, the sequence of the 3'UTR does not contain the sequence shown in SEQ ID NO: 11 or SEQ ID NO: 12; When the 5'UTR contains a sequence as shown in SEQ ID NO: 6, the 3'UTR does not contain a sequence as shown in SEQ ID NO: 8 or SEQ ID NO: 10-12.

9. The UTR combination as described in claim 8, characterized in that, The 5'UTR is selected from sequences shown as SEQ ID NO: 1-3, and the 3'UTR contains sequences selected from sequences shown as SEQ ID NO: 7-10 or SEQ ID NO:

12.

10. The UTR combination as described in claim 7 or 8, characterized in that, The sequence of the 5'UTR is shown in SEQ ID NO: 1, and the sequence of the 3'UTR is shown in SEQ ID NO: 8; or, The sequence of the 5'UTR is shown in SEQ ID NO: 1, and the sequence of the 3'UTR is shown in SEQ ID NO: 10; or, The sequence of the 5'UTR is shown in SEQ ID NO: 1, and the sequence of the 3'UTR is shown in SEQ ID NO: 12; or, The sequence of the 5'UTR is shown in SEQ ID NO: 2, and the sequence of the 3'UTR is shown in SEQ ID NO: 7; or, The sequence of the 5'UTR is shown in SEQ ID NO: 2, and the sequence of the 3'UTR is shown in SEQ ID NO: 9; or, The sequence of the 5'UTR is shown in SEQ ID NO: 2, and the sequence of the 3'UTR is shown in SEQ ID NO: 10; or, The sequence of the 5'UTR is shown in SEQ ID NO: 3, and the sequence of the 3'UTR is shown in SEQ ID NO:

9.

11. An expression element, characterized in that, The expression element comprises a promoter and a UTR combination as described in any one of claims 8-10.

12. A gene expression cassette, characterized in that, The gene expression cassette includes the expression element as described in claim 11; Preferably, the gene expression also includes a target gene and a terminator.

13. A recombinant expression vector, characterized in that, The recombinant expression vector includes the UTR combination as described in any one of claims 8-10, the expression element as described in claim 11, or the expression cassette as described in claim 12.

14. A transformant, characterized in that, The transformant includes the expression cassette as described in claim 12, or the recombinant expression vector as described in claim 13; Preferably, the host cell of the transformant is a mammalian cell, such as 293T cells, HeLa cells, A549 cells, Vero cells, HuH-7 cells, or DC2.4 cells.

15. The use of the UTR combination as described in any one of claims 8-10, the expression element as described in claim 11, the gene expression cassette as described in claim 12, the recombinant expression vector as described in claim 13, or the transformant as described in claim 14 in improving protein translation efficiency.

Citation Information

Patent Citations

  • Method and kit for detecting specific DNA methylation modification site in plant flower organ

    CN104293890A

  • Nucleic acid molecules encoding leptin, compositions and uses

    CN117089548A

  • Discover biological features using composite images

    US20110110569A1

  • UTR molecule for increasing protein expression level

    WO2024109866A1