A method and apparatus for processing an mRNA vaccine design task

By combining a pre-trained Uni-RNA model and a vaccine characteristic prediction model trained on an antigen vaccine dataset with the NSGA-II algorithm and in vitro and in vivo experimental optimization, the problems of long design cycles and low efficiency of mRNA vaccines are solved, and a high-efficiency and quality-assured design process is realized.

CN119479809BActive Publication Date: 2025-11-18SHANGHAI ALGORITHM INNOVATION RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411478355.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-11-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

The current mRNA vaccine design process is characterized by long design cycles, low efficiency, extensive experimental analysis, and reliance on expert experience, resulting in high costs.

Method used

A pre-trained Uni-RNA model was used as the backbone for feature extraction. A vaccine characteristic prediction model was trained using an antigen vaccine dataset. The NSGA-II algorithm was used for multi-objective optimization. The mRNA sequence was optimized by combining in vitro and in vivo experiments, which reduced the number of experiments and improved design efficiency.

Benefits of technology

By using model prediction and process optimization, we can reduce the number of experimental analyses, shorten the design cycle, improve the efficiency of mRNA vaccine design, ensure sequence quality, and overcome the limitations of expert experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479809B_ABST
    Figure CN119479809B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to a kind of mRNA vaccine design task processing method and device, the method comprises: with Uni-RNA model as main stem design a vaccine characteristic prediction model for predicting three characteristics of mRNA vaccine, and train model based on antigen vaccine data set, and based on the maximum correlation between codon usage frequency, secondary structure and GC content and vaccine stability, translation efficiency and immunogenicity design three objective functions;And according to the antigen sequence input by user, mRNA sequence initialization is carried out, and according to the multi-objective optimization mode of NSGA-II algorithm, mRNA sequence set is iteratively optimized by means of vaccine characteristic prediction model and three objective functions and corresponding in vitro / in vivo experimental means, and the corresponding optimized mRNA sequence set is fed back to user. By the application, the sequence richness can be improved, the design quality can be ensured, the design efficiency can be improved, and the design cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a processing method and device for mRNA vaccine design tasks. BACKGROUND

[0002] mRNA vaccines have shown great potential and advantages in the field of immunotherapy, especially in dealing with emerging infectious diseases such as the novel coronavirus (COVID-19). Compared with traditional vaccines, mRNA vaccines have the characteristics of rapid response, flexibility and efficient production. Therefore, mRNA vaccines have received widespread attention and application during the COVID-19 pandemic. However, although mRNA vaccine technology has many advantages, there are still many challenges in designing mRNA sequences with optimal stability (usually characterized by the half-life of mRNA in the body), translation efficiency (usually characterized by the expression amount of target antigen protein) and immunogenicity (usually characterized by antibody titer). The most common challenge is the cost and efficiency problem: each mRNA sequence planned for initial design needs to go through in vivo / in vitro experimental analysis process before proceeding to the next step of screening and optimization, and each mRNA sequence optimized at each step also needs to go through in vivo / in vitro experimental analysis before moving on to the next step; such iterative cycle not only requires a large amount of manpower and material resources to conduct multiple repetitive experiments, but also causes long design cycle and low design efficiency. SUMMARY

[0003] The present application aims at the defects of the prior art, and provides an mRNA vaccine design task processing method and device, electronic equipment and computer readable storage medium. The present application pre-designs a working model for predicting three vaccine properties (stability, translation efficiency and immunogenicity) of mRNA sequences based on a pre-trained Uni-RNA model as a feature extraction backbone, and trains the model based on an antigen vaccine dataset constructed by big data collection. The present application identifies the maximum correlation between the three basic features (codon usage frequency, RNA secondary structure and GC content) and the three vaccine properties (stability, translation efficiency and immunogenicity) that the model can predict, and determines three objective functions (stability objective function, translation efficiency objective function and immunogenicity objective function) based on the identification results. Then in the specific mRNA vaccine design task processing process: the present application first uses an mRNA sequence generation tool to generate the transcription mRNA sequence of the target antigen protein in batches to obtain an initialized mRNA sequence set; then the vaccine property prediction model and the three objective functions are used to iteratively optimize and screen the mRNA sequence set according to the multi-objective optimization method of the NSGA-II algorithm to obtain an optimized set with reduced sequence number; then the stability and translation efficiency of the optimized set are analyzed based on in vitro experiments; if the error between the in vitro experimental analysis result and the model prediction result is large, the model is automatically optimized based on the experimental data, and the mRNA sequence set is iteratively optimized after optimization; if the error between the in vitro experimental analysis result and the model prediction result is controllable, the immunogenicity of the optimized set is further analyzed based on in vivo experiments; if the error between the in vivo experimental analysis result and the model prediction result is large, the model is automatically optimized again based on the experimental data, and the mRNA sequence set is iteratively optimized after optimization; if the error between the in vivo experimental analysis result and the model prediction result is controllable, the current latest optimized set is output as the final design candidate set. The mRNA sequence initialization method of the present application can get rid of the limitations of expert experience and give rich and diverse sequence structures. The vaccine property prediction model of the present application can reduce the number of experimental analysis times, shorten the design cycle and improve the design efficiency. The model training method and multi-objective optimization process of the present application can effectively guarantee the quality of the designed mRNA sequence.

[0004] To achieve the above object, the first aspect of the embodiment of the present application provides an mRNA vaccine design task processing method, which comprises:

[0005] A working model for predicting three vaccine properties of mRNA sequences is designed with a Uni-RNA model as a feature extraction backbone, and is denoted as a corresponding vaccine property prediction model; the three vaccine properties include stability, translation efficiency and immunogenicity; the vaccine property prediction model includes the Uni-RNA model, three basic feature prediction heads, a feature fusion module and three vaccine property prediction heads;

[0006] A large amount of antigen protein and corresponding mRNA vaccine information is collected through a big data collection method to construct a corresponding antigen vaccine data set; and the vaccine property prediction model is trained based on the antigen vaccine data set; and after the model training is completed, the maximum correlation relationship between the three output features of the three basic feature prediction heads and the three output properties of the three vaccine property prediction heads is identified, and the corresponding three objective functions are determined based on the identification results; the three objective functions include a stability objective function, a translation efficiency objective function and an immunogenicity objective function;

[0007] The molecular sequence of a target antigen protein input by a user is received as a corresponding first antigen sequence; and mRNA sequence initialization is performed according to the first antigen sequence to obtain a corresponding first mRNA sequence set; and the first mRNA sequence set is iteratively optimized by means of the vaccine property prediction model, the three objective functions and corresponding in vitro / in vivo experimental methods in a multi-objective optimization manner of the NSGA-II algorithm to obtain a corresponding optimized mRNA sequence set, which is fed back to the user; the first mRNA sequence set is composed of multiple first mRNA sequences; and the optimized mRNA sequence set is composed of one or more optimized mRNA sequences.

[0008] Preferably, the vaccine property prediction model is used to predict the stability, translation efficiency and immunogenicity features of the input mRNA sequence and output corresponding stability prediction data, translation efficiency prediction data and immunogenicity prediction data;

[0009] The vaccine property prediction model includes the Uni-RNA model, the three basic feature prediction heads, the feature fusion module and the three vaccine property prediction heads; the three basic feature prediction heads include a codon usage frequency prediction head, an RNA secondary structure prediction head and a GC content prediction head; and the three vaccine property prediction heads include a stability prediction head, a translation efficiency prediction head and an immunogenicity prediction head;

[0010] The model input end of the vaccine property prediction model is used to receive the corresponding mRNA sequence, and the first model output end, the second model output end and the third model output end are respectively used to output the corresponding stability prediction data, the translation efficiency prediction data and the immunogenicity prediction data;

[0011] The input end of the Uni-RNA model is connected with the model input end, and the output end is connected with the input ends of the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head respectively; the output ends of the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head are connected with the first, second and third input ends of the feature fusion module; the output end of the feature fusion module is connected with the input ends of the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head respectively; the output ends of the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head are connected with the first, second and third model output ends respectively;

[0012] The Uni-RNA model has completed model pre-training; the Uni-RNA model is used for feature coding processing of the input mRNA sequence to obtain a corresponding feature coding vector, and the feature coding vector is sent to the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head;

[0013] The codon usage frequency prediction head is implemented based on a type of nonlinear regression model or a type of neural network model; the codon usage frequency prediction head is used for predicting the codon usage frequency of each base position on the mRNA sequence according to the feature coding vector to obtain a corresponding codon usage frequency prediction vector, and the codon usage frequency prediction vector is sent to the feature fusion module; the nonlinear regression model at least includes MLP model, NNR model, GBDT model and XGBoost model; the neural network model at least includes MLP model, CNN model, ResNet model, GNN model;

[0014] The RNA secondary structure prediction head is implemented based on another type of neural network model; the RNA secondary structure prediction head is used for predicting the secondary structure features of each base position on the mRNA sequence according to the feature coding vector to obtain a corresponding secondary structure prediction vector, and the secondary structure prediction vector is sent to the feature fusion module;

[0015] The GC content prediction head is implemented based on another type of nonlinear regression model or another type of neural network model; the GC content prediction head is used for predicting the GC content of each base position on the mRNA sequence according to the feature coding vector to obtain a corresponding GC content prediction vector, and the GC content prediction vector is sent to the feature fusion module;

[0016] The feature fusion module is used for vector splicing of the codon usage frequency prediction vector, the secondary structure prediction vector and the GC content prediction vector to obtain a corresponding splicing vector, and the splicing vector is sent to the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head;

[0017] The stability prediction head is implemented based on another type of the nonlinear regression model or another type of the neural network model; the stability prediction head is configured to predict the stability feature of the mRNA sequence according to the splicing vector and output corresponding stability prediction data;

[0018] The translation efficiency prediction head is implemented based on another type of the nonlinear regression model or another type of the neural network model; the translation efficiency prediction head is configured to predict the translation feature of the mRNA sequence according to the splicing vector and output corresponding translation efficiency prediction data;

[0019] The immunogenicity prediction head is implemented based on another type of the nonlinear regression model or another type of the neural network model; the immunogenicity prediction head is configured to predict the immunogenicity feature of the mRNA sequence according to the splicing vector and output corresponding immunogenicity prediction data.

[0020] Preferably, the antigen vaccine data set includes a plurality of antigen vaccine data records; each antigen vaccine data record includes an antigen molecule sequence, an mRNA molecule sequence, a codon usage frequency label vector, a secondary structure label vector, a GC content label vector, stability label data, translation efficiency label data, and immunogenicity label data.

[0021] Preferably, the model training of the vaccine property prediction model based on the antigen vaccine data set specifically includes:

[0022] The mRNA molecule sequence, the codon usage frequency label vector, the secondary structure label vector, and the GC content label vector of each antigen vaccine data record in the antigen vaccine data set are extracted to form a corresponding first data record; and all the first data records obtained form a corresponding first data set;

[0023] The mRNA molecule sequence, the stability label data, the translation efficiency label data, and the immunogenicity label data of each antigen vaccine data record in the antigen vaccine data set are extracted to form a corresponding second data record; and all the second data records obtained form a corresponding second data set;

[0024] The first-stage model training of the Uni-RNA model and the three basic feature prediction heads of the vaccine property prediction model is performed based on the first data set; after the first round of model training is completed, the second-stage model training of the three vaccine property prediction heads of the vaccine property prediction model is performed based on the second data set; and after the second round of model training is completed, the current model training is confirmed to be completed.

[0025] Further, the first-stage model training of the Uni-RNA model and the three basic feature prediction heads of the vaccine property prediction model based on the first data set specifically comprises:

[0026] Step 51, based on a preset first segmentation ratio, randomly segmenting the first data set into two sub-data sets, denoted as a corresponding first training set and a first evaluation set;

[0027] Wherein, the first training set and the first evaluation set are both composed of a plurality of first data records; the ratio of the total number of records of the first training set to the total number of records of the first evaluation set satisfies the first segmentation ratio;

[0028] Step 52, extracting the first first data record of the first training set as a corresponding current training record;

[0029] Step 53, taking the codon usage frequency label vector, the secondary structure label vector and the GC content label vector of the current training record as the corresponding first label vector P tag1 , second label vector P tag2 and third label vector P tag3 ;

[0030] Step 54, inputting the mRNA molecule sequence of the current training record into the Uni-RNA model to obtain a corresponding first feature encoding vector through feature encoding processing; and inputting the first feature encoding vector into the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head respectively to obtain a corresponding first prediction vector P pre1 , second prediction vector P pre2 and third prediction vector P pre3 ;

[0031] Step 55, bringing the first prediction vector P pre1 and the corresponding first label vector P tag1 into a preset first loss function L a , and bringing the second prediction vector P pre2 and the corresponding second label vector P tag2 into a preset second loss function L b , and bringing the third prediction vector P pre3 and the corresponding third label vector P tag3 into a preset third loss function L c , and calculating the first loss function L a , the second loss function L b and the third loss function Lc The first model loss function L M1 ; and the function values of the first, second, third loss functions and the first model loss function are calculated and the calculation results are taken as the corresponding first loss value, second loss value, third loss value and first overall loss value.

[0032] The first model loss function L M1 is:

[0033] L M1 = L a (P pre1 , P tag1 ) + L b (P pre2 , P tag2 ) + L c (P pre3 , P tag3 ).

[0034] The first loss function L a at least includes an L1 loss function and an L2 loss function; the second loss function L b is composed of a base type loss function and a base structure loss function, the base type loss function at least includes a cross-entropy loss function, and the base structure loss function at least includes an L1 loss function and an L2 loss function; and the third loss function L c at least includes an L1 loss function and an L2 loss function.

[0035] Step 56, whether the first, second, third loss values and the first overall loss value all satisfy the respective corresponding first loss value range, second loss value range, third loss value range and first overall loss value range is identified; if the first, second, third loss values and the first overall loss value all satisfy the respective corresponding loss value range, go to step 57; if the first loss value does not satisfy the first loss value range, based on the preset first model optimizer, the model parameters of the Uni-RNA model and the codon usage frequency prediction head are modulated in the direction of making the first loss function L a reach the minimum value, and return to step 54 at the end of this round of parameter modulation; if the second loss value does not satisfy the second loss value range, based on the preset second model optimizer, the model parameters of the Uni-RNA model and the RNA secondary structure prediction head are modulated in the direction of making the second loss function L b reach the minimum value, and return to step 54 at the end of this round of parameter modulation; if the third loss value does not satisfy the third loss value range, based on the preset third model optimizer, the model parameters of the Uni-RNA model and the RNA secondary structure prediction head are modulated in the direction of making the third loss function L cThe direction reaching the minimum value modulates the model parameters of the Uni-RNA model and the GC content prediction head for one round, and returns to step 54 at the end of the current round of parameter modulation; if the first overall loss value does not meet the first overall loss value range, the fourth preset model optimizer is used to modulate the model parameters of the Uni-RNA model and the three basic feature prediction heads in the direction reaching the minimum value of the first model loss function L M1 The direction reaching the minimum value modulates the model parameters of the Uni-RNA model and the three basic feature prediction heads for one round, and returns to step 54 at the end of the current round of parameter modulation;

[0036] The first, second, third and fourth model optimizers at least include an SGD optimizer and an ADAM optimizer;

[0037] Step 57, identify whether the current training record is the last first data record of the first training set; if yes, go to step 58; if no, extract the next first data record of the first training set as a new current training record and return to step 53;

[0038] Step 58, traverse all the first data records of the first evaluation set for one round; and in the current round of traversal, the first data record currently traversed is taken as a corresponding current evaluation record; the mRNA molecule sequence of the current evaluation record is input into the Uni-RNA model for feature encoding processing to obtain a corresponding second feature encoding vector; the second feature encoding vector is input into the three basic feature prediction heads respectively for prediction processing to obtain a corresponding fourth, fifth and sixth prediction vector to form a corresponding first prediction vector sequence; the codon usage frequency label vector, the secondary structure label vector and the GC content label vector of the current evaluation record are taken as a corresponding fourth, fifth and sixth label vector to form a corresponding first label vector sequence; and the first prediction vector sequence and the first label vector sequence form a corresponding first prediction-label pair; and at the end of the current round of traversal, all the first prediction-label pairs obtained are input into a preset first model evaluation function to obtain a corresponding first evaluation value;

[0039] The first model evaluation function at least includes an RMSE function;

[0040] Step 59, identify whether the first evaluation value meets a preset first evaluation value range; if not, return to step 51 for further training; if yes, confirm that the first stage model training is completed.

[0041] Further, the second stage model training of the three vaccine property prediction heads of the vaccine property prediction model based on the second data set comprises:

[0042] Step 61, based on a preset second split ratio, randomly split the second data set into two sub-data sets, denoted as a corresponding second training set and a second evaluation set;

[0043] Wherein, the second training set and the second evaluation set are both composed of a plurality of second data records; the ratio of the total number of records of the second training set to the total number of records of the second evaluation set meets the second split ratio;

[0044] Step 62, extract the first second data record of the second training set as a corresponding current training record;

[0045] Step 63, record the stability label data, translation efficiency label data and immunogenicity label data of the current training record as corresponding first label data D tag1 , second label data D tag2 and third label data D tag3 ;

[0046] Step 64, input the mRNA molecule sequence of the current training record into the vaccine property prediction model for prediction processing to obtain corresponding first prediction data D pre1 , second prediction data D pre2 and third prediction vector D pre3 ;

[0047] Step 65, bring the first prediction data D pre1 and the corresponding first label data D tag1 into a preset fourth loss function L e , and bring the second prediction data D pre2 and the corresponding second label data D tag2 into a preset fifth loss function L f , and bring the third prediction vector D pre3 and the corresponding third label data D tag3 into a preset sixth loss function L g , and add the fourth loss function L e , the fifth loss function L f and the sixth loss function L g to form a corresponding second model loss function L M2 ; and calculate the function values of the fourth, fifth, sixth loss functions and the second model loss function, and take the calculation results as the corresponding fourth loss value, fifth loss value, sixth loss value and second overall loss value;

[0048] Wherein, the second model loss function L M2 is:

[0049]

[0049] L M2 = L e (D pre1 , D tag2 ) + L f (D pre2 , D tag2 ) + L g (D pre3 , D tag3 );

[0050] The fourth loss function L e includes at least an L1 loss function, an L2 loss function, and a cross-entropy loss function; the fifth loss function L f includes at least an L1 loss function and an L2 loss function; and the sixth loss function L g includes at least an L1 loss function, an L2 loss function, and a cross-entropy loss function.

[0051] In step 66, it is identified whether the fourth, fifth, and sixth loss values and the second overall loss value all satisfy the respective fourth loss value range, fifth loss value range, sixth loss value range, and second overall loss value range. If the fourth, fifth, and sixth loss values and the second overall loss value all satisfy the respective loss value range, the process proceeds to step 67. If the fourth loss value does not satisfy the fourth loss value range, the model parameters of the stability prediction head are modulated in a direction that minimizes the fourth loss function L e based on a preset fifth model optimizer, and the process returns to step 64 at the end of the current round of parameter modulation. If the fifth loss value does not satisfy the fifth loss value range, the model parameters of the translation efficiency prediction head are modulated in a direction that minimizes the fifth loss function L f based on a preset sixth model optimizer, and the process returns to step 64 at the end of the current round of parameter modulation. If the sixth loss value does not satisfy the sixth loss value range, the model parameters of the immunogenicity prediction head are modulated in a direction that minimizes the sixth loss function L g based on a preset seventh model optimizer, and the process returns to step 64 at the end of the current round of parameter modulation. If the second overall loss value does not satisfy the second overall loss value range, the model parameters of the three-vaccine-property prediction head are modulated in a direction that minimizes the second model loss function L M2 based on a preset eighth model optimizer, and the process returns to step 64 at the end of the current round of parameter modulation.

[0052] The fifth, sixth, seventh, and eighth model optimizers all include at least an SGD optimizer and an ADAM optimizer.

[0053] Step 67, identifying whether the current training record is the last second data record of the second training set; if yes, going to step 68; if no, extracting the next second data record of the second training set as a new current training record and returning to step 63;

[0054] Step 68, performing a round of traversal on all the second data records of the second evaluation set; and in the round of traversal, taking the second data record currently traversed as a corresponding current evaluation record; inputting the mRNA molecule sequence of the current evaluation record into the vaccine property prediction model for prediction processing to obtain a corresponding first prediction data sequence composed of fourth, fifth and sixth prediction data; and taking the stability label data, the translation efficiency label data and the immunogenicity label data of the current evaluation record as a corresponding first label data sequence composed of fourth, fifth and sixth label data; and taking the first prediction data sequence and the first label data sequence as a corresponding second prediction-label pair; and at the end of the round of traversal, inputting all the second prediction-label pairs obtained into a preset second model evaluation function to obtain a corresponding second evaluation value;

[0055] Among them, the second model evaluation function at least includes the RMSE function;

[0056] Step 69, identifying whether the second evaluation value meets a preset second evaluation value range; if not, returning to step 61 to continue training; if yes, confirming that the second stage model training is ended.

[0057] Preferably, after the model training is ended, the maximum correlation relationship between the three output features of the three basic feature prediction heads and the three output properties of the three vaccine property prediction heads is identified and based on the identification result, three target functions are determined, specifically including:

[0058] Step 71, after the model training is ended, taking the mapping relationship between the three types of basic feature prediction vectors generated by the vaccine property prediction model in the model prediction process and the three types of property prediction data finally output by the model as a corresponding stability mapping relationship F A , translation efficiency mapping relationship F B and immunogenicity mapping relationship F C ;

[0059] Among them, y a =F A (x a ,x b ,x c ), y b =F B (xa x b x c y c = F C (x a x b x c );

[0060] independent variables x a , x b , x c are the codon usage frequency prediction vector, the secondary structure prediction vector and the GC content prediction vector generated in the model prediction process respectively, and dependent variables y a , y b , y c are the stability prediction data, the translation efficiency prediction data and the immunogenicity prediction data output by the model finally;

[0061] Step 72, each of the antigen vaccine data records of the antigen vaccine data set is recorded as record R i , 1≤i≤N, N is the total number of records of the antigen vaccine data set; and the mRNA molecule sequence of each record R i is input into the vaccine property prediction model for prediction processing to obtain six prediction results corresponding to x a,i , x b,i , x c,i , y a,i , y b,i , y c,i ; and a corresponding first independent / dependent variable data group is composed of the prediction results x a,i , x b,i , x c,i , y a,i corresponding to each index i, and N first independent / dependent variable data groups are composed of the first data group set corresponding to the N first independent / dependent variable data groups; and a corresponding second independent / dependent variable data group is composed of the prediction results x a,i , x b,i , x c,i , y b,i corresponding to each index i, and N second independent / dependent variable data groups are composed of the second data group set corresponding to the N second independent / dependent variable data groups; and a corresponding third independent / dependent variable data group is composed of the prediction results x a,i , x b,i , x c,i , y c,i corresponding to each index i, and N third independent / dependent variable data groups are composed of the third data group set corresponding to the N third independent / dependent variable data groups;

[0062] Step 73, based on the SHAP evaluation algorithm, evaluating the contribution coefficients of the independent variables x a , b , c to the dependent variable y a , and taking the independent variable x a , b or x c with the largest contribution coefficient as the maximum correlation variable x a of the dependent variable y ma ; and based on the SHAP evaluation algorithm, evaluating the contribution coefficients of the independent variables x a , b , c to the dependent variable y b , and taking the independent variable x a , b or x c with the largest contribution coefficient as the maximum correlation variable x b of the dependent variable y mb ; and based on the SHAP evaluation algorithm, evaluating the contribution coefficients of the independent variables x a , b , c to the dependent variable y c , and taking the independent variable x a , b or x c with the largest contribution coefficient as the maximum correlation variable x c of the dependent variable y mc ;

[0063] Step 74, taking the dependent variable y a and its corresponding maximum correlation variable x ma into the pre-set correlation function F corr to obtain the corresponding stability objective function F corr (x ma ,y a ); and taking the dependent variable y b and its corresponding maximum correlation variable x mb into the correlation function F corr to obtain the corresponding translation efficiency objective function F corr (x mb ,y b ); and taking the dependent variable y c and its corresponding maximum correlation variable x mc into the correlation function F corr to obtain the corresponding immunogenicity objective function F corr (x mc ,y c) ; and the three corresponding target functions are composed of the stability target function, the translation efficiency target function and the immunogenicity target function; the correlation function F corr The correlation function F includes at least a Pearson correlation function, a Spearman correlation function and a Kendall correlation function.

[0064] Preferably, the mRNA sequence initialization according to the first antigen sequence to obtain a corresponding first mRNA sequence set specifically includes:

[0065] The transcription mRNA sequence of the first antigen sequence is subjected to sequence generation processing using a preset mRNA sequence generation tool to obtain a plurality of corresponding first mRNA sequences, thereby forming a corresponding first mRNA sequence set; the mRNA sequence generation tool is an mRNA sequence reverse translation tool or an mRNA sequence reverse decoding model; the mRNA sequence reverse translation tool at least includes a Translate Tool tool of ExPASy, which is used for reverse translation and output of the transcription mRNA sequence corresponding to a specified antigen sequence; the mRNA sequence reverse decoding model at least includes a reverse decoding model based on a GANs model, which is used for reverse decoding and output of the transcription mRNA sequence corresponding to a specified antigen sequence.

[0066] Preferably, the multi-objective optimization method according to the NSGA-II algorithm is used to iteratively optimize the first mRNA sequence set by means of the vaccine property prediction model, the three target functions and corresponding in vitro / in vivo experimental methods to obtain a corresponding optimized mRNA sequence set, which is fed back to the user, and specifically includes:

[0067] Step 91, initializing a first iteration counter to 1;

[0068] Step 92, inputting each first mRNA sequence of the first mRNA sequence set into the vaccine property prediction model for prediction, and bringing the prediction result into the three target functions for calculation to obtain corresponding first target function values, second target function values and third target function values; and sorting all the first target function values, all the second target function values and all the third target function values obtained according to the function values from high to low to obtain corresponding first, second and third function value sequences;

[0069] Step 93, based on the first, second and third function value sequences, the non-dominated level and the crowded distance of each of the first mRNA sequences are identified according to the non-dominated level and crowded distance identification rules in the NSGA-II algorithm; and based on the tournament selection strategy in the NSGA-II algorithm, the mRNA sequence individuals are selected according to the non-dominated level and crowded distance of each of the first mRNA sequences, and a corresponding first individual set is formed by the selected all first mRNA sequences; and based on the preset individual crossover rule, all first mRNA sequences in the first individual set are subjected to individual crossover processing, and each first mRNA sequence after crossover is subjected to individual mutation processing; and after all first mRNA sequences in the first individual set complete individual mutation, the first individual set and the first mRNA sequence set are merged to obtain a new first mRNA sequence set;

[0070] The individual crossover rule includes single-point crossover and multi-point crossover.

[0071] Step 94, the first iteration counter is incremented by 1; and whether the first iteration counter after incrementing by 1 exceeds the preset iteration counter threshold is identified; if yes, go to step 95; if no, return to step 92.

[0072] Step 95, each of the first mRNA sequences of the latest first mRNA sequence set is input into the vaccine property prediction model for prediction to obtain corresponding first stability prediction data, first translation efficiency prediction data and first immunogenicity prediction data; and based on the preset stability and translation efficiency sorting rule, all first stability prediction data and all first translation efficiency prediction data obtained are sorted respectively, and one or more first mRNA sequences with stability and translation efficiency sorting both before the specified sorting position are extracted to form a corresponding preferred set.

[0073] Step 96, the stability and translation efficiency of each of the first mRNA sequences in the preferred set are analyzed by in vitro experimental means to obtain corresponding first stability experimental data and first translation efficiency experimental data; and the relative error of the first stability prediction data and the first stability experimental data of each of the first mRNA sequences in the preferred set is calculated to obtain corresponding first relative error, and the relative error of the first translation efficiency prediction data and the first translation efficiency experimental data of each of the first mRNA sequences in the preferred set is calculated to obtain corresponding second relative error.

[0074] Step 97, identify whether all the obtained first and second relative errors are less than the corresponding first and second relative error thresholds; if yes, go to step 98; if no, perform one round of parameter optimization processing on the stability prediction head and the translation efficiency prediction head of the vaccine property prediction model according to the supervised model training mode with the obtained all the first stability experiment data and the first translation efficiency experiment data as label data, and return to step 91 when the round of optimization processing is completed;

[0075] Step 98, obtain the corresponding first immunogenicity experiment data by experimentally analyzing the immunogenicity properties of each of the first mRNA sequences in the preferred set through in vivo experiment means; and calculate the relative error of the first immunogenicity prediction data and the first immunogenicity experiment data of each of the first mRNA sequences in the preferred set to obtain the corresponding third relative error;

[0076] Step 99, identify whether all the obtained third relative errors are less than the corresponding third relative error threshold; if no, perform one round of parameter optimization processing on the immunogenicity prediction head of the vaccine property prediction model according to the supervised model training mode with the obtained all the first immunogenicity experiment data as label data, and return to step 91 when the round of optimization processing is completed; if yes, take each of the first mRNA sequences in the preferred set as the corresponding optimized mRNA sequence, and form the corresponding optimized mRNA sequence set by all the obtained optimized mRNA sequences and output.

[0077] The second aspect of the embodiment of the present application provides a device for implementing the processing method of the mRNA vaccine design task of the first aspect described above, and the device comprises a working model construction module, a working model training module and a vaccine design task processing module.

[0078] The working model construction module is used to design a working model for predicting three vaccine properties of mRNA sequences with Uni-RNA model as a feature extraction backbone, which is recorded as a corresponding vaccine property prediction model; the three vaccine properties include stability, translation efficiency and immunogenicity; the vaccine property prediction model comprises the Uni-RNA model, three basic feature prediction heads, a feature fusion module and three vaccine property prediction heads.

[0079] The working model training module is configured to collect a large amount of antigen protein and corresponding mRNA vaccine information by a big data collection method to construct a corresponding antigen vaccine data set, and train a vaccine characteristic prediction model based on the antigen vaccine data set, and identify the maximum correlation between three output characteristics of the three basic feature prediction heads and three output characteristics of the three vaccine characteristic prediction heads after the model training is completed, and determine corresponding three objective functions based on the identification result; the three objective functions include a stability objective function, a translation efficiency objective function and an immunogenicity objective function.

[0080] The vaccine design task processing module is configured to receive a molecular sequence of a target antigen protein input by a user as a corresponding first antigen sequence, initialize an mRNA sequence according to the first antigen sequence to obtain a corresponding first mRNA sequence set, and perform iterative optimization processing on the first mRNA sequence set by means of the vaccine characteristic prediction model, the three objective functions and corresponding in vitro / in vivo experimental means according to a multi-objective optimization method of the NSGA-II algorithm to obtain a corresponding optimized mRNA sequence set and feed back to the user; the first mRNA sequence set is composed of a plurality of first mRNA sequences; and the optimized mRNA sequence set is composed of one or more optimized mRNA sequences.

[0081] The third aspect of the embodiment of the present application provides an electronic device, comprising a memory, a processor and a transceiver.

[0082] The processor is configured to be coupled with the memory, read and execute instructions in the memory to realize the method steps of the first aspect described above;

[0083] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transceiving.

[0084] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores computer instructions, when the computer instructions are executed by the computer, the computer instructions make the computer execute the instructions of the method of the first aspect described above.

[0085] The embodiment of the application provides an mRNA vaccine design task processing method, device, electronic equipment and computer readable storage medium. From the above content, it can be known that the embodiment of the application pre-designs a working model for predicting three vaccine characteristics (stability, translation efficiency and immunogenicity) of an mRNA sequence by taking a pre-trained Uni-RNA model as a feature extraction backbone, and the working model is recorded as a vaccine characteristic prediction model, trains the model based on an antigen vaccine data set constructed by a big data collection method, identifies the maximum correlation between three basic characteristics (codon usage frequency, RNA secondary structure and GC content) and three vaccine characteristics (stability, translation efficiency and immunogenicity) that can be predicted by the model, and determines three objective functions (stability objective function, translation efficiency objective function and immunogenicity objective function) based on the identification result. In the specific mRNA vaccine design task processing process: the embodiment of the application first generates the transcription mRNA sequence of the target antigen protein in batches by using an mRNA sequence generation tool to obtain an initialized mRNA sequence set; then the vaccine characteristic prediction model and the three objective functions are used to iteratively optimize and screen the mRNA sequence set according to the multi-objective optimization mode of the NSGA-II algorithm, so as to obtain an optimized set with reduced sequence number; then the stability and translation efficiency of the optimized set are analyzed based on an in-vitro experiment method; if the error between the in-vitro experiment analysis result and the model prediction result is large, the model is automatically optimized based on the experimental data, and the mRNA sequence set is iteratively optimized again after optimization; if the error between the in-vitro experiment analysis result and the model prediction result is controllable, the immunogenicity of the optimized set is further analyzed based on an in-vivo experiment method; if the error between the in-vivo experiment analysis result and the model prediction result is large, the model is automatically optimized again based on the experimental data, and the mRNA sequence set is iteratively optimized again after optimization; if the error between the in-vivo experiment analysis result and the model prediction result is controllable, the current latest optimized set is output as a final design candidate set. The mRNA sequence initialization mode of the embodiment of the application gets rid of the limitation of expert experience, improves the structural diversity of the mRNA sequence; the vaccine characteristic prediction model of the embodiment of the application reduces the number of experimental analysis times, shortens the design period and improves the design efficiency; and the model training method and the multi-objective optimization process effectively guarantee the design quality of the mRNA sequence. BRIEF DESCRIPTION OF DRAWINGS

[0086] Figure 1 A schematic diagram of an mRNA vaccine design task processing method provided for the embodiment one of the application;

[0087] Figure 2 A module structure diagram of the vaccine characteristic prediction model provided for the embodiment one of the application;

[0088] Figure 3 A module structure diagram of a processing device for an mRNA vaccine design task is provided for the second embodiment of the present application.

[0089] Figure 4 A structural schematic diagram of an electronic device is provided for the third embodiment of the present application. DETAILED DESCRIPTION

[0090] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0091] The first embodiment of the present application provides a processing method for an mRNA vaccine design task, as shown in a flowchart of the processing method for an mRNA vaccine design task provided for the first embodiment of the present application. Figure 1 The first embodiment of the present application provides a processing method for an mRNA vaccine design task, as shown in a flowchart of the processing method for an mRNA vaccine design task provided for the first embodiment of the present application.

[0092] Step 1. Design a working model for predicting three vaccine properties of mRNA sequences with Uni-RNA model as a feature extraction backbone, denoted as a corresponding vaccine property prediction model.

[0093] Here, the three vaccine properties mentioned in the embodiments of the present application include stability, translation efficiency and immunogenicity; the stability is usually characterized by the half-life of mRNA in vivo, the translation efficiency is usually characterized by the expression amount of target antigen protein, and the immunogenicity is usually characterized by antibody titer.

[0094] The vaccine property prediction model of the embodiments of the present application is used to predict the stability, translation efficiency and immunogenicity features of the input mRNA sequence and output corresponding stability prediction data, translation efficiency prediction data and immunogenicity prediction data; here, the stability prediction data, translation efficiency prediction data and immunogenicity prediction data are actually half-life prediction data, antigen protein expression amount prediction data and antibody titer prediction data.

[0095] As shown in the flowchart of the processing method for an mRNA vaccine design task provided for the first embodiment of the present application, the method mainly includes the following steps: Figure 2 As shown in the module structure diagram of the vaccine property prediction model provided for the first embodiment of the present application, the model components of the vaccine property prediction model include: Uni-RNA model, three basic feature prediction heads, feature fusion module and three vaccine property prediction heads; wherein the three basic feature prediction heads include codon usage frequency prediction head, RNA secondary structure prediction head and GC content prediction head; the three vaccine property prediction heads include stability prediction head, translation efficiency prediction head and immunogenicity prediction head.

[0096] The input / output end of the vaccine property prediction model is explained as follows: the model input end is used to receive the corresponding mRNA sequence, and the first model output end, the second model output end and the third model output end are used to output the corresponding stability prediction data, translation efficiency prediction data and immunogenicity prediction data respectively.

[0097] The connection relationship of each model component is as follows: the input end of the Uni-RNA model is connected with the model input end, and the output end is connected with the input end of the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head respectively; the output end of the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head is connected with the first, second and third input end of the feature fusion module; the output end of the feature fusion module is connected with the input end of the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head respectively; the output end of the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head is connected with the corresponding first, second and third model output end respectively.

[0098] The component functions of each model component are as follows:

[0099] 1) Uni-RNA model:

[0100] The Uni-RNA model of the embodiment of the application has completed model pre-training; the Uni-RNA model is used for feature coding processing of the input mRNA sequence to obtain the corresponding feature coding vector, and sends the feature coding vector to the codon usage frequency prediction head, the RNA secondary structure prediction head and the GC content prediction head;

[0101] Here, the Uni-RNA model of the embodiment of the application is a feature coding model based on a transformer encoder structure, and the Uni-RNA model has completed pre-training; the model structure, function and pre-training scheme of this model have been disclosed in the technical document "UNI-RNA: UNIVERSAL PRE-TRAINED MODELS REVOLUTIONIZE RNA RESEARCH"; it can be known from the technical document that the Uni-RNA model can be used for base-level deep learning of an RNA sequence.

[0102] 2) Codon usage frequency prediction head:

[0103] The codon usage frequency prediction head of the embodiment of the application is implemented based on a type of nonlinear regression model or a type of neural network model; the codon usage frequency prediction head is used for predicting the codon usage frequency of each base position on the mRNA sequence according to the feature coding vector to obtain the corresponding codon usage frequency prediction vector, and sends the codon usage frequency prediction vector to the feature fusion module;

[0104] Here, the nonlinear regression model mentioned by the embodiment of the application at least includes an MLP model, an NNR model, a GBDT model and an XGBoost model; the neural network model at least includes an MLP model, a CNN model, a ResNet model and a GNN model;

[0105] 3) RNA secondary structure prediction head:

[0106] The RNA secondary structure prediction head of the embodiment of the application is implemented based on another type of neural network model; the RNA secondary structure prediction head is used for predicting the secondary structure features of each base position on the mRNA sequence according to the feature coding vector to obtain a corresponding secondary structure prediction vector and sending the feature fusion module;

[0107] 4) GC content prediction head:

[0108] The GC content prediction head of the embodiment of the application is implemented based on another type of nonlinear regression model or another type of neural network model; the GC content prediction head is used for predicting the GC content of each base position on the mRNA sequence according to the feature coding vector to obtain a corresponding GC content prediction vector and sending the feature fusion module;

[0109] 5) Feature fusion module:

[0110] The feature fusion module of the embodiment of the application is used for vector splicing the codon usage frequency prediction vector, the secondary structure prediction vector and the GC content prediction vector to obtain a corresponding splicing vector and sending the splicing vector to the stability prediction head, the translation efficiency prediction head and the immunogenicity prediction head;

[0111] 6) Stability prediction head:

[0112] The stability prediction head of the embodiment of the application is implemented based on another type of nonlinear regression model or another type of neural network model; the stability prediction head is used for predicting the stability features of the mRNA sequence according to the splicing vector and outputting corresponding stability prediction data;

[0113] 7) Translation efficiency prediction head:

[0114] The translation efficiency prediction head of the embodiment of the application is implemented based on another type of nonlinear regression model or another type of neural network model; the translation efficiency prediction head is used for predicting the translation features of the mRNA sequence according to the splicing vector and outputting corresponding translation efficiency prediction data;

[0115] 8) Immunogenicity prediction head:

[0116] The immunogenicity prediction head of the embodiment of the application is implemented based on another type of nonlinear regression model or another type of neural network model; the immunogenicity prediction head is used to predict the immunogenicity characteristics of the mRNA sequence according to the splicing vector and output corresponding immunogenicity prediction data.

[0117] Step 2, a large amount of antigen protein and corresponding mRNA vaccine information is collected by a big data collection method to construct a corresponding antigen vaccine data set; and a vaccine characteristic prediction model is trained based on the antigen vaccine data set; and after the model training is completed, the maximum correlation relationship between the three output characteristics of the three basic characteristic prediction heads and the three output characteristics of the three vaccine characteristic prediction heads is identified, and the corresponding three objective functions are determined based on the identification results;

[0118] Specifically, step 21, a large amount of antigen protein and corresponding mRNA vaccine information is collected by a big data collection method to construct a corresponding antigen vaccine data set;

[0119] The antigen vaccine data set includes a plurality of antigen vaccine data records; the antigen vaccine data record includes an antigen molecule sequence, an mRNA molecule sequence, a codon usage frequency label vector, a secondary structure label vector, a GC content label vector, a stability label data, a translation efficiency label data and an immunogenicity label data.

[0120] Here, the embodiment of the application collects data from a plurality of public mRNA vaccine databases and a plurality of public technical literature / materials / magazines / papers by a big data collection method, the collection objects include various antigen proteins and corresponding mRNA vaccine information, the mRNA vaccine information here at least includes corresponding mRNA molecule sequence, codon usage frequency, secondary structure, GC content, stability parameter, i.e., mRNA half-life, translation efficiency parameter, i.e., expression amount of target antigen protein, immunogenicity parameter, i.e., antibody titer of corresponding antibody; it should be noted that if it is found that the characteristic / characteristic parameter (codon usage frequency, secondary structure, GC content, stability parameter, translation efficiency parameter, immunogenicity parameter) of a certain mRNA vaccine is missing during data collection, it will be calculated and filled in by a preset biochemical information tool.

[0121] Step 22, and based on the antigen vaccine data set, a vaccine characteristic prediction model is trained;

[0122] Specifically, step 221, the mRNA molecule sequence, the codon usage frequency label vector, the secondary structure label vector and the GC content label vector of each antigen vaccine data record in the antigen vaccine data set are extracted to form a corresponding first data record; and all the first data records obtained form a corresponding first data set;

[0123] Here, the first dataset consists of multiple first data records; each first data record consists of an mRNA molecule sequence, a codon usage frequency tag vector, a secondary structure tag vector, and a GC content tag vector.

[0124] Step 222: Extract the mRNA molecular sequence, stability tag data, translation efficiency tag data, and immunogenicity tag data from each antigen vaccine data record in the antigen vaccine dataset to form a corresponding second data record; and combine all the obtained second data records to form the corresponding second dataset;

[0125] Here, the resulting second dataset consists of multiple second data records; each second data record consists of mRNA molecule sequence, stability tag data, translation efficiency tag data, and immunogenicity tag data.

[0126] Step 223, and based on the first dataset, perform the first stage of model training on the Uni-RNA model of the vaccine characteristic prediction model and the prediction heads of the three basic features;

[0127] Specifically, it includes: step 2231, randomly dividing the first dataset into two subsets based on a preset first segmentation ratio, denoted as the first training set and the first evaluation set;

[0128] Wherein, the first segmentation ratio is a preset ratio value, such as 8:2; both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio;

[0129] Step 2232: Extract the first data record of the first training set as the corresponding current training record;

[0130] Step 2233: Denote the codon usage frequency label vector, secondary structure label vector, and GC content label vector of the current training record as the corresponding first label vector P. tag1 The second label vector P tag2 and the third label vector P tag3 ;

[0131] Step 2234: Input the currently recorded mRNA molecular sequence into the Uni-RNA model for feature encoding processing to obtain the corresponding first feature encoding vector; then input the first feature encoding vector into the codon usage frequency prediction head, RNA secondary structure prediction head, and GC content prediction head respectively for prediction processing to obtain the corresponding first prediction vector P. pre1 Second prediction vector P pre2 and the third prediction vector P pre3 ;

[0132] Step 2235, the first prediction vector P pre1 and the corresponding first label vector P tag1 Substitute the preset first loss function L a and the second prediction vector P pre2 and the corresponding second label vector P tag2 Substitute the preset second loss function L b And the third prediction vector P pre3 and the corresponding third label vector P tag3 Substitute the preset third loss function L c And by the first loss function L a Second loss function L b and the third loss function L c The summation constitutes the corresponding first model loss function L. M1 The function values ​​of the first, second, and third loss functions, as well as the first model loss function, are calculated and the calculation results are used as the corresponding first loss value, second loss value, third loss value, and first overall loss value.

[0133] Wherein, the first model loss function L M1 for:

[0134] L M1 =L a (P pre1 ,P tag1 )+L b (P pre2 ,P tag2 )+L c (P pre3 ,P tag3 );

[0135] First loss function L a It includes at least the L1 loss function and the L2 loss function; a second loss function L b It consists of a base type loss function and a base structure loss function. The base type loss function includes at least a cross-entropy loss function, and the base structure loss function includes at least an L1 loss function and an L2 loss function; a third loss function L... c It should include at least the L1 loss function and the L2 loss function;

[0136] Step 2236: Identify whether the first, second, and third loss values ​​and the first overall loss value all satisfy their respective first loss value range, second loss value range, third loss value range, and first overall loss value range; if the first, second, and third loss values ​​and the first overall loss value all satisfy their respective loss value ranges, proceed to step 2237; if the first loss value does not satisfy the first loss value range, then based on the preset first model optimizer, move towards making the first loss function L... a The Uni-RNA model and the model parameters using the frequency prediction head for codons are modulated in one round to reach the minimum value, and the process returns to step 2234 at the end of this round of parameter modulation; if the second loss value does not meet the second loss value range, the optimization is performed based on the preset second model optimizer in a direction that makes the second loss function L... b The model parameters of the Uni-RNA model and the RNA secondary structure prediction head are modulated in one round to reach the minimum value, and the process returns to step 2234 at the end of this round of parameter modulation; if the third loss value does not meet the range of the third loss value, the optimization is performed based on the preset third model optimizer to move towards the direction that makes the third loss function L c The model parameters of the Uni-RNA model and the GC content prediction head are modulated in one round to reach the minimum value, and the process returns to step 2234 at the end of this round of parameter modulation; if the first overall loss value does not meet the first overall loss value range, then the first model loss function L is optimized based on the preset fourth model optimizer. M1 The direction that reaches the minimum value modulates the model parameters of the Uni-RNA model and the three basic feature prediction heads in one round, and returns to step 2234 when the parameter modulation of this round ends;

[0137] Among them, the first, second, and third loss value ranges and the first overall loss value range are four pre-set loss value ranges; the first, second, third, and fourth model optimizers all include at least the SGD optimizer and the ADAM optimizer;

[0138] Step 2237: Identify whether the current training record is the last first data record of the first training set; if yes, proceed to step 2238; if no, extract the next first data record of the first training set as the new current training record and return to step 2233.

[0139] Step 2238: Perform a traversal of all first data records in the first evaluation set; during this traversal, the currently traversed first data record is taken as the corresponding current evaluation record; the mRNA molecular sequence of the current evaluation record is input into the Uni-RNA model for feature encoding processing to obtain the corresponding second feature encoding vector; the second feature encoding vector is input into the three basic feature prediction heads for prediction processing to obtain the corresponding fourth, fifth, and sixth prediction vectors, forming the corresponding first prediction vector sequence; the codon usage frequency tag vector, secondary structure tag vector, and GC content tag vector of the current evaluation record are recorded as the corresponding fourth, fifth, and sixth tag vectors, forming the corresponding first tag vector sequence; the first prediction vector sequence and the first tag vector sequence form the corresponding first prediction-tag pair; and at the end of this traversal, all the obtained first prediction-tag pairs are input into the preset first model evaluation function to calculate the corresponding first evaluation value;

[0140] The first model evaluation function includes at least the RMSE function;

[0141] Step 2239: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 2231 to continue training; if it meets the range, confirm that the first stage of model training has ended.

[0142] The first evaluation value range is a pre-set evaluation value range;

[0143] Step 224, and after the first round of model training, the second stage of model training is carried out on the three vaccine characteristic prediction heads of the vaccine characteristic prediction model based on the second dataset;

[0144] Specifically, it includes: step 2241, randomly dividing the second dataset into two subsets based on a preset second segmentation ratio, denoted as the corresponding second training set and second evaluation set;

[0145] The second split ratio is a pre-set ratio, such as 8:2; both the second training set and the second evaluation set consist of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second split ratio.

[0146] Step 2242: Extract the first second data record of the second training set as the corresponding current training record;

[0147] Step 2243: Record the stability label data, translation efficiency label data, and immunogenicity label data of the current training record as the corresponding first label data D. tag1 Second label data D tag2 and third label data D tag3 ;

[0148] Step 2244: Input the currently recorded mRNA molecular sequence into the vaccine characteristic prediction model for prediction processing to obtain the corresponding first prediction data D. pre1 Second prediction data D pre2 and the third prediction vector D pre3 ;

[0149] Step 2245, the first prediction data D pre1 and the corresponding first label data D tag1 Substitute the preset fourth loss function L e And the second prediction data D pre2 and the corresponding second label data D tag2 Substitute the preset fifth loss function L f And the third prediction vector D pre3 and the corresponding third label data D tag3 Substitute the preset sixth loss function L g And by the fourth loss function L e The fifth loss function L f and the sixth loss function L g The summation constitutes the corresponding second model loss function L. M2 The function values ​​of the fourth, fifth, and sixth loss functions, as well as the second model loss function, are calculated and the results are used as the corresponding fourth loss value, fifth loss value, sixth loss value, and second overall loss value.

[0150] Among them, the loss function of the second model L M2 for:

[0151] L M2 =L e (D pre1 D tag2 )+L f (D pre2 D tag2 )+L g (D pre3 D tag3 );

[0152] Fourth loss function L e It includes at least the L1 loss function, L2 loss function, and cross-entropy loss function; and a fifth loss function L. f It includes at least the L1 loss function and the L2 loss function; a sixth loss function L g It should include at least the L1 loss function, the L2 loss function, and the cross-entropy loss function;

[0153] Step 2246: Identify whether the fourth, fifth, and sixth loss values, as well as the second overall loss value, all satisfy their respective corresponding loss value ranges; if the fourth, fifth, and sixth loss values, as well as the second overall loss value, satisfy their respective loss value ranges, proceed to step 2247; if the fourth loss value does not satisfy the fourth loss value range, then based on the preset fifth model optimizer, move towards making the fourth loss function L... e The model parameters of the stability prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 2244 when the parameter modulation ends; if the fifth loss value does not meet the range of the fifth loss value, then the preset sixth model optimizer is used to move towards the direction that makes the fifth loss function L f The model parameters of the translation efficiency prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 2244 when this round of parameter modulation ends; if the sixth loss value does not meet the range of the sixth loss value, then the preset seventh model optimizer is used to move towards making the sixth loss function L g The model parameters of the immunogenicity prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 2244 when this round of parameter modulation ends; if the second overall loss value does not meet the range of the second overall loss value, then the optimization is performed based on the preset eighth model optimizer in the direction that makes the second model loss function L M2 The direction that reaches the minimum value modulates the model parameters of the three vaccine characteristic prediction heads in one round, and returns to step 2244 when the parameter modulation of this round ends;

[0154] Among them, the fourth, fifth, and sixth loss value ranges and the second overall loss value range are four pre-set loss value ranges; the fifth, sixth, seventh, and eighth model optimizers all include at least the SGD optimizer and the ADAM optimizer;

[0155] Step 2247: Identify whether the current training record is the last second data record of the second training set; if yes, proceed to step 2248; if no, extract the next second data record of the second training set as the new current training record and return to step 2243.

[0156] Step 2248: Perform a traversal of all second data records in the second evaluation set; during this traversal, the currently traversed second data record is taken as the corresponding current evaluation record; the mRNA molecular sequence of the current evaluation record is input into the vaccine characteristic prediction model for prediction processing to obtain the corresponding fourth, fifth, and sixth prediction data, which form the corresponding first prediction data sequence; the stability tag data, translation efficiency tag data, and immunogenicity tag data of the current evaluation record are recorded as the corresponding fourth, fifth, and sixth tag data, which form the corresponding first tag data sequence; the first prediction data sequence and the first tag data sequence form the corresponding second prediction-tag pair; and at the end of this traversal, all the obtained second prediction-tag pairs are input into the preset second model evaluation function to calculate the corresponding second evaluation value;

[0157] The second model evaluation function includes at least the RMSE function;

[0158] Step 2249: Identify whether the second evaluation value meets the preset range of the second evaluation value; if not, return to step 2241 to continue training; if it meets the range, confirm that the second stage of model training has ended.

[0159] The second evaluation value range is a pre-set evaluation value range;

[0160] Step 225, and confirm the end of this model training after the second round of model training;

[0161] Step 23: After the model training is completed, identify the maximum correlation between the three output features of the three basic feature prediction heads and the three output features of the three vaccine characteristic prediction heads, and determine the corresponding three objective functions based on the identification results.

[0162] The three objective functions include a stability objective function, a translation efficiency objective function, and an immunogenicity objective function;

[0163] Specifically, this includes: Step 231, after the model training is completed, recording the mapping relationship between the three basic feature prediction vectors generated by the vaccine characteristic prediction model during the model prediction process and the three characteristic prediction data finally output by the model as the corresponding stability mapping relationship F. A Translation efficiency mapping relationship F B Mapping relationship with immunogenicity F C ;

[0164] Among them, y a =F A (x a ,x b ,x c ), y b =F B (xa ,x b ,x c ), y c =F C (x a ,x b ,x c );

[0165] Independent variable x a x b x c These are the codon usage frequency prediction vector, secondary structure prediction vector, and GC content prediction vector generated during the model prediction process, respectively, and the dependent variable y. a y b y c These are the stability prediction data, translation efficiency prediction data, and immunogenicity prediction data that are the final outputs of the model, respectively.

[0166] Step 232: Denote each antigen vaccine data record in the antigen vaccine dataset as record R. i , 1≤i≤N, where N is the total number of records in the antigen vaccine dataset; and each record R i The mRNA molecular sequence was input into the vaccine characteristic prediction model for prediction processing, resulting in six corresponding prediction results, x. a,i x b,i x c,i y a,i y b,i y c,i And the prediction result x corresponding to each index i. a,i x b,i x c,i y a,i Form a corresponding first independent / dependent variable data set, and form a corresponding first data set set from the N obtained first independent / dependent variable data sets; and form a prediction result x corresponding to each index i. a,i x b,i x c,i y b,i This forms a corresponding second independent / dependent variable data set, and the resulting N second independent / dependent variable data sets form a corresponding second data set set; and the prediction result x corresponding to each index i is... a,i x b,i x c,i y c,i Form a corresponding third independent / dependent variable data set, and form a corresponding third data set set from the N obtained third independent / dependent variable data sets;

[0167] Step 233: Based on the SHAP evaluation algorithm, evaluate the independent variable x according to the first data set. a xb x c For the dependent variable y a The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y a The largest correlation variable x ma Based on the SHAP evaluation algorithm, the independent variable x is evaluated according to the second data set. a x b x c For the dependent variable y b The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y b The largest correlation variable x mb Based on the SHAP evaluation algorithm, the independent variable x is evaluated according to the third dataset. a x b x c For the dependent variable y c The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y c The largest correlation variable x mc ;

[0168] Step 234, the dependent variable y a Its corresponding largest correlation variable x ma Substitute the preset correlation function F corr The corresponding stability objective function F is obtained. corr (x ma ,y a ); and the dependent variable y b Its corresponding largest correlation variable x mb Substitute the correlation function F corr The corresponding translation efficiency objective function F is obtained. corr (x mb ,y b ); and the dependent variable y c Its corresponding largest correlation variable x mc Substitute the correlation function F corr The corresponding immunogenicity objective function F is obtained. corr (x mc ,y c The objective function consists of three corresponding objective functions: a stability objective function, a translation efficiency objective function, and an immunogenicity objective function.

[0169] Among them, the correlation function F corr It should include at least the Pearson correlation function, the Spearman correlation function, and the Kendall correlation function.

[0170] Step 3: Receive the molecular sequence of the target antigen protein input by the user as the corresponding first antigen sequence; initialize the mRNA sequence according to the first antigen sequence to obtain the corresponding first mRNA sequence set; and iteratively optimize the first mRNA sequence set according to the multi-objective optimization method of the NSGA-II algorithm, using the vaccine characteristic prediction model, three objective functions and corresponding in vitro / in vivo experimental methods to obtain the corresponding optimized mRNA sequence set and feed it back to the user.

[0171] Specifically, this includes: Step 31, receiving the molecular sequence of the target antigen protein input by the user as the corresponding first antigen sequence;

[0172] Step 32, and initialize the mRNA sequence according to the first antigen sequence to obtain the corresponding first mRNA sequence set;

[0173] The first mRNA sequence set consists of multiple first mRNA sequences;

[0174] Specifically, this includes: using a preset mRNA sequence generation tool to generate a sequence from the transcribed mRNA sequence of the first antigen sequence to obtain multiple corresponding first mRNA sequences that form a corresponding first mRNA sequence set;

[0175] Here, the mRNA sequence generation tool in this embodiment of the invention is either an mRNA sequence reverse translation tool or an mRNA sequence reverse decoding model; wherein, the mRNA sequence reverse translation tool includes at least the ExPASy Translate Tool, which is used to reverse translate and output the transcribed mRNA sequence corresponding to a specified antigen sequence; the mRNA sequence reverse decoding model includes at least a reverse decoding model based on the GANs model, which is used to reverse decode and output the transcribed mRNA sequence corresponding to a specified antigen sequence.

[0176] Step 323, and following the multi-objective optimization method of the NSGA-II algorithm, the first mRNA sequence set is iteratively optimized using the vaccine characteristic prediction model, three objective functions, and corresponding in vitro / in vivo experimental methods to obtain the corresponding optimized mRNA sequence set, which is then fed back to the user.

[0177] The optimized mRNA sequence set consists of one or more optimized mRNA sequences.

[0178] Specifically, this includes: Step 3231, initializing the first iteration counter to 1;

[0179] Step 3232: Input each of the first mRNA sequences in the first mRNA sequence set into the vaccine characteristic prediction model for prediction, and input the prediction results into three objective functions to calculate the corresponding first objective function value, second objective function value, and third objective function value; and sort all the obtained first objective function values, all the second objective function values, and all the third objective function values ​​from high to low to obtain the corresponding first, second, and third function value sequences.

[0180] Step 3233: According to the non-dominated level and crowding distance identification rules in the NSGA-II algorithm, the non-dominated level and crowding distance corresponding to each first mRNA sequence are confirmed based on the first, second, and third function value sequences; and based on the tournament selection strategy in the NSGA-II algorithm, individual mRNA sequences are selected according to the non-dominated level and crowding distance of each first mRNA sequence, and the selected first mRNA sequences form the corresponding first body set; and based on the preset individual crossover rules, individual crossover processing is performed on all first mRNA sequences in the first body set, and individual mutation processing is performed on each first mRNA sequence that has completed the crossover; and after all first mRNA sequences in the first body set have completed individual mutation, the first body set and the first mRNA sequence set are merged to obtain a new first mRNA sequence set;

[0181] Among them, individual crossover rules include single-point crossover and multi-point crossover;

[0182] Step 3234: Increment the first iteration counter by 1; and identify whether the incremented first iteration counter exceeds the preset iteration counter threshold; if yes, proceed to step 3235; if no, return to step 3232.

[0183] Here, the iteration counter threshold is a pre-set threshold parameter;

[0184] Step 3235: Input each of the first mRNA sequences in the latest first mRNA sequence set into the vaccine characteristic prediction model to obtain the corresponding first stability prediction data, first translation efficiency prediction data and first immunogenicity prediction data; and sort all the obtained first stability prediction data and all the first translation efficiency prediction data according to the preset stability and translation efficiency ranking rules, and extract one or more first mRNA sequences that are ranked before the specified ranking to form the corresponding preferred set;

[0185] Here, the stability and translation efficiency ranking rule in this embodiment of the invention is a pre-set ranking rule, which is sorted from best to worst by default. The higher the ranking, the better the corresponding stability / translation efficiency, and vice versa. The specified ranking is a pre-set ranking value. The stability and translation efficiency ranking can use the same value as their respective specified ranking, or different values ​​can be defined for stability and translation efficiency as their respective specified rankings.

[0186] Step 3236: The stability and translation efficiency of each first mRNA sequence in the preferred set are analyzed by in vitro experiments to obtain the corresponding first stability experimental data and first translation efficiency experimental data; the relative error between the first stability prediction data and the first stability experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding first relative error; and the relative error between the first translation efficiency prediction data and the first translation efficiency experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding second relative error.

[0187] Here, the in vitro experimental methods mentioned in the embodiments of the present invention are the in vitro experimental methods used in the conventional mRNA vaccine design process. These experimental methods synthesize mRNA in vitro through transcription in a cell-free system and verify its expression efficiency and stability in an in vitro expression system. For example, T7 RNA polymerase is used to perform in vitro transcription and synthesize candidate mRNA sequences, and quantitative real-time polymerase chain reaction (qPCR) and enzyme-linked immunosorbent assay (ELISA) are used to detect the translation efficiency / stability characterization data of the mRNA.

[0188] Step 3237: Identify whether all the obtained first and second relative errors are less than the corresponding first and second relative error thresholds; if yes, proceed to step 3238; if no, use all the first stability experimental data and first translation efficiency experimental data obtained this time as label data, perform a round of parameter optimization on the stability prediction head and translation efficiency prediction head of the vaccine characteristic prediction model according to the supervised model training method, and return to step 3231 when the optimization process ends.

[0189] Here, the first and second relative error thresholds in this embodiment of the invention are two preset error thresholds;

[0190] Step 3238: The immunogenicity characteristics of each first mRNA sequence in the preferred set are analyzed by in vivo experimental methods to obtain the corresponding first immunogenicity experimental data; and the relative error between the first immunogenicity prediction data and the first immunogenicity experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding third relative error.

[0191] Here, the in vivo experimental methods mentioned in the embodiments of the present invention are the in vivo experimental methods used in the conventional mRNA vaccine design process. These experimental methods evaluate the immunogenicity of candidate mRNA vaccines in animal models. For example, mice or rabbits are selected as animal models, and the candidate mRNA vaccine is administered to the animal models using a standard immunization procedure. Then, the antibody titers corresponding to the injected vaccine are evaluated using methods such as flow cytometry (FACS), enzyme-linked immunospot assay (ELISPOT), and neutralizing antibody assay.

[0192] Step 3239: Identify whether all obtained third relative errors are less than the corresponding third relative error threshold; if not, use all obtained first immunogenicity experimental data as label data, perform a round of parameter optimization on the immunogenicity prediction head of the vaccine characteristic prediction model according to the supervised model training method, and return to step 3231 when the optimization process ends; if yes, use each first mRNA sequence in the preferred set as the corresponding optimized mRNA sequence, and output the corresponding optimized mRNA sequence set composed of all obtained optimized mRNA sequences.

[0193] Here, the third relative error threshold in this embodiment of the invention is a pre-set error threshold.

[0194] Figure 3 This is a module structure diagram of a processing device for an mRNA vaccine design task provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a working model construction module 201, a working model training module 202, and a vaccine design task processing module 203.

[0195] The working model construction module 201 is used to design a working model for predicting three vaccine characteristics of mRNA sequences, denoted as the corresponding vaccine characteristic prediction model, with the Uni-RNA model as the feature extraction backbone. The three vaccine characteristics include stability, translation efficiency, and immunogenicity. The vaccine characteristic prediction model includes the Uni-RNA model, three basic feature prediction heads, a feature fusion module, and three vaccine characteristic prediction heads.

[0196] The working model training module 202 is used to collect a large amount of antigen protein and its corresponding mRNA vaccine information through big data collection to construct the corresponding antigen vaccine dataset; and to train the vaccine characteristic prediction model based on the antigen vaccine dataset; and after the model training is completed, to identify the maximum correlation between the three output features of the three basic feature prediction heads and the three output features of the three vaccine characteristic prediction heads, and to determine the corresponding three objective functions based on the identification results; the three objective functions include the stability objective function, the translation efficiency objective function, and the immunogenicity objective function.

[0197] The vaccine design task processing module 203 receives the molecular sequence of the target antigen protein input by the user as the corresponding first antigen sequence; initializes the mRNA sequence according to the first antigen sequence to obtain the corresponding first mRNA sequence set; and iteratively optimizes the first mRNA sequence set according to the multi-objective optimization method of the NSGA-II algorithm, using a vaccine characteristic prediction model, three objective functions, and corresponding in vitro / in vivo experimental methods to obtain the corresponding optimized mRNA sequence set and feeds it back to the user; the first mRNA sequence set consists of multiple first mRNA sequences; the optimized mRNA sequence set consists of one or more optimized mRNA sequences.

[0198] The present invention provides a processing device for mRNA vaccine design tasks, which can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0199] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the working model construction module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and called and executed by a processing element of the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0200] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).

[0201] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0202] Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 4 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.

[0203] exist Figure 4The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0204] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0205] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.

[0206] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing mRNA vaccine design tasks. As described above, this invention pre-designs a working model (denoted as the vaccine characteristic prediction model) for predicting three vaccine characteristics (stability, translation efficiency, and immunogenicity) of mRNA sequences, using a pre-trained Uni-RNA model as the feature extraction backbone. This model is trained on an antigen vaccine dataset constructed through large-scale data collection. The maximum correlation between the three basic features (codon usage frequency, RNA secondary structure, and GC content) that the model can predict and the three vaccine characteristics (stability, translation efficiency, and immunogenicity) is identified, and three objective functions (stability objective function, translation efficiency objective function, and immunogenicity objective function) are determined based on the identification results. In the specific process of mRNA vaccine design: This embodiment of the invention first uses an mRNA sequence generation tool to generate batches of transcribed mRNA sequences of the target antigen protein to obtain an initialized mRNA sequence set; then, using a vaccine characteristic prediction model and three objective functions, the mRNA sequence set is iteratively optimized and screened according to the NSGA-II algorithm's multi-objective optimization method to obtain a preferred set with a reduced number of sequences; next, the stability and translation efficiency of the preferred set are experimentally analyzed based on in vitro experimental methods; if the error between the in vitro experimental analysis results and the model prediction results is large, the model is automatically optimized based on the experimental data, and the mRNA sequence set is iteratively optimized again after optimization; if the error between the in vitro experimental analysis results and the model prediction results is controllable, the immunogenicity of the preferred set is further analyzed based on in vivo experimental methods; if the error between the in vivo experimental analysis results and the model prediction results is large, the model is automatically optimized again based on the experimental data, and the mRNA sequence set is iteratively optimized again after optimization; if the error between the in vivo experimental analysis results and the model prediction results is controllable, the latest preferred set is output as the final design candidate set. The mRNA sequence initialization method of this invention overcomes the limitations of expert experience and improves the structural diversity of mRNA sequences; the vaccine characteristic prediction model of this invention reduces the number of experimental analyses, shortens the design cycle, and improves design efficiency; the model training method and multi-objective optimization process of this invention effectively ensure the design quality of mRNA sequences.

[0207] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0208] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for processing mRNA vaccine design tasks, characterized in that, The method includes: A working model for predicting three vaccine characteristics of mRNA sequences is designed using the Uni-RNA model as the feature extraction backbone, denoted as the corresponding vaccine characteristic prediction model. The three vaccine characteristics include stability, translation efficiency, and immunogenicity. The vaccine characteristic prediction model includes the Uni-RNA model, three basic feature prediction heads, a feature fusion module, and three vaccine characteristic prediction heads. A large amount of antigen protein and its corresponding mRNA vaccine information was collected through big data acquisition methods to construct a corresponding antigen vaccine dataset. The vaccine characteristic prediction model was trained based on the antigen vaccine dataset. After the model training was completed, the maximum correlation between the three output features of the three basic feature prediction heads and the three output features of the three vaccine characteristic prediction heads was identified, and three corresponding objective functions were determined based on the identification results. The three objective functions include a stability objective function, a translation efficiency objective function, and an immunogenicity objective function. The system receives the molecular sequence of the target antigen protein input by the user as the corresponding first antigen sequence; and initializes the mRNA sequence based on the first antigen sequence to obtain the corresponding first mRNA sequence set; then, using the multi-objective optimization method of the NSGA-II algorithm, iteratively optimizes the first mRNA sequence set with the aid of the vaccine characteristic prediction model, the three objective functions, and the corresponding in vitro / in vivo experimental methods to obtain the corresponding optimized mRNA sequence set, which is then fed back to the user; the first mRNA sequence set consists of multiple first mRNA sequences; and the optimized mRNA sequence set consists of one or more optimized mRNA sequences. The vaccine characteristic prediction model is used to predict the stability, translation efficiency, and immunogenicity characteristics of the input mRNA sequence and output the corresponding stability prediction data, translation efficiency prediction data, and immunogenicity prediction data. The three basic feature prediction heads include a codon usage frequency prediction head, an RNA secondary structure prediction head, and a GC content prediction head; the three vaccine characteristic prediction heads include a stability prediction head, a translation efficiency prediction head, and an immunogenicity prediction head. The Uni-RNA model has completed model pre-training; the Uni-RNA model is used to perform feature encoding processing on the input mRNA sequence to obtain the corresponding feature encoding vector, which is then sent to the codon usage frequency prediction head, the RNA secondary structure prediction head, and the GC content prediction head; The codon usage frequency prediction head is used to predict the codon usage frequency at each base position on the mRNA sequence based on the feature encoding vector, and then send the corresponding codon usage frequency prediction vector to the feature fusion module. The RNA secondary structure prediction head is used to predict the secondary structure features of each base position on the mRNA sequence based on the feature encoding vector, obtain the corresponding secondary structure prediction vector, and send it to the feature fusion module. The GC content prediction head is used to predict the GC content at each base position on the mRNA sequence based on the feature encoding vector, obtain the corresponding GC content prediction vector, and send it to the feature fusion module. The feature fusion module is used to concatenate the codon usage frequency prediction vector, the secondary structure prediction vector, and the GC content prediction vector to obtain a corresponding concatenated vector, which is then sent to the stability prediction head, the translation efficiency prediction head, and the immunogenicity prediction head. The stability prediction head is used to predict the stability characteristics of the mRNA sequence based on the splicing vector and output the corresponding stability prediction data. The translation efficiency prediction head is used to predict the translation features of the mRNA sequence based on the splicing vector and output the corresponding translation efficiency prediction data. The immunogenicity prediction head is used to predict the immunogenicity characteristics of the mRNA sequence based on the splicing vector and output the corresponding immunogenicity prediction data.

2. The method for processing mRNA vaccine design tasks according to claim 1, characterized in that, The model input terminal of the vaccine characteristic prediction model is used to receive the corresponding mRNA sequence, and the first model output terminal, the second model output terminal, and the third model output terminal are used to output the corresponding stability prediction data, the translation efficiency prediction data, and the immunogenicity prediction data, respectively. The input terminal of the Uni-RNA model is connected to the model input terminal, and the output terminals are respectively connected to the input terminals of the codon usage frequency prediction head, the RNA secondary structure prediction head, and the GC content prediction head; the output terminals of the codon usage frequency prediction head, the RNA secondary structure prediction head, and the GC content prediction head are connected to the first, second, and third input terminals of the feature fusion module; the output terminal of the feature fusion module is respectively connected to the input terminals of the stability prediction head, the translation efficiency prediction head, and the immunogenicity prediction head; the output terminals of the stability prediction head, the translation efficiency prediction head, and the immunogenicity prediction head are respectively connected to the corresponding first, second, and third model output terminals; The codons are implemented using a frequency prediction head based on a nonlinear regression model or a neural network model; the nonlinear regression model includes at least an MLP model, an NNR model, a GBDT model, and an XGBoost model; the neural network model includes at least an MLP model, a CNN model, a ResNet model, and a GNN model. The RNA secondary structure prediction head is implemented based on another type of neural network model described above; The GC content prediction head is implemented based on another type of nonlinear regression model or another type of neural network model; The stability prediction head is implemented based on another type of nonlinear regression model or another type of neural network model; The translation efficiency prediction head is implemented based on another type of nonlinear regression model or another type of neural network model; The immunogenicity prediction head is implemented based on another type of nonlinear regression model or another type of neural network model.

3. The method for processing mRNA vaccine design tasks according to claim 2, characterized in that, The antigen vaccine dataset includes multiple antigen vaccine data records; the antigen vaccine data records include antigen molecule sequences, mRNA molecule sequences, codon usage frequency tag vectors, secondary structure tag vectors, GC content tag vectors, stability tag data, translation efficiency tag data, and immunogenicity tag data.

4. The method for processing mRNA vaccine design tasks according to claim 3, characterized in that, The training of the vaccine characteristic prediction model based on the antigen vaccine dataset specifically includes: The mRNA molecule sequence, codon usage frequency tag vector, secondary structure tag vector, and GC content tag vector of each antigen vaccine data record in the antigen vaccine dataset are extracted to form a corresponding first data record; and all the obtained first data records are combined to form a corresponding first dataset. The mRNA molecule sequence, stability tag data, translation efficiency tag data, and immunogenicity tag data of each antigen vaccine data record in the antigen vaccine dataset are extracted to form a corresponding second data record; and all the obtained second data records form a corresponding second dataset; The first stage of model training is performed on the Uni-RNA model and the three basic feature prediction heads of the vaccine characteristic prediction model based on the first dataset; after the first round of model training, the second stage of model training is performed on the three vaccine characteristic prediction heads of the vaccine characteristic prediction model based on the second dataset; and the model training is confirmed to be completed after the second round of model training.

5. The method for processing mRNA vaccine design tasks according to claim 4, characterized in that, The first-stage model training based on the Uni-RNA model and the three basic feature prediction heads of the vaccine characteristic prediction model using the first dataset specifically includes: Step 51: Based on a preset first segmentation ratio, the first dataset is randomly divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 52: Extract the first data record of the first training set as the corresponding current training record; Step 53: Denote the codon usage frequency label vector, the secondary structure label vector, and the GC content label vector of the current training record as the corresponding first label vector P. tag1 The second label vector P tag2 and the third label vector P tag3 ; Step 54: Input the currently recorded mRNA molecular sequence into the Uni-RNA model for feature encoding processing to obtain the corresponding first feature encoding vector; and input the first feature encoding vector into the codon usage frequency prediction head, the RNA secondary structure prediction head, and the GC content prediction head for prediction processing to obtain the corresponding first prediction vector P. pre1 Second prediction vector P pre2 and the third prediction vector P pre3 ; Step 55, the first prediction vector P pre1 and the corresponding first label vector P tag1 Substitute the preset first loss function L a and the second prediction vector P pre2 and the corresponding second label vector P tag2 Substitute the preset second loss function L b and the third prediction vector P pre3 and the corresponding third label vector P tag3 Substitute the preset third loss function L c And by the first loss function L a The second loss function L b and the third loss function L c The summation constitutes the corresponding first model loss function L. M1 The function values ​​of the first, second, and third loss functions, as well as the first model loss function, are calculated and the calculation results are used as the corresponding first loss value, second loss value, third loss value, and first overall loss value. Wherein, the first model loss function L M1 for: ; The first loss function L a It includes at least the L1 loss function and the L2 loss function; the second loss function L b It consists of a base type loss function and a base structure loss function, wherein the base type loss function includes at least a cross-entropy loss function, and the base structure loss function includes at least an L1 loss function and an L2 loss function; the third loss function L... c It should include at least the L1 loss function and the L2 loss function; Step 56: Identify whether the first, second, and third loss values ​​and the first overall loss value all satisfy their respective first loss value range, second loss value range, third loss value range, and first overall loss value range; if the first, second, and third loss values ​​and the first overall loss value all satisfy their respective loss value ranges, proceed to step 57; if the first loss value does not satisfy the first loss value range, then based on the preset first model optimizer, move towards making the first loss function L... a The Uni-RNA model and the model parameters of the codon frequency prediction head are modulated in one round to reach the minimum value, and the process returns to step 54 at the end of this round of parameter modulation; if the second loss value does not meet the second loss value range, then based on the preset second model optimizer, the process moves towards making the second loss function L... b The model parameters of the Uni-RNA model and the RNA secondary structure prediction head are modulated in one round to reach the minimum value, and the process returns to step 54 at the end of this round of parameter modulation; if the third loss value does not meet the range of the third loss value, then based on the preset third model optimizer, the process moves towards making the third loss function L... c The model parameters of the Uni-RNA model and the GC content prediction head are modulated in one round to reach the minimum value, and the process returns to step 54 at the end of this round of parameter modulation; if the first overall loss value does not meet the first overall loss value range, then based on the preset fourth model optimizer, the process moves towards making the first model loss function L... M1 The model parameters of the Uni-RNA model and the three basic feature prediction heads are modulated in one round in the direction of reaching the minimum value, and the process returns to step 54 when the parameter modulation ends. Among them, the first, second, third and fourth model optimizers all include at least the SGD optimizer and the ADAM optimizer; Step 57: Identify whether the current training record is the last first data record in the first training set; if yes, proceed to step 58; if no, extract the next first data record in the first training set as the new current training record and return to step 53. Step 58: Perform a traversal of all the first data records in the first evaluation set; during this traversal, the currently traversed first data record is taken as the corresponding current evaluation record; the mRNA molecular sequence of the current evaluation record is input into the Uni-RNA model for feature encoding processing to obtain the corresponding second feature encoding vector; the second feature encoding vector is input into the three basic feature prediction heads for prediction processing to obtain the corresponding fourth, fifth, and sixth prediction vectors to form the corresponding first prediction vector sequence; the codon usage frequency tag vector, the secondary structure tag vector, and the GC content tag vector of the current evaluation record are recorded as the corresponding fourth, fifth, and sixth tag vectors to form the corresponding first tag vector sequence; the first prediction vector sequence and the first tag vector sequence form the corresponding first prediction-tag pair; and at the end of this traversal, all the obtained first prediction-tag pairs are input into a preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function includes at least the RMSE function; Step 59: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 51 to continue training; if it meets the range, confirm that the first stage of model training has ended.

6. The method for processing mRNA vaccine design tasks according to claim 4, characterized in that, The second-stage model training based on the second dataset for the three vaccine characteristic prediction heads of the vaccine characteristic prediction model specifically includes: Step 61: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set; The second training set and the second evaluation set are both composed of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second segmentation ratio. Step 62: Extract the first second data record of the second training set as the corresponding current training record; Step 63: Record the stability label data, translation efficiency label data, and immunogenicity label data of the current training record as the corresponding first label data D. tag1 Second label data D tag2 and third label data D tag3 ; Step 64: Input the mRNA molecular sequence recorded in the current training record into the vaccine characteristic prediction model for prediction processing to obtain the corresponding first prediction data D. pre1 Second prediction data D pre2 and the third prediction vector D pre3 ; Step 65, the first predicted data D pre1 and the corresponding first tag data D tag1 Substitute the preset fourth loss function L e and the second predicted data D pre2 and the corresponding second tag data D tag2 Substitute the preset fifth loss function L f and the third prediction vector D pre3 and the corresponding third tag data D tag3 Substitute the preset sixth loss function L g And by the fourth loss function L e The fifth loss function L f and the sixth loss function L g The summation constitutes the corresponding second model loss function L. M2 The function values ​​of the fourth, fifth, and sixth loss functions, as well as the second model loss function, are calculated and the results are used as the corresponding fourth loss value, fifth loss value, sixth loss value, and second overall loss value. Wherein, the second model loss function L M2 for: ; The fourth loss function L e It includes at least the L1 loss function, the L2 loss function, and the cross-entropy loss function; the fifth loss function L f It includes at least the L1 loss function and the L2 loss function; the sixth loss function L g It should include at least the L1 loss function, the L2 loss function, and the cross-entropy loss function; Step 66: Identify whether the fourth, fifth, and sixth loss values ​​and the second overall loss value all satisfy their respective corresponding ranges for the fourth, fifth, sixth, and second overall loss values; if the fourth, fifth, and sixth loss values ​​and the second overall loss value all satisfy their respective ranges for the loss value, proceed to step 67; if the fourth loss value does not satisfy the fourth loss value range, then based on the preset fifth model optimizer, move towards making the fourth loss function L... e The model parameters of the stability prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 64 when this round of parameter modulation ends; if the fifth loss value does not meet the range of the fifth loss value, then the preset sixth model optimizer is used to move towards the direction that makes the fifth loss function L f The model parameters of the translation efficiency prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 64 at the end of this round of parameter modulation; if the sixth loss value does not meet the range of the sixth loss value, then the preset seventh model optimizer is used to optimize the sixth loss function L. g The model parameters of the immunogenicity prediction head are modulated once in the direction that reaches the minimum value, and the process returns to step 64 when this round of parameter modulation ends; if the second overall loss value does not meet the range of the second overall loss value, then based on the preset eighth model optimizer, the process moves towards making the second model loss function L... M2 The model parameters of the three vaccine characteristic prediction heads are modulated once in the direction that reaches the minimum value, and the process returns to step 64 when the parameter modulation ends. The fifth, sixth, seventh, and eighth model optimizers all include at least an SGD optimizer and an ADAM optimizer; Step 67: Identify whether the current training record is the last second data record of the second training set; if yes, proceed to step 68; if no, extract the next second data record of the second training set as the new current training record and return to step 63. Step 68: Perform a traversal of all the second data records in the second evaluation set; during this traversal, the currently traversed second data record is taken as the corresponding current evaluation record; the mRNA molecular sequence of the current evaluation record is input into the vaccine characteristic prediction model for prediction processing to obtain the corresponding fourth, fifth, and sixth prediction data to form the corresponding first prediction data sequence; the stability tag data, translation efficiency tag data, and immunogenicity tag data of the current evaluation record are recorded as the corresponding fourth, fifth, and sixth tag data to form the corresponding first tag data sequence; the first prediction data sequence and the first tag data sequence form the corresponding second prediction-tag pair; and at the end of this traversal, all the obtained second prediction-tag pairs are input into a preset second model evaluation function to calculate the corresponding second evaluation value; The second model evaluation function includes at least the RMSE function; Step 69: Identify whether the second evaluation value meets the preset second evaluation value range; if not, return to step 61 to continue training; if it does, confirm that the second stage of model training has ended.

7. The method for processing mRNA vaccine design tasks according to claim 3, characterized in that, After model training, the maximum correlation between the three output features of the three basic feature prediction heads and the three output features of the three vaccine characteristic prediction heads is identified, and three corresponding objective functions are determined based on the identification results. Specifically, this includes: Step 71: After the model training is completed, the mapping relationship between the three basic feature prediction vectors generated by the vaccine characteristic prediction model during the model prediction process and the three characteristic prediction data finally output by the model is recorded as the corresponding stability mapping relationship F. A Translation efficiency mapping relationship F B Mapping relationship with immunogenicity F C ; where y a = F A (x a , x b , x c ), y b = F B (x a , x b , x c ), y c = F C (x a , x b , x c ); Independent variable x a x b x c These are the codon usage frequency prediction vector, the secondary structure prediction vector, and the GC content prediction vector generated during the model prediction process, respectively, with the dependent variable y. a y b y c These are the stability prediction data, translation efficiency prediction data, and immunogenicity prediction data that are ultimately output by the model, respectively. Step 72, denot each of the antigen vaccine data records in the antigen vaccine dataset as record R. i 1≤i≤N, where N is the total number of records in the antigen vaccine dataset; and each of the records R i The mRNA molecule sequence is input into the vaccine characteristic prediction model for prediction processing, resulting in six corresponding prediction results, x. a,i x b,i x c,i y a,i y b,i y c,i And the prediction result x corresponding to each index i. a,i x b,i x c,i y a,i A corresponding first independent / dependent variable data set is formed, and the resulting N first independent / dependent variable data sets form a corresponding first data set set; and the prediction result x corresponding to each index i is... a,i x b,i x c,i y b,i A corresponding second independent / dependent variable data set is formed, and the resulting N second independent / dependent variable data sets form a corresponding second data set set; and the prediction result x corresponding to each index i is... a,i x b,i x c,i y c,i A corresponding third independent / dependent variable data set is formed, and the N obtained third independent / dependent variable data sets form a corresponding third data set set; Step 73: Based on the SHAP evaluation algorithm, evaluate the independent variable x according to the first data set. a x b x c For the dependent variable y a The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y a The largest correlation variable x ma Based on the SHAP evaluation algorithm, the independent variable x is evaluated according to the second data set. a x b x c For the dependent variable y b The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y b The largest correlation variable x mb Based on the SHAP evaluation algorithm, the independent variable x is evaluated according to the third data set. a x b x c For the dependent variable y c The contribution coefficients and the independent variable x with the largest contribution coefficient. a x b or x c As the dependent variable y c The largest correlation variable x mc ; Step 74, the dependent variable y a Its corresponding largest correlation variable x ma Substitute the preset correlation function F corr The corresponding stability objective function F is obtained. corr (x ma ,y a ); and the dependent variable y b Its corresponding largest correlation variable x mb Substitute the correlation function F corr The corresponding translation efficiency objective function F is obtained. corr (x mb ,y b ); and the dependent variable y c Its corresponding largest correlation variable x mc Substitute the correlation function F corr The corresponding immunogenicity objective function F is obtained. corr (x mc ,y c The three objective functions are composed of the stability objective function, the translation efficiency objective function, and the immunogenicity objective function; the correlation function F corr It should include at least the Pearson correlation function, the Spearman correlation function, and the Kendall correlation function.

8. The method for processing mRNA vaccine design tasks according to claim 1, characterized in that, The step of initializing the mRNA sequence based on the first antigen sequence to obtain the corresponding first mRNA sequence set specifically includes: The transcribed mRNA sequence of the first antigen sequence is processed using a preset mRNA sequence generation tool to generate multiple corresponding first mRNA sequences, forming a corresponding first mRNA sequence set. The mRNA sequence generation tool is either an mRNA sequence reverse translation tool or an mRNA sequence reverse decoding model. The mRNA sequence reverse translation tool includes at least the ExPASy Translate Tool, used for reverse translating and outputting the transcribed mRNA sequence corresponding to a specified antigen sequence. The mRNA sequence reverse decoding model includes at least a reverse decoding model based on GANs models, used for reverse decoding and outputting the transcribed mRNA sequence corresponding to a specified antigen sequence.

9. The method for processing mRNA vaccine design tasks according to claim 1, characterized in that, The multi-objective optimization method based on the NSGA-II algorithm, using the vaccine characteristic prediction model, the three objective functions, and corresponding in vitro / in vivo experimental methods, iteratively optimizes the first mRNA sequence set to obtain an optimized mRNA sequence set, which is then fed back to the user. Specifically, this includes: Step 91: Initialize the first iteration counter to 1; Step 92: Input each of the first mRNA sequences in the first mRNA sequence set into the vaccine characteristic prediction model for prediction, and input the prediction results into the three objective functions to calculate the corresponding first objective function value, second objective function value, and third objective function value; and sort all the obtained first objective function values, all the second objective function values, and all the third objective function values ​​from high to low to obtain the corresponding first, second, and third function value sequences. Step 93: According to the non-dominated level and crowding distance identification rules in the NSGA-II algorithm, the non-dominated level and crowding distance corresponding to each first mRNA sequence are confirmed based on the first, second, and third function value sequences; and based on the tournament selection strategy in the NSGA-II algorithm, individual mRNA sequences are selected according to the non-dominated level and crowding distance of each first mRNA sequence, and all selected first mRNA sequences form the corresponding first body set; and based on the preset individual crossover rules, individual crossover processing is performed on all first mRNA sequences in the first body set, and individual mutation processing is performed on each first mRNA sequence that has completed the crossover; and after all first mRNA sequences in the first body set have completed individual mutation, the first body set and the first mRNA sequence set are merged to obtain a new first mRNA sequence set; The individual crossover rules include single-point crossover and multi-point crossover; Step 94: Increment the first iteration counter by 1; and identify whether the incremented first iteration counter exceeds a preset iteration counter threshold; if yes, proceed to step 95; if no, return to step 92. Step 95: Input each of the first mRNA sequences in the latest first mRNA sequence set into the vaccine characteristic prediction model to obtain the corresponding first stability prediction data, first translation efficiency prediction data and first immunogenicity prediction data; and sort all the obtained first stability prediction data and all the first translation efficiency prediction data according to the preset stability and translation efficiency ranking rules, and extract one or more first mRNA sequences whose stability and translation efficiency rankings are both above the specified ranking to form the corresponding preferred set. Step 96: The stability and translation efficiency of each first mRNA sequence in the preferred set are analyzed using in vitro experiments to obtain corresponding first stability experimental data and first translation efficiency experimental data; the relative error between the first stability prediction data and the first stability experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding first relative error; and the relative error between the first translation efficiency prediction data and the first translation efficiency experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding second relative error. Step 97: Identify whether all the obtained first and second relative errors are less than the corresponding first and second relative error thresholds; if yes, proceed to step 98; if no, use all the first stability experimental data and the first translation efficiency experimental data obtained this time as label data, perform a round of parameter optimization on the stability prediction head and the translation efficiency prediction head of the vaccine characteristic prediction model according to the supervised model training method, and return to step 91 when the end of this round of optimization. Step 98: The immunogenicity characteristics of each first mRNA sequence in the preferred set are analyzed by in vivo experimental methods to obtain the corresponding first immunogenicity experimental data; and the relative error between the first immunogenicity prediction data and the first immunogenicity experimental data of each first mRNA sequence in the preferred set is calculated to obtain the corresponding third relative error. Step 99: Identify whether all the obtained third relative errors are less than the corresponding third relative error threshold; if not, use all the first immunogenicity experimental data obtained this time as label data, perform a round of parameter optimization on the immunogenicity prediction head of the vaccine characteristic prediction model according to the supervised model training method, and return to step 91 when the optimization process ends; if yes, use each of the first mRNA sequences in the preferred set as the corresponding optimized mRNA sequences, and output the corresponding optimized mRNA sequence set composed of all the obtained optimized mRNA sequences.

10. An apparatus for performing a processing method for designing an mRNA vaccine according to any one of claims 1-9, characterized in that, The device includes: a working model construction module, a working model training module, and a vaccine design task processing module; The working model construction module is used to design a working model for predicting three vaccine characteristics of mRNA sequences, denoted as the corresponding vaccine characteristic prediction model, with the Uni-RNA model as the feature extraction backbone. The three vaccine characteristics include stability, translation efficiency, and immunogenicity. The vaccine characteristic prediction model includes the Uni-RNA model, three basic feature prediction heads, a feature fusion module, and three vaccine characteristic prediction heads. The working model training module is used to collect a large amount of antigen protein and its corresponding mRNA vaccine information through big data collection to construct a corresponding antigen vaccine dataset; and to train the vaccine characteristic prediction model based on the antigen vaccine dataset; and after the model training is completed, to identify the maximum correlation between the three output features of the three basic feature prediction heads and the three output features of the three vaccine characteristic prediction heads, and to determine the corresponding three objective functions based on the identification results; the three objective functions include a stability objective function, a translation efficiency objective function, and an immunogenicity objective function; The vaccine design task processing module receives the molecular sequence of the target antigen protein input by the user as the corresponding first antigen sequence; initializes the mRNA sequence according to the first antigen sequence to obtain the corresponding first mRNA sequence set; and iteratively optimizes the first mRNA sequence set using the multi-objective optimization method of the NSGA-II algorithm, with the help of the vaccine characteristic prediction model, the three objective functions, and the corresponding in vitro / in vivo experimental methods, to obtain the corresponding optimized mRNA sequence set and feed it back to the user; the first mRNA sequence set consists of multiple first mRNA sequences; and the optimized mRNA sequence set consists of one or more optimized mRNA sequences.

11. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Tumor neoantigen prediction method and system, electronic equipment and storage medium

    CN117690495A

  • MRNA vaccine sequence optimization method and system based on mRVBERT model

    CN118248217A