A processing method and device for a tRNA adaptation index weight prediction model
By constructing an adaptive index weight prediction model and a downstream task model, the limitations and instability of tRNA adaptive index weights among different species are solved, and the accurate prediction of tRNA adaptive index weights and gene expression levels under different conditions is achieved, which improves the universality and stability of the model.
Patent Information
- Application Number
- CN202411470052.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-10-21
AI Technical Summary
The existing tRNA adaptation index weight derivation methods show great limitations and instability among different species, which is difficult to accurately reflect the degree of matching codons and tRNA abundance, and affect the study of gene expression regulation and protein synthesis efficiency.
The adaptation index weight prediction model and three downstream task models were constructed. By collecting and training the genome sequences of the target species, an end-to-end prediction model was established to predict the tRNA adaptation index weight under different conditions and further predict the gene expression level.
Reduces processing complexity, improves model universality and species adaptability, enhances prediction stability, and can adapt to different species without changing the model structure when switching target species.
Smart Images

Figure CN119207560B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a processing method and device for a tRNA adaptation index weight prediction model. Background Art
[0002] The tRNA adaptation index (tAI) weight is used to reflect the matching degree between each codon and its corresponding tRNA abundance in a given species. The tRNA adaptation index weight has important applications in the research of gene expression regulation, protein synthesis efficiency, and evolutionary biology. However, deriving the tRNA adaptation index weight for each codon has always been a difficult point in bioinformatics research. Most of the existing derivation methods rely on experimental data or are based on simplified theoretical models, and these methods often show great limitations and instabilities when facing different species. Summary of the Invention
[0003] The objective of the present invention is to provide a processing method, device, electronic device, and computer-readable storage medium for a tRNA adaptation index weight prediction model in view of the defects of the prior art. The present invention constructs an adaptation index weight prediction model for processing tRNA adaptation index weight prediction tasks and three downstream task models (transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model); and collects genomic sequences of a certain type of target species under different conditions and constructs a dataset based on the collected data to train the adaptation index weight prediction model and the three downstream task models; after the training is completed, the adaptation index weight prediction model is used to predict the tRNA adaptation index weight of the current target species under the current given conditions, and the gene expression levels (transcription efficiency, translation efficiency, gene abundance) of a given target gene (DNA gene, synthetic gene) are further predicted according to the tRNA adaptation index weight through the three downstream task models. The four types of models provided by the present invention (adaptation index weight prediction model, transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model) are all end-to-end prediction models. Predicting the tRNA adaptation index weight and gene expression levels based on these four types of models can reduce the processing complexity and improve the model generality; when switching the target species for the four types of models of the present invention, there is no need to change the model structure, and only one round of training is required using the dataset of the current species, thereby not only improving the species adaptability of the model, but also further improving the prediction stability of each type of species through training.
[0004] To achieve the above objective, in the first aspect of the embodiments of the present invention, a processing method for a tRNA adaptation index weight prediction model is provided, and the method includes:
[0005] Construct an adaptation index weight prediction model for processing the tRNA adaptation index weight prediction task; and design three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model;
[0006] Collect genomic sequences of the target species under different conditions and construct a first data set based on the collected data; and perform model training on the adaptation index weight prediction model and the three downstream task models based on the first data set;
[0007] After the model training is completed, receive the current experimental conditions, the current genomic sequence, and the current target gene input by the user; and preprocess the current experimental conditions, the current genomic sequence, and the current target gene to obtain corresponding current experimental condition encoding vectors, current genomic sequence encoding vectors, current tRNA gene copy number vectors, and current target gene codon index sets; the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence;
[0008] The adaptation index weight prediction model performs tRNA adaptation index weight prediction processing on the current genomic sequence encoding vector, the current experimental condition encoding vector, and the current tRNA gene copy number vector to obtain a corresponding current adaptation index weight vector; and the three downstream task models respectively perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing on the current adaptation index weight vector and the current target gene codon index set to obtain corresponding current transcription efficiency, current translation efficiency, and current gene abundance, and form a corresponding task report to feedback to the user.
[0009] Preferably, the adaptation index weight prediction model is used to perform tRNA adaptation index weight prediction processing on the input genomic sequence encoding vector A, experimental condition encoding vector B, and tRNA gene copy number vector C and output a corresponding adaptation index weight vector W;
[0010] The genomic sequence encoding vector A is composed of a codon encoding vector A C and an anticodon encoding vector A AC spliced together; the codon encoding vector A C is formed by sequentially sorting the codon encodings c i of all codon types of the corresponding genomic sequence; the anticodon encoding vector A AC is formed by sequentially sorting the anticodon encodings ac j of all anticodon types of the corresponding genomic sequence; the codon encoding c iThe total number is denoted as N C , and the anticodon encodes ac j The total number is denoted as N AC , 1 ≤ i ≤ N C , 1 ≤ j ≤ N AC ; the codon encodes c i and the anticodon encodes ac j Each is a triple-base encoding;
[0011] The experimental condition encoding vector B includes at least temperature encoding, humidity encoding, light encoding, nutrient condition encoding, physiological / culture stage encoding;
[0012] The tRNA gene copy number vector C consists of N AC gene copy numbers tGCN j ; the gene copy number tGCN j corresponds one-to-one with the anticodon encoding ac j ;
[0013] The adaptation index weight vector W consists of N C adaptation index weights w i ; the adaptation index weight w i corresponds one-to-one with the codon encoding c i ;
[0014] The first model input end, the second model input end, and the third model input end of the adaptation index weight prediction model are respectively used to receive the corresponding genome sequence encoding vector A, the experimental condition encoding vector B, and the tRNA gene copy number vector C, and the model output end is used to output the corresponding adaptation index weight vector W;
[0015] The adaptation index weight prediction model includes a Uni-RNA model, a first MLP model, a second MLP model, a feature fusion module, a third MLP model, and an adaptation index weight estimation module;
[0016] The input end of the Uni-RNA model is connected to the first model input end, and the output ends are respectively connected to the input ends of the first and second MLP models; the output ends of the first and second MLP models are respectively connected to the second and third input ends of the feature fusion module; the first input end of the feature fusion module is connected to the second model input end, and the output end is connected to the input end of the third MLP model; the output end of the third MLP model is connected to the second input end of the adaptation index weight estimation module; the first input end of the adaptation index weight estimation module is connected to the third model input end, and the output end is connected to the model output end;
[0017] The Uni-RNA model has completed pre-training; the Uni-RNA model is used to perform feature encoding on the genomic sequence encoding vector A to obtain the corresponding encoded feature X and send it to the first and second MLP models;
[0018] The first MLP model is used to predict the intracellular concentration of tRNA genes corresponding to each anticodon type according to the encoded feature X to obtain the corresponding concentration prediction vector P and send it to the feature fusion module; the concentration prediction vector P is composed of N AC predicted concentrations p j ; the predicted concentration p j corresponds one-to-one with the anticodon encoding ac j ;
[0019] The second MLP model is used to predict the pairing probability of each codon-anticodon type pair according to the encoded feature X to obtain the corresponding pairing probability matrix M and send it to the feature fusion module; the shape of the pairing probability matrix M is N C ×N AC ; it is composed of N C ×N AC pairing probabilities m i,j ; each pairing probability m i,j corresponds to a group of the codon encoding c i and the anticodon encoding ac j ;
[0020] The feature fusion module is used to perform dimensionality reduction on the pairing probability matrix M to obtain the corresponding pairing probability vector V M ; and perform vector concatenation on the pairing probability vector V M , the concentration prediction vector P and the experimental condition encoding vector B to obtain the corresponding concatenated vector R and send it to the third MLP model;
[0021] The third MLP model is used to predict the coupling coefficient of each codon-anticodon type pair according to the concatenated vector R to obtain the corresponding coupling coefficient tensor S and send it to the adaptation index weight estimation module; the shape of the coupling coefficient tensor S is N C ×N AC ; it is composed of N C ×N AC coupling coefficients s i,j ; each coupling coefficient s i,j corresponds to a group of the codon encoding c i and the anticodon encoding ac j ;
[0022] The adaptation index weight estimation module is used to estimate the adaptation index weights of each codon type according to the tRNA gene copy number vector C and the coupling coefficient tensor S to obtain the corresponding adaptation index weight vector W; the adaptation index weight vector W consists of N C adaptation index weights w i ; the adaptation index weight w i corresponds one-to-one with the codon encoding c i ; the estimation formula of the adaptation index weight w i is:
[0023] Preferably, the transcription efficiency prediction model is used to predict the transcription efficiency of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding transcription efficiency TPE;
[0024] The target gene codon index set G consists of codon indexes corresponding to all codon types of the specified target gene;
[0025] The transcription efficiency prediction model includes a first screening module and a fourth MLP model; the input end of the first screening module is connected to the input end of the transcription efficiency prediction model, and the output end is connected to the input end of the fourth MLP model; the output end of the fourth MLP model is connected to the output end of the transcription efficiency prediction model;
[0026] The first screening module is used to extract the adaptation index weights w i in the adaptation index weight vector W whose index i matches each of the codon indexes in the target gene codon index set G to form a corresponding first weight vector and send it to the fourth MLP model; the fourth MLP model is used to predict the transcription efficiency according to the first weight vector to obtain the corresponding transcription efficiency TPE and output it;
[0027] The translation efficiency prediction model is used to predict the translation efficiency of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding translation efficiency TLE;
[0028] The translation efficiency prediction model includes a second screening module and a fifth MLP model; the input end of the second screening module is connected to the input end of the translation efficiency prediction model, and the output end is connected to the input end of the fifth MLP model; the output end of the fifth MLP model is connected to the output end of the translation efficiency prediction model;
[0029] The second sub-screening module is used to extract the adaptation index weights w where the index i in the adaptation index weight vector W matches each of the codon indices in the target gene codon index set G, and form a corresponding second weight vector to send to the fifth MLP model; the fifth MLP model is used to perform translation efficiency prediction according to the second weight vector to obtain the corresponding gene translation efficiency TLE and output it; i Extract and form a corresponding second weight vector to send to the fifth MLP model; the fifth MLP model is used to perform translation efficiency prediction according to the second weight vector to obtain the corresponding gene translation efficiency TLE and output it;
[0030] The gene abundance prediction model is used to perform prediction processing on the gene abundance of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G, and output the corresponding gene abundance TGA;
[0031] The gene abundance prediction model includes a third sub-screening module and a sixth MLP model; the input end of the third sub-screening module is connected to the input end of the gene abundance prediction model, and the output end is connected to the input end of the sixth MLP model; the output end of the sixth MLP model is connected to the output end of the gene abundance prediction model;
[0032] The third sub-screening module is used to extract the adaptation index weights w where the index i in the adaptation index weight vector W matches each of the codon indices in the target gene codon index set G, i Extract and form a corresponding third weight vector to send to the sixth MLP model; the sixth MLP model is used to perform gene abundance prediction according to the third weight vector to obtain the corresponding gene abundance TGA and output it.
[0033] Preferably, the first data set includes a plurality of first data records;
[0034] Each of the first data records corresponds to a genomic sequence obtained by an observer through manual or machine observation of a type of cell of the target species under a set condition;
[0035] The first data record includes a first genomic sequence, a first set condition, a first codon data table, a first anticodon data table, and a first gene expression data table;
[0036] The data format of the first genomic sequence is a general genomic sequence, genomic map, or gene expression profile;
[0037] The first set condition is the observation / experimental set condition corresponding to the current genomic sequence, including at least temperature, humidity, light, nutritional condition, physiological / culture stage;
[0038] The first codon data table, the first anticodon data table, and the first gene expression data table are all analysis and measurement data obtained by an observer using a preset bioinformatics tool to perform directional analysis and measurement on the first genomic sequence;
[0039] Among them, the bioinformatics tool at least includes tRNAscan-SE software, Galaxy software, SAMtools software, and GATK software;
[0040] The first codon data table includes multiple first codon data; each first codon data corresponds to a codon type in the first genomic sequence; the first codon data includes a codon identifier, a codon triple-base encoding, and a codon usage bias; the codon identifier is the unique identifier of the current codon type; the codon triple-base encoding is the triple-base encoding of the current codon type; the codon usage bias is the usage bias scalar of the current codon type in the first genomic sequence;
[0041] The first anticodon data table includes multiple first anticodon data; each first anticodon data corresponds to an anticodon type in the first genomic sequence; the first anticodon data includes an anticodon identifier, an anticodon triple-base encoding, and the tRNA gene copy number; the anticodon identifier is the unique identifier of the current anticodon type; the anticodon triple-base encoding is the triple-base encoding of the current anticodon type; the tRNA gene copy number is the total number of tRNA genes corresponding to the current anticodon type in the first genomic sequence;
[0042] The first gene expression data table includes multiple first gene expression data; each first gene expression data corresponds to a DNA gene or a synthetic gene in the first genomic sequence; the first gene expression data includes a gene identifier, a gene type, a gene base encoding sequence, a gene codon sequence, a gene transcription efficiency, a gene translation efficiency, and a gene abundance; the gene identifier is the unique identifier of the current gene; the gene type at least includes a DNA gene and a synthetic gene; the gene base encoding sequence is the base encoding sequence of the current gene; the gene codon sequence is composed of the codon identifiers of all codon types of the current gene; the gene transcription efficiency, the gene translation efficiency, and the gene abundance are the transcription efficiency scalar, translation efficiency scalar, and abundance scalar of the current gene under the first set conditions.
[0043] Preferably, the model training of the adaptation index weight prediction model and the three downstream task models based on the first dataset specifically includes:
[0044] Step 501: Prepare training data based on the first dataset to obtain a corresponding second dataset; extract the first training genome sequence encoding vector, the first training experimental condition encoding vector, the first training gene copy number vector, and the first codon usage bias vector of each second data record in the second dataset to form a corresponding third data record; remove duplicates from all the obtained third data records; and form a corresponding third dataset from all the deduplicated third data records.
[0045] Among them, the second dataset includes multiple second data records; the second data record includes a first training genome sequence encoding vector, a first training experimental condition encoding vector, a first training gene copy number vector, a first codon usage bias vector, a first training codon index set, a first label transcription efficiency, a first label translation efficiency, and a first label gene abundance; the third dataset is composed of multiple third data records, and each third data record is composed of the corresponding first training genome sequence encoding vector, the first training experimental condition encoding vector, the first training gene copy number vector, and the first codon usage bias vector.
[0046] Step 502: Randomly split the third dataset based on a preset first splitting ratio to obtain a corresponding first training set and a first evaluation set.
[0047] Among them, both the first training set and the first evaluation set are composed of multiple third data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio.
[0048] Step 503: Take the first third data record in the first training set as the corresponding current training record.
[0049] Step 504: Input the first training genome sequence encoding vector, the first training experimental condition encoding vector, and the first training gene copy number vector of the current training record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain a corresponding first predicted adaptation index weight vector.
[0050] Step 505: Substitute the first predicted adaptation index weight vector and the first codon usage deviation vector of the current training record into a preset first correlation function; and based on a preset first model optimization algorithm, perform one round of parameter modulation on the adaptation index weight prediction model in the direction that maximizes the positive correlation of the first correlation function; and at the end of this round of parameter modulation, identify whether the current training record is the last third data record of the first training set; if so, go to Step 506; if not, extract the next third data record of the first training set as the new current training record and return to Step 504 to continue training;
[0051] Among them, the first correlation function includes at least Pearson correlation function, Spearman correlation function, and Kendall correlation function; the first model optimization algorithm includes at least hill climbing algorithm and genetic algorithm;
[0052] Step 506: Perform one round of traversal on all the third data records of the first evaluation set; and during this round of traversal, use the currently traversed third data record as the corresponding current evaluation record; and input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current evaluation record into the adaptation index weight prediction model to perform tRNA adaptation index weight prediction processing to obtain the corresponding second predicted adaptation index weight vector; and substitute the second predicted adaptation index weight vector and the first codon usage deviation vector of the current training record into the first correlation function for calculation to obtain the corresponding first correlation degree; and at the end of this round of traversal, input all the obtained first correlation degrees into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value;
[0053] Among them, the first model evaluation function includes at least MAE function, MSE function, and RMSE function;
[0054] Step 507: Identify whether the first evaluation value meets a preset first evaluation value range; if the first evaluation value meets the first evaluation value range, identify whether all the obtained first correlation degrees meet a preset first correlation degree range, if not, return to Step 501 to continue training, if so, solidify the model parameters of the adaptation index weight prediction model and go to Step 508; if the first evaluation value does not meet the first evaluation value range, return to Step 501 to continue training;
[0055] Step 508: Randomly divide the second data set based on a preset second segmentation ratio to obtain the corresponding second training set and second evaluation set;
[0056] Among them, both the second training set and the second evaluation set are composed of a plurality of the second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second segmentation ratio;
[0057] Step 509, use the first second data record in the second training set as the corresponding current training record;
[0058] Step 510, input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current training record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding third predicted adaptation index weight vector;
[0059] Step 511, input the third predicted adaptation index weight vector and the first training codon index set of the current training record into the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model respectively for prediction to obtain the corresponding first predicted transcription efficiency, first predicted translation efficiency, and first predicted gene abundance;
[0060] Step 512, substitute the first predicted transcription efficiency and the first labeled transcription efficiency of the current training record into a preset first transcription loss function for calculation to obtain the corresponding first loss value; and substitute the first predicted translation efficiency and the first labeled translation efficiency of the current training record into a preset first translation loss function for calculation to obtain the corresponding second loss value; and substitute the first predicted gene abundance and the first labeled gene abundance of the current training record into a preset first abundance loss function for calculation to obtain the corresponding third loss value;
[0061] Among them, the first transcription loss function, the first translation loss function, and the first abundance loss function all at least include the L1 loss function and the L2 loss function;
[0062] Step 513: Identify whether the first, second, and third loss values meet the preset first, second, and third loss value ranges. If the first, second, and third loss values all meet their corresponding first, second, and third loss value ranges, proceed to Step 514. If the first loss value does not meet the first loss value range, perform one round of parameter modulation on the transcription efficiency prediction model based on the preset first model optimizer in the direction of minimizing the function value of the first transcription loss function, and return to Step 511 at the end of this round of parameter modulation. If the second loss value does not meet the second loss value range, perform one round of parameter modulation on the translation efficiency prediction model based on the preset second model optimizer in the direction of minimizing the function value of the first translation loss function, and return to Step 511 at the end of this round of parameter modulation. If the third loss value does not meet the third loss value range, perform one round of parameter modulation on the gene abundance prediction model based on the preset third model optimizer in the direction of minimizing the function value of the first abundance loss function, and return to Step 511 at the end of this round of parameter modulation.
[0063] Among them, the first, second, and third model optimizers all at least include SGD series optimizers and ADAM series optimizers.
[0064] Step 514: Conduct one round of traversal of all the second data records in the second evaluation set. During this round of traversal, use the currently traversed second data record as the corresponding current evaluation record. Input the first training genomic sequence encoding vector, the first training experimental condition encoding vector, and the first training gene copy number vector of the current evaluation record into the adaptability index weight prediction model to perform tRNA adaptability index weight prediction processing to obtain the corresponding fourth predicted adaptability index weight vector. Input the fourth predicted adaptability index weight vector and the first training codon index set of the current evaluation record into the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model respectively to perform predictions to obtain the corresponding second predicted transcription efficiency, second predicted translation efficiency, and second predicted gene abundance. Combine the second predicted transcription efficiency, the second predicted translation efficiency, and the second predicted gene abundance to form a corresponding first predicted vector, combine the first labeled transcription efficiency, the first labeled translation efficiency, and the first labeled gene abundance of the current evaluation record to form a corresponding first labeled vector, and combine the first predicted vector and the first labeled vector to form a corresponding first prediction-label pair. At the end of this round of traversal, input all the obtained first prediction-label pairs into the preset second model evaluation function for calculation to obtain the corresponding second evaluation value.
[0065] Among them, the second model evaluation function at least includes the MAE function and the RMSE function;
[0066] Step 515, identify whether the second evaluation value meets a preset second evaluation value range; if the second evaluation value meets the second evaluation value range, solidify the model parameters of the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model, and go to step 516; if the second evaluation value does not meet the second evaluation value range, return to step 508 to continue training;
[0067] Step 516, stop model training and confirm that the model training of the adaptation index weight prediction model and the three downstream task models is completed.
[0068] Further, the preparation of training data based on the first data set to obtain a corresponding second data set specifically includes:
[0069] Step 61, use each of the first data records in the first data set as a corresponding current data record;
[0070] Step 62, use the first set conditions, the first codon data table, the first anticodon data table, and the first gene expression data table in the current data record as the corresponding current set conditions, current codon data table, current anticodon data table, and current gene expression data table;
[0071] Step 63, extract all the codon triplet base codes in the current codon data table and sort them sequentially to form a corresponding first codon coding vector; extract all the anticodon triplet base codes in the current anticodon data table and sort them sequentially to form a corresponding first anticodon coding vector; and splice the first codon coding vector and the first anticodon coding vector to obtain a corresponding first training genome sequence coding vector;
[0072] Step 64, based on a preset conditional coding rule, encode the temperature, humidity, light, nutrient conditions, and physiological / culture stage in the current set conditions to obtain corresponding temperature coding, humidity coding, light coding, nutrient condition coding, and physiological / culture stage coding, which together form a corresponding first training experimental condition coding vector;
[0073] Step 65, extract all the tRNA gene copy numbers in the current anticodon data table and sort them sequentially to form a corresponding first training gene copy number vector;
[0074] Step 66: Extract all the codon usage biases in the current codon data table and sort them in order to form a corresponding first codon usage bias vector.
[0075] Step 67: Conduct a round of traversal on all the first gene expression data in the current gene expression data table; during this round of traversal, regard the currently traversed first gene expression data as the corresponding current gene expression data; and regard the gene codon sequence, gene transcription efficiency, gene translation efficiency, and gene abundance of the current gene expression data as the corresponding first training codon index set, the first label transcription efficiency, the first label translation efficiency, and the first label gene abundance; and form a corresponding second data record from the first training genome sequence coding vector, the first training experimental condition coding vector, the first training gene copy number vector, the first codon usage bias vector, and the first training codon index set, the first label transcription efficiency, the first label translation efficiency, and the first label gene abundance corresponding to the current gene expression data; and at the end of this round of traversal, form a first record set corresponding to the current data record from all the second data records obtained in this round of traversal.
[0076] Step 68: Merge the data records of the first record set corresponding to all the first data records to obtain a corresponding second data set.
[0077] Preferably, the preprocessing of the current experimental condition, the current genome sequence, and the current target gene to obtain a corresponding current experimental condition coding vector, current genome sequence coding vector, current tRNA gene copy number vector, and current target gene codon index set specifically includes:
[0078] Step 71: Based on a preset condition coding rule, code the temperature, humidity, light, nutrient condition, and physiological / culture stage in the current experimental condition respectively to obtain corresponding temperature coding, humidity coding, light coding, nutrient condition coding, and physiological / culture stage coding to form the corresponding current experimental condition coding vector.
[0079] Step 72: Use a preset bioinformatics tool to analyze and count the codon types of all mRNA genes in the current genome sequence, and sort all the counted codon types in order to obtain a corresponding first type sequence; and sort the triple-base codings of all the codon types in the first type sequence in order to obtain a corresponding current codon coding vector.
[0080] Step 73: Use the bioinformatics tool to analyze and count the anticodon types of all tRNA genes within the current genomic sequence, and sort all the counted anticodon types in order to obtain a corresponding second type sequence; and sort the triplet base codings of all anticodon types in the second type sequence in order to obtain a corresponding current anticodon coding vector.
[0081] Step 74: Concatenate the current codon coding vector and the current anticodon coding vector to obtain a corresponding current genomic sequence coding vector.
[0082] Step 75: Use the bioinformatics tool to statistically measure the total number of gene replications of each tRNA gene within the current genomic sequence and use the measurement result as the corresponding tRNA gene copy number; and sort all the obtained tRNA gene copy numbers in order to obtain a corresponding current tRNA gene copy number vector.
[0083] Step 76: Use the bioinformatics tool to analyze and count all codon types of the current target gene to obtain a corresponding third type sequence; and extract the sorting index of each codon type in the third type sequence in the first type sequence as a corresponding first codon index; and form a corresponding current target gene codon index set from all the obtained first codon indexes.
[0084] Step 77: Output the current experimental condition coding vector, the current genomic sequence coding vector, the current tRNA gene copy number vector, and the current target gene codon index set obtained this time as the processing result of this preprocessing.
[0085] In the second aspect of the embodiments of the present invention, there is provided a device for implementing the processing method of the tRNA adaptation index weight prediction model described in the first aspect above. The device includes: a model construction module, a model training module, a data preparation module, and a model application module.
[0086] The model construction module is used to construct an adaptation index weight prediction model for processing tRNA adaptation index weight prediction tasks; and design three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model.
[0087] The model training module is used to collect data on genomic sequences of a target species under different conditions and construct a first data set based on the collected data; and perform model training on the adaptation index weight prediction model and the three downstream task models based on the first data set.
[0088] The data preparation module is used to receive the current experimental conditions, the current genomic sequence, and the current target gene input by the user after the model training is completed; and preprocess the current experimental conditions, the current genomic sequence, and the current target gene to obtain the corresponding current experimental condition encoding vector, current genomic sequence encoding vector, current tRNA gene copy number vector, and current target gene codon index set; the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence.
[0089] The model application module is used to perform tRNA adaptation index weight prediction processing on the current genomic sequence encoding vector, the current experimental condition encoding vector, and the current tRNA gene copy number vector by the adaptation index weight prediction model to obtain the corresponding current adaptation index weight vector; and perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing on the current adaptation index weight vector and the current target gene codon index set by three downstream task models respectively to obtain the corresponding current transcription efficiency, current translation efficiency, and current gene abundance, and form a corresponding task report to feedback to the user.
[0090] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0091] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;
[0092] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
[0093] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0094] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for a tRNA adaptation index weight prediction model. As can be seen from the above, the embodiment of the present invention constructs an adaptation index weight prediction model for processing tRNA adaptation index weight prediction tasks and three downstream task models (transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model); and collects genomic sequences of a certain type of target species under different conditions and constructs a dataset based on the collected data to train the adaptation index weight prediction model and the three downstream task models; after the training is completed, the adaptation index weight prediction model is used to predict the tRNA adaptation index weight of the current target species under the current given conditions, and the gene expression levels (transcription efficiency, translation efficiency, gene abundance) of a given target gene (DNA gene, synthetic gene) are further predicted according to the tRNA adaptation index weight through the three downstream task models. The four types of models provided by the embodiment of the present invention (adaptation index weight prediction model, transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model) are all end-to-end prediction models. Predicting the tRNA adaptation index weight and gene expression levels based on these four types of models not only reduces the processing complexity but also improves the model versatility; when switching the target species, the four types of models in the embodiment of the present invention do not need to change the model structure, and only need to be trained once using the dataset of the current species. In this way, not only the species adaptability of the model is improved, but also the prediction stability of the model for each type of species is further improved through training. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 Schematic diagram of a processing method for a tRNA adaptation index weight prediction model provided in Embodiment 1 of the present invention;
[0096] Figure 2 Module structure diagram of the adaptation index weight prediction model provided in Embodiment 1 of the present invention;
[0097] Figure 3 Module structure diagrams of the transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model provided in Embodiment 1 of the present invention;
[0098] Figure 4 Module structure diagram of a processing apparatus for a tRNA adaptation index weight prediction model provided in Embodiment 2 of the present invention;
[0099] Figure 5 Schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0100] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0101] Embodiment 1 of the present invention provides a processing method for a tRNA adaptation index weight prediction model, as Figure 1 shown in the schematic diagram of a processing method for a tRNA adaptation index weight prediction model provided in Embodiment 1 of the present invention. The method mainly includes the following steps:
[0102] Step 1, construct an adaptation index weight prediction model for processing the tRNA adaptation index weight prediction task; and design three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model.
[0103] Here, the adaptation index weight prediction model of the embodiment of the present invention is used to perform tRNA adaptation index weight prediction processing based on the input genomic sequence coding vector A, experimental condition coding vector B, and tRNA gene copy number vector C, and output the corresponding adaptation index weight vector W.
[0104] The above genomic sequence coding vector A is composed of a codon coding vector A C and an anticodon coding vector A AC spliced together; wherein, the codon coding vector A C is composed of the codon codings c i of all codon types of the corresponding genomic sequence sorted in order; the anticodon coding vector A AC is composed of the anticodon codings ac j of all anticodon types of the corresponding genomic sequence sorted in order; the total number of codon codings c i is denoted as N C , the total number of anticodon codings ac j is denoted as N AC , 1 ≤ i ≤ N C , 1 ≤ j ≤ N AC ; the codon coding c i and the anticodon coding ac j are each a triple-base coding.
[0105] The above experimental condition coding vector B includes at least temperature coding, humidity coding, light coding, nutrient condition coding, physiological / culture stage coding.
[0106] The above tRNA gene copy number vector C is composed of N ACThe gene copy number tGCN j Composition; among them, the gene copy number tGCN j Corresponds one-to-one with the anticodon encoding ac j One-to-one correspondence.
[0107] The above adaptation index weight vector W consists of N C Adaptation index weights w i Composition; among them, the adaptation index weight w i Corresponds one-to-one with the codon encoding c i One-to-one correspondence.
[0108] Such as Figure 2 As shown in the module structure diagram of the adaptation index weight prediction model provided in Embodiment 1 of the present invention, the first model input end, the second model input end, and the third model input end of the adaptation index weight prediction model are respectively used to receive the corresponding genomic sequence coding vector A, experimental condition coding vector B, and tRNA gene copy number vector C, and the model output end is used to output the corresponding adaptation index weight vector W.
[0109] The model components of the adaptation index weight prediction model include: Uni-RNA model, first MLP model, second MLP model, feature fusion module, third MLP model, and adaptation index weight estimation module. The connection relationships of each model component are as follows: the input end of the Uni-RNA model is connected to the first model input end, and the output ends are respectively connected to the input ends of the first and second MLP models; the output ends of the first and second MLP models are respectively connected to the second and third input ends of the feature fusion module; the first input end of the feature fusion module is connected to the second model input end, and the output end is connected to the input end of the third MLP model; the output end of the third MLP model is connected to the second input end of the adaptation index weight estimation module; the first input end of the adaptation index weight estimation module is connected to the third model input end, and the output end is connected to the model output end.
[0110] The functions of each model component of the adaptation index weight prediction model are as follows.
[0111] 1) Uni-RNA model:
[0112] The Uni-RNA model of the embodiment of the present invention has been pre-trained; this Uni-RNA model is used to perform feature encoding on the genomic sequence coding vector A to obtain the corresponding coding feature X and send it to the first and second MLP models.
[0113] It should be noted that the model structure, functions, and pre-training scheme of the Uni-RNA model in the embodiments of the present invention have been publicly disclosed in the technical literature "UNI-RNA: UNIVERSAL PRE-TRAINED MODELS REVOLUTIONIZE RNA RESEARCH". It can be known from this technical literature that the Uni-RNA model can perform deep learning on the association relationships of each codon+codon encoding pair, codon+anticodon encoding pair, and anticodon+anticodon encoding pair in a genomic sequence encoding vector A and output a corresponding encoding feature, namely encoding feature X.
[0114] 2) First MLP model:
[0115] The first MLP model in the embodiments of the present invention is used to predict the intracellular concentration of tRNA genes corresponding to each anticodon type based on the encoding feature X and send the corresponding concentration prediction vector P to the feature fusion module.
[0116] Here, the concentration prediction vector P consists of N AC predicted concentrations p j ; the predicted concentration p j corresponds one-to-one with the anticodon encoding ac j .
[0117] 3) Second MLP model:
[0118] The second MLP model in the embodiments of the present invention is used to predict the pairing probability of each codon-anticodon type pair based on the encoding feature X and send the corresponding pairing probability matrix M to the feature fusion module.
[0119] Here, the shape of the pairing probability matrix M is N C ×N AC , and it consists of N C ×N AC pairing probabilities m i,j ; each pairing probability m i,j corresponds to a group of codon encoding c i and anticodon encoding ac j .
[0120] 4) Feature fusion module:
[0121] The feature fusion module in the embodiments of the present invention is used to perform dimensionality reduction processing on the pairing probability matrix M to obtain the corresponding pairing probability vector V M ; and perform vector splicing on the pairing probability vector V M , the concentration prediction vector P, and the experimental condition encoding vector B to obtain the corresponding splicing vector R and send it to the third MLP model.
[0122] 5) Third MLP model:
[0123] The third MLP model in the embodiments of the present invention is used to predict the coupling coefficient for each group of codon-anticodon type pairs according to the splicing vector R to obtain the corresponding coupling coefficient tensor S and send it to the adaptation index weight estimation module.
[0124] Here, the shape of the coupling coefficient tensor S is N C ×N AC , and it consists of N C ×N AC coupling coefficients s i,j ; each coupling coefficient s i,j corresponds to a group of codon encoding c i and anticodon encoding ac j . The coupling coefficient s i,j mentioned here is known from the following as a calculation parameter for calculating the adaptation index weight.
[0125] It should be noted that the splicing vector R incorporates three types of characteristic information that affect the coupling coefficient (intracellular concentration of tRNA genes, codon-anticodon pairing probability, experimental conditions). The third MLP model can learn the mapping relationship between these three types of characteristic information and the coupling coefficient, and the generalization and accuracy of the obtained coupling coefficient will be better.
[0126] 6) Adaptation index weight estimation module:
[0127] The adaptation index weight estimation module in the embodiments of the present invention is used to estimate the adaptation index weight for each codon type according to the tRNA gene copy number vector C and the coupling coefficient tensor S to obtain the corresponding adaptation index weight vector W.
[0128] Here, the adaptation index weight vector W consists of N C adaptation index weights w i ; the adaptation index weight w i corresponds one-to-one with the codon encoding c i ; and the estimation formula for the adaptation index weight w i is:
[0129] As can be seen from the functions of the model components of the above adaptation index weight prediction model, the adaptation index weight prediction model is used to predict the adaptation index weights of all codons within a given genomic sequence. To predict the gene expression levels (such as transcription efficiency, translation efficiency, gene abundance, etc.) of a certain type of target gene (DNA gene, synthetic gene) in this given genomic sequence, further analysis is required through the three downstream task models (transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model) of the embodiments of the present invention based on the known adaptation index weight vector W and the known target gene.
[0130] The model structures of the three downstream task models of the embodiments of the present invention are as Figure 3 shown in the module structure diagrams of the transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model provided in Embodiment 1 of the present invention. A brief description of the three downstream task models is given below.
[0131] 1) Transcription efficiency prediction model:
[0132] As Figure 3 shown, the transcription efficiency prediction model of the embodiments of the present invention is used to predict the transcription efficiency of a specified target gene based on the input adaptation index weight vector W and the target gene codon index set G and output the corresponding transcription efficiency TPE.
[0133] The target gene codon index set G mentioned here is composed of codon indexes corresponding to all codon types of the specified target gene.
[0134] The model components of this transcription efficiency prediction model include: a first screening module and a fourth MLP model. The connection relationship of each component is: the input end of the first screening module is connected to the input end of the transcription efficiency prediction model, and the output end is connected to the input end of the fourth MLP model; the output end of the fourth MLP model is connected to the output end of the transcription efficiency prediction model. The component functions of each component are: the first screening module is used to extract the adaptation index weights w i in the adaptation index weight vector W whose index i matches each codon index of the target gene codon index set G to form a corresponding first weight vector and send it to the fourth MLP model; the fourth MLP model is used to predict the transcription efficiency based on the first weight vector to obtain the corresponding transcription efficiency TPE and output it.
[0135] It should be noted that, as can be seen from the foregoing, the adaptation index weight vector W contains the adaptation index weights of all codons within the entire genomic sequence; while predicting the transcription efficiency of a certain type of target gene (DNA gene, synthetic gene), only the adaptation index weights related to the current target gene need to be used for prediction. Therefore, the transcription efficiency prediction model of the present invention embodiment will first select the adaptation index weights related to the current target gene from the adaptation index weight vector W based on the first screening module to form the first weight vector, and then the fourth MLP model will perform transcription efficiency prediction according to the first weight vector.
[0136] 2) Translation efficiency prediction model:
[0137] As Figure 3 shown, the translation efficiency prediction model of the present invention embodiment is used to perform prediction processing on the translation efficiency of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding translation efficiency TLE.
[0138] The model components of this translation efficiency prediction model include: a second screening module and a fifth MLP model. The connection relationship of each component is: the input end of the second screening module is connected to the input end of the translation efficiency prediction model, and the output end is connected to the input end of the fifth MLP model; the output end of the fifth MLP model is connected to the output end of the translation efficiency prediction model. The component functions of each component are: the second screening module is used to extract the adaptation index weights w i whose index i in the adaptation index weight vector W matches each codon index in the target gene codon index set G to form the corresponding second weight vector and send it to the fifth MLP model; the fifth MLP model is used to perform translation efficiency prediction according to the second weight vector to obtain the corresponding gene translation efficiency TLE and output it.
[0139] Similar to the transcription efficiency prediction model, when predicting the translation efficiency of a certain type of target gene (DNA gene, synthetic gene), only the adaptation index weights related to the current target gene need to be used for prediction. Therefore, the translation efficiency prediction model of the present invention embodiment will first select the adaptation index weights related to the current target gene from the adaptation index weight vector W based on the second screening module to form the second weight vector, and then the fifth MLP model will perform translation efficiency prediction according to the second weight vector.
[0140] 3) Gene abundance prediction model:
[0141] As Figure 3 shown, the gene abundance prediction model of the present invention embodiment is used to perform prediction processing on the gene abundance of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding gene abundance TGA.
[0142] The model components of the gene abundance prediction model include: a third screening module and a sixth MLP model. The connection relationship of each component is as follows: the input end of the third screening module is connected to the input end of the gene abundance prediction model, and the output end is connected to the input end of the sixth MLP model; the output end of the sixth MLP model is connected to the output end of the gene abundance prediction model. The component functions of each component are as follows: the third screening module is used to extract the adaptation index weights w i in the adaptation index weight vector W that match the codon indexes in the target gene codon index set G with index i to form a corresponding third weight vector and send it to the sixth MLP model; the sixth MLP model is used to predict the gene abundance based on the third weight vector to obtain the corresponding gene abundance TGA and output it.
[0143] Similar to the transcription efficiency prediction model, when predicting the gene abundance of a certain type of target gene (DNA gene, synthetic gene), it is only necessary to predict according to the codon adaptation index weights related to the current target gene. Therefore, the gene abundance prediction model in the embodiments of the present invention will first select the codon adaptation index weights related to the current target gene from the adaptation index weight vector W based on the third screening module to form a third weight vector, and then the sixth MLP model will predict the gene abundance according to the third weight vector.
[0144] Step 2: Collect data on the genomic sequences of the target species under different conditions and construct a first data set based on the collected data; and train the adaptation index weight prediction model and three downstream task models based on the first data set;
[0145] Specifically, it includes: Step 21: Collect data on the genomic sequences of the target species under different conditions and construct a first data set based on the collected data;
[0146] Here, the first data set constructed in the embodiments of the present invention corresponds to a type of target species; the first data set includes multiple first data records; each first data record corresponds to a genomic sequence obtained by an observer through manual or machine observation of a type of cell of the corresponding target species under a set condition; and each first data record is composed of the following data: a first genomic sequence, a first set condition, a first codon data table, a first anticodon data table, and a first gene expression data table;
[0147] Among them, the data format of the first genomic sequence is a general genomic sequence, genomic map, or gene expression profile; the first set of conditions is the observation / experimental set of conditions corresponding to the current genomic sequence, including at least temperature, humidity, light, nutritional conditions, physiological / culture stage; the first codon data table, the first anticodon data table, and the first gene expression data table are all analysis and measurement data obtained by the observer using preset bioinformatics tools to perform directional analysis and measurement on the first genomic sequence; the bioinformatics tools mentioned here include at least tRNAscan-SE software, Galaxy software, SAMtools software, and GATK software;
[0148] The following further explains the data content of the first codon data table, the first anticodon data table, and the first gene expression data table:
[0149] 1) The first codon data table:
[0150] The first codon data table includes multiple first codon data; each first codon data corresponds to a type of codon in the first genomic sequence; that is to say, the first codon data table covers all types of codons in the first genomic sequence;
[0151] The first codon data includes a codon identifier, a codon triplet base encoding, and a codon usage bias; among them, the codon identifier is the unique identifier of the current codon type; the codon triplet base encoding is the triplet base encoding of the current codon type; the codon usage bias is the usage bias scalar of the current codon type in the first genomic sequence;
[0152] 2) The first anticodon data table:
[0153] The first anticodon data table includes multiple first anticodon data; each first anticodon data corresponds to a type of anticodon in the first genomic sequence; that is to say, the first anticodon data table covers all types of anticodons in the first genomic sequence;
[0154] The first anticodon data includes an anticodon identifier, an anticodon triplet base encoding, and the tRNA gene copy number; among them, the anticodon identifier is the unique identifier of the current anticodon type; the anticodon triplet base encoding is the triplet base encoding of the current anticodon type; the tRNA gene copy number is the total number of tRNA genes corresponding to the current anticodon type in the first genomic sequence;
[0155] 3) The first gene expression data table:
[0156] The first gene expression data table includes multiple first gene expression data; each first gene expression data corresponds to a DNA gene or a synthetic gene in the first genomic sequence; that is to say, the first gene expression data table covers all DNA genes and synthetic genes (if any) in the first genomic sequence;
[0157] The first gene expression data includes gene identification, gene type, gene base coding sequence, gene codon sequence, gene transcription efficiency, gene translation efficiency, and gene abundance; among them, the gene identification is the unique identifier of the current gene; the gene type includes at least DNA genes and synthetic genes; the gene base coding sequence is the base coding sequence of the current gene; the gene codon sequence is composed of codon identifiers of all codon types of the current gene; the gene transcription efficiency, gene translation efficiency, and gene abundance are the transcription efficiency scalar, translation efficiency scalar, and abundance scalar of the current gene under the first set conditions;
[0158] Step 22, and based on the first data set, model training is performed on the adaptation index weight prediction model and three downstream task models;
[0159] Specifically, it includes: Step 2201, training data preparation is performed based on the first data set to obtain a corresponding second data set; and the first training genomic sequence coding vector, the first training experimental condition coding vector, the first training gene copy number vector, and the first codon usage bias vector of each second data record in the second data set are extracted to form a corresponding third data record; and all the obtained third data records are de-duplicated; and all the de-duplicated third data records form the corresponding third data set;
[0160] Specifically, it includes: Step 22011, training data preparation is performed based on the first data set to obtain a corresponding second data set;
[0161] Among them, the second data set includes multiple second data records; the second data record includes the first training genomic sequence coding vector, the first training experimental condition coding vector, the first training gene copy number vector, the first codon usage bias vector, the first training codon index set, the first label transcription efficiency, the first label translation efficiency, and the first label gene abundance;
[0162] Specifically, it includes: Step 220111, each first data record of the first data set is used as the corresponding current data record;
[0163] Step 220112, the first set conditions, the first codon data table, the first anticodon data table, and the first gene expression data table of the current data record are used as the corresponding current set conditions, current codon data table, current anticodon data table, and current gene expression data table;
[0164] Step 220113: Extract all codon triple-base encodings in the current codon data table, sort them in order, and form a corresponding first codon encoding vector; extract all anticodon triple-base encodings in the current anticodon data table, sort them in order, and form a corresponding first anticodon encoding vector; and splice the first codon encoding vector and the first anticodon encoding vector to obtain a corresponding first training genome sequence encoding vector.
[0165] Step 220114: Based on the preset conditional encoding rules, encode the temperature, humidity, light, nutrient conditions, and physiological / culture stage in the current set conditions respectively to obtain corresponding temperature encoding, humidity encoding, light encoding, nutrient condition encoding, and physiological / culture stage encoding, and form a corresponding first training experimental condition encoding vector.
[0166] Here, the conditional encoding rules are a set of pre-specified encoding rules used to encode data for conditional factors such as temperature, humidity, light, nutrient conditions, and physiological / culture stage. The encoding rules used can be customized according to actual application requirements. For example, encode temperature, humidity, and light in the form of numerical values or normalized numerical values with positive and negative relationships, encode different classifications of nutrient conditions in the form of positive integers, and encode the degree of different classifications in the form of decimals / percentages. The corresponding nutrient condition encoding is composed of classification encoding + degree encoding, and encode different stages of the physiological / culture stage in the form of positive integers.
[0167] Step 220115: Extract all tRNA gene copy numbers in the current anticodon data table, sort them in order, and form a corresponding first training gene copy number vector.
[0168] Step 220116: Extract all codon usage biases in the current codon data table, sort them in order, and form a corresponding first codon usage bias vector.
[0169] Step 220117: Perform a round of traversal on all the first gene expression data in the current gene expression data table; during this round of traversal, use the currently traversed first gene expression data as the corresponding current gene expression data; use the gene codon sequence, gene transcription efficiency, gene translation efficiency, and gene abundance of the current gene expression data as the corresponding first training codon index set, first label transcription efficiency, first label translation efficiency, and first label gene abundance; and form a corresponding second data record from the first training genome sequence coding vector, first training experimental condition coding vector, first training gene copy number vector, first codon usage bias vector, and the first training codon index set, first label transcription efficiency, first label translation efficiency, and first label gene abundance corresponding to the current gene expression data; at the end of this round of traversal, form a first record set corresponding to the current data record from all the second data records obtained in this round of traversal.
[0170] Step 220118: Perform data record merging on the first record sets corresponding to all the first data records to obtain a corresponding second data set.
[0171] Here, as can be seen from the above steps 22011[220111 - 220118], the first training codon index set, first label transcription efficiency, first label translation efficiency, and first label gene abundance of each second data record in the second data set are personalized, but the first training genome sequence coding vector, first training experimental condition coding vector, first training gene copy number vector, and first codon usage bias vector of all the second data records generated from the same first data record coincide.
[0172] It should be noted that as can be seen from the following, the embodiments of the present invention will train the adaptation index weight prediction model based on the first training genome sequence coding vector, first training experimental condition coding vector, first training gene copy number vector, and first codon usage bias vector of the second data record. If the second data set is used as the reference data set to train the adaptation index weight prediction model, there will inevitably be a repeated training scenario where the training data is the same multiple times, and the training cycle will also be relatively long. To shorten the training cycle and improve the training efficiency, the embodiments of the present invention hereby construct a data set dedicated to training the adaptation index weight prediction model and without duplicate records, that is, the third data set, according to the second data set through the following step 22012.
[0173] Step 22012, extract the first training genome sequence encoding vectors, first training experimental condition encoding vectors, first training gene copy number vectors, and first codon usage bias vectors of each second data record in the second dataset to form a corresponding third data record; de-duplicate all the obtained third data records; and form a corresponding third dataset from all the de-duplicated third data records;
[0174] Among them, the third dataset consists of multiple third data records, and each third data record consists of the corresponding first training genome sequence encoding vector, first training experimental condition encoding vector, first training gene copy number vector, and first codon usage bias vector;
[0175] Step 2202, randomly split the third dataset based on a preset first splitting ratio to obtain a corresponding first training set and first evaluation set;
[0176] Among them, the first splitting ratio is a preset ratio value, such as 9:1; both the first training set and the first evaluation set consist of multiple third data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio;
[0177] Step 2203, take the first third data record in the first training set as the corresponding current training record;
[0178] Step 2204, input the first training genome sequence encoding vector, first training experimental condition encoding vector, and first training gene copy number vector of the current training record into the tRNA adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain a corresponding first predicted adaptation index weight vector;
[0179] Step 2205, input the first predicted adaptation index weight vector and the first codon usage bias vector of the current training record into a preset first correlation function; and based on a preset first model optimization algorithm, modulate the parameters of the adaptation index weight prediction model in the direction of maximizing the positive correlation of the first correlation function for one round; and at the end of this round of parameter modulation, identify whether the current training record is the last third data record in the first training set; if so, go to Step 2206; if not, extract the next third data record in the first training set as the new current training record and return to Step 2204 to continue training;
[0180] Among them, the first correlation function includes at least the Pearson correlation function, Spearman correlation function, and Kendall correlation function; the first model optimization algorithm includes at least the hill climbing algorithm and genetic algorithm;
[0181] Step 2206, perform a round of traversal on all the third data records in the first evaluation set; during this round of traversal, use the currently traversed third data record as the corresponding current evaluation record; input the first training genome sequence encoding vector, the first training experimental condition encoding vector, and the first training gene copy number vector of the current evaluation record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding second predicted adaptation index weight vector; bring the second predicted adaptation index weight vector and the first codon usage bias vector of the current training record into the first correlation function for calculation to obtain the corresponding first correlation degree; at the end of this round of traversal, bring all the obtained first correlation degrees into the preset first model evaluation function for calculation to obtain the corresponding first evaluation value;
[0182] Among them, the first model evaluation function at least includes the MAE function, the MSE function, and the RMSE function;
[0183] Step 2207, identify whether the first evaluation value meets the preset first evaluation value range; if the first evaluation value meets the first evaluation value range, identify whether all the obtained first correlation degrees meet the preset first correlation degree range, if not, return to step 2201 to continue training, if so, solidify the model parameters of the adaptation index weight prediction model and go to step 2208; if the first evaluation value does not meet the first evaluation value range, return to step 2201 to continue training;
[0184] Among them, the first evaluation value range is a preset evaluation value range; the first correlation degree range is a preset correlation degree range;
[0185] Step 2208, randomly split the second data set based on the preset second split ratio to obtain the corresponding second training set and second evaluation set;
[0186] Among them, the second split ratio is a preset ratio value, such as 8:2: both the second training set and the second evaluation set are composed of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set meets the second split ratio;
[0187] Step 2209, use the first second data record in the second training set as the corresponding current training record;
[0188] Step 2210, input the first training genome sequence encoding vector, the first training experimental condition encoding vector, and the first training gene copy number vector of the current training record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding third predicted adaptation index weight vector;
[0189] Step 2211: Input the third predicted adaptation index weight vector and the first training codon index set of the current training record into the transcriptional efficiency prediction model, the translational efficiency prediction model, and the gene abundance prediction model respectively for prediction to obtain the corresponding first predicted transcriptional efficiency, first predicted translational efficiency, and first predicted gene abundance;
[0190] Step 2212: Substitute the first predicted transcriptional efficiency and the first labeled transcriptional efficiency of the current training record into a preset first transcriptional loss function for calculation to obtain the corresponding first loss value; substitute the first predicted translational efficiency and the first labeled translational efficiency of the current training record into a preset first translational loss function for calculation to obtain the corresponding second loss value; substitute the first predicted gene abundance and the first labeled gene abundance of the current training record into a preset first abundance loss function for calculation to obtain the corresponding third loss value;
[0191] Among them, the first transcriptional loss function, the first translational loss function, and the first abundance loss function all at least include the L1 loss function and the L2 loss function;
[0192] Step 2213: Identify whether the first, second, and third loss values meet the preset first, second, and third loss value ranges; if the first, second, and third loss values all meet the corresponding first, second, and third loss value ranges, then go to Step 2214; if the first loss value does not meet the first loss value range, then based on a preset first model optimizer, perform one round of parameter modulation on the transcriptional efficiency prediction model in the direction of minimizing the function value of the first transcriptional loss function, and return to Step 2211 at the end of this round of parameter modulation; if the second loss value does not meet the second loss value range, then based on a preset second model optimizer, perform one round of parameter modulation on the translational efficiency prediction model in the direction of minimizing the function value of the first translational loss function, and return to Step 2211 at the end of this round of parameter modulation; if the third loss value does not meet the third loss value range, then based on a preset third model optimizer, perform one round of parameter modulation on the gene abundance prediction model in the direction of minimizing the function value of the first abundance loss function, and return to Step 2211 at the end of this round of parameter modulation;
[0193] Among them, the first, second, and third loss value ranges are three preset loss value ranges; the first, second, and third model optimizers all at least include SGD series optimizers and ADAM series optimizers;
[0194] Step 2214: Perform a round of traversal on all the second data records in the second evaluation set; during this round of traversal, use the currently traversed second data record as the corresponding current evaluation record; input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current evaluation record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding fourth predicted adaptation index weight vector; input the fourth predicted adaptation index weight vector and the first training codon index set of the current evaluation record into the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model respectively for prediction to obtain the corresponding second predicted transcription efficiency, second predicted translation efficiency, and second predicted gene abundance; form a corresponding first predicted vector from the second predicted transcription efficiency, second predicted translation efficiency, and second predicted gene abundance, form a corresponding first labeled vector from the first labeled transcription efficiency, first labeled translation efficiency, and first labeled gene abundance of the current evaluation record, and form a corresponding first predicted-labeled pair from the first predicted vector and the first labeled vector; at the end of this round of traversal, input all the obtained first predicted-labeled pairs into the preset second model evaluation function for calculation to obtain the corresponding second evaluation value.
[0195] Among them, the second model evaluation function at least includes the MAE function and the RMSE function.
[0196] Step 2215: Identify whether the second evaluation value meets the preset second evaluation value range; if the second evaluation value meets the second evaluation value range, solidify the model parameters of the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model and go to Step 2216; if the second evaluation value does not meet the second evaluation value range, return to Step 2208 to continue training.
[0197] Among them, the second evaluation value range is a pre-set evaluation value range.
[0198] Step 2216: Stop the model training and confirm that the training of the adaptation index weight prediction model and the three downstream task models is completed.
[0199] Step 3: After the model training is completed, receive the current experimental conditions, the current genome sequence, and the current target gene input by the user; preprocess the current experimental conditions, the current genome sequence, and the current target gene to obtain the corresponding current experimental condition coding vector, the current genome sequence coding vector, the current tRNA gene copy number vector, and the current target gene codon index set.
[0200] Specifically, it includes: Step 31: After the model training is completed, receive the current experimental conditions, the current genome sequence, and the current target gene input by the user.
[0201] Here, the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species corresponding to the most recent model training under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence;
[0202] Step 32, and preprocess the current experimental conditions, the current genomic sequence, and the current target gene to obtain the corresponding current experimental condition encoding vector, current genomic sequence encoding vector, current tRNA gene copy number vector, and current target gene codon index set;
[0203] Specifically, it includes: Step 321, based on the condition encoding rule, encode the temperature, humidity, light, nutrient conditions, and physiological / culture stage in the current experimental conditions respectively to obtain the corresponding temperature encoding, humidity encoding, light encoding, nutrient condition encoding, and physiological / culture stage encoding, which together form the corresponding current experimental condition encoding vector;
[0204] Step 322, use bioinformatics tools to analyze and count the codon types of all mRNA genes in the current genomic sequence, and sort the counted codon types in order to obtain the corresponding first type sequence; and sort the triple-base encodings of all codon types in the first type sequence in order to obtain a corresponding current codon encoding vector;
[0205] Step 323, use bioinformatics tools to analyze and count the anticodon types of all tRNA genes in the current genomic sequence, and sort the counted anticodon types in order to obtain the corresponding second type sequence; and sort the triple-base encodings of all anticodon types in the second type sequence in order to obtain a corresponding current anticodon encoding vector;
[0206] Step 324, splice the current codon encoding vector and the current anticodon encoding vector to obtain the corresponding current genomic sequence encoding vector;
[0207] Step 325, use bioinformatics tools to statistically measure the total number of gene replications of each tRNA gene in the current genomic sequence and use the measurement result as the corresponding tRNA gene copy number; and sort all the obtained tRNA gene copy numbers in order to obtain the corresponding current tRNA gene copy number vector;
[0208] Step 326: Use bioinformatics tools to analyze and count all codon types of the current target gene to obtain the corresponding third-type sequence; extract the sorting index of each codon type of the third-type sequence in the first-type sequence as a corresponding first codon index; and form a corresponding current target gene codon index set from all the obtained first codon indexes.
[0209] Step 327: Output the current experimental condition coding vector, the current genomic sequence coding vector, the current tRNA gene copy number vector, and the current target gene codon index set obtained this time as the processing result of this preprocessing.
[0210] Step 4: The adaptation index weight prediction model performs tRNA adaptation index weight prediction processing based on the current genomic sequence coding vector, the current experimental condition coding vector, and the current tRNA gene copy number vector to obtain the corresponding current adaptation index weight vector; and the three downstream task models respectively perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing based on the current adaptation index weight vector and the current target gene codon index set to obtain the corresponding current transcription efficiency, current translation efficiency, and current gene abundance, and form a corresponding task report to feedback to the user.
[0211] Finally, it should also be noted that based on the four types of models in the embodiments of the present invention, multiple model application scenarios can be extended. Three typical scenarios are given below for illustration.
[0212] Scenario 1: Under specified experimental conditions (i.e., the experimental condition coding vector remains unchanged), identify the genomic sequence coding vectors, tRNA gene copy number vectors, and target gene codon index sets corresponding to each DNA / synthetic gene of a genomic sequence, and thus obtain multiple sets of gene identification data (experimental condition coding vector + genomic sequence coding vector + tRNA gene copy number vector + target gene codon index set). Then, input the experimental condition coding vector + genomic sequence coding vector + tRNA gene copy number vector of each set of gene identification data into the adaptation index weight prediction model to obtain the adaptation index weight vectors of each DNA / synthetic gene; then, substitute the adaptation index weight vectors of each DNA / synthetic gene into the fixed gene-level tRNA adaptation index calculation formula to obtain the tRNA adaptation indexes of each DNA / synthetic gene; then, substitute the adaptation index weight vectors of each DNA / synthetic gene + target gene codon index set into the three downstream task models (transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model) to obtain the transcription efficiency, translation efficiency, and gene abundance of each DNA / synthetic gene. Thus, the gene expression levels of different genes in a genomic sequence can be horizontally compared from multiple data dimensions (tRNA adaptation index, transcription efficiency, translation efficiency, gene abundance) based on the prediction data.
[0213] Scenario 2: Switch the experimental conditions and obtain a corresponding genomic sequence under different experimental conditions; and predict the multi-dimensional gene expression data (tRNA adaptation index, transcription efficiency, translation efficiency, gene abundance) of each DNA / synthetic gene under each experimental condition through an operation method similar to that in Scenario 1; thus, the gene expression levels of each DNA / synthetic gene under different experimental conditions can be compared longitudinally based on the predicted data.
[0214] Scenario 3: Predict the multi-dimensional gene expression data (tRNA adaptation index, transcription efficiency, translation efficiency, gene abundance) of each DNA / synthetic gene under each experimental condition through an operation method similar to that in Scenario 1 under a specified experimental condition; obtain relevant experimental detection data (such as transcription efficiency, translation efficiency, gene abundance, etc.) through experimental means under the specified experimental condition, and ensure the accuracy of the experimental detection data through other experimental or computational means; then judge the accuracy of the four types of models by comparing the predicted and experimental data of each gene expression data item (transcription efficiency, translation efficiency, gene abundance, etc.); when the accuracy drops or does not meet the standard, use the known predicted-experimental data pairs as a reference and optimize the parameters of some or all of the models that cause the prediction to fail to meet the standard in the four types of models through the model parameter optimizer in the manner of model training; subsequently, the accuracy of other experimental data that has not been verified for accuracy can be evaluated using the predicted data of the four types of models after optimization.
[0215] Figure 4 The figure is a module structure diagram of a processing device for a tRNA adaptation index weight prediction model provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or can be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device can be a device or a chip system of the foregoing terminal device or server. As Figure 4 shown, the device includes: a model construction module 201, a model training module 202, a data preparation module 203, and a model application module 204.
[0216] The model construction module 201 is used to construct an adaptation index weight prediction model for processing the tRNA adaptation index weight prediction task; and design three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model.
[0217] The model training module 202 is used to collect data on the genomic sequences of the target species under different conditions and construct a first data set based on the collected data; and train the adaptation index weight prediction model and the three downstream task models based on the first data set.
[0218] The data preparation module 203 is used to receive the current experimental conditions, the current genomic sequence, and the current target gene input by the user after the model training is completed; and preprocess the current experimental conditions, the current genomic sequence, and the current target gene to obtain the corresponding current experimental condition encoding vector, the current genomic sequence encoding vector, the current tRNA gene copy number vector, and the current target gene codon index set; the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence.
[0219] The model application module 204 is used to perform tRNA adaptation index weight prediction processing by the adaptation index weight prediction model according to the current genomic sequence encoding vector, the current experimental condition encoding vector, and the current tRNA gene copy number vector to obtain the corresponding current adaptation index weight vector; and perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing by three downstream task models respectively according to the current adaptation index weight vector and the current target gene codon index set to obtain the corresponding current transcription efficiency, the current translation efficiency, and the current gene abundance, and form a corresponding task report to feedback to the user.
[0220] The processing device of a tRNA adaptation index weight prediction model provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0221] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element in hardware or in the form of instructions in software.
[0222] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain above module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a System-on-a-chip (SOC).
[0223] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The above available medium may be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0224] Figure 5 The structural schematic diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server that implements the method of the foregoing embodiments, or may be a terminal device or a server that is connected to the foregoing terminal device or server and implements the method of the foregoing embodiments. As Figure 5As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the method of the foregoing embodiments. Preferably, the electronic device according to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0225] In Figure 5 The system bus 305 mentioned may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0226] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0227] It should be noted that the embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the methods and processing procedures provided in the above embodiments.
[0228] An embodiment of the present invention provides a processing method, device, electronic device, and computer-readable storage medium for a tRNA adaptation index weight prediction model. As can be seen from the above, the embodiment of the present invention constructs an adaptation index weight prediction model for processing tRNA adaptation index weight prediction tasks and three downstream task models (transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model); and collects genomic sequences of a certain type of target species under different conditions and constructs a data set based on the collected data to train the adaptation index weight prediction model and the three downstream task models; after the training is completed, the adaptation index weight prediction model is used to predict the tRNA adaptation index weight of the current target species under the current given conditions, and the gene expression levels (transcription efficiency, translation efficiency, gene abundance) of a given target gene (DNA gene, synthetic gene) are further predicted based on the tRNA adaptation index weight through the three downstream task models. The four types of models provided by the embodiment of the present invention (adaptation index weight prediction model, transcription efficiency prediction model, translation efficiency prediction model, and gene abundance prediction model) are all end-to-end prediction models. Predicting the tRNA adaptation index weight and gene expression levels based on these four types of models not only reduces the processing complexity but also improves the model versatility; when switching the target species, the four types of models in the embodiment of the present invention do not need to change the model structure, and only need to be trained once using the data set of the current species. In this way, not only the species adaptability of the model is improved, but also the prediction stability of the model for each type of species is further improved through training.
[0229] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0230] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A processing method for a tRNA adaptation index weight prediction model, characterized in that The method includes: Constructing an adaptation index weight prediction model for processing the tRNA adaptation index weight prediction task; and designing three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model; Collecting data on the genomic sequences of the target species under different conditions and constructing a first data set based on the collected data; and training the adaptation index weight prediction model and the three downstream task models based on the first data set; After the model training is completed, receiving the current experimental conditions, the current genomic sequence, and the current target gene input by the user; and preprocessing the current experimental conditions, the current genomic sequence, and the current target gene to obtain corresponding current experimental condition encoding vectors, current genomic sequence encoding vectors, current tRNA gene copy number vectors, and current target gene codon index sets; the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence; The adaptation index weight prediction model performs tRNA adaptation index weight prediction processing based on the current genomic sequence encoding vector, the current experimental condition encoding vector, and the current tRNA gene copy number vector to obtain a corresponding current adaptation index weight vector; and the three downstream task models respectively perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing based on the current adaptation index weight vector and the current target gene codon index set to obtain a corresponding current transcription efficiency, current translation efficiency, and current gene abundance, and form a corresponding task report to feedback to the user; Among them, the adaptation index weight prediction model is used to perform tRNA adaptation index weight prediction processing according to the input genomic sequence encoding vector A, experimental condition encoding vector B, and tRNA gene copy number vector C and output a corresponding adaptation index weight vector W; The genomic sequence encoding vector A is encoded by a codon encoding vector A C and an anticodon encoding vector A AC spliced together; the codon encoding vector A C is composed of codons c i of all codon types of the corresponding genomic sequence arranged in sequence; the anticodon encoding vector A AC is composed of anticodons ac j of all anticodon types of the corresponding genomic sequence arranged in sequence; the total number of the codons c i is denoted as N C , and the total number of the anticodons ac j is denoted as N AC , where 1 ≤ i ≤ N C , and 1 ≤ j ≤ N AC ; the codon c i and the anticodon ac j are each a triplet base encoding; The experimental condition encoding vector B at least includes temperature encoding, humidity encoding, light encoding, nutrient condition encoding, physiological / culture stage encoding; The tRNA gene copy number vector C consists of N AC gene copy numbers tGCN j ; the gene copy number tGCN j corresponds one-to-one with the anticodon encoding ac j ; The adaptation index weight vector W consists of N C adaptation index weights w i ; the adaptation index weights w i correspond one-to-one with the codon encoding c i ; The first model input end, the second model input end, and the third model input end of the adaptation index weight prediction model are respectively used to receive the corresponding genomic sequence encoding vector A, the experimental condition encoding vector B, and the tRNA gene copy number vector C, and the model output end is used to output the corresponding adaptation index weight vector W; The adaptation index weight prediction model includes a Uni-RNA model, a first MLP model, a second MLP model, a feature fusion module, a third MLP model, and an adaptation index weight estimation module; The input end of the Uni-RNA model is connected to the input end of the first model, and the output end is respectively connected to the input ends of the first and second MLP models; the output ends of the first and second MLP models are respectively connected to the second and third input ends of the feature fusion module; the first input end of the feature fusion module is connected to the input end of the second model, and the output end is connected to the input end of the third MLP model; the output end of the third MLP model is connected to the second input end of the adaptation index weight estimation module; the first input end of the adaptation index weight estimation module is connected to the input end of the third model, and the output end is connected to the model output end; The Uni-RNA model has completed pre-training; the Uni-RNA model is used to perform feature encoding on the genomic sequence encoding vector A to obtain the corresponding encoding feature X and send it to the first and second MLP models; The first MLP model is used to predict the intracellular concentration of tRNA genes corresponding to each anticodon type based on the encoded feature X, obtain the corresponding concentration prediction vector P, and send it to the feature fusion module; the concentration prediction vector P is composed of N AC predicted concentrations p j ; the predicted concentration p j corresponds one-to-one with the anticodon encoding ac j ; The second MLP model is used to predict the pairing probability of each group of codon-anticodon type pairs based on the encoded feature X to obtain a corresponding pairing probability matrix M and send it to the feature fusion module; the shape of the pairing probability matrix M is N C ×N AC , which consists of N C ×N AC pairing probabilities m i,j ; each of the pairing probabilities m i,j corresponds to a group of the codon encoding c i and the anticodon encoding ac j . The feature fusion module is used to perform dimensionality reduction on the paired probability matrix M to obtain a corresponding paired probability vector V M ; and perform vector concatenation on the paired probability vector V M , the concentration prediction vector P, and the experimental condition encoding vector B to obtain a corresponding concatenated vector R and send it to the third MLP model; The third MLP model is used to predict the coupling coefficient of each group of codon-anticodon type pairs according to the spliced vector R, and send the corresponding coupling coefficient tensor S to the adaptation index weight estimation module; the shape of the coupling coefficient tensor S is N C ×N AC , which consists of N C ×N AC coupling coefficients s i,j ; each of the coupling coefficients s i,j corresponds to a group of the codon encoding c i and the anticodon encoding ac j . The adaptation index weight estimation module is used to estimate the adaptation index weights of each codon type according to the tRNA gene copy number vector C and the coupling coefficient tensor S to obtain the corresponding adaptation index weight vector W; the adaptation index weight vector W consists of N C adaptation index weights w i ; the adaptation index weight w i corresponds one-to-one with the codon encoding c i ; the estimation formula of the adaptation index weight w i is: ; The transcription efficiency prediction model is used to perform prediction processing on the transcription efficiency of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding transcription efficiency TPE; The target gene codon index set G consists of codon indexes corresponding to all codon types of the specified target gene; The transcription efficiency prediction model includes a first screening module and a fourth MLP model; the input end of the first screening module is connected to the input end of the transcription efficiency prediction model, and the output end is connected to the input end of the fourth MLP model; the output end of the fourth MLP model is connected to the output end of the transcription efficiency prediction model; The first sub-screening module is used to extract the adaptation index weights w in which the index i in the adaptation index weight vector W matches each of the codon indexes in the target gene codon index set G, and form a corresponding first weight vector to send to the fourth MLP model; the fourth MLP model is used to perform transcription efficiency prediction according to the first weight vector to obtain the corresponding transcription efficiency TPE and output it; i Extract and form a corresponding first weight vector and send it to the fourth MLP model; the fourth MLP model is used to perform transcription efficiency prediction according to the first weight vector to obtain the corresponding transcription efficiency TPE and output it; The translation efficiency prediction model is used to perform prediction processing on the translation efficiency of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding translation efficiency TLE; The translation efficiency prediction model includes a second screening module and a fifth MLP model; the input end of the second screening module is connected to the input end of the translation efficiency prediction model, and the output end is connected to the input end of the fifth MLP model; the output end of the fifth MLP model is connected to the output end of the translation efficiency prediction model; The second sub-screening module is used to extract the adaptation index weights w where the index i in the adaptation index weight vector W matches each of the codon indexes in the target gene codon index set G, and form a corresponding second weight vector to send to the fifth MLP model; the fifth MLP model is used to perform translation efficiency prediction according to the second weight vector to obtain the corresponding gene translation efficiency TLE and output it; i Extract them to form a corresponding second weight vector and send it to the fifth MLP model; the fifth MLP model is used to predict the translation efficiency according to the second weight vector to obtain the corresponding gene translation efficiency TLE and output it; The gene abundance prediction model is used to perform prediction processing on the gene abundance of a specified target gene according to the input adaptation index weight vector W and the target gene codon index set G and output the corresponding gene abundance TGA; The gene abundance prediction model includes a third screening module and a sixth MLP model; the input end of the third screening module is connected to the input end of the gene abundance prediction model, and the output end is connected to the input end of the sixth MLP model; the output end of the sixth MLP model is connected to the output end of the gene abundance prediction model; The third sub-screening module is used to extract the adaptation index weights w where the index i in the adaptation index weight vector W matches each of the codon indices in the target gene codon index set G, and form a corresponding third weight vector to send to the sixth MLP model; the sixth MLP model is used to perform gene abundance prediction according to the third weight vector to obtain the corresponding gene abundance TGA and output it. i Extract and form a corresponding third weight vector to send to the sixth MLP model; the sixth MLP model is used to perform gene abundance prediction according to the third weight vector to obtain the corresponding gene abundance TGA and output it.
2. The processing method of the tRNA adaptation index weight prediction model according to claim 1, wherein, The first data set includes a plurality of first data records; Each of the first data records corresponds to a genomic sequence obtained by an observer through manual or machine observation of a type of cell of the target species under a set condition; The first data record includes a first genomic sequence, a first set of conditions, a first codon data table, a first anticodon data table, and a first gene expression data table; The data format of the first genomic sequence is a general genomic sequence, genomic map, or gene expression profile; The first set of conditions is the observation / experimental set of conditions corresponding to the current genomic sequence, including at least temperature, humidity, light, nutritional conditions, physiological / culture stage; The first codon data table, the first anticodon data table, and the first gene expression data table are all analysis and measurement data obtained by an observer using a preset bioinformatics tool to perform directional analysis and measurement on the first genomic sequence; Among them, the bioinformatics tool includes at least tRNAscan-SE software, Galaxy software, SAMtools software, GATK software; The first codon data table includes multiple first codon data; each first codon data corresponds to a type of codon in the first genomic sequence; the first codon data includes a codon identifier, a codon triplet base encoding, and a codon usage bias; the codon identifier is the unique identifier of the current codon type; the codon triplet base encoding is the triplet base encoding of the current codon type; the codon usage bias is the usage bias scalar of the current codon type in the first genomic sequence; The first anticodon data table includes multiple first anticodon data; each first anticodon data corresponds to a type of anticodon in the first genomic sequence; the first anticodon data includes an anticodon identifier, an anticodon triplet base encoding, and the tRNA gene copy number; the anticodon identifier is the unique identifier of the current anticodon type; the anticodon triplet base encoding is the triplet base encoding of the current anticodon type; the tRNA gene copy number is the total number of tRNA genes corresponding to the current anticodon type in the first genomic sequence; The first gene expression data table includes multiple first gene expression data; each first gene expression data corresponds to a DNA gene or synthetic gene in the first genomic sequence; the first gene expression data includes a gene identifier, a gene type, a gene base encoding sequence, a gene codon sequence, a gene transcription efficiency, a gene translation efficiency, and a gene abundance; the gene identifier is the unique identifier of the current gene; the gene type includes at least DNA genes and synthetic genes; the gene base encoding sequence is the base encoding sequence of the current gene; the gene codon sequence is composed of the codon identifiers of all codon types of the current gene; the gene transcription efficiency, the gene translation efficiency, and the gene abundance are the transcription efficiency scalar, translation efficiency scalar, and abundance scalar of the current gene under the first set of conditions.
3. The processing method of the tRNA adaptation index weight prediction model according to claim 2, wherein, The model training of the adaptation index weight prediction model and the three downstream task models based on the first data set specifically includes: Step 501: Prepare training data based on the first dataset to obtain a corresponding second dataset; extract the first training genome sequence encoding vector, the first training experimental condition encoding vector, the first training gene copy number vector, and the first codon usage bias vector of each second data record in the second dataset to form a corresponding third data record; remove duplicates from all the obtained third data records; and form a corresponding third dataset from all the deduplicated third data records. Among them, the second dataset includes multiple second data records; the second data record includes a first training genome sequence encoding vector, a first training experimental condition encoding vector, a first training gene copy number vector, a first codon usage bias vector, a first training codon index set, a first label transcription efficiency, a first label translation efficiency, and a first label gene abundance; the third dataset is composed of multiple third data records, and each third data record is composed of the corresponding first training genome sequence encoding vector, the first training experimental condition encoding vector, the first training gene copy number vector, and the first codon usage bias vector. Step 502: Randomly split the third dataset based on a preset first splitting ratio to obtain a corresponding first training set and a first evaluation set. Among them, both the first training set and the first evaluation set are composed of multiple third data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio. Step 503: Take the first third data record in the first training set as the corresponding current training record. Step 504: Input the first training genome sequence encoding vector, the first training experimental condition encoding vector, and the first training gene copy number vector of the current training record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain a corresponding first predicted adaptation index weight vector. Step 505: Substitute the first predicted adaptation index weight vector and the first codon usage bias vector of the current training record into a preset first correlation function; and based on a preset first model optimization algorithm, perform one round of parameter modulation on the adaptation index weight prediction model in the direction of maximizing the positive correlation of the first correlation function; and at the end of this round of parameter modulation, identify whether the current training record is the last third data record in the first training set; if so, go to Step 506; if not, extract the next third data record in the first training set as the new current training record and return to Step 504 to continue training. Among them, the first correlation function includes at least the Pearson correlation function, the Spearman correlation function, and the Kendall correlation function; the first model optimization algorithm includes at least the hill climbing algorithm and the genetic algorithm. Step 506, perform a round of traversal on all the third data records in the first evaluation set; during this round of traversal, take the currently traversed third data record as the corresponding current evaluation record; input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current evaluation record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding second predicted adaptation index weight vector; substitute the second predicted adaptation index weight vector and the first codon usage deviation vector of the current training record into the first correlation function for calculation to obtain the corresponding first correlation degree; at the end of this round of traversal, substitute all the obtained first correlation degrees into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value; Among them, the first model evaluation function at least includes the MAE function, the MSE function, and the RMSE function; Step 507, identify whether the first evaluation value meets a preset first evaluation value range; if the first evaluation value meets the first evaluation value range, then identify whether all the obtained first correlation degrees meet a preset first correlation degree range, if not, return to step 501 to continue training, if so, solidify the model parameters of the adaptation index weight prediction model and go to step 508; if the first evaluation value does not meet the first evaluation value range, return to step 501 to continue training; Step 508, randomly split the second data set based on a preset second splitting ratio to obtain the corresponding second training set and second evaluation set; Among them, both the second training set and the second evaluation set are composed of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set meets the second splitting ratio; Step 509, take the first second data record in the second training set as the corresponding current training record; Step 510, input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current training record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding third predicted adaptation index weight vector; Step 511, input the third predicted adaptation index weight vector and the first training codon index set of the current training record into the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model respectively for prediction to obtain the corresponding first predicted transcription efficiency, first predicted translation efficiency, and first predicted gene abundance; Step 512: Substitute the first predicted transcription efficiency and the first labeled transcription efficiency of the current training record into a preset first transcription loss function for calculation to obtain a corresponding first loss value; substitute the first predicted translation efficiency and the first labeled translation efficiency of the current training record into a preset first translation loss function for calculation to obtain a corresponding second loss value; substitute the first predicted gene abundance and the first labeled gene abundance of the current training record into a preset first abundance loss function for calculation to obtain a corresponding third loss value; Among them, the first transcription loss function, the first translation loss function, and the first abundance loss function all at least include the L1 loss function and the L2 loss function; Step 513: Identify whether the first, second, and third loss values meet the preset first, second, and third loss value ranges; if the first, second, and third loss values all meet the corresponding first, second, and third loss value ranges, go to Step 514; if the first loss value does not meet the first loss value range, based on a preset first model optimizer, perform one round of parameter modulation on the transcription efficiency prediction model in the direction of minimizing the function value of the first transcription loss function, and return to Step 511 at the end of this round of parameter modulation; if the second loss value does not meet the second loss value range, based on a preset second model optimizer, perform one round of parameter modulation on the translation efficiency prediction model in the direction of minimizing the function value of the first translation loss function, and return to Step 511 at the end of this round of parameter modulation; if the third loss value does not meet the third loss value range, based on a preset third model optimizer, perform one round of parameter modulation on the gene abundance prediction model in the direction of minimizing the function value of the first abundance loss function, and return to Step 511 at the end of this round of parameter modulation; Among them, the first, second, and third model optimizers all at least include SGD series optimizers and ADAM series optimizers; Step 514, perform a round of traversal on all the second data records in the second evaluation set; during this round of traversal, take the currently traversed second data record as the corresponding current evaluation record; input the first training genome sequence coding vector, the first training experimental condition coding vector, and the first training gene copy number vector of the current evaluation record into the adaptation index weight prediction model for tRNA adaptation index weight prediction processing to obtain the corresponding fourth predicted adaptation index weight vector; input the fourth predicted adaptation index weight vector and the first training codon index set of the current evaluation record into the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model respectively for prediction to obtain the corresponding second predicted transcription efficiency, second predicted translation efficiency, and second predicted gene abundance; form a corresponding first predicted vector from the second predicted transcription efficiency, the second predicted translation efficiency, and the second predicted gene abundance, form a corresponding first labeled vector from the first labeled transcription efficiency, the first labeled translation efficiency, and the first labeled gene abundance of the current evaluation record, and form a corresponding first predicted-label pair from the first predicted vector and the first labeled vector; at the end of this round of traversal, input all the obtained first predicted-label pairs into a preset second model evaluation function for calculation to obtain the corresponding second evaluation value; Among them, the second model evaluation function at least includes the MAE function and the RMSE function; Step 515, identify whether the second evaluation value meets a preset second evaluation value range; if the second evaluation value meets the second evaluation value range, solidify the model parameters of the transcription efficiency prediction model, the translation efficiency prediction model, and the gene abundance prediction model and go to Step 516; if the second evaluation value does not meet the second evaluation value range, return to Step 508 to continue training; Step 516, stop model training and confirm that the model training of the adaptation index weight prediction model and the three downstream task models is completed.
4. The processing method of the tRNA adaptation index weight prediction model according to claim 3, wherein The preparation of training data based on the first data set to obtain the corresponding second data set specifically includes: Step 61, take each first data record in the first data set as the corresponding current data record; Step 62, take the first set conditions, the first codon data table, the first anticodon data table, and the first gene expression data table of the current data record as the corresponding current set conditions, current codon data table, current anticodon data table, and current gene expression data table; Step 63: Extract all the codon triplet base codings in the current codon data table, sort them in sequence, and form a corresponding first codon coding vector; extract all the anticodon triplet base codings in the current anticodon data table, sort them in sequence, and form a corresponding first anticodon coding vector; perform vector splicing on the first codon coding vector and the first anticodon coding vector to obtain a corresponding first training genome sequence coding vector; Step 64: Based on a preset conditional coding rule, encode the temperature, humidity, light, nutrient condition, and physiological / culture stage in the current set conditions respectively to obtain corresponding temperature coding, humidity coding, light coding, nutrient condition coding, and physiological / culture stage coding, and form a corresponding first training experimental condition coding vector; Step 65: Extract all the tRNA gene copy numbers in the current anticodon data table, sort them in sequence, and form a corresponding first training gene copy number vector; Step 66: Extract all the codon usage biases in the current codon data table, sort them in sequence, and form a corresponding first codon usage bias vector; Step 67: Perform a round of traversal on all the first gene expression data in the current gene expression data table; during this round of traversal, use the currently traversed first gene expression data as the corresponding current gene expression data; use the gene codon sequence, gene transcription efficiency, gene translation efficiency, and gene abundance of the current gene expression data as the corresponding first training codon index set, the first label transcription efficiency, the first label translation efficiency, and the first label gene abundance; form a corresponding second data record from the first training genome sequence coding vector, the first training experimental condition coding vector, the first training gene copy number vector, the first codon usage bias vector, and the first training codon index set, the first label transcription efficiency, the first label translation efficiency, and the first label gene abundance corresponding to the current gene expression data; at the end of this round of traversal, form a first record set corresponding to the current data record from all the second data records obtained in this round of traversal; Step 68: Perform data record merging on the first record set corresponding to all the first data records to obtain a corresponding second data set.
5. The processing method of the tRNA adaptation index weight prediction model according to claim 1, characterized in that The preprocessing of the current experimental conditions, the current genome sequence, and the current target gene to obtain a corresponding current experimental condition coding vector, current genome sequence coding vector, current tRNA gene copy number vector, and current target gene codon index set specifically includes: Step 71, based on a preset conditional encoding rule, encode the temperature, humidity, light, nutrient condition, and physiological / culture stage in the current experimental condition respectively to obtain corresponding temperature encoding, humidity encoding, light encoding, nutrient condition encoding, and physiological / culture stage encoding, which form the corresponding current experimental condition encoding vector; Step 72, use a preset bioinformatics tool to analyze and count the codon types of all mRNA genes in the current genomic sequence, and sort the counted codon types in order to obtain a corresponding first type sequence; and sort the triplet base encodings of all codon types in the first type sequence in order to obtain a corresponding current codon encoding vector; Step 73, use the bioinformatics tool to analyze and count the anticodon types of all tRNA genes in the current genomic sequence, and sort the counted anticodon types in order to obtain a corresponding second type sequence; and sort the triplet base encodings of all anticodon types in the second type sequence in order to obtain a corresponding current anticodon encoding vector; Step 74, splice the current codon encoding vector and the current anticodon encoding vector to obtain the corresponding current genomic sequence encoding vector; Step 75, use the bioinformatics tool to statistically calculate the total number of gene replications of each tRNA gene in the current genomic sequence and use the calculation result as the corresponding tRNA gene copy number; and sort all the obtained tRNA gene copy numbers in order to obtain the corresponding current tRNA gene copy number vector; Step 76, use the bioinformatics tool to analyze and count all codon types of the current target gene to obtain a corresponding third type sequence; and extract the sorting index of each codon type in the third type sequence in the first type sequence as a corresponding first codon index; and form the corresponding current target gene codon index set from all the obtained first codon indexes; Step 77, output the current experimental condition encoding vector, the current genomic sequence encoding vector, the current tRNA gene copy number vector, and the current target gene codon index set obtained this time as the processing result of this preprocessing.
6. An apparatus for implementing a processing method of the tRNA adaptation index weight prediction model according to any one of claims 1-5, characterized in that, The device includes: a model construction module, a model training module, a data preparation module, and a model application module; The model construction module is used to construct an adaptation index weight prediction model for processing the tRNA adaptation index weight prediction task; and design three downstream task models: a transcription efficiency prediction model, a translation efficiency prediction model, and a gene abundance prediction model; The model training module is used to collect data on the genomic sequences of the target species under different conditions and construct a first data set based on the collected data; and perform model training on the adaptation index weight prediction model and the three downstream task models based on the first data set; The data preparation module is used to receive the current experimental conditions, the current genomic sequence, and the current target gene input by the user after the model training is completed; and preprocess the current experimental conditions, the current genomic sequence, and the current target gene to obtain the corresponding current experimental condition encoding vector, current genomic sequence encoding vector, current tRNA gene copy number vector, and current target gene codon index set; the current genomic sequence is the genomic sequence obtained by an observer through manual or machine observation / experimental means for a type of cell of the target species under the current experimental conditions; the current target gene is a DNA gene or a synthetic gene in the current genomic sequence. The model application module is used to perform tRNA adaptation index weight prediction processing on the current genomic sequence encoding vector, the current experimental condition encoding vector, and the current tRNA gene copy number vector by the adaptation index weight prediction model to obtain the corresponding current adaptation index weight vector; and perform corresponding transcription efficiency, translation efficiency, and gene abundance prediction processing on the current adaptation index weight vector and the current target gene codon index set by three downstream task models respectively to obtain the corresponding current transcription efficiency, current translation efficiency, and current gene abundance, and form a corresponding task report to feedback to the user.
7. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-5; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is made to execute the method according to any one of claims 1-5.