Method and device for determining prokaryotic protein translation intensity prediction model

By designing the protein translation intensity prediction model of the convolution kernel module, the problem of large data sets, high cost and poor generalization capabilities in the existing technology is solved, and more efficient and accurate prediction of prokaryotic protein translation intensity is achieved.

CN120164530APending Publication Date: 2025-06-17SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146100.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the machine learning model for the translation efficiency of prokaryotic proteins requires a large data set, high acquisition cost and average generalization ability.

Method used

A protein translation intensity prediction model including the first convolution kernel module and the second convolution kernel module is designed. The first convolution kernel module is used to convolutionize data on non-translated area, and the second convolution kernel module is used to convolutionize data on the translation area. By collecting multiple experimental data of multiple prokaryotes and training models, the prediction accuracy and efficiency are improved.

Benefits of technology

It improves the accuracy and efficiency of the prediction of translation intensity of prokaryotic proteins, reduces the demand for training sample size, reduces the cost of data acquisition, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164530A_ABST
    Figure CN120164530A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biological information, and provides a prokaryote protein translation intensity prediction model determination method and device, and the method comprises the steps: collecting experimental data of a plurality of prokaryotes in advance; encoding the experimental data to obtain encoded data identified by a computer; dividing the coded data into a training set and a test set; multiple preset protein translation intensity prediction models are trained by using the training set, the multiple preset protein translation intensity prediction models comprise a convolution layer, the convolution layer at least comprises a first convolution kernel module and a second convolution kernel module, the first convolution kernel module is used for performing convolution processing on untranslated region data, and the second convolution kernel module is used for performing convolution processing on the untranslated region data; the second convolution kernel module is used for performing convolution processing on the translation area data; and measuring the prediction precision of each protein translation intensity prediction model by using the test set, and determining a final protein translation intensity prediction model according to the prediction precision of each protein translation intensity prediction model. According to the invention, the generalization ability, prediction precision and efficiency of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and in particular to a method and device for determining a prokaryotic protein translation intensity prediction model. Background Art

[0002] Translation is a fundamental process in biology, which refers to the synthesis of proteins through the codon sequence on mRNA (Messenger RNA). In microbial engineering, it is usually necessary to finely control the translation and expression of proteins. The translation regulation of proteins involves many molecular mechanisms and factors, which can affect the translation rate and thus regulate protein synthesis. The translation process of prokaryotes includes initiation, elongation, and termination stages, and each stage is affected by complex regulatory mechanisms. Usually, a gene regulatory factor controls another regulatory factor, and so on, forming a gene regulatory network. These mechanisms ensure that cells can precisely regulate protein synthesis under different environmental conditions to meet different needs.

[0003] Researchers usually achieve stable and reliable cell behavior by constructing some simple gene circuits. In the construction of small-scale gene circuits, the trial-and-error optimization method can still meet the requirements. However, with the continuous increase in the system scale and complexity, it becomes extremely inefficient to optimize engineered gene circuits and metabolic pathways by the trial-and-error method. Especially when synthesizing longer DNA and constructing larger-scale genomes, the limitations of the trial-and-error optimization method are more obvious.

[0004] Ribosome, as a complex molecular machine composed of RNA and proteins, plays a key role in protein synthesis.

[0005] In the prior art, the following methods are usually used to study protein expression levels: RBS Library, thermodynamic model, and machine learning model.

[0006] In the above methods, RBS Library requires a large number of experiments and screening, which has the problem of low efficiency; the thermodynamic model cannot distinguish the factors affecting protein translation intensity in different bacteria, that is, the predicted protein translation intensity is the same regardless of the bacteria; the machine learning model requires a large dataset, with high acquisition cost, general generalization ability, and low prediction accuracy. Summary of the Invention

[0007] The present invention is used to solve the technical problem in the prior art that the machine learning model for prokaryotic protein translation efficiency requires a large dataset, has high acquisition cost, and has general generalization ability.

[0008] To solve the above technical problem, the first aspect of the present invention provides a method for determining a prokaryotic protein translation intensity prediction model, including:

[0009] Pre-collect multiple experimental data of various prokaryotes, and each experimental data includes the translation intensity of the prokaryote at each sequence data;

[0010] Encode the experimental data to obtain encoded data recognizable by a computer;

[0011] Divide the encoded data of multiple groups of experimental data into a training set and a test set;

[0012] Use the sequence data as input and the translation intensity as output, and use the training set to train a plurality of preset protein translation intensity prediction models. Among them, the protein translation intensity prediction model includes a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer includes at least a first convolutional kernel module and a second convolutional kernel module. The first convolutional kernel module is used to perform convolutional processing on non-translated region data, and the second convolutional kernel module is used to perform convolutional processing on translated region data;

[0013] Use the test set to measure the prediction accuracy of each protein translation intensity prediction model, and determine the final protein translation intensity prediction model according to the prediction accuracy of each protein translation intensity prediction model.

[0014] As a further embodiment of the present invention, collecting multiple groups of experimental data includes:

[0015] For each prokaryote, make a first sequence containing a green fluorescent protein gene, measure the fluorescence intensity when expressing the first sequence, and determine the translation intensity according to the measured fluorescence intensity; or

[0016] For each prokaryote, make a second sequence containing a specific sequence and a gene, perform transcriptome sequencing and ribosome footprint sequencing on the plasmid of the second sequence, and obtain the translation intensity by dividing the normalized ribosome footprint sequencing result by the normalized transcriptome sequencing result.

[0017] As a further embodiment of the present invention, encoding the experimental data to obtain encoded data recognizable by a computer includes:

[0018] Digitally encode the translation intensity;

[0019] Perform binary one-hot encoding on the sequence data to obtain sequence encoding data of a fixed length.

[0020] As a further embodiment of the present invention, performing binary one-hot encoding on the sequence data to obtain sequence encoding data of a fixed length includes:

[0021] Judge whether the non-translated region in the sequence data conforms to the first preset base length;

[0022] If the non-translated region does not meet the first preset base length, adjust the non-translated region to the first preset base length;

[0023] After the non-translated region meets the first preset base length and the non-translated region is adjusted to the first preset base length, determine whether the length of the translated region in the sequence data meets the second preset base length;

[0024] If the translated region does not meet the second preset base length, adjust the translated region to the second preset base length;

[0025] If the translated region meets the second preset base length and the translated region is adjusted to the second preset base length, encode the sequence data.

[0026] In a further embodiment of the present invention, adjusting the non-translated region to the first preset base length includes:

[0027] When the length of the non-translated region is greater than the first preset base length, delete the N bases on the left side of the non-translated region;

[0028] When the length of the non-translated region is less than the first preset base length, supplement N bases on the left side of the non-translated region so that the length of the non-translated region is equal to the first preset base length.

[0029] In a further embodiment of the present invention, adjusting the translated region to the second preset base length includes:

[0030] When the length of the translated region is greater than the second preset base length, delete the N bases on the right side of the translated region;

[0031] When the length of the translated region is less than the second preset base length, supplement N bases on the right side of the translated region so that the length of the translated region is equal to the second preset base length.

[0032] In a further embodiment of the present invention, the process of determining the first preset base length includes:

[0033] Set the first preset base length to multiple first candidate lengths; for each first candidate length, train a prediction model and calculate the accuracy of the prediction model; screen out the prediction models that meet the preset accuracy; from the first candidate lengths related to the screened prediction models, select the smallest first candidate length as the first preset base length;

[0034] The process of determining the second preset base length includes:

[0035] Set the second preset base length to multiple second candidate lengths, where each second candidate length is a multiple of three; for each second candidate length, train a prediction model and calculate the accuracy of the prediction model; set the second candidate length associated with the prediction model with the maximum accuracy as the second preset base length.

[0036] As a further embodiment of the present invention, the convolutional layer further includes: a splicing module and a third convolutional kernel module;

[0037] The sizes and / or convolutional strides of the first convolutional kernels in the first convolutional kernel modules of the protein translation intensity prediction models are not the same;

[0038] The sizes and / or convolutional strides of the second convolutional kernels in the second convolutional kernel modules of the protein translation intensity prediction models are not the same;

[0039] The splicing module is used to perform splicing processing on the convolutional results of the first convolutional kernel module and the convolutional results of the second convolutional kernel module;

[0040] The third convolutional kernel module is used to perform convolutional processing on the splicing results of the splicing module.

[0041] As a further embodiment of the present invention, the first convolutional kernel module includes multiple first convolutional kernels, the size of the first convolutional kernel is [kernel_sizes, 4 + 2 * int(kernel_sizes / 2)], and the kernel_sizes of each first convolutional kernel are different;

[0042] The size of the second convolutional kernel in the second convolutional kernel module is [3, 4];

[0043] The size of the third convolutional kernel in the third convolutional kernel module is [kernel_sizes, 1 + 2 * int(kernel_sizes / 2)].

[0044] The second aspect of the present invention provides a method for predicting the protein translation intensity of prokaryotes, including:

[0045] Obtain multiple sequence data of the prokaryote to be measured;

[0046] Input each sequence data of the prokaryote to be measured into the protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the prokaryote to be measured;

[0047] Construct a genetic circuit according to the protein translation intensity of each sequence data of the prokaryote to be measured.

[0048] The third aspect of the present invention provides a device for determining a protein translation intensity prediction model of prokaryotes, including:

[0049] A data collection unit for pre-collecting multiple experimental data of various prokaryotes, where each experimental data includes the translation intensity of the prokaryote at each sequence data;

[0050] An encoding unit for encoding the experimental data to obtain encoded data recognizable by a computer;

[0051] A data partitioning unit for partitioning the encoded data of multiple groups of experimental data into a training set and a test set;

[0052] A training unit for using the sequence data as input and the translation intensity as output, and training a plurality of preset protein translation intensity prediction models using the training set, where the protein translation intensity prediction model includes a convolutional layer, a pooling layer, and a fully connected layer, and the convolutional layer at least includes a first convolutional kernel module and a second convolutional kernel module, and the first convolutional kernel module is used for performing convolutional processing on non-translated region data, and the second convolutional kernel module is used for performing convolutional processing on translated region data;

[0053] A testing unit for measuring the prediction accuracy of each protein translation intensity prediction model using the test set, and determining the final protein translation intensity prediction model according to the prediction accuracy of each protein translation intensity prediction model.

[0054] A fourth aspect of the present invention provides a prokaryotic protein translation intensity prediction device, including:

[0055] A data acquisition unit for acquiring multiple sequence data of a prokaryote to be measured;

[0056] A prediction unit for inputting each sequence data of the prokaryote to be measured into the protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the prokaryote to be measured;

[0057] An application unit for constructing a genetic circuit according to the protein translation intensity of each sequence data of the prokaryote to be measured.

[0058] A fifth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the method described in any one of the foregoing embodiments.

[0059] A sixth aspect of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of a computer device, it implements the method described in any one of the foregoing embodiments.

[0060] A seventh aspect of the present invention provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor of a computer device, it implements the method described in any one of the foregoing embodiments.

[0061] The method and device for determining a prokaryotic protein translation intensity prediction model provided by the present invention design a protein translation intensity prediction model including a first convolution kernel module and a second convolution kernel module. The first convolution kernel module is used to perform convolution processing on non-translated region data, and the second convolution kernel module is used to perform convolution processing on translated region data. It can comprehensively analyze the influence of non-translated regions and translated regions on protein translation intensity, improve the prediction accuracy and efficiency of prokaryotic protein translation intensity. At the same time, through the setting of the first convolution module and the second convolution module, the requirement for the amount of training samples can also be reduced, and the data acquisition cost can be reduced. By collecting multiple experimental data of various prokaryotes and training a protein translation intensity prediction model based on the multiple experimental data of various prokaryotes, the generalization ability of the model can be improved.

[0062] To make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 Shows the flowchart of the method for determining a prokaryotic protein translation intensity prediction model according to an embodiment of the present invention;

[0065] Figure 2 Shows the flowchart of the process of encoding sequence data according to an embodiment of the present invention;

[0066] Figure 3 Shows the structural diagram of a prokaryotic protein translation intensity prediction model according to an embodiment of the present invention;

[0067] Figure 4 Shows the flowchart of the method for predicting prokaryotic protein translation intensity according to an embodiment of the present invention;

[0068] Figure 5 Shows the structural diagram of the device for determining a prokaryotic protein translation intensity prediction model according to an embodiment of the present invention;

[0069] Figure 6 Shows the structural diagram of the device for predicting prokaryotic protein translation intensity according to an embodiment of the present invention;

[0070] Figure 7 Shows the schematic diagram of the verification result according to an embodiment of the present invention;

[0071] Figure 8 Shows the structural diagram of the computer device according to an embodiment of the present invention.

[0072] Explanation of the reference numerals in the drawings:

[0073] 301, Convolutional layer;

[0074] 302, Pooling layer;

[0075] 303, Fully connected layer;

[0076] 501, Data collection unit;

[0077] 502, Encoding unit;

[0078] 503, Data partitioning unit;

[0079] 504, Training unit;

[0080] 505, Testing unit;

[0081] 601, Data acquisition unit;

[0082] 602, Prediction unit;

[0083] 603, Application unit;

[0084] 802, Computer device;

[0085] 804, Processor;

[0086] 806, Memory;

[0087] 808, Driving mechanism;

[0088] 810, Input / output module;

[0089] 812, Input device;

[0090] 814, Output device;

[0091] 816, Rendering device;

[0092] 818, Graphical user interface;

[0093] 820, Network interface;

[0094] 822, Communication link;

[0095] 824, Communication bus. Detailed implementation manners

[0096] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0097] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0098] This specification provides the method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, it may include more or fewer operation steps. The step sequence listed in the embodiments is only one way among the execution sequences of numerous steps, and does not represent the only execution sequence. When the actual system or device product is executed, it can be executed in the order shown in the embodiments or the drawings or executed in parallel.

[0099] It should be noted that the data involved in the present invention (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by users or fully authorized by all parties, and the acquisition, transmission, storage, use and processing of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0100] It should be noted that in the embodiments of the present invention, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present invention, but it does not mean that the applicant has already or necessarily used this solution.

[0101] As a complex molecular machine composed of RNA and proteins, ribosomes play a crucial role in protein synthesis. For example, specific sequences and secondary structures in the 5’untranslated region (5’UTR) can affect the binding efficiency of translation initiation factors, thereby influencing the translation intensity; the codon usage bias in the coding sequence affects the binding rate of tRNA, and some codons may be more efficient than others, thus increasing the translation speed; miRNA inhibits translation or promotes degradation by binding to the 3'UTR (3’untranslated region, 3’UTR) or 5'UTR of mRNA, thereby reducing the translation intensity, and so on.

[0102] In the prior art, the following methods are usually adopted to study protein expression levels: RBS Library, thermodynamic models, and machine learning models.

[0103] The RBS Library method mainly relies on introducing a large number of mutations at specific sites of DNA sequences and screening the desired sequences from all mutants. Specifically, for example, researchers utilize the structural dynamics and sequence specificity of RNA to achieve post-transcriptional regulation of gene expression by designing different riboregulators and establishing a library for screening. After transcription, there is a ribosome binding site on mRNA, which can be used for ribosome docking. For cis-inhibition, researchers insert a complementary cis-repressive RNA (crRNA) upstream of the ribosome binding site of the target gene. This crRNA forms a stem-loop structure, physically hindering the binding of ribosomes to the RBS, thereby inhibiting gene expression. For trans-activation, a trans-activating RNA (taRNA) is expressed. After taRNA binds to crRNA, the inhibitory structure of crRNA is disrupted, the RBS is exposed, and ribosomes can rebind to the RBS to initiate gene translation, thereby activating the expression of the target gene. Another example is that researchers designed a DNA cassette system, inserting synthetic DNA oligonucleotides between the transcription and translation initiation sites of a gene to create various 5′ hairpin structures, establishing a library and screening. The strengths and complexities of these hairpin structures are different, thus affecting the half-life of mRNA. The research results show that the impact of 5′ hairpin structures on mRNA stability is significant and diverse. Some hairpin structures have a strong dependence on mRNA stability, while other structures have no significant effect. These findings indicate that 5′ hairpin structures play an important role in the stability of Escherichia coli mRNA, and gene expression can be regulated by designing specific hairpin structures.

[0104] The thermodynamic model method mainly studies protein translation based on thermodynamic models. For example, based on thermodynamic models, combined with key molecular interactions in the translation initiation process, and through optimization algorithms, it realizes the prediction and control of the target translation initiation rate. Another example is to construct a four-parameter free energy model using thermodynamic principles to quantify the binding free energy between ribosomes and mRNA.

[0105] The machine learning model method mainly includes the following three types:

[0106] 1. Adopt a method of unbiased design experiments to systematically analyze the effects of nucleotide sequences, secondary structures, codons, and amino acid properties on translation efficiency. Specifically, the researchers conducted a large-scale design experiment, designed 244,000 DNA sequences, used full factorial design to evaluate the combinations of nucleotides, secondary structures, codons, and amino acid properties, and measured the abundance and degradation of reporter gene transcripts, polysome profiles, protein yields, and growth rates for each sequence. Use Multivariate Analysis of Variance (ANOVA) to analyze the effects of sequence characteristics (such as secondary structure, codon composition, amino acid properties) and their first-order interactions on translation efficiency. Through this method, the independent contribution of each characteristic to translation efficiency can be quantified, and a linear regression model can be constructed using sequence characteristics (such as secondary structure free energy, codon composition, amino acid properties) as independent variables and protein yield or translation efficiency as the dependent variable to predict translation efficiency.

[0107] 2. A method combining massively parallel reporter assay (MPRA) and deep learning was developed to predict the impact of human 5'UTR sequences on translation efficiency. This method holds that cis-regulatory elements in many human 5'UTRs have been individually characterized, but there is currently a lack of a method to accurately predict protein expression solely based on 5'UTR sequences. This limits the ability to estimate the impact of genomic coding variations and the ability to engineer 5'UTRs for precise translation control. The method first designed an MPRA experiment, creating a library containing 280,000 gene sequences, each sequence containing a random 5'UTR and a constant region (including the CDS and 3'UTR of enhanced green fluorescent protein eGFP). In vitro transcription and transfection of HEK293T cells were used, the polysome fraction was collected and RNA sequencing was performed, and the translation efficiency was determined by measuring the average ribosome loading (MRL) of each 5'UTR. Next, a convolutional neural network (CNN) was used to train the model. The input to the model was 260,000 5'UTR sequences, and the test set was 20,000 sequences. Grid search was used to optimize the hyperparameters. The model identified sequences with upstream start codons (uAUGs) or upstream open reading frames (uORFs) with significantly reduced ribosome loading, and at the same time, the usage rates of CUG and GUG as alternative start codons were lower than that of AUG, which was consistent with prior knowledge. At the same time, by calculating the minimum free energy (MFE) of the 5'UTR and comparing it with the MRL, the inhibitory effect of secondary structure on ribosome loading was quantified. In addition, the model predicted 45 single nucleotide variations (SNVs) related to human diseases and explained 77% of these variations. By combining MPRA and deep learning methods, this method established a model capable of predicting translation efficiency from 5'UTR sequences.

[0108] 3. The accuracy and data efficiency of deep learning models in protein expression prediction under different dataset scales and sequence diversity conditions. By using strategies to control sequence diversity, methods to achieve high-precision prediction on smaller datasets were explored. This method used a fluorescence dataset containing more than 228,000 E. coli strains, with a 96nt variable region in front of the sfGFP coding gene of these strains. Randomization was carried out by designing experimental methods to ensure balanced coverage of the sequence space and controlled diversity of variations. This method trained a variety of machine learning models, including traditional non-deep learning models (such as ridge regression, multi-layer perceptron, support vector machine regression, and random forest regression) and convolutional neural networks (CNN). Different DNA encoding strategies (such as global biophysical properties, overlapping k-mers, one-hot encoding at single-base resolution) were used to characterize the sequences, and finally R 2Evaluate the prediction accuracy of the model on the test set. The study found that non-deep learning models perform poorly on small datasets, but the prediction accuracy can be significantly improved by increasing the size of the training set. CNN has a higher prediction accuracy than non-deep models with the same amount of data. Even on small datasets, CNN can achieve good performance. At the same time, controlling sequence diversity can significantly improve data efficiency and the coverage of the model. By introducing diverse sequences into the training set, the model can achieve accurate predictions within a larger range of sequence space.

[0109] In summary, the RBS Libarary requires a large number of experiments and screening, which has the problem of low efficiency; the thermodynamic model cannot distinguish the factors affecting protein translation intensity in different bacteria, that is, the predicted protein translation intensity is the same regardless of the bacteria; the machine learning model requires a large dataset, has a high acquisition cost, average generalization ability, and has the problem of low model accuracy.

[0110] In an embodiment of the present invention, to solve the technical problem in the prior art that the machine learning model for the protein translation efficiency of prokaryotes requires a large dataset, has a high acquisition cost, and average generalization ability, a method for determining a prokaryotic protein translation intensity prediction model is provided, as Figure 1 shown, including:

[0111] Step 101, pre-collect multiple experimental data of various prokaryotes, and each experimental data includes the translation intensity of the prokaryote at each sequence data.

[0112] In some specific embodiments, collecting multiple experimental data of various prokaryotes includes:

[0113] For each prokaryote, make a first sequence containing the green fluorescent protein gene, measure the fluorescence intensity when expressing the first sequence, and determine the translation intensity according to the measured fluorescence intensity.

[0114] In some specific embodiments, collecting multiple experimental data of various prokaryotes includes:

[0115] For each prokaryote, make a second sequence containing a specific sequence and gene, perform transcriptome sequencing and ribosome footprint sequencing on the plasmid of the second sequence, and obtain the translation intensity by dividing the normalized ribosome footprint sequencing result by the transcriptome sequencing result.

[0116] Step 102, encode the experimental data to obtain encoded data recognizable by a computer.

[0117] Specifically, this step includes: digitally encoding the translation intensity; performing binary one-hot encoding on the sequence data to obtain sequence encoded data of a fixed length.

[0118] Step 103: Divide the encoded data of multiple groups of experimental data into a training set and a test set.

[0119] To ensure the accuracy of the test, the distribution of experimental data in the training set and the test set is the same.

[0120] Step 104: Use the sequence data as the input and the translation intensity as the output, and train a plurality of preset protein translation intensity prediction models using the training set.

[0121] Among them, the plurality of preset protein translation intensity prediction models include a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer includes at least a first convolutional kernel module and a second convolutional kernel module. The first convolutional kernel module is used to perform convolutional processing on untranslated region data, and the second convolutional kernel module is used to perform convolutional processing on translated region data.

[0122] Among them, the untranslated region (UTR) specifically refers to a series of base sequences upstream of the start codon (the first three bases of the coding region - that is, CDS -, usually ATG). The coding region (CDS) is the region where the gene is located (that is, the part that is truly expressed as a protein during translation).

[0123] Step 105: Measure the prediction accuracy of each protein translation intensity prediction model using the test set, and determine the final protein translation intensity prediction model according to the prediction accuracy of each protein translation intensity prediction model.

[0124] In this embodiment, by designing a protein translation intensity prediction model including a first convolutional kernel module and a second convolutional kernel module, where the first convolutional kernel module is used to perform convolutional processing on untranslated region data, and the second convolutional kernel module is used to perform convolutional processing on translated region data, it is possible to comprehensively analyze the influence of the untranslated region and the translated region on the protein translation intensity, improve the prediction accuracy and efficiency of prokaryotic protein translation intensity, and at the same time, through the settings of the first convolutional module and the second convolutional module, the requirement for the training sample size can also be reduced, and the data acquisition cost can be reduced. The present invention can improve the generalization ability of the model by collecting multiple experimental data of multiple prokaryotes and training a protein translation intensity prediction model based on the multiple experimental data of multiple prokaryotes. In addition, the protein translation intensity prediction model established based on this embodiment can design and optimize gene circuits more precisely and efficiently, and achieve fine control of protein expression.

[0125] In some embodiments of the present invention, as Figure 2 shown, perform binary one-hot encoding on the sequence data to obtain sequence encoded data of a fixed length, including:

[0126] Step 201: Determine whether the untranslated region in the sequence data meets a first preset base length.

[0127] Step 202, if the untranslated region does not meet the first preset base length, adjust the untranslated region to the first preset base length.

[0128] Specifically, adjusting the untranslated region to the first preset base length includes:

[0129] When the length of the untranslated region is greater than the first preset base length, delete the N bases on the left side of the untranslated region;

[0130] When the length of the untranslated region is less than the first preset base length, supplement N bases on the left side of the untranslated region so that the length of the untranslated region is equal to the first preset base length.

[0131] The basis for the above adjustment method is as follows: The structure of a segment of DNA is usually an untranslated region + a translated region. For example, a 100-base untranslated region is followed by a 100-base translated region. What this invention aims to explore is the initiation intensity of protein translation, that is, the influence of the base pairs near the start codon on protein translation. The start codon is the first three bases of the translated region. Therefore, what this invention explores is the part of the untranslated region closest to the start codon and the part of the translated region closest to the start codon. If the position of the start codon is assumed to be the 0 position, what this method explores is the influence of the untranslated region [-the first preset base length, -1] and the translated region [0, the second preset base length - 1] on the translation intensity. Therefore, for the extra bases (the extra bases in the untranslated region are (-∞, -the first preset base length - 1], and the extra bases in the translated region are [the second preset base length, +∞)), their influence on the model accuracy should be excluded.

[0132] For example, if the first preset base length is 50 and the second preset base length is 36, then what this invention explores is the influence of the untranslated region [-50, -1] and the translated region [0, 35] on the translation intensity. Therefore, the extra bases (-∞, -51] and [36, +∞) should be deleted.

[0133] Step 203, after the untranslated region meets the first preset base length and the untranslated region is adjusted to the first preset base length, determine whether the length of the translated region in the sequence data meets the second preset base length. That is, if the untranslated region meets the first preset base length, or after the untranslated region is adjusted to the first preset base length, determine whether the length of the translated region in the sequence data meets the second preset base length.

[0134] Step 204, if the translated region does not meet the second preset base length, adjust the translated region to the second preset base length.

[0135] Specifically, adjusting the translated region to the second preset base length includes:

[0136] When the length of the translation region is greater than the second preset base length, delete the N bases on the right side of the translation region;

[0137] When the length of the translation region is less than the second preset base length, supplement N bases on the right side of the translation region so that the length of the translation region is equal to the second preset base length.

[0138] For the reason of the adjustment method in this step, reference can be made to the description of step 202, which will not be elaborated here.

[0139] Step 205, if the translation region meets the second preset base length and after adjusting the translation region to the second preset base length, encode the sequence data. That is, if the translation region meets the second preset base length or after adjusting the translation region to the second preset base length, encode the sequence data.

[0140] In an embodiment of the present invention, the process of determining the first preset base length includes:

[0141] Set the first preset base length to multiple first candidate lengths; for each first candidate length, train a prediction model and calculate the accuracy of the prediction model; screen out the prediction models that meet the preset accuracy; from the first candidate lengths related to the screened-out prediction models, select the smallest first candidate length as the first preset base length;

[0142] The process of determining the second preset base length includes:

[0143] Set the second preset base length to multiple second candidate lengths, where each second candidate length is a multiple of three; for each second candidate length, train a prediction model and calculate the accuracy of the prediction model; set the second candidate length related to the prediction model with the maximum accuracy as the second preset base length.

[0144] In a specific embodiment of the present invention, the first preset base length is 50 and the second preset base length is 36.

[0145] Specifically, for the non-translation region: First, set the length of the non-translation region to 0 to explore the separate influence of the non-translation region by excluding the influence of the translation region on translation. In previous experiments, several different lengths were set for testing. For shorter lengths (such as 30, etc.), they lost some influencing factors resulting in lower translation accuracy than relatively longer lengths. And for longer lengths (such as 80, 90, etc.), their prediction accuracy did not show further improvement in general cases, and in some cases, due to the need for a large number of N bases (specified as empty, encoded as [0, 0, 0, 0] to exclude their influence on the model) for the original short sequences, data redundancy and sparsity problems occurred. After the above analysis, the first preset base length is set to 50.

[0146] For the translation region, the reading of the translation region is in units of three bases. Combining prior knowledge (usually the first 35 base pairs have a greater impact on the translation intensity) and experimental results (experiments were conducted using 10, 12, 30, and 36 respectively, and the best results were obtained when the length was 36), 36 is therefore selected as the second preset base length.

[0147] In one embodiment of the present invention, as Figure 3 shown, the preset multiple protein translation intensity prediction models include a convolutional layer 301, a pooling layer 302, and a fully connected layer 303. Among them, the fully connected layer is a network structure.

[0148] The convolutional layer 301 includes a first convolutional kernel module, a second convolutional kernel module, a splicing module, and a third convolutional kernel module.

[0149] The first convolutional kernel module is used to perform convolutional processing on non-translation region data, the second convolutional kernel module is used to perform convolutional processing on translation region data, the splicing module is used to splice the convolutional results of the first convolutional kernel module and the convolutional results of the second convolutional kernel module, and the third convolutional kernel module is used to perform convolutional processing on the splicing results of the splicing module.

[0150] The sizes and / or convolutional strides of the first convolutional kernels in the first convolutional kernel modules of each protein translation intensity prediction model are not the same.

[0151] The sizes and / or convolutional strides of the second convolutional kernels in the second convolutional kernel modules of each protein translation intensity prediction model are not the same.

[0152] In a specific embodiment of the present invention, the first convolutional kernel module includes multiple first convolutional kernels, the size of the first convolutional kernel is [kernel_sizes, 4 + 2 * int(kernel_sizes / 2)], and the kernel_sizes of each first convolutional kernel are different. For example, kernel_sizes are set to 3, 5, 7, etc. Specifically, during implementation, in order to avoid missing information, the convolutional stride of the convolutional kernels in the first convolutional kernel module is 1.

[0153] The size of the second convolutional kernel in the second convolutional kernel module is [3, 4]. Specifically, considering the characteristics of the DNA sequence, the translation region sequence starts with the start codon (usually ATG, and in a few cases GTG, TTG, CTG), and each reading during the translation process is in groups of 3 bases. Therefore, the convolutional kernel specification for the translation region is set to [3, 4], and the convolutional stride during convolution is 3.

[0154] The size of the third convolutional kernel in the third convolutional kernel module is [kernel_sizes, 1 + 2 * int(kernel_sizes2 / 2)].

[0155] During specific implementation, multiple first convolutional kernels in the first convolutional kernel module can be used to perform secondary convolution on non-translated region data. After splicing the convolution results of the first convolutional kernel module and the convolution results of the second convolutional module through a splicing module, a single convolution is performed using a convolutional kernel [kernel_sizes, 1 + 2 * int(kernel_sizes2 / 2)].

[0156] In this embodiment, considering the problem that the influence of bases at different positions in the non-translated region and different bases at the same position on translation is not known to people, multiple convolutional kernels are set for the non-translated region. Finally, the results of convolutional kernels of different lengths are spliced together, and the data to be sampled is selected in the fully connected layer, thereby improving the prediction accuracy of the model.

[0157] At the same time, this embodiment also takes into account that during translation, the ribosome moves along the 5' end to the 3' end of the mRNA transcribed from DNA, reads three bases each time, that is, one codon, and then recruits the corresponding aminoacyl-tRNA to the ribosome for peptide chain synthesis according to the amino acid corresponding to the codon. This reading method is continuous, without overlap and interval, ensuring the accuracy and efficiency of protein synthesis. According to the characteristics of the DNA sequence, the translation region sequence starts from the start codon (usually ATG, and in a few cases GTG, TTG, CTG), and each reading during translation is in groups of 3 bases. Therefore, the convolutional kernel specification of the translation region is set to [3, 4]. This way of setting the convolutional kernel of the translation region can improve the prediction accuracy of the translation region.

[0158] In an embodiment of the present invention, a method for predicting the protein translation intensity of prokaryotes is also provided, as Figure 4 shown, including:

[0159] Step 401, obtaining multiple sequence data of the prokaryote to be measured.

[0160] Step 402, inputting each sequence data of the prokaryote to be measured into the protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the prokaryote to be measured.

[0161] Step 403, constructing a genetic circuit according to the protein translation intensity of each sequence data of the prokaryote to be measured.

[0162] Protein expression regulation is the key to constructing genetic circuits with complex logical functions. By reasonably combining and designing different protein expression regulatory elements (such as promoters, enhancers, transcription factor binding sites, etc.), genetic circuits with logical functions such as AND gates, OR gates, and NOT gates can be constructed. These genetic circuits can integrate and process multiple signals to achieve precise control of cell behavior.

[0163] Protein expression regulation helps to simulate and reconstruct complex signal pathways and regulatory networks in biological systems. By designing and constructing artificial gene circuits, we can gain in-depth understanding of the working principles of biological systems, providing a theoretical basis and technical support for the development of bioengineering and biotechnology. For example, constructing gene circuits that can sense changes in the concentration of specific substances in the environment and respond accordingly for the development of biosensors.

[0164] Based on the same inventive concept, the present invention also provides an apparatus for determining a prokaryotic protein translation intensity prediction model, as described in the following embodiments. Since the principle of the apparatus for determining a prokaryotic protein translation intensity prediction model to solve problems is similar to that of the method for determining a prokaryotic protein translation intensity prediction model, therefore, the implementation of the apparatus for determining a prokaryotic protein translation intensity prediction model can refer to the method for determining a prokaryotic protein translation intensity prediction model, and the repeated parts will not be elaborated.

[0165] Specifically, as Figure 5 shown, the apparatus for determining a prokaryotic protein translation intensity prediction model includes:

[0166] A data collection unit 501, configured to pre-collect a plurality of experimental data of a variety of prokaryotes, and each experimental data includes the translation intensity of the prokaryote at each sequence data.

[0167] An encoding unit 502, configured to encode the experimental data to obtain encoded data recognizable by a computer.

[0168] A data partitioning unit 503, configured to partition the encoded data of multiple groups of experimental data into a training set and a test set.

[0169] A training unit 504, configured to use the sequence data as input and the translation intensity as output, and train a plurality of preset protein translation intensity prediction models using the training set.

[0170] Among them, the plurality of preset protein translation intensity prediction models include a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer at least includes a first convolutional kernel module and a second convolutional kernel module. The first convolutional kernel module is used to perform convolutional processing on non-translated region data, and the second convolutional kernel module is used to perform convolutional processing on translated region data.

[0171] A testing unit 505, configured to measure the prediction accuracy of each protein translation intensity prediction model using the test set, and determine the final protein translation intensity prediction model according to the prediction accuracy of each protein translation intensity prediction model.

[0172] The present invention designs a protein translation intensity prediction model including a first convolution kernel module and a second convolution kernel module. The first convolution kernel module is used to perform convolution processing on untranslated region data, and the second convolution kernel module is used to perform convolution processing on translated region data, which can comprehensively analyze the influence of untranslated regions and translated regions on protein translation intensity, improve the prediction accuracy and efficiency of prokaryotic protein translation intensity. At the same time, through the setting of the first convolution module and the second convolution module, the requirement for the amount of training samples can also be reduced, and the data acquisition cost can be reduced. The present invention collects multiple experimental data of various prokaryotes and trains a protein translation intensity prediction model based on the multiple experimental data of various prokaryotes, which can improve the generalization ability of the model.

[0173] In an embodiment of the present invention, a device for predicting prokaryotic protein translation intensity is further provided, as Figure 6 shown, including:

[0174] A data acquisition unit 601, configured to acquire multiple sequence data of a to-be-detected prokaryote;

[0175] A prediction unit 602, configured to input each sequence data of the to-be-detected prokaryote into the protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the to-be-detected prokaryote;

[0176] An application unit 603, configured to construct a genetic circuit according to the protein translation intensity of each sequence data of the to-be-detected prokaryote.

[0177] In a specific embodiment of the present invention, the above technical solution is verified with the experimental data of Corynebacterium glutamicum and Bacillus subtilis, and the verification results are as Figure 7 shown. Figure 7 In, the abscissa is the observed value (experimental measurement value, that is, the corresponding translation intensity measured by experiment), and the ordinate is the model prediction value (the translation intensity predicted by the model). All values are calculated using log10. From Figure 7 it can be seen that the model has high accuracy and feasibility in the verification of laboratory data.

[0178] In an embodiment of the present invention, a computer device is further provided, as Figure 8As shown, the computer device 802 may include one or more processors 804, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 802 may also include any memory 806 for storing any kind of information such as code, settings, data, etc. By way of non-limiting example, the memory 806 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any memory may store information using any technology. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 802. In one case, when the processor 804 executes the associated instructions stored in any memory or combination of memories, the computer device 802 may perform any operation of the associated instructions. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.

[0179] The computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via the input device 812) and for providing various outputs (via the output device 814). A particular output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input device 812, and the output device 814 may not be included and the computer device 802 may be only a computer device in a network. The computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.

[0180] The communication link 822 may be implemented in any way, for example, via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 822 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.

[0181] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is run by a processor, it executes the steps of the above method.

[0182] An embodiment of the present invention also provides a computer-readable instruction, and when the processor executes the instruction, the program therein causes the processor to execute the steps of the above method.

[0183] It should be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0184] It should also be understood that in the embodiments of the present invention, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the front and rear associated objects.

[0185] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0186] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0187] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be a connection in electrical, mechanical, or other forms.

[0188] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0189] In addition, in each embodiment of the present invention, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0190] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0191] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for determining a prokaryotic protein translation intensity prediction model, characterized in that: include: Collect multiple experimental data of various prokaryotes in advance, each experimental data includes the translation intensity of the prokaryotes at each sequence data; Encoding the experimental data to obtain coded data recognized by a computer; Divide the encoded data of multiple groups of experimental data into training sets and test sets; The sequence data is used as input and the translation strength is used as output, and the training set is used to train a plurality of preset protein translation strength prediction models, wherein the protein translation strength prediction model comprises a convolution layer, a pooling layer and a fully connected layer, and the convolution layer comprises at least a first convolution kernel module and a second convolution kernel module, the first convolution kernel module is used to perform convolution processing on the non-translation region data, and the second convolution kernel module is used to perform convolution processing on the translation region data; The test set is used to measure the prediction accuracy of each protein translation strength prediction model, and the final protein translation strength prediction model is determined according to the prediction accuracy of each protein translation strength prediction model.

2. The method according to claim 1, characterized in that Collect multiple experimental data of various prokaryotes including: For each prokaryotic organism, preparing a first sequence containing a green fluorescent protein gene, measuring the fluorescence intensity when the first sequence is expressed, and determining the translation intensity based on the measured fluorescence intensity; or For each prokaryotic organism, a second sequence comprising a specific sequence and a gene is prepared, transcriptome sequencing and ribosomal imprint sequencing are performed on the plasmid of the second sequence, and the translation intensity is obtained by dividing the normalized ribosomal imprint sequencing result by the normalized transcriptome sequencing result.

3. The method according to claim 1, characterized in that The experimental data is encoded to obtain computer-recognizable encoded data, including: Digital coding of translation intensity; The sequence data is binary one-hot encoded to obtain sequence encoding data of fixed length.

4. The method according to claim 3, characterized in that Perform binary one-hot encoding on the sequence data to obtain fixed-length sequence encoding data, including: Determining whether the untranslated region in the sequence data meets a first preset base length; If the untranslated region does not conform to the first preset base length, adjusting the untranslated region to the first preset base length; If the untranslated region meets the first preset base length and after the untranslated region is adjusted to the first preset base length, determining whether the length of the translated region in the sequence data meets the second preset base length; If the translation region does not meet the second preset base length, adjusting the translation region to the second preset base length; If the translation region meets the second preset base length and after the translation region is adjusted to the second preset base length, the sequence data is encoded.

5. The method according to claim 4, characterized in that Adjusting the untranslated region to a first predetermined base length comprises: When the length of the untranslated region is greater than the first preset base length, the N bases on the left side of the untranslated region are deleted; When the length of the untranslated region is less than the first preset base length, N bases are added to the left side of the untranslated region so that the length of the untranslated region is equal to the first preset base length.

6. The method according to claim 4, characterized in that Adjusting the translation region to a second predetermined base length comprises: When the length of the translation region is greater than the second preset base length, the N bases on the right side of the translation region are deleted; When the length of the translation region is less than the second preset base length, N bases are added to the right side of the translation region so that the length of the translation region is equal to the second preset base length.

7. The method according to claim 4, characterized in that The first preset base length determination process includes: The first preset base length is set to a plurality of first candidate lengths; for each first candidate length, a prediction model is trained and the accuracy of the prediction model is calculated; a prediction model that satisfies the preset accuracy is selected; and from the first candidate lengths related to the selected prediction models, the smallest first candidate length is selected as the first preset base length; The second preset base length determination process includes: The second preset base length is set to multiple second candidate lengths, wherein each second candidate length is a multiple of three; for each second candidate length, a prediction model is trained, and the accuracy of the prediction model is calculated; and the second candidate length associated with the prediction model with the highest accuracy is set to the second preset base length.

8. The method according to claim 1, characterized in that The convolution layer also includes: a splicing module and a third convolution kernel module; The size and / or convolution step of the first convolution kernel in the first convolution kernel module of each protein translation intensity prediction model are different; The size and / or convolution step of the second convolution kernel in the second convolution kernel module of each protein translation intensity prediction model are different; The splicing module is used to splice the convolution result of the first convolution kernel module and the convolution result of the second convolution kernel module; The third convolution kernel module is used to perform convolution processing on the splicing result of the splicing module.

9. The method according to claim 8, characterized in that The first convolution kernel module includes multiple first convolution kernels, the size of the first convolution kernel is [kernel_sizes, 4+2*int(kernel_sizes / 2)], and the kernel_sizes of each first convolution kernel is different; The size of the second convolution kernel in the second convolution kernel module is [3, 4]; The size of the third convolution kernel in the third convolution kernel module is [kernel_sizes,1+2*int(kernel_sizes2 / 2)].

10. A method for predicting the translation strength of prokaryotic proteins, characterized in that: include: Acquire multiple sequence data of the prokaryotic organism to be tested; Inputting each sequence data of the prokaryotic organism to be tested into a protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the prokaryotic organism to be tested; According to the protein translation intensity of each sequence data of the prokaryotic organism to be tested, a gene circuit is constructed.

11. A device for determining a prokaryotic protein translation intensity prediction model, characterized in that: include: A data collection unit, used for pre-collecting a plurality of experimental data of a plurality of prokaryotes, each experimental data including the translation intensity of the prokaryotes at each sequence data; A coding unit, used for coding the experimental data to obtain coded data recognized by a computer; A data division unit, used for dividing the coded data of multiple groups of experimental data into a training set and a test set; A training unit, used to take the sequence data as input and the translation strength as output, and use the training set to train a plurality of preset protein translation strength prediction models, wherein the protein translation strength prediction model comprises a convolution layer, a pooling layer and a fully connected layer, the convolution layer comprises at least a first convolution kernel module and a second convolution kernel module, the first convolution kernel module is used to perform convolution processing on the non-translation region data, and the second convolution kernel module is used to perform convolution processing on the translation region data; The testing unit is used to measure the prediction accuracy of each protein translation strength prediction model using the test set, and determine the final protein translation strength prediction model according to the prediction accuracy of each protein translation strength prediction model.

12. A prokaryotic protein translation intensity prediction device, characterized in that: include: A data acquisition unit, used to acquire multiple sequence data of the prokaryotic organism to be tested; A prediction unit, used for inputting each sequence data of the prokaryotic organism to be tested into a protein translation intensity prediction model to obtain the protein translation intensity of each sequence data of the prokaryotic organism to be tested; The application unit is used to construct a gene circuit according to the protein translation intensity of each sequence data of the prokaryotic organism to be tested.

13. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor of a computer device, the method according to any one of claims 1 to 9 is implemented.

15. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor of a computer device, the method according to any one of claims 1 to 9 is implemented.