Gene sequence optimizing and screening method based on deep learning
By constructing a gene sequence optimization and screening model using deep learning methods, the problem of unknown host action mechanisms in recombinant protein expression was solved, enabling efficient screening of gene sequences with high expression levels in E. coli and significantly improving the soluble expression of α-glucan phosphorylase.
Patent Information
- Application Number
- CN202410619650.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for recombinant protein expression suffer from low expression levels of heterologous proteins or protein folding errors due to unknown host mechanisms. Furthermore, deep learning methods lack reliability and broad applicability, making it difficult to effectively optimize gene sequences.
A deep learning-based gene sequence optimization and screening method was adopted. By combining a synonymous codon generation model and a gene expression level prediction model with multiple machine learning algorithms, a mapping relationship between gene sequences and soluble expression levels was constructed to generate gene sequences with high expression levels.
This method achieves efficient screening of gene sequences with optimal soluble expression levels in E. coli with an accuracy of 82%, and improves protein expression levels compared to existing methods, especially increasing the soluble expression level of α-glucan phosphorylase by 20.52 times.
Smart Images

Figure CN120998306A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of protein expression, and particularly relates to a gene sequence optimization and screening method based on deep learning. BACKGROUND
[0002] High-efficiency soluble expression of recombinant proteins is a crucial link in the fields of modern biological science and bioengineering. With the development of synthetic biology technology and industry, high-efficiency expression of target genes through optimization can reduce the cost of catalysts, improve production efficiency, and reduce production costs.
[0003] Most of the existing optimization strategies mainly optimize sequences based on biological indicators and empirical indicators. By optimizing gene synthesis, regulating transcription and translation efficiency, and considering protein folding and other factors, the optimization results indicators such as codon adaptation index (CAI) and GC content are adjusted. However, these factors are not completely independent, but rather interact with each other, collectively affecting the expression level of recombinant proteins. Moreover, due to the presence of more unknown mechanisms within the host, optimization of the above factors may still result in low expression of heterologous proteins or protein folding errors forming inclusion bodies.
[0004] With the rapid development of deep learning technology, it has shown great potential in the field of natural language processing, such as ChatGPT generating dialogues and Al pahFold2 predicting protein three-dimensional structures. Therefore, in addition to the above codon optimization by adjusting factors affecting protein expression, several deep learning technologies have been used to optimize sequences, such as BiLSTM-CRF, ICOR, and CO-T5. However, there are still problems, such as BiLSTM-CRF and ICOR not introducing the relationship between gene sequences and their soluble expression levels, and CO-T5 being the latest research achievement, but it only fine-tunes on a few similar proteins and lacks reliable experimental verification.
[0005] Therefore, there is an urgent need to find a codon optimization strategy that can avoid incomplete understanding of the mechanisms within the target host and is suitable for a wide range of species-derived proteins. In addition, although traditional codon optimization techniques generally consider multiple factors affecting protein expression, these optimization methods are limited by unknown mechanisms within the target host cell, resulting in unreliable and incorrect optimization. Deep learning technology, as the most promising solution, can translate protein sequence language into gene sequence language based on the codon usage patterns and distribution characteristics of the target host expression system. However, due to the uncertainty in generating gene sequences, the generated codon sequences still need to be screened. SUMMARY
[0006] The technical problem to be solved by the present application is to provide a gene sequence optimization and screening method based on deep learning.
[0007] A gene sequence optimization and screening method based on deep learning, the steps are as follows:
[0008] (1) Set the original amino acid sequence or nucleotide sequence of the recombinant protein to be optimized;
[0009] (2) The operator sets the synonymous codon generation SCG model decoding parameters to determine the number of gene sequences generated by synonymous codons;
[0010] (3) The operator selects a machine learning algorithm to evaluate the expression level of the gene sequence.
[0011] Preferably, in the gene sequence optimization and screening method based on deep learning, the specific operation method of step (1) is: setting the sequence parameter to be optimized as the original amino acid sequence or nucleotide sequence of the recombinant protein, the model will check the sequence type of the input content before accepting the sequence parameter to be optimized, if it is an amino acid sequence, the model will add a termination symbol "*" at the end of the sequence, if it is a nucleotide sequence, the model will convert the nucleotide sequence into an amino acid sequence according to the standard codon table.
[0012] Preferably, in the gene sequence optimization and screening method based on deep learning, the specific operation method of step (2) is: setting the synonymous codon generation model decoding parameters for executing the beam search decoding strategy, by setting multiple sets of model decoding parameters, the total number of generated gene sequences is the sum of the beam size of multiple parameters.
[0013] Preferably, in the gene sequence optimization and screening method based on deep learning, the model decoding parameters are batch size (BatchSize), beam search width (BeamWidth), and beam search result size (BeamSize).
[0014] Preferably, in the gene sequence optimization and screening method based on deep learning, in step (2), the SCG model receives the amino acid sequence and decoding parameters, extracts the amino acid sequence information, combines the host intracellular codon usage mode information learned in the model, and iteratively generates the next nucleotide through beam search and pruning operations until all gene sequences are generated, and then outputs to a table file after removing duplicates.
[0015] Preferably, in the above-mentioned gene sequence optimization and screening method based on deep learning, the specific operation method of step (3) is: providing four machine learning algorithms: support vector machine (SVM), logistic regression (LR), deep perception network (MLP) and light gradient boosting machine (LightGBM), selecting several or all of the algorithms to evaluate the gene sequence, evaluating each gene sequence by the algorithm, obtaining a probability score of the solubility expression level ranging from 0 to 1, and the closer to 1, the more likely the gene sequence is a high solubility expression level gene sequence; and calculating the average probability score of each gene, and the highest comprehensive ranking is marked as the final codon optimization result.
[0016] Preferably, in the above-mentioned gene sequence optimization and screening method based on deep learning, in step (3), the gene sequence file is generated based on the SCG model, 18 feature descriptors are used for constructing a 1024-dimensional sequence feature vector for each gene sequence, including codon composition and various physical and chemical properties, and 768-dimensional sequence features extracted by a transfer learning language model are used, including high-dimensional gene sequence information. The machine learning algorithm inputs the probability that the gene sequence belongs to a high expression level based on the above-mentioned total of 1792-dimensional feature vectors, and saves the gene sequence to a file; calculates the mean probability that each gene sequence belongs to a high expression level sequence, saves the mean probability to a file, and takes the sequence with the highest mean probability as the final optimization result.
[0017] Preferably, in the above-mentioned gene sequence optimization and screening method based on deep learning, the recombinant protein is alpha-glucan phosphorylase.
[0018] The beneficial effects of the present application are:
[0019] The gene sequence optimization and screening method based on deep learning is a method for optimizing and screening a gene (DNA) sequence of a to-be-expressed protein through a deep learning algorithm. An end-to-end deep learning algorithm is adopted to learn the codon usage pattern in Escherichia coli, so that sequence feature information can be extracted from a recombinant protein amino acid sequence, and the codon usage pattern information of Escherichia coli is combined to realize the design of a large number of gene sequences similar to the genome sequence of Escherichia coli and having the codon distribution pattern of Escherichia coli, and the situation that the result generated by the deep learning model falls into a local optimal solution is avoided. Meanwhile, these gene sequences have the characteristics of easy translation expression and possible high soluble expression level in Escherichia coli. In order to find out the gene sequence with the best expression level, the soluble expression level information of the gene sequence is introduced, a classic machine learning algorithm is trained through biological feature engineering and transfer learning feature engineering, and a mapping relationship between the gene sequence and the soluble expression level is constructed, so that the gene sequence with the best soluble expression level is selected from a large number of synonymous codon sequences of the target recombinant protein. Specifically, compared with the prior art, the method has the following advantages:
[0020] (1) The deep learning language model trained by using the largest current non-redundant Escherichia coli gene sequence data of more than 150,000 pieces can effectively avoid the influence of unknown mechanisms of the target host, and meanwhile, based on a widely-sourced gene soluble expression level data set, a mapping relationship between the gene sequence and the soluble expression level is constructed by using feature engineering combined with a machine learning algorithm, and the accuracy rate of identifying high expression level sequences is as high as 82%;
[0021] (2) Compared with the impossible global search of synonymous codons and the greedy search prone to generate local optimal solutions, the beam search strategy is adopted to generate the gene sequence with the highest global score, and the pruning operation is adopted, that is, when the model decoding gradually generates the next codon, the codon not matching the corresponding amino acid is removed, so that the amino acid sequence converted from all the synonymous codon sequences according to the standard codon table is consistent with the original amino acid sequence;
[0022] (3) The gene expression level data set is derived from 171 organisms, including 2 eukaryotes, 18 archaea and 151 bacteria, and 6348 genes, and meanwhile, feature engineering is improved, 18 kinds of effective feature descriptors are combined to form a 1024-dimensional biological feature vector, and 768-dimensional sequence features extracted by a transfer learning model are adopted, and a plurality of machine learning algorithms are used for common evaluation, and the accuracy rate is high. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1An implementation flowchart of the deep learning-based gene sequence optimization and screening method according to the present application.
[0024] Figure 2 A comparison chart of soluble expression of different gene sequences of alpha-glucan phosphorylase from Thermotoga maritima, wherein Opt is the sequence optimized based on the method according to the present application, GenScript is the sequence optimized by the codon tool of GenScript Corporation, and Wild-type is the wild-type sequence. DETAILED DESCRIPTION
[0025] In order for those skilled in the art to better understand the technical solutions of the present application, the technical solutions of the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0026] The following deep learning-based gene sequence optimization and screening method takes Escherichia coli as an example of a chassis microorganism to construct a synonymous codon generation model (SCG) and a gene expression level prediction model (GELP).
[0027] The SCG model component is a classic sequence-to-sequence model in the field of natural language processing, and its network structure adopts the famous Encoder-Decoder architecture, which is composed of an amino acid sequence encoding embedding layer, an encoder, a nucleotide sequence encoding embedding layer, a decoder, and a linear layer for generating prediction results. The input size of the amino acid sequence encoding embedding layer of the SCG model is 25, the output size is 512, the number of Encoder layers is 3, the input size of the nucleotide sequence encoding embedding layer is 68, the output size is 512, the number of Decoder layers is 3, the number of attention heads is 8, the size of the hidden layer is 512, the optimizer is Adam, the learning rate is 0.001, and the loss function is KLDivLoss. The SCG model is iteratively trained on a non-redundant Escherichia coli dataset of more than 150,000 collected from the NCBI Reference Sequence database. The encoding method is: a dictionary is created for protein sequences and gene sequences respectively, both of which contain 4 special mappings, and numbers 0-3 represent <unk>unknown character, <pad>padding character, <bos>start character, <eos>Stop character. Protein sequence is assigned a unique and distinct value with 21 numbers (4-24) for 20 amino acids and a stop symbol *, where the number assigned to each amino acid is assigned according to the frequency of each amino acid in the dataset; Similarly, DNA sequence is assigned a unique and distinct value with 64 numbers (4-67) for 64 codons, where the number assigned to each codon is assigned according to the frequency of each codon in the dataset. The encoder and decoder of the SCG model generate the probability distribution of the next base by collecting the input protein sequence and the generated gene sequence features, and then generate a large number of synonymous codon sequences that meet the codon usage pattern of E. coli by the model decoding method of beam search and pruning operation, which retains several bases with higher probability and splices them into the existing gene sequence, and then iterates in turn to generate a large number of synonymous codon sequences that meet the codon usage pattern of E. coli. The SCG model optimizes the sequence to use higher preference codons to design and generate gene sequences, and the overall peak value of the optimized sequence CAI is leveled at 0.9, and the model tends to use preferred codons while not choosing all codons with the highest preference codon, which generates valuable rare codons with low frequency, and the average identity measured in the test set reaches 79% compared with the original gene.
[0028] The GELP model component is composed of multiple model algorithms (Support Vector Machine (SVM), Logistic Regression (LR), Deep Perceptron Network (MLP) and Light Gradient Boosting Machine (LightGBM)), the purpose is to build the mapping relationship between gene sequence and gene expression level, including using biological feature engineering and transfer learning feature engineering to extract sequence features, and then through machine learning algorithm for prediction evaluation, these model algorithms are all on a large-scale unified expression and purification of recombinant protein gene expression data in Escherichia coli (from NESG), and after pretreatment, only the gene expression data set containing low expression low solubility, high expression high solubility two types is iteratively trained. Using iDNA class in iFeatureOmega-CLI tool (https: / / github.com / Superzchen / iFeatureOmega-CLI) from 7 categories of 37 feature descriptors including nucleotide composition, nucleotide position specificity, electron-ion interaction pseudopotential, autocorrelation and cross-covariance, physicochemical properties, mutual information and pseudo-nucleotide composition (Table 1) to select 18 effective feature descriptors (Table 1) to form a 1024-dimensional biological feature vector. At the same time, using transfer learning technology, fine-tuning large language model (DNABERT-2: https: / / github.com / MAGICS-LAB / DNABERT_2), and using DNABERT-2 model to extract 768-dimensional feature vector from gene sequence. After splicing two kinds of features, multiple machine learning algorithms are used for ten-fold cross-validation, among which support vector machine (SVM), logistic regression (LR), deep perception network (MLP) and light gradient boosting machine (LightGBM) perform best, the accuracy is more than 75%, the highest is 82.5%, the area under the ROC curve is more than 0.83, the Matthew correlation coefficient is more than 0.5, and the prediction performance is strong.
[0029] Table 1
[0030]
[0031]
[0032] Example 1
[0033] As Figure 1 As shown, a deep learning-based alpha-glucan phosphorylase (alphaGP, the protein to be optimized) gene sequence optimization and screening method, alphaGP itself is a key enzyme for decomposing alpha-glucan, and can reversibly catalyze the phosphorolysis reaction of alpha-glucan. AlphaGP has important significance in many in vitro synthetic biology systems with alpha-glucan as substrate, such as hydrogen production from sugar, electricity production from sugar, and inositol production from starch. The alphaGP protein used in the experiment is derived from Thermotoga maritima, and the Uniport ID is O33831 (https: / / www.uniprot.org / uniprotkb / O33831 / entry), the sequence length is 822, the molecular weight is 96.139kDa, the amino acid sequence is shown in the sequence table SEQ ID NO: 1, and the wild-type gene sequence is named Wild-type, and the nucleotide sequence is shown in the sequence table SEQ ID NO: 2. The specific steps are as follows:
[0034] Step 1, set the sequence parameter to be optimized:
[0035] The sequence parameter to be optimized is set as the amino acid sequence of the alphaGP protein, and a termination symbol is added after the amino acid sequence of the alphaGP protein, and the SCG model is input;
[0036] Step 2: Set the synonymous codon generation model decoding parameter:
[0037] As shown in Table 2, a total of 6 sets of decoding parameters are set, and the parameters are input into the SCG model. The SCG model encoder captures the protein amino acid sequence information, and the decoder combines the learning of the E. coli codon usage pattern and the gradually generated nucleotide sequence information, and searches for the optimal sequence in the sequence solution space according to the decoding scheme parameter setting, and a total of 1260 gene sequences are expected to be generated;
[0038] Table 2
[0039]
[0040]
[0041] The parameters in step 2 above can be set in countless sets, the batch size is the number of model calculations at a time (generally an exponential power of 2, conducive to computer calculation), the beam search width is the number of results retained under one calculation (ranging from 1 to 6, corresponding to the maximum in amino acids, leucine L, which has 6 synonymous codons), and the beam search result size is the total output result data amount (ranging from 1 to infinity), the three parameters together determine the time consumption and the computing resources required for the SCG model optimization. The SCG model encoder will capture the protein amino acid sequence information, and the decoder will search for the optimal sequence in the sequence solution space according to the context information, combined with the learned Escherichia coli codon usage pattern and the gradually generated nucleotide sequence information, according to the decoding scheme parameter setting. A total of 1260 gene sequences are expected to be generated.
[0042] Step 3: Select machine learning algorithms to evaluate the expression levels of gene sequences
[0043] SVM, LR, MLP and LightGBM are selected as evaluation algorithms to evaluate 1260 gene sequences. At the same time, based on the αGP protein amino acid sequence, the nucleotide sequence obtained by optimization using the online codon optimization tool of Genscript (https: / / www.genscript.com.cn / tools / gensmart%2dcodon%2doptimization) is marked as GenScript, and the nucleotide sequence is shown in the sequence table SEQ ID NO: 3. The sequences Wild-type and GenScript are also predicted using the above selected evaluation algorithms, and the probability value between 0 and 1 is finally obtained. The sequence with the maximum probability mean value is taken as the final optimization result, which is named Opt, and the nucleotide sequence is shown in the sequence table SEQ ID NO: 4. The prediction results are shown in Table 3 below:
[0044] Table 3
[0045]
[0046]
[0047] Example 2
[0048] The Wild-type, GenScript and Opt sequences described in Table 2 of Example 1 were constructed into E. coli pET28a vector to obtain the corresponding plasmids, and the plasmids were transformed into E. coli BL21 (DE3) for expression. The colonies on the TB solid medium plate were picked with a gun head and placed in TB liquid medium containing Kana resistance at a concentration of 50 μg / mL, and the medium volume was 5 mL. Then, the culture was incubated at 37°C in a shaker to OD600=0.6-0.8, IPTG was added to a final concentration of 0.1 mM for induction, and the shaker temperature was changed to 16°C for incubation for 20-22 hours. After centrifugation at 6000 rpm for 5 minutes, the cells were collected, resuspended and precipitated with buffer A (50 mM HEPES, 50 mM NaCl, pH 7.5). The cell suspension was subjected to ultrasonic disruption, and the cell lysate was centrifuged at 8000 rpm for 20 min to collect the supernatant. SDS-PAGE was performed with 12% protein gel to detect the protein purity in the whole cell lysate and the supernatant. The gel result is shown in Figure 2 Figure 1, and the data results are shown in Table 4 below. The sequence Opt is highly overexpressed and completely soluble, with the whole cell lysate protein content increased by 15.55 times relative to the sequence Wild-type and 5.76 times relative to the sequence GenScript; the supernatant soluble protein content is increased by 20.52 times relative to the sequence Wild-type and 6.78 times relative to the sequence GenScript.
[0049] Table 4
[0050]
[0051]
[0052] In summary, the SCG model constructed by the Transformer with multi-head attention mechanism and the Encoder-Decoder architecture can design and generate gene sequences conforming to the codon usage pattern of the chassis microorganism from the protein amino acid sequence in batches, and then introduce a gene solubility expression level information dataset with a wide range of protein sources. The sequence features extracted by extensive biological characteristics and large language model transfer learning are used to train a classic machine learning model, and the mapping relationship between the gene sequence and the solubility expression level is established from multiple angles, so as to realize the screening of the gene sequence with the best solubility expression level in the target host from the massive reverse translation DNA sequence. Taking Escherichia coli as the chassis microorganism, based on the above deep learning strategy, the amino acid sequence of alpha-glucan phosphorylase (alphaGP) is predicted to have a gene sequence with a high solubility expression level in Escherichia coli, and related protein expression experiments are carried out in Escherichia coli. Compared with the wild type original sequence and the sequence optimized by the Kings River company, the SDS-PAGE gel results show that the gene sequence provided by our deep learning algorithm has a higher solubility expression level than the wild type original sequence and the sequence optimized by the Kings River, and the supernatant protein content is 20.52 times that of the wild type sequence and 6.78 times that of the sequence optimized by the Kings River.
[0053] The above-described embodiments are merely preferred embodiments of the present application and are not intended to limit the scope of the present application. Various modifications and improvements to the technical solutions of the present application made by those of ordinary skill in the art without departing from the design spirit of the present application shall fall within the scope of protection of the present application as defined by the claims.< / eos> < / bos> < / pad> < / unk>
Claims
1. A gene sequence optimization and screening method based on deep learning, characterized in that: The specific steps are as follows: (1) Set the original amino acid sequence or nucleotide sequence of the recombinant protein to be optimized; (2) The operator sets the synonymous codon generation model (SCG) decoding parameters to determine the number of gene sequences generated by synonymous codons; (3) The operator selects a machine learning algorithm to evaluate the expression level of the gene sequence. 2.The deep learning based genetic sequence optimization and screening method according to claim 1, wherein: The specific operation method of step (1) is: set the sequence parameter to be optimized as the original amino acid sequence or nucleotide sequence of the recombinant protein. The model will check the sequence type of the input content before accepting the sequence parameter to be optimized. If it is an amino acid sequence, the model will add a termination symbol "*" at the end of the sequence. If it is a nucleotide sequence, the model will convert the nucleotide sequence into an amino acid sequence according to the standard codon table. 3.The deep learning based genetic sequence optimization and screening method according to claim 1, wherein: The specific operation method of step (2) is: set the synonymous codon generation model decoding parameters to execute the beam search decoding strategy. By setting multiple sets of model decoding parameters, the total number of gene sequences generated is the sum of the beam sizes of the multiple parameters. 4.The deep learning based genetic sequence optimization and screening method according to claim 3, wherein: The model decoding parameters are batch size, beam search width, and beam search result size.
5. The deep learning-based genetic sequence optimization and screening method according to claim 1 or 3 or 4, characterized in that: In step (2), the SCG model receives the amino acid sequence and decoding parameters, extracts the amino acid sequence information, combines the host's codon usage pattern information learned by the model, and iteratively generates the next nucleotide through beam search and pruning operations, until all gene sequences are generated. After removing duplicates, output to a table file. 6.The deep learning based genetic sequence optimization and screening method according to claim 1, wherein: The specific operation method of step (3) is: provide four machine learning algorithms: support vector machine, logistic regression, deep perception network, and light gradient boosting machine. Select several or all of these algorithms to evaluate the gene sequence. Through algorithm evaluation of each gene sequence, a probability score of the solubility expression level ranging from 0 to 1 is obtained. The closer to 1, the more likely it is that the gene sequence has a high solubility expression level. At the same time, the average probability score of each gene is calculated, and the highest comprehensive ranking is marked as the final codon optimization result. 7.The deep learning-based genetic sequence optimization and screening method according to claim 1 or 6, wherein: In step (3), based on the SCG model, a gene sequence file is generated. Feature engineering uses 18 feature descriptors to construct a 1024-dimensional sequence feature vector for each gene sequence, including codon composition and various physical and chemical properties. At the same time, a 768-dimensional sequence feature extracted by a transfer learning language model is used, which contains high-dimensional gene sequence information. Machine learning algorithms based on the above 1792-dimensional feature vectors input the probability of the gene sequence belonging to a high expression level, and save it to a file; calculate the mean probability of each gene sequence belonging to a high expression level sequence, save it to a file, and the sequence with the highest mean probability is the final optimization result. 8.The deep learning based genetic sequence optimization and screening method of claim 1, wherein: Using Escherichia coli as the chassis microorganism, the recombinant protein is alpha-glucan phosphorylase.