Microbial strain DNA mixed coding and alternative Transformer depth prediction method for gene expression of microbial strain DNA mixed coding
By combining the alternating Transformer deep learning framework with CNN and multi-head attention mechanism, the problems of weak DNA sequence feature extraction and model generalization ability in microbial engineering are solved, achieving high-precision gene expression prediction and improving the stability and efficiency of prediction.
Patent Information
- Application Number
- CN202510957567.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies struggle to efficiently capture complex local semantic patterns and global long-range dependencies in DNA sequences during microbial engineering. They also suffer from insufficient fusion of multi-source information and weak model generalization ability, resulting in low accuracy in gene expression prediction and reliance on costly trial-and-error experiments.
We employ the alternating Transformer deep learning framework, combining CNN and multi-head attention mechanisms to integrate local semantic perception with global contextual information. Through multi-layer nonlinear transformation and regularization mechanisms, we construct an end-to-end deep learning network model to extract high-dimensional semantic information and global contextual dependencies from long DNA sequences.
It significantly improves the accuracy and stability of gene expression prediction, solves the problems of difficult feature extraction of long DNA sequences and weak model generalization ability, achieves high-precision gene expression prediction, and reduces experimental costs and cycle time.
Smart Images

Figure CN120808868A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of bioinformatics and artificial intelligence, and particularly relates to a microorganism strain DNA hybrid coding and an alternating Transformer deep prediction method for gene expression thereof. BACKGROUND
[0002] Microbial engineering plays a strategic role in the fields of biopharmaceuticals, industrial enzyme production, genetic engineering, and synthetic biology, and the core challenge lies in efficient expression of target genes. Traditional strain development processes require the introduction of exogenous genes into host cells to improve expression efficiency by optimizing genetic elements such as promoters, ribosome binding sites (RBS), etc. However, the existing technology has three bottlenecks: 1. Insufficient sequence feature extraction: Conventional machine learning models (such as linear regression, random forest, support vector machine) are difficult to capture complex local semantic patterns (such as codon bias, base combination effects) and global long-range dependencies (such as cross-region interactions formed by mRNA secondary structure) in DNA sequences; 2. Lack of multi-source information fusion: Existing methods rely on single sequence data or simple statistical features (such as GC content), and fail to integrate multi-dimensional biophysical features such as nucleotide composition, codon adaptation index (CAI), codon bottleneck position / intensity, etc., resulting in insufficient prediction information dimension; 3. Weak model generalization ability: Prediction tools based on heuristic rules or shallow networks (such as Codon Optimization Algorithms) have significant errors (MAE>30%) in complex expression systems, forcing research and development to rely on high-cost trial and error experiments, with a single strain construction-validation cycle lasting several weeks.
[0003] With the development of artificial intelligence, deep learning-driven gene expression prediction is becoming a new cornerstone of synthetic biology, and its value lies not only in the improvement of prediction accuracy, but also in the reconstruction of the "design-build-test-learn" closed loop.
[0004] However, traditional deep learning networks (such as CNN) models can extract local base patterns, but are limited by fixed receptive fields and cannot model global dependencies of long sequences (such as the synergistic effect of distant promoters and translation initiation sites); single-layer Transformers can capture global correlations, but they weaken key local features (such as RBS core regions) by focusing too much on irrelevant positions, and are prone to overfitting with limited samples. SUMMARY
[0005] The application aims to provide an alternating Transformer deep prediction method for DNA mixed coding of microbial strain and gene expression. Compared with traditional artificial experimental methods, the application has significant improvement in stability and repeatability, and is not affected by operation environment, reagent quality, personnel skills and other factors. Compared with classic machine learning methods such as decision tree and support vector machine, the alternating Transformer deep learning framework proposed by the application solves the problems of long DNA sequence feature extraction difficulty and weak model generalization ability through multi-layer nonlinear transformation and regularization mechanism.
[0006] To solve the above technical problems, the application adopts the following technical solutions: S1: Microbial strain DNA sequence data loading preprocessing, including missing value processing, 4 base binary coding, biological characteristic data normalization coding, and splicing the coded data into a 12-row n-column combined feature matrix; S2: Read the combined feature matrix of step S1, build an end-to-end deep learning network model, design an alternating layer structure of CNN-Transformer-CNN-Transformer-CNN-MLP, efficiently extract DNA long sequence high-dimensional semantic information and global context dependency relationship, and finally output the gene expression level; S3: Use 2 times 5-fold cross-validation to evaluate the model, and measure the prediction performance of the model based on the determination coefficient (R 2 ), mean absolute error (MAE), and mean square error (MSE).
[0007] Further, a preferred implementation method is provided, and the step S1 includes: S1.1: Preprocess the data containing missing values. If the DNA sequence is missing, remove the sequence; if the DNA sequence is complete and only some biological feature statistics are missing, use the nearest neighbor algorithm to estimate the expected value of the biological feature from the DNA sequence with similar structure to fill in.
[0008] S1.2: Standardize the DNA sequence of the target microbial strain, and cut the DNA sequence length to n; identify and replace irregular bases, and map them to a zero vector [0, 0, 0, 0] to ensure matrix integrity.
[0009] S1.3: Map the 4 regular bases of the DNA sequence of length n using 0-1 coding, map base A to [1, 0, 0, 0], map base T to [0, 1, 0, 0], map base C to [0, 0, 1, 0], and map base G to [0, 0, 0, 1], and construct a 4-row n-column binary matrix.
[0010] S1.4: Normalize the 8 biological characteristics of the DNA sequence, i.e., nucleotide composition, codon adaptation index, codon slope bottleneck position, codon slope bottleneck strength, polypeptide hydrophobicity, and mRNA secondary structure stability (5' end region, middle coding region, 3' end region), and construct an 8-row n-column biological feature matrix, with the same normalized value for each row of the matrix. The normalized value ranges from 0 to 1, and its calculation expression is: where x is the original value of the DNA sequence biological feature, x' is the normalized value, is the minimum value of the biological feature, is the maximum value of the biological feature.
[0011] S1.5: Concatenate the two encoding matrices constructed by S1.3 and S1.4 to form a 12-row n-column combined feature matrix for input into the deep learning model. The first 4 rows of the matrix are the binary encoding of the DNA sequence bases, and the last 8 rows are the normalized encoding of the sequence biological features, with the matrix element value ranging from 0 to 1.
[0012] The step S2 is specifically: S2.1: Construct a deep network architecture that combines CNN and alternating Transformer encoders, with an alternating stacking structure of CNN-Transformer-CNN-Transformer-CNN-Multilayer Perceptron (MLP). By alternating CNN layers and Transformer layers, effective fusion of local feature extraction and global context modeling across multiple levels of features is achieved. CNN layers perform basic to abstract feature level construction and provide spatial invariance, while Transformer layers model long-range global dependencies on these feature levels. The final MLP layer projects all the features extracted and fused by the previous layers to the output space.
[0013] S2.2: In each Transformer layer, a multi-head attention mechanism is used to project each element of the CNN feature map to multiple subspaces and assign different attention weights. This mechanism calculates the weights by modeling the interdependence between elements in the feature map, thereby capturing the spatial dependence of DNA base sequences and their features.
[0014] The multi-head attention mechanism is implemented through the following steps: First, the DNA combination feature matrix is extracted through CNN feature to obtain the feature map , after linear transformation, three matrices of query (Q), key (K) and value (V) are generated. The calculation expression is: ; ; ; in 、 、 is a learnable weight matrix.
[0015] Then, the generated Q, K, V matrices are divided into h heads, and each head calculates the attention score independently, and its calculation expression is: in 、 、 It is The query, key and value matrices of each head, is the dimension of the key vector, used to scale the dot product to avoid vanishing gradients.
[0016] The output of multi-head attention is the concatenation and linear transformation of the attention results of each head, and its calculation expression is: in , is the weight matrix of the output linear transformation.
[0017] The step S3 is specifically as follows: S3.1: Construct a model based on the coefficient of determination (R 2 ), mean absolute error (MAE), and mean squared error (MSE) are used to measure the multi-dimensional prediction performance of deep learning models. R² measures the goodness of fit of deep learning models for regression tasks, MAE measures the average of the absolute differences between predicted and true gene expression values, and MSE measures the average of the squared differences between predicted and true gene expression values.
[0018] S3.2: Use 2-fold 5-fold cross-validation, which is divided into three steps: 1. First 5-fold cross-validation: Perform a first random permutation operation on the original dataset D. Divide the permuted dataset into five mutually exclusive subsets, each containing approximately equal number of data samples. Perform the following operations five times in turn: select one of the subsets as the test set, and the remaining four subsets as the training set; train the model using the training set and evaluate the model performance using the test set, generating the corresponding performance score (denoted as S1_1, S1_2, S1_3, S1_4, S1_5). Based on the performance scores S1_1 to S1_5 generated by the five evaluations, calculate the average performance score μ1 and its standard deviation σ1 of the first round of 5-fold cross-validation.
[0019] 2. Second 5-fold cross-validation: Perform a second random permutation operation on the original dataset D, where the second permutation operation uses a different random seed than the first permutation operation to ensure that the subsets obtained by the two divisions are mutually exclusive. Divide the second permuted dataset into five new mutually exclusive subsets, each containing approximately equal number of data samples. Perform the following operations five times in turn: select one of the subsets as the validation / test set, and the remaining four subsets as the training set; train the model using the training set and evaluate the model performance using the validation / test set, generating the corresponding performance score (denoted as S2_1, S2_2, S2_3, S2_4, S2_5). Based on the performance scores S2_1 to S2_5 generated by the five evaluations, calculate the average performance score μ2 and its standard deviation σ2 of the second round of 5-fold cross-validation.
[0020] 3. Performance summary: By performing steps 1 and 2, a total of ten independent model training and evaluation are completed using ten mutually exclusive data subset divisions, generating ten performance scores (S1_1, S1_2, S1_3, S1_4, S1_5, S2_1, S2_2, S2_3, S2_4, S2_5). Summarize the ten performance scores to calculate their total average performance score μ_total and total standard deviation σ_total as the comprehensive performance evaluation indicator of the deep network.
[0021] The beneficial effects of the present application are: 1. The present application deeply integrates multi-layer CNN and Transformer-encoder architecture, which screens local features of DNA encoding matrix through convolution and pooling layers of CNN, and then establishes long-distance dependency between bases and cross-modal association between DNA sequence and other biological features using multi-head attention mechanism of Transformer-encoder, combined with residual connection and layer normalization design to ensure model stability. This collaborative mechanism breaks through the limitations of traditional models for long sequence modeling, and realizes efficient capture of key driving factors of gene expression.
[0022] 2. Compared to existing technologies, this invention successfully addresses three major bottlenecks in scenarios involving long, multi-base sequences and multi-feature inputs: first, it uses a multi-head attention mechanism to analyze complex positional relationships between bases and cross-sequence similarities; second, it establishes an interaction model between DNA sequences and auxiliary biological indicators; and third, it significantly improves the accuracy of gene expression prediction (the embodiment achieves a 15% improvement over the CNN baseline model). Especially for high-throughput genomic data, the DNA hybrid coding matrix and CNN pre-feature map effectively reduce the computational complexity of the Transformer, balancing prediction accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 Flowchart of the alternating Transformer deep prediction method for mixed DNA coding of microbial strains and their gene expression in an embodiment of the present invention.
[0025] Figure 2 Flowchart of the 2-fold 5-fold cross validation method in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] like Figure 1 As shown in the embodiment of the present application, an embodiment of an alternate Transformer deep prediction method for mixed DNA coding of microbial strains and their gene expression is shown, and the specific steps are as follows: S1: Loading and preprocessing of DNA sequence data of microbial strains, including missing value processing, binary encoding of four bases, normalization encoding of biological characteristic data, and splicing the encoded data into a 12-row n-column combined feature matrix; S2: read the combined feature matrix of step S1, build a deep learning network model based on alternating layer Transformer, design the alternating layer structure of CNN-Transformer-CNN-Transformer-CNN-MLP, efficiently extract high-dimensional semantic information and global context dependence of DNA sequence, and finally output the gene expression level; S3: evaluate the model by adopting 2 times 5-fold cross-validation, and measure the prediction performance of the model based on the coefficient of determination (R 2 ), mean absolute error (MAE), and mean square error (MSE).
[0028] Further, a preferred implementation method is provided, and the step S1 comprises: S1.1: pre-process the data containing missing values. If the DNA sequence is missing, delete the sequence; if the DNA sequence is complete and only some biological feature statistics are missing, use the nearest neighbor algorithm to estimate the expected value of the biological feature from the DNA sequence with similar structure to fill in the missing values.
[0029] S1.2: standardize the DNA sequence of the target microorganism strain, and cut the DNA sequence length to n; identify and replace the irregular bases, and map them to the all-zero vector [0, 0, 0, 0] to ensure the integrity of the matrix.
[0030] S1.3: encode each base of the screened DNA sequence using 0-1 encoding, map base A to [1, 0, 0, 0], base T to [0, 1, 0, 0], base C to [0, 0, 1, 0], and base G to [0, 0, 0, 1], and construct a 4-row n-column binary matrix.
[0031] S1.4: normalize the 8 biological features of the DNA sequence, i.e. nucleotide composition, codon adaptation index, codon slope bottleneck position, codon slope bottleneck strength, polypeptide hydrophobicity, mRNA secondary structure stability (5' end region, middle coding region, 3' end region), and construct an 8-row n-column biological feature matrix, each row of the matrix takes the same normalized value. The normalized value range is 0 to 1, and the calculation expression is: where x is the original value of the DNA sequence biological feature, x' is the normalized value, is the minimum value of the biological feature, is the maximum value of the biological feature.
[0032] S1.5: Concatenate the two encoding matrices constructed in S1.3 and S1.4 to form a 12-row, n-column combined feature matrix for input to the deep learning model. The first four rows of this matrix contain the binary encoding of the DNA sequence, and the last eight rows contain the normalized encoding of the sequence's biometric features. The values of the matrix elements range from 0 to 1.
[0033] The step S2 is specifically as follows: S2.1: Construct a deep network architecture that deeply fuses CNNs and alternating Transformer encoders, employing an alternating stacking structure of CNN-Transformer-CNN-Transformer-CNN-Multi-Layer Perceptron (MLP). By alternating CNN and Transformer layers, we effectively combine local feature extraction and global context modeling across multiple feature levels. CNN layers construct feature hierarchies from base to abstract and provide spatial invariance, while Transformer layers model long-range global dependencies across these feature hierarchies. The final MLP layer projects the features extracted and fused by all the previous layers into the output space.
[0034] S2.2: In each Transformer layer, a multi-head attention mechanism is used to project each element of the CNN feature map into multiple subspaces and assign different attention weights. This mechanism calculates weights by modeling the interdependencies between elements in the feature map, thereby capturing the spatial dependencies of DNA base sequences and their features.
[0035] The multi-head attention mechanism is implemented through the following steps: First, the DNA combination feature matrix is extracted through CNN feature to obtain the feature map , after linear transformation, three matrices of query (Q), key (K) and value (V) are generated. The calculation expression is: ; ; ; in 、 、 is a learnable weight matrix.
[0036] Then, the generated Q, K, V matrices are divided into h heads, and each head calculates the attention score independently, and its calculation expression is: in 、 、 It is The query, key and value matrices of each head, is the dimension of the key vector, used to scale the dot product to avoid gradient vanishing.
[0037] The output of multi-head attention is the concatenation and linear transformation of the results of each head of attention, whose computational expression is: wherein , is the weight matrix of the output linear transformation.
[0038] The step S3 is specifically: S3.1: Constructing a multi-dimensional prediction performance index of the deep learning model based on the determination coefficient (R 2 ), mean absolute error (MAE), and mean square error (MSE). R² measures the goodness of fit of the deep learning model for the regression task, MAE measures the average of the absolute values of the differences between the predicted values and the true values of gene expression, and MSE measures the average of the squares of the differences between the predicted values and the true values of gene expression.
[0039] S3.2: Using 2 times 5-fold cross-validation, which is specifically divided into three steps: 1. First 5-fold cross-validation: performing a first random permutation operation on the original data set D. The permuted data set is divided into five mutually exclusive subsets, each containing approximately the same number of data samples. The following operations are performed five times in turn: selecting one of the subsets as the test set, and the remaining four subsets as the training set; training the model using the training set and evaluating the model performance using the test set to generate the corresponding performance score (denoted as S1_1, S1_2, S1_3, S1_4, S1_5). Based on the performance scores S1_1 to S1_5 generated by the five evaluations, the average performance score μ1 of the first round of 5-fold cross-validation and its standard deviation σ1 are calculated.
[0040] 2. Second 5-fold cross-validation: performing a second random permutation operation on the original data set D, wherein the random seed used in the second permutation operation is different from that used in the first permutation operation to ensure that the subsets obtained by the two divisions are mutually exclusive. The second permuted data set is divided into five new mutually exclusive subsets, each containing approximately the same number of data samples. The following operations are performed five times in turn: selecting one of the subsets as the validation / test set, and the remaining four subsets as the training set; training the model using the training set and evaluating the model performance using the validation / test set to generate the corresponding performance score (denoted as S2_1, S2_2, S2_3, S2_4, S2_5). Based on the performance scores S2_1 to S2_5 generated by the five evaluations, the average performance score μ2 of the second round of 5-fold cross-validation and its standard deviation σ2 are calculated.
[0041] 3. Performance aggregation: By performing step 1 and step 2, ten independent model training and evaluation are completed with ten mutually exclusive data subset divisions, generating ten performance scores (S1_1, S1_2, S1_3, S1_4, S1_5, S2_1, S2_2, S2_3, S2_4, S2_5). The ten performance scores are aggregated to calculate the total average performance score μ_total and its total standard deviation σ_total as the comprehensive performance evaluation indicator of the deep network.
[0042] The embodiments of the present application are described in detail above with reference to the drawings, but the present application is not limited to the above-described embodiments, and for those skilled in the art, after learning the contents described in the present application, a number of equivalent transformations and substitutions can be made without departing from the principles of the present application, and these equivalent transformations and substitutions should also be considered to belong to the protection scope of the present application.
Claims
1. A method for deep prediction of mixed DNA coding of microbial strains and their gene expression using an alternating Transformer, characterized in that: The following steps are involved: S1: Loading and preprocessing of DNA sequence data of microbial strains, including missing value processing, four base binary encoding, normalized encoding of biological characteristic data, and splicing the encoded data into a 12-row n-column combined feature matrix; S2: Read the combined feature matrix from step S1, build an end-to-end deep learning network model, and design an alternating stacking structure of CNN-Transformer-CNN-Transformer-CNN-MLP to efficiently extract high-dimensional semantic information and global context dependencies of long DNA sequences, and finally output gene expression levels; S3: Cross-validate the deep learning model and measure the prediction performance of the model based on the coefficient of determination (R2), mean absolute error (MAE), and mean square error (MSE).
2. The method for deep prediction of mixed DNA coding of microbial strains and their gene expression using an alternating Transformer according to claim 1, characterized in that: Step S1 is specifically as follows: S1-1: Use 0-1 encoding to encode each base of a DNA sequence fragment of length n, mapping base A to [1, 0, 0, 0], base T to [0, 1, 0, 0], base C to [0, 0, 1, 0], and base G to [0, 0, 0,1], to construct a binary matrix with 4 rows and n columns; S1-2: Normalize the eight biological characteristics of the DNA sequence (nucleotide composition, codon adaptation index, codon ramp bottleneck position, codon ramp bottleneck strength, peptide hydrophilicity, and mRNA secondary structure stability) to construct an 8-row, n-column biological characteristic matrix, with each row of the matrix taking the same normalized value; S1-3: The two encoding matrices generated by S1-1 and S1-2 are cascaded along the row direction to form a combined feature matrix with a dimension of 12×n, which serves as the input of the deep learning model.
3. The method for deep prediction of mixed DNA coding of microbial strains and their gene expression using an alternating Transformer according to claim 1, characterized in that: Step S2 is specifically as follows: S2-1: Build a deep learning network that combines multiple convolutional layers (CNN) and alternating layers of Transformer encoders, and design an alternating layer structure of CNN-Transformer-CNN-Transformer-CNN -MLP; S2-2: In each Transformer layer, a multi-head attention mechanism is used to assign different attention weights to each element of the CNN feature map from multiple subspaces. The weight of each element is determined by calculating the correlation between each element and other elements in the feature map, that is, the spatial correlation of the DNA base sequence and its features is obtained.
4. The method for deep prediction of mixed DNA coding of microbial strains and their gene expression using an alternating Transformer according to claim 1, characterized in that: Step S3 is specifically as follows: S3-1: Construct a 2 ), mean absolute error (MAE), and mean square error (MSE) are used to measure the prediction performance indicators of deep learning models; S3-2: Two 5-fold cross-validation runs were performed. Before each run, the training data was randomly shuffled and divided into five mutually exclusive subsets. Each subset was used as the test set, and the remaining four subsets were combined into the training set. Final model performance was assessed by averaging the performance metrics from the two independent validation runs.