Probiotic screening method based on convolutional neural network
Through the method based on convolutional neural network, the problems of cumbersome operation, high computational complexity and limited classification accuracy of traditional probiotic screening methods are solved, and efficient and accurate probiotic screening is achieved to adapt to different data environments and strain characteristics.
Patent Information
- Application Number
- CN202510657232.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-29
AI Technical Summary
The existing probiotic screening methods rely on traditional biological experiments and sequence comparisons, and have problems such as cumbersome operation, high cost, high computational complexity, and insufficient processing capabilities for complex data. Especially in the era of high-throughput sequencing, it is difficult to effectively identify unknown strains, and traditional machine learning methods rely on artificial features to lead to limited classification accuracy.
The method based on convolutional neural network is adopted to process DNA sequences through regular expressions, and k-mer fragments are generated using sliding window method. Combined with multi-phase and multi-scale feature extraction, k-mer frequency vectors are constructed, and classified through convolutional neural networks. Combined with multi-task learning and attention mechanisms, the model structure is optimized to improve screening efficiency and accuracy.
It significantly improves the efficiency and accuracy of probiotic screening, reduces computing resource consumption, enhances the processing ability of complex data, improves the adaptability to unknown strains and generalization capabilities of models, and reduces the dependence on reference databases.
Smart Images

Figure CN120564844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biological probiotic screening, and specifically relates to a probiotic screening method based on convolutional neural network. Background Art
[0002] In the current field of probiotic screening technology, it mainly relies on traditional biological experimental methods and computational methods based on sequence alignment. Although traditional biological experimental methods have a certain degree of reliability, they are cumbersome to operate, long in cycle and expensive, making it difficult to meet the needs of large-scale, high-throughput screening. Computational methods based on sequence alignment, such as BLAST, achieve classification by comparing the target sequence with known sequences in the database. However, with the continuous deepening of microbial research and the explosive growth of data volume, this type of method has exposed many problems. On the one hand, it is heavily dependent on the quality and completeness of the reference database. When the database lacks unknown probiotic strains, especially new or mutant strains, the accuracy of the comparison results will drop significantly, and these unknown strains cannot be effectively identified. On the other hand, as the scale of data continues to expand, the computational complexity increases sharply. In the era of high-throughput sequencing, processing large amounts of genomic data will consume a lot of computing resources and the cost is extremely high.
[0003] Furthermore, traditional machine learning methods, such as support vector machines (SVMs) and random forests, have been widely used in probiotic genome classification tasks. However, these methods have significant limitations. They rely heavily on manually extracted features, such as GC content and sequence length, and are unable to automatically learn deep-level patterns from sequences. Especially when dealing with complex strain sequences, manually extracted features often fail to fully capture key information in the sequence, resulting in limited classification accuracy.
[0004] In summary, existing probiotic screening methods have shortcomings in accuracy, efficiency, and the ability to process complex data. An innovative method is urgently needed to solve these problems in order to promote the development of probiotic screening technology. Summary of the Invention
[0005] In order to effectively overcome the drawbacks of traditional sequence alignment methods that rely on the integrity of the reference database and have high computational complexity, and at the same time break through the limitations of traditional machine learning methods in artificial feature engineering, the present invention proposes a probiotic screening method based on a convolutional neural network, which includes: reading the original probiotic-related DNA sequence file, matching the DNA sequence file using a regular expression, and removing all non-ATCG base characters to obtain a normalized DNA sequence; screening the normalized DNA sequence; segmenting the screened DNA sequence using a sliding window method to generate a set of k-mer fragments; constructing a k-mer frequency vector based on the frequency of each k-mer fragment in the sequence; normalizing the k-mer frequency vector; inputting the normalized k-mer frequency vector into the trained optimized convolutional neural network to obtain a DNA sequence classification result; and determining whether the strain corresponding to the sequence is a probiotic based on the DNA sequence classification result.
[0006] Beneficial effects of the present invention:
[0007] 1) This paper uses a sliding window to extract k-mer frequency features, combines multi-scale and multi-phase processing, and simultaneously calculates the contribution value of each dimension to select some dimensions that are more critical to classification to construct vectors, further reducing the computational complexity of traditional sequence alignment methods from O(n 2 ) is reduced to O(n), significantly reducing computing resource consumption. When processing large-scale genomic data generated by high-throughput sequencing, screening tasks can be completed more quickly, effectively improving the efficiency of probiotic screening.
[0008] 2) This method, based on a deep learning architecture based on convolutional neural networks, combines multi-scale k-mer feature extraction, positional encoding, and an attention mechanism, resulting in powerful automatic learning capabilities. By calculating dimensional contribution values to screen key dimensions, it can precisely focus on the distribution patterns of k-mers and deeply explore key features in local and global sequences. Compared to manually designed features or unfiltered dimensions, this method reduces subjectivity and information loss, significantly improving the classification accuracy of probiotic screening.
[0009] 3) This invention uses maximum-minimum normalization to reduce the impact of sequence length differences on the model. Combining residual connections and a dropout mechanism (which can be added appropriately during training) effectively suppresses overfitting caused by high-dimensional sparse data. Furthermore, because the selected dimensions are more representative, the application of improved focal loss and contrast loss can be more effective, further enhancing the model's generalization and stability, and ensuring the accuracy of probiotic screening in different data environments.
[0010] 4) This invention supports dynamic adjustment of the k-mer window length (3≤k≤6) and can flexibly optimize the network structure according to actual needs, such as increasing the depth of the convolutional layer and adjusting the parameters of the attention module. When calculating the dimension contribution value, the evaluation method and threshold can also be adjusted according to different situations. This flexible and customized screening strategy can more efficiently screen for different bacterial species characteristics or genetic variations of mutant strains, reduce dependence on reference databases, and improve the predictive adaptability of unknown sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is the overall flow chart of the present invention;
[0012] Figure 2 This is the convolutional network structure diagram of the present invention. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0014] A probiotic screening method based on convolutional neural network, such as Figure 1 and Figure 2 As shown, the method includes: reading an original probiotic-related DNA sequence file, matching the DNA sequence file using a regular expression, and removing all non-ATCG base characters to obtain a normalized DNA sequence; screening the normalized DNA sequence; segmenting the screened DNA sequence using a sliding window method to generate a set of k-mer fragments; constructing a k-mer frequency vector based on the frequency of each k-mer fragment in the sequence; normalizing the k-mer frequency vector; inputting the normalized k-mer frequency vector into a trained optimized convolutional neural network to obtain a DNA sequence classification result; and determining whether the strain corresponding to the sequence is a probiotic based on the DNA sequence classification result.
[0015] In this example, before performing probiotic screening, the original probiotic-related DNA sequence files were first read. Specifically, a certain amount of DNA sequence data was obtained from a public dataset, such as 237 complete probiotic gene sequences and 470 non-probiotic gene sequences. This data was then divided into a training set and a test set at a ratio of 9:1.
[0016] After acquiring the data, special processing is performed. The raw DNA sequence from each FASTA-formatted file in the dataset is first read. A regular expression (e.g., [^ATCG]) is used to precisely match and remove all non-ATCG base characters, converting the raw sequence into a normalized DNA sequence. For example, the sequence ATGNRCT becomes ATGCT after cleaning. The normalized DNA sequence is then divided into three subsequences (phase 0, phase 1, and phase 2) based on the reading frame phase, using the start codon (ATG) as a reference. If the sequence lacks a clear start codon, a sliding window method is used to traverse the entire sequence, generating three phased subsequences with a step size of 1.
[0017] Next, each phase subsequence was subjected to k-mer dynamic optimization segmentation, using windows of k = 3, k = 4, and k = 5. The frequencies of 3-mers, 4-mers, and 5-mers in each phase were counted to construct a multi-phase, multi-scale feature vector, with a higher weight assigned to the phase 0 subsequence. To reduce the dimensionality of the feature vector and retain key information, statistical methods such as information gain and chi-square tests were used to evaluate the correlation between each k-mer dimension and the target probiotic classification. Taking information gain as an example, its formula is:
[0018]
[0019] Where S is the dataset, A is a feature (i.e., a k-mer dimension), is the entropy of the dataset S, and is the subset when feature A takes the value v. Based on the calculated contribution value, an appropriate threshold is set, and only k-mer dimensions with contribution values above the threshold are retained to construct a new feature vector.
[0020] Using the functional region database pre-trained by the hidden Markov model (HMM), the normalized DNA sequence was scanned to locate potential functional regions (such as the pks and nrps gene cluster regions), and the k-mer frequency in the functional region was given a 2-fold weight, while the non-functional region retained the original frequency. Similarly, the contribution value of the k-mer dimension related to the functional region and the target probiotic classification was calculated, and an appropriate threshold was set to screen out the dimensions with high contribution values. These screened dimensions were integrated with the dimensions screened out in the previous phase-specific k-mer feature extraction to further optimize the feature vector.
[0021] In this embodiment, screening the normalized DNA sequence includes: dividing the normalized DNA sequence into three phase subsequences; performing k-mer dynamic optimization segmentation on each phase subsequence; counting the frequencies of each segmented subsequence to construct a multi-phase multi-scale feature vector; calculating the contribution value of each k-mer dimension in the multi-phase multi-scale feature vector, and evaluating the correlation between each k-mer dimension and the target probiotic classification; setting a first threshold, comparing the calculated contribution value with the set threshold, and constructing a new feature vector for the k-mer dimension with a contribution value higher than the threshold; using a functional region database pre-trained with a hidden Markov model to scan the normalized DNA sequence and locate the functional region of the normalized DNA sequence; assigning a weighted coefficient to the k-mer frequency within the functional region, maintaining the original frequency of the non-functional region, calculating the contribution value of the k-mer dimension related to the functional region and the target probiotic classification, setting a second threshold, and screening out dimensions with high contribution values based on the second threshold; integrating the screened dimensions with the new feature vector to obtain a screened DNA sequence.
[0022] The above-mentioned screening approach based on multi-phase k-mer feature extraction, contribution value screening, and functional region weighting effectively reduces redundant features and enhances the expression of key features. This not only improves the model's ability to focus on probiotic sequence features, but also reduces the input feature dimensionality while maintaining classification accuracy, thereby improving model training and inference efficiency. Furthermore, the HMM-based functional region database scanning approach highlights features in regions related to probiotic function, enhancing model interpretability while improving adaptability and generalization to complex or unknown strains, ultimately improving the accuracy, robustness, and computational efficiency of the probiotic classification model. This data processing pipeline is unconventional in the field. Traditional machine learning methods for probiotic classification rely heavily on sequence alignment or fixed artificial features and do not incorporate phase information and k-mer contribution value screening strategies. Furthermore, the innovative integration of hidden Markov models for functional region localization combined with feature weighting is a rare practice in traditional techniques that utilizes structural annotation information for deep learning feature screening. Therefore, this pipeline, through statistical evaluation and a feature enhancement strategy guided by biological functional knowledge, achieves a more in-depth and accurate modeling of probiotic sequence features, demonstrating significant innovation and non-obviousness.
[0023] In this embodiment, a sliding window method is used to segment the normalized DNA sequence after special processing (including dimensional screening) to generate a set of k-mer segments. The frequency of each k-mer segment in the dimension determined after screening is counted to construct a k-mer frequency vector. Then, a maximum-minimum normalization method is used to map the values of each dimension of the vector to the interval [0,1] to obtain a standardized feature vector.
[0024] A convolutional neural network (CNN) was designed and trained based on preprocessed data for probiotic screening and classification. The model uses a fully one-dimensional convolutional architecture and takes a normalized, dimensionally filtered DNA sequence vector as input. Sinusoidal positional encoding is added after the input layer to preserve sequence positional information. The network architecture is as follows: the first layer has 32 1D convolution kernels of size 3; the second layer has 64 1D convolution kernels of size 5; and the third layer has 128 1D convolution kernels of size 7. Residual connections are introduced to enhance feature transfer. A channel-spatial dual attention module is introduced between convolutional layers to improve the model's focus on key features. The network effectively extracts and classifies probiotic features through the collaborative work of multiple convolutional layers, adaptively selected pooling layers, and fully connected layers. During training, a modified Focal Loss loss function and the Adam algorithm (adjustable parameters) are used as the optimizer. Training epochs are set to 200 to ensure full model convergence. At the same time, multi-task branch auxiliary training was enabled, with training conducted in a 7:2:1 ratio between the main task and the auxiliary task, guiding the model to learn richer feature representations. This architecture, combining attention mechanisms with multi-task learning, significantly improved the classification accuracy of probiotic screening.
[0025] In this embodiment, multi-task learning includes two auxiliary tasks, functional region prediction and phase distribution classification, in addition to the primary classification task. The output layer of the functional region prediction task adds one neuron to predict whether the input sequence contains a functional region characteristic of probiotics; the output layer of the phase distribution classification task adds three neurons to predict the dominant phase in the sequence. By setting the weight ratio of the primary task to the auxiliary tasks at 7:2:1, the model is guided to learn a richer set of features. This multi-task learning approach enables the model to analyze and learn from sequences from multiple perspectives, thereby improving its overall performance.
[0026] The trained model is evaluated using the test set, with metrics such as accuracy, recall, and F1 value calculated to assess its performance. In practice, after the aforementioned data preprocessing (including dimension screening) and feature extraction steps are performed on the DNA sequence of the test strain, it is fed into the trained model. The model outputs a probability score for whether the sequence belongs to probiotics or non-probiotics. Based on a pre-set score threshold, the target strain's category can be accurately determined, enabling efficient screening of probiotics.
[0027] In this example, normalizing the k-mer frequency vector involves counting the absolute frequencies of 256 4-mer segments in the sequence and constructing a 256-dimensional frequency vector based on the counts. To eliminate the impact of varying sequence lengths on the analysis results, a Min-Max Normalization method is used to map the values of each dimension of the vector to the interval [0, 1]. The calculation formula is:
[0028]
[0029] where x max 、x min are the maximum and minimum 4-mer frequencies in the vector, respectively.
[0030] In this embodiment, a channel-space dual attention module is introduced between the convolutional layers. Among them, the channel attention (SE module) generates channel weights by performing global average pooling and full connection operations on the output of the convolutional layer, thereby adaptively adjusting the importance of different channel features. The spatial attention (SA module) performs spatial dimension attention weighting on the feature map, focusing on the key sequence area of probiotics. Channel attention and spatial attention are connected in series to form a dual attention mechanism. This mechanism can significantly enhance the model's ability to focus on high-value features, allowing the model to more accurately focus on key information related to probiotic screening among many features, thereby improving the accuracy of screening.
[0031] When training convolutional neural networks, an improved FocalLoss is used to address the class imbalance problem. By introducing adaptive weights for sample difficulty, the weights are dynamically adjusted according to the predicted probability of the samples during training. The final loss function is:
[0032]
[0033] Among them, γ is the focusing parameter, which is set to 2; n is the total amount of data, α i is the category weight coefficient, p i is the probability of belonging to the positive class. This increases the penalty for difficult samples. This allows the model to focus more on minority samples, improves the classification ability of imbalanced data, and thus improves the accuracy of probiotic screening.
[0034] In this embodiment, a contrast loss branch is added after the fully connected layer to measure the difference between samples by calculating the cosine similarity of the feature vectors of paired samples. The contrast loss and the main classification loss are weighted in a ratio of 2:8, and the total loss is L total =0.8L focal +0.2L contrastThis approach can effectively improve the separability of the feature space, making the features learned by the model more discriminative and further optimizing the performance of the model.
[0035] The method of this invention forms a complete logical closed loop from data processing, feature extraction, model design, to loss optimization. In particular, the method of calculating dimensional contribution values during the data processing phase to screen key dimensions, synergistically working with other steps, further effectively improves the accuracy and efficiency of probiotic screening, providing stronger support for probiotic research and application.
[0036] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A probiotic screening method based on convolutional neural network, characterized in that: include: The original probiotic-related DNA sequence file is read, and the DNA sequence file is matched using regular expressions. All non-ATCG base characters are removed to obtain a normalized DNA sequence; the normalized DNA sequence is screened; the screened DNA sequence is segmented using the sliding window method to generate a set of k-mer fragments; a k-mer frequency vector is constructed based on the frequency of occurrence of each k-mer fragment in the sequence; the k-mer frequency vector is normalized; the normalized k-mer frequency vector is input into the trained optimized convolutional neural network to obtain the DNA sequence classification result; based on the DNA sequence classification result, it is determined whether the strain corresponding to the sequence is a probiotic.
2. A probiotic screening method based on convolutional neural network according to claim 1, characterized in that: Removing all non-ATCG base characters includes: constructing a regular expression [^ATCG], matching the regular expression with the DNA sequence, and if the match is successful, retaining the successfully matched DNA fragment; if the match is unsuccessful, removing the unsuccessfully matched DNA fragment.
3. A probiotic screening method based on convolutional neural network according to claim 1, characterized in that: The screening of normalized DNA sequences includes: dividing the normalized DNA sequence into three phase subsequences; performing k-mer dynamic optimization segmentation on each phase subsequence; counting the frequencies of each segmented subsequence to construct a multi-phase multi-scale feature vector; calculating the contribution value of each k-mer dimension in the multi-phase multi-scale feature vector, and evaluating the correlation between each k-mer dimension and the target probiotic classification; setting a first threshold, comparing the calculated contribution value with the set threshold, and constructing a new feature vector for the k-mer dimension with a contribution value higher than the threshold; using a functional region database pre-trained with a hidden Markov model to scan the normalized DNA sequence and locate the functional region of the normalized DNA sequence; assigning a weighted coefficient to the k-mer frequency in the functional region, keeping the original frequency in the non-functional region, calculating the contribution value of the k-mer dimension related to the functional region and the target probiotic classification, setting a second threshold, and screening out dimensions with high contribution values according to the second threshold; integrating the screened dimensions with the new feature vector to obtain the screened DNA sequence.
4. A probiotic screening method based on convolutional neural network according to claim 3, characterized in that: The k-mer frequencies within the functional regions were given a weighting factor of 2.
5. The probiotic screening method based on convolutional neural network according to claim 3, characterized in that: The calculation formula for dimension contribution value is: Among them, S is the data set, A is the feature, H(S) is the entropy of the data set S, S v It is the subset when feature A takes the value v.
6. A probiotic screening method based on convolutional neural network according to claim 1, characterized in that: The sliding window length in the sliding window method is set to k=4 and the step size is s=1.
7. The probiotic screening method based on convolutional neural network according to claim 1, characterized in that: The optimized convolutional neural network includes the first convolutional layer, the second convolutional layer, the third convolutional layer, the channel-space dual attention module, the fully connected layer and the output layer, and each layer is connected in sequence; the first convolutional layer is configured with 32 1D convolution kernels with a kernel size of 3; the second convolutional layer is configured with 64 1D convolution kernels with a kernel size of 5; the third convolutional layer is equipped with 128 1D convolution kernels with a kernel size of 7, and is combined with residual connections; the fully connected layer contains 128 neurons and uses ReLU activation; the output layer has two neurons and uses the Softmax function for classification.
8. The probiotic screening method based on convolutional neural network according to claim 7, characterized in that: Training the optimized convolutional neural network includes: adding sinusoidal position encoding to the input feature vector and inputting it into the network structure, extracting local and global sequence features through multiple layers of one-dimensional convolutional layers in sequence, and introducing channel attention and spatial attention modules between the convolutional layers to enhance the model's perception of key information; in the training phase, a multi-task learning strategy is adopted to simultaneously perform the main classification task, functional area prediction task and phase distribution classification task, and guide the model to extract richer sequence representations through weight ratios; in the optimization process, the Adam optimizer is selected, combined with the improved Focal Loss to deal with the category imbalance problem, and contrast loss is introduced to improve feature distinguishability.
9. A probiotic screening method based on convolutional neural network according to claim 8, characterized in that: The loss function of the model is: Among them, γ is the focusing parameter, n is the total amount of data, α i is the category weight coefficient, p i is the probability of belonging to the positive class.
10. The probiotic screening method based on convolutional neural network according to claim 6, characterized in that: The initial learning list of the Adam optimizer is set to 1×10-3 and the training round epoch=200.
Citation Information
Patent Citations
Gene sequence digital realizing method based on information entropy and system thereof
CN109903812A
Phage host prediction method, apparatus and device, and storage medium
CN113658633A
Method for accurately identifying soil pathogenic bacterium pollution by using sequence k-mer frequency optimization characteristics
CN114842908A
Method for accurately identifying beneficial bacteria according to whole genome sequence characteristics
CN117668693A
Screening method and system for synthetic biological probiotics
CN119252334A
Cited By
Synthetic probiotic phenotypic characteristic prediction method and system
CN121171363A
Probiotic tolerance characteristic prediction method and device, electronic equipment and storage medium
CN122117042A
Probiotic tolerance property prediction method, device, electronic equipment and storage medium
CN122117042B