RNA sequence classification method based on Mangbar model and semi-supervised learning
Through the RNA sequence classification method of the Mamba model and semi-supervised learning, the problems of poor processing of long sequence data and high annotation costs in RNA sequence classification are solved, and efficient feature extraction and classification performance improvement are achieved, especially excellent performance on long sequence data.
Patent Information
- Application Number
- CN202511255952.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing technologies in RNA sequence classification have problems such as poor processing of high-throughput long sequence data, high annotation costs, low computational efficiency and high memory usage due to data sparsity, making it difficult to effectively utilize unlabeled data and capture long sequence dependencies.
An RNA sequence classification method based on the Mamba model and semi-supervised learning is adopted. By constructing an encoder-decoder structure and combining it with the Mamba module of the selective state space model, feature extraction and unsupervised latent representation learning are achieved. Unlabeled data is used to improve classification performance, and the training process is optimized through a dual-path structure and dynamic loss function.
It improves the generalization ability and computational efficiency of RNA sequence classification, reduces memory usage, and improves classification performance and feature utilization, especially in processing long sequence data.
Smart Images

Figure CN120748508A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and specifically relates to an RNA sequence classification method based on the Mamba model and semi-supervised learning, which can be used for the identification and classification of non-coding RNAs and messenger RNAs such as circRNA, IncRNA and mRNA, including technologies such as deep learning, semi-supervised learning, and state-space models. Background Art
[0002] RNA sequence classification is of great significance in bioinformatics and clinical medicine research, playing a key role in identifying non-coding RNA functions, screening disease-associated RNA markers, and discovering drug target cell sites. While RNA sequences are relatively easy to obtain, high-quality annotated data typically requires specialized experiments and expert interpretation, resulting in high annotation costs and limited availability, severely hindering the performance of fully supervised learning classification models.
[0003] At present, mainstream deep learning methods such as the attention-based Transformer have excellent performance in sequence modeling, but the training complexity is high, resulting in poor processing of high-throughput long-sequence RNA data. In addition, RNA sequences are often represented by k-mer features, and the k-mer dimension increases exponentially with the growth of k value, resulting in high and sparse input feature dimensions. Traditional dense storage and processing methods have high memory usage and low computational efficiency, which reduces the utilization of features in deep model training. Therefore, there is an urgent need for a new RNA sequence classification method that can fully utilize unlabeled data, effectively capture long sequence dependencies, and take into account the optimization of computing resources. Summary of the Invention
[0004] To overcome the problems of poor sequence modeling capabilities and low utilization of unlabeled RNA in existing technologies, this paper proposes an RNA sequence classification method based on the Mamba model and semi-supervised learning. By constructing a semi-supervised framework based on an encoder-decoder structure and integrating the Mamba module of the selective state space model, it effectively realizes the extraction of long-term dependency features of RNA sequences and unsupervised latent representation learning, improving the performance and generalization ability of classification tasks. The specific technical solution of this invention includes four steps: Step 1: Feature Extraction and Latent Space Encoding. K-mer frequency features are extracted from the RNA sequence to generate a high-dimensional sparse feature vector. The K-mer features are stored in the Compressed Sparse Row (CSR) format. An encoder network is then used to compress the 1344-dimensional high-dimensional input into a 256-dimensional latent space vector. The latent space features are then L2-normalized to enhance class discrimination.
[0005] Step 2: Fusion of the Mamba module and the attention mechanism. The Mamba module is introduced into the residual block. The traditional multi-head attention module and the state-space Mamba module are dynamically selected in the third and fourth residual blocks of the encoder. When the sequence length is less than 300, the traditional multi-head attention module is still used. When the sequence length is greater than 300, the state transfer matrix (with complexity O(L)) is used to capture long-range dependencies. Based on the Selective State Space Model (SSM) theory, the module generates discretized parameters A, B, and C through linear projection. A gating mechanism is used to implement selective state updates, effectively capturing long-sequence contextual information with low complexity. Mechanisms such as residual connections, LayerNorm, and Dropout are retained to ensure the stability and generalization of network training.
[0006] Step 3: Semi-supervised joint training and loss design. A dual-path architecture is constructed, with an encoder-decoder path implemented for unsupervised feature learning. The encoder maps the input to a 256-dimensional latent space, and the decoder reconstructs the original input from the latent space. The classification path is used for supervised learning tasks and includes a two-layer perceptron classification head. The encoder serves both reconstruction and classification tasks, and the latent space features are L2-normalized and used for both reconstruction and classification. The loss function is designed as the weighted sum of reconstruction loss and classification loss:
[0007] in, is the cross entropy loss, is the mean square error loss, and To adjust the hyperparameters of the weights, Take 5, Take 1 to balance the classification performance and feature reconstruction quality; each batch has 128 samples, with labeled samples and unlabeled samples accounting for 20% and 80% respectively. Dynamically adjust batch construction proportions; during training, the gradients of labeled and unlabeled data are separated, Gaussian noise is added to the unlabeled data to enhance reconstruction strength, and consistency regularization is used to calculate the mean squared error of two different dropouts for the same sample; a double early stopping mechanism is introduced, with the primary stopping condition being that the F1 score on the validation set does not improve for 15 consecutive rounds, and the auxiliary condition being that the ratio of reconstruction loss to classification loss deviates from the baseline for more than 10 rounds; strict training monitoring ensures the stability of semi-supervised learning.
[0008] Step 4: Classification prediction and feature application. The RNA sequence to be predicted is also subjected to k-mer feature extraction and sparse coding before being input into the trained model. The model outputs multi-class classification probabilities and a 256-dimensional latent space vector to achieve functional prediction, which can be used for downstream functional analysis, clustering, and visualization to assist in biological research. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 This is a schematic diagram of the overall process of the RNA sequence classification method based on the Mamba model and semi-supervised learning.
[0010] Figure 2 This is a semi-supervised learning network framework diagram of the encoder-decoder structure.
[0011] Figure 3 It is a schematic diagram of the network structure of the residual block integrated Mamba module.
[0012] Figure 4 Flowchart of data flow processing for model training.
[0013] Figure 5 This is the data flow processing flowchart for the model inference phase. DETAILED DESCRIPTION
[0014] The present invention is described in detail below with reference to the accompanying drawings and examples.
[0015] Feature extraction and latent space encoding. Figure 2 This paper demonstrates the process of RNA sequence data preprocessing, k-mer feature construction, and latent space vector generation. Data preparation: Three types of RNA data from public databases were selected: mRNA, lncRNA, and circRNA, with a total of approximately 61,888 samples. The RNA sequence length ranged from 200 to 2000 nt, and sequences containing N ratios greater than 10% were removed. The data was then cleaned by replacing uracil in the sequence with thymine, using a unified DNA representation. Sequences containing more than three consecutive unknown bases and sequences with N base ratios exceeding 10% were then removed. Sampling stratification was used to divide the training set, validation set, and test set into a 14:3:3 ratio, with labeled samples accounting for 15% of the training set and the rest as unlabeled data. K-mer feature extraction: Setting k = 4, the sequence is divided into 4-mer segments with a step size of 1 and a non-overlapping sliding window. possibilities; for each k , count all possible k-mer occurrence frequencies, and then divide the frequencies by the effective sequence length for normalization; to enhance feature richness, extract k =3 and k=5 features, and the horizontal splicing operation of sparse matrices is used to splice the three types of k-mer features into a sparse matrix of dimension 1344; the generated sparse features are stored in CSR format, and the memory usage of a single batch is reduced by about 70%; the latent space encoding adopts the improved UnsupervisedResNet architecture. First, the input layer receives the 1344-dimensional k-mer sparse feature vector and performs dense mapping; secondly, the 1344-dimensional features are compressed to a 1024-dimensional dense vector through the input projection layer; then, the data passes through 4 residual blocks composed of multi-head attention and Mamba, of which the first two blocks are composed of multi-head attention, and the last two blocks are composed of Mamba modules containing linear layers, Batch, and GELU activation functions; then, the 1024-dimensional features are compressed to a 256-dimensional latent space vector; finally, the output is connected with the downstream tasks, which are used for classification tasks and reconstruction tasks respectively.
[0016] Mamba module ensemble and long sequence modeling. Figure 3 Demonstrates how to integrate the Mamba module into the encoder to improve long sequence modeling capabilities and optimize video memory usage.
[0017] Mamba module position and structure: In the third and fourth residual blocks of the encoder, when the sequence length is greater than 300, the Mamba module is used to dynamically replace the traditional multi-head attention. The Mamba module first receives the residual block from the previous layer. The feature vector is input into the normalization layer unit for normalization; then, the normalized features are mapped into the parameters required for state update distribution and branch through linear projection, and one-dimensional causal convolution is performed on the input branch to capture context information; then, the state is updated according to the selected state space model, the time evolution of the hidden state is realized through the state transfer matrix, and the new input is injected in combination with the gate control mechanism; then, the updated state is output through the output projection module to generate the output feature; finally, the output and input features are superimposed through the residual link, and Dropout is performed to enhance the stability and generalization ability of network training to obtain the output of the Mamba module; parameter configuration: input / output dimension Degree: 1024, state dimension: 16, convolution kernel size: 4, expansion factor: 2; The Mamba module is based on the selective state space model SSM, and captures long-range dependencies through the state transfer matrix. The input projection maps the input x into two branches (x, z) through a linear layer, and performs one-dimensional causal convolution on the x branch: kernel_size=4, padding=3, and generates state parameters B: 16 dimensions and C: 32 dimensions through x_proj. The time step parameter dt: 1024 dimensions is generated through dt_proj, and then discretization calculation is performed. The state transfer matrix A is the core component of the Mamba module, which is used to control the time evolution characteristics of the hidden state. Its initialization time is ; The state update is realized by the following discretized equation:
[0018] in Corresponding to memory decay, represents new input injection, , t It is dynamically generated from input features. To optimize video memory, gradient checkpointing technology is used to reduce activation value storage and mixed precision training is adopted. Each residual block contains a convolution unit, LayerNorm, 0.1 Dropout and residual connection to ensure gradient stability.
[0019] Semi-supervised joint training and loss design. Figure 4 Demonstrate how to improve model performance under small sample conditions by jointly optimizing classification and reconstruction tasks.
[0020] Decoder and classifier design: The decoder uses a symmetric 4-layer convolutional network to reconstruct the 256-dimensional latent vector into a 1344-dimensional dense k-mer vector layer by layer for unsupervised reconstruction tasks; the classification head is a two-layer fully connected network that outputs a three-category probability distribution.
[0021] Loss function design: The overall loss function is , in is the cross entropy loss; is the mean square error loss; the weight is , ensuring that classification performance is prioritized; training strategy: the optimizer uses AdamW, the initial learning rate is 1e-5, and the weight decay is 1e-5, Parameters (0.9, 0.999); batch size 128, each batch contains approximately 26 labeled samples, accounting for 20% of the 1400 randomly sampled samples, and 102 unlabeled samples. The labeled samples are used to calculate the classification loss and reconstruction loss, while the unlabeled samples only participate in the reconstruction loss calculation. By setting the classification loss weight and reconstruction loss weight separately, the model prioritizes classification performance while using unlabeled data to learn more generalized potential feature representations; training for 150 epochs; at the learning rate scheduling level, the loss monitoring method is adopted. When the loss stagnates for 4 epochs, the learning rate is halved.
[0022] Classification prediction and downstream feature application. Figure 5 Demonstrate the inference process: the input RNA sequence undergoes the same k-mer feature extraction and sparse coding; only the encoder and classification head are used for forward inference, without the decoder; multi-class classification probabilities and 256-dimensional latent space vectors are output; the input RNA sequence is first converted from uracil to thymine, and then sequences containing more than 3 consecutive N bases or N ratios > 10% are filtered; a sliding window with a step size of 1 is used to extract k=3, 4, and 5 k-mer features, and using the horizontal concatenation operation of sparse matrices to splice the three types of k-mer features into a sparse matrix of dimension 1344, load the pre-trained model, and disable gradient calculation for model inference; the output result probs represents the probability distribution of mRNA, lncRNA, and circRNA, and latent is 256-dimensional floating-point data with norm = 1; downstream applications of feature vectors include t-SNE visualization of latent vectors to realize RNA family clustering; high-confidence prediction results are used for disease marker screening or functional annotation research; latent vectors can be fused with other multi-omics data for multimodal bioinformatics analysis.
[0023] The present invention uses the low-coverage benchmark dataset AttenRNA to test the performance of a proposed RNA sequence classification method based on the Mamba model and semi-supervised learning. The dataset contains 61,888 different types of biological sequence RNA. After the samples of the dataset undergo k-mer feature extraction, compressed sparse row storage, encoder network compression and L2 normalization in step 1, the features are distributed on the same hypersphere. The normalized features are then input into the residual block stacking structure built in step 2. The first two residuals retain the traditional multi-head attention to extract contextual features. The third and fourth residual blocks are replaced by the Mamba module, which processes long sequence dependencies through a selective state space mechanism. Finally, a hybrid representation that fuses local and global features is output, providing a more discriminative feature basis for downstream tasks. The model is trained on the training set and validation set of the sample dataset, tested on the test set, and finally the f1 score of the model is calculated. Based on the training and testing of the method proposed in the present invention, the f1 score value on the AttenRNA dataset was 0.9192, which is 2% higher than the current best-performing target accuracy prediction model AttenRNA. The present invention builds a hybrid framework that integrates convolution and Mamba and designs a semi-supervised loss function based on the characteristics of biological sequences. Therefore, the performance is higher than other existing methods and can effectively improve the classification ability of different RNA biological sequences.
[0024] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A RNA sequence classification method based on the Mamba model and semi-supervised learning, characterized in that: Constructing a semi-supervised framework based on an encoder-decoder structure and integrating the selective state space model Mamba module includes the following steps: Step 1: Preprocess the RNA sequence by replacing uracil with thymine and removing sequences containing more than three consecutive N bases or N accounting for more than 10%; extract k-mer features with k=3, 4, and 5 respectively using a sliding window with a step size of 1, and use the horizontal concatenation operation of the sparse matrix to concatenate the three types of k-mer features into a sparse matrix of dimension 1344. The encoder compresses them into a 256-dimensional latent vector and performs L2 normalization; Step 2: In the third and fourth residual blocks of the encoder, when the sequence length is greater than 300, the multi-head attention module is dynamically replaced with a Mamba module based on a selective state space model. The state transfer matrix captures long-range dependencies, and the state is updated in combination with a gating mechanism. Linear projection is used to generate discretized parameters A, B, C and the time step Δt. Residual connections, LayerNorm, and Dropout are retained to ensure training stability. Step 3: Construct an encoder-decoder to implement semi-supervised joint training. The classification head uses a multi-layer perceptron to output three-category results for the gradient separation of labeled data and unlabeled data. A double early stopping mechanism is introduced. The total loss is the weighted sum of cross entropy and mean square error. The model uses labeled samples for supervised training and unlabeled samples for reconstruction. Step 4: Input the RNA sequence to be classified and after the above processing, output multi-class classification probability and 256-dimensional potential features for RNA family clustering, function prediction and downstream visualization analysis.
2. The RNA sequence classification method based on the Mamba model and semi-supervised learning according to claim 1, characterized in that: The dual early stopping mechanism in step 3 includes: the main stopping condition is that the F1 score of the validation set does not improve for 15 consecutive rounds, and the auxiliary condition is that the ratio of the reconstruction loss to the classification loss deviates from the baseline for more than 10 rounds.
3. The RNA sequence classification method based on the Mamba model and semi-supervised learning according to claim 1, characterized in that The state update of the Mamba module is realized by the following discretization equation: The state transfer matrix A is the core component of the Mamba module, which is used to control the time evolution characteristics of the hidden state. Its initialization time is ; The state update is realized by the following discretized equation: ;in Corresponding to memory decay, represents new input injection, , t Dynamically generated from input features; parameter configuration: input / output dimension: 1024, state dimension: 16, convolution kernel size: 4, expansion factor:
2.
4. The RNA sequence classification method based on the Mamba model and semi-supervised learning according to claim 1, characterized in that: The semi-supervised joint training loss function is designed as the weighted sum of reconstruction loss and classification loss: ;in, is the cross entropy loss, is the mean square error loss, and To adjust the hyperparameters of the weights, Take 5, Set to 1 to balance classification performance and feature reconstruction quality.
5. The RNA sequence classification method based on the Mamba model and semi-supervised learning according to claim 1, characterized in that Horizontal concatenation of sparse matrices: RNA sequences are often represented by k-mer features. The k-mer dimension increases exponentially with the k value, resulting in high-dimensional and sparse input features. Traditional dense storage and processing methods have high memory usage and low computational efficiency, reducing the utilization of features in deep model training. k-mer frequency features are extracted from RNA sequences to generate high-dimensional sparse feature vectors, and the k-mer features are stored using the Compressed Sparse Row format.
Citation Information
Patent Citations
Supervised learning method for discriminating mRNA and lncRNA
CN108595913A
Biological sequence feature extraction method based on word embedding and auto-encoder fusion
CN113392929A
Metagenome contigs classification method based on self-supervised learning
CN113393898A
RNA-protein binding site prediction method and system based on self-attention mechanism
CN114023376A
Single-cell RNA (Ribonucleic Acid) sequence gene regulation and control inference method based on deep learning
CN116825204A
Cited By
Method and system for identifying coding potential of lincRNA small peptide
CN121838875A