A base recognition method, device and storage medium based on the Transformer architecture for nanopore sequencing
Through the autoregressive decoding method based on the Transformer architecture, combined with autoregressive decoding and beam search strategies, the problem of insufficient base recognition accuracy in nanopore sequencing is solved, and high-precision and robust base sequence recognition are achieved.
Patent Information
- Application Number
- CN202311838741.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-12-28
AI Technical Summary
The base recognition accuracy in existing nanopore sequencing technologies needs to be further improved, especially when dealing with long-term dependencies and noise interference, and existing methods are difficult to effectively capture the long-term dependencies and accurately decode the sequence.
The base recognition method based on the Transformer architecture is adopted, and the convolution module is downsampled and feature extraction is performed, and the encoding module performs global context modeling. The decoding module adopts autoregressive decoding and beam search strategies, and combines the limiting weight control decoder output to generate high-precision base sequences.
It improves the accuracy and robustness of base recognition, and is suitable for nanopore sequencing data with long read length and high error rate, achieving high-precision base sequence recognition.
Smart Images

Figure CN118038972B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information and communication technologies for nucleic acid sequence analysis, and particularly to a base recognition method, apparatus, and storage medium for nanopore sequencing based on the Transformer architecture. Background Art
[0002] Nanopore sequencing is a high-throughput sequencing method based on nanopore technology, which realizes base recognition by measuring the change in the electrical signal when a DNA or RNA molecule passes through a nanopore. Nanopore sequencing technology has the characteristics of being fast, real-time, and at the single-molecule level, and thus has broad application prospects in the fields of genomics research, medical diagnosis, and personalized medicine.
[0003] The principle of nanopore sequencing is based on the interaction between ion channels and nucleic acid molecules. An ion channel is a membrane protein with a tiny pore diameter, usually in the nanoscale range. When a nucleic acid molecule passes through an ion channel, it blocks the ion channel, thereby causing a change in the ion current and resulting in a change in the electrical signal. The change in the electrical signal is closely related to the base sequence of the nucleic acid molecule because different bases have different blocking effects on the ion channel, thus causing a specific electrical signal pattern. Therefore, nanopore sequencing technology uses this change in the electrical signal to identify the base sequence of nucleic acid molecules.
[0004] The ability to accurately decode the base sequence is a key factor determining the downstream analysis of nanopore sequencing. However, the original current signal is usually interfered by various factors, such as noise, electrode drift, and the interaction between bases. Therefore, high-precision signal decoding in nanopore sequencing is a challenging problem and is also crucial for the popularization and application of nanopore sequencing.
[0005] In the past decade, several base recognition methods for nanopore sequencing have been proposed. Its iterative process can be divided into three stages: The first stage is mainly based on the Hidden Markov Model (HMM). Since the sequencing electrical signal data corresponds to the change in the measured ion current from one nucleotide sequence of 5 or 6 bases (k-mer) to another sequence (k-mer) during the translocation of nucleic acid molecules through the nanopore. The base recognition method based on HMM regards the ion current signal as a chain of observable events, while k-mers are regarded as states within the HMM. Since the first nucleotide of each state overlaps with the last nucleotide of the previous state, the joint probability of the nucleotide sequence can be calculated, and the path with the largest total joint probability represents the final predicted sequence. However, there are the following disadvantages in base recognition through HMM: HMM can only model short-term dependencies, but there are long-term dependencies in nanopore sequencing data, and it is difficult for HMM to capture long-term dependencies; HMM is a prior model of known DNA sequences, which will lead to too high error rates when dealing with unknown base sequences. Typical practices are such as Metrichor and Nanocall. The algorithm basis adopted in the second stage is the Recurrent Neural Network (RNN), but its input is still the k-mer events segmented according to the sharp change of the electrical signal. Due to the changes in translocation speed and noise signals, the segmentation is prone to errors. Typical recognition tools in this stage are such as DeepNano and BasecRAWller. The third stage is the currently mainstream base recognition method, which performs base recognition by constructing an end-to-end deep learning neural network. Currently, the main model frameworks in this stage are to perform downsampling and feature extraction through Convolutional Neural Network (CNN), perform context modeling through RNN, and perform many-to-one mapping decoding through Connectionist Temporal Classification (CTC). However, the CNN model cannot well capture the temporal information in the sequence; although RNN can model long-range dependencies, there are problems of gradient vanishing and gradient explosion; CTC has the defect of assuming independence between events. Typical algorithms in this stage are such as Chiron and Bonito.
[0006] The Transformer architecture was initially proposed to solve the machine translation task. It effectively captures the long-range dependencies in the sequence through the self-attention mechanism and the multi-head attention mechanism, and models the correlation between the source sequence and the target sequence through the autoregressive decoding mode. It has achieved milestone results in fields such as natural language processing, image recognition, and automatic speech recognition. At the same time, in the biological field, this architecture has also demonstrated powerful capabilities in protein generation, structure prediction, drug design, and biological sequence recognition.
[0007] The literature "Improving basecalling accuracy with transformer" discloses the use of FishNChips to identify base sequences. CN116312791A discloses a nanopore sequence recognition network structure based on Transformer. First, a nanopore sequence is input, and the input nanopore sequence is converted into an image of a fixed size through a sequence-to-image module; then the converted image is input into the sequence-to-image network for training, and finally the category of the sequence is output. Among them, the sequence-to-image network includes a Transformer structure and a classification layer, and the Transformer structure includes an encoder, a decoder, and a multi-head attention mechanism module. However, the above method does not adopt an autoregressive decoding method in the decoding process, and directly performs greedy search on the probability matrix output by the model. This decoding strategy is extremely likely to obtain a local optimal solution and lacks the ability to backtrack. Therefore, once the choice made leads to entering an unfavorable state, greedy search cannot correct the error. Summary of the Invention
[0008] Aiming at the deficiencies in the prior art, the present invention provides a base recognition method, device, and storage medium for nanopore sequencing based on the Transformer architecture. It solves the problem that the base recognition accuracy in the prior art needs to be further improved.
[0009] In the first aspect of the present invention, a base recognition method for nanopore sequencing based on the Transformer architecture is provided, including the following steps:
[0010] Step S1: The convolution module receives the original current signal of nanopore sequencing, performs downsampling and feature information extraction, and performs position encoding to obtain a fused vector;
[0011] Step S2: Input the fused vector of Step S1 into the encoding module to perform global context modeling and characterization on the current signal to generate a high-dimensional intermediate hidden vector;
[0012] Step S3: The decoding module simultaneously receives the hidden vector of Step S2 and the partial base sequence already generated by the decoding module, predicts the candidate base at the next time step, and when the decoder generates the end decoding character, the base recognition ends, that is, the base sequence corresponding to the original current signal is generated.
[0013] In a specific embodiment of the present invention, in Step S1, the original current signal undergoes preprocessing and segmentation processing steps;
[0014] The preprocessing includes: identifying the valid signal and noise signal of the original current signal, removing the noise signal, and segmenting the step signal;
[0015] The segmented processing includes: intercepting a plurality of current signal segments of a preset length from the current signal according to a preset overlapping rate;
[0016] Preferably, the length of each current signal segment is not greater than 10,000, and the overlapping rate of the head and tail of each segment of the signal is one-tenth of the signal length.
[0017] In a specific embodiment of the present invention, in step S1, the original current signal is a long-read current signal obtained by nanopore sequencing;
[0018] Preferably, the convolution module includes a plurality of groups of convolutional layers, and the convolutional layer includes a one-dimensional convolutional layer, a batch normalization layer, a swish activation function, and a scaling layer;
[0019] More preferably, the specific calculation principle of the swish activation function is:
[0020] f(x) = x * sigmoid(x)
[0021]
[0022] where x is the input;
[0023] More preferably, the convolution module includes three groups of convolutional layers. The output dimension of the one-dimensional convolutional layer in the first group is 4, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolutional layer in the second group is 16, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolutional layer in the third group is 384, the kernel size is 19, and the stride is 5;
[0024] More preferably, the scaling layer maps the eigenvalue to between -0.5 and 3.5.
[0025] In a specific embodiment of the present invention, in step S2, the encoder includes a plurality of groups of encoding layers, and the encoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, and a linear layer;
[0026] Preferably, the calculation principle of the multi-head self-attention layer is:
[0027] (Q, K, V) = X(W Q , W K , W V )
[0028]
[0029] where X is the input matrix, (W Q , W K , W V) are three trainable parameter matrices; Q, K, and V are the query, key, and value matrices respectively, obtained by linearly transforming the input matrix, and d k is the dimension of the key;
[0030] NultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0031]
[0032] where W O is a trainable parameter matrix.
[0033] In a specific embodiment of the present invention, in step S3, the decoding module includes a plurality of groups of decoding layers, and each decoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, a cross-attention layer, a residual connection, a normalization layer, and a linear layer;
[0034] Preferably, the method for the decoding module to predict candidate bases at the next time step includes a method combining autoregression and beam search;
[0035] More preferably, the method combining autoregression and beam search includes: the decoding module generates K candidate bases at the next time step through a beam search strategy according to the input vector composed of the hidden vector generated by the encoding module and the partial base sequence already generated by the decoding module and the constraint weight W;
[0036] More preferably, the method of inputting the partial base sequence already generated by the decoding module into the decoding module includes: inputting the partial base sequence already generated by the decoding module into the word embedding layer to obtain word vectors, performing positional encoding on the word vectors, and then inputting them into the decoding module, where the initial character is [BOS];
[0037] Even more preferably, the value range of the constraint weight W is 0 ≤ W ≤ 1; preferably 0.1 ≤ W ≤ 0.7;
[0038] Even more preferably, the value range of the number of candidates K for beam search is 1 ≤ K ≤ 4, and K is an integer.
[0039] In a specific embodiment of the present invention, step S3 further includes combining and processing multiple decoding results. Except for retaining all bases in the first segment, the head of each of the remaining base sequences removes 0.1275 * the length of the electrical signal overlapping part of bases.
[0040] In a specific embodiment of the present invention, in steps S1 - S3, the optimizer used during model training is AdamW.
[0041] In the second aspect of the present invention, a base recognition system for nanopore sequencing based on the Transformer architecture is provided, including: a convolutional module, a high-dimensional intermediate hidden vector generation module, and a base sequence generation module;
[0042] The convolutional module is used to downsample and extract features from the original current signal;
[0043] The high-dimensional intermediate hidden vector generation module is used to perform global context modeling and characterization on the current signal to be sequenced to generate a high-dimensional intermediate hidden vector;
[0044] The base sequence generation module is used to receive the hidden vector and the generated base sequence, predict the candidate base at the next time step, and when the decoding character is generated, the base recognition ends, that is, the base sequence corresponding to the original current signal is generated.
[0045] Preferably, the convolutional module includes a plurality of groups of convolutional layers, and each convolutional layer includes a one-dimensional convolutional layer, a batch normalization layer, a swish activation function, and a scaling layer;
[0046] The high-dimensional intermediate hidden vector generation module includes a plurality of groups of encoding layers, and each encoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, and a linear layer;
[0047] The base sequence generation module includes a plurality of groups of decoding layers, and each decoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, a cross-attention layer, a residual connection, a normalization layer, and a linear layer;
[0048] Preferably, the base recognition device further includes a current signal preprocessing module and a current signal segmentation processing module;
[0049] The current signal preprocessing module is used to identify the effective signal and noise signal of the original current signal, remove the noise signal, and segment the step signal;
[0050] The current signal segmentation processing module is used to intercept a plurality of current signal segments with a preset length from the current signal according to a preset overlap rate.
[0051] In the third aspect of the present invention, a base recognition device is provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above-mentioned base recognition method is implemented.
[0052] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned base recognition method.
[0053] The technical principle of the present invention is:
[0054] The correlation between the source sequence (such as the original current signal) and the target sequence (such as the base sequence) is modeled through the autoregressive decoding mode. Specifically, the present invention adopts a beam search strategy based on the autoregressive method in the decoding stage. The autoregressive decoding is based on the generated part and gradually generates the next output. This method allows the model to make decisions based on the previously generated content at each step, thereby obtaining a more accurate and coherent output sequence, which is very effective for processing sequence tasks with long-term dependencies. The beam search considers multiple candidate solutions at each decoding time step to avoid falling into the local optimal solution, and backtracks at the last time step to obtain the optimal solution among the candidate solutions, and has error correction capabilities. At the same time, the present invention innovatively proposes to limit the weights and control the diversity of the output results of the decoder, which can further improve the prediction accuracy.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] First, for the first time, a Transformer-based autoregressive decoding method is provided for nanopore sequencing base recognition with high accuracy and robustness. Second, the accuracy of base recognition is improved by modeling the contextual information of the sequence. Third, it is suitable for processing nanopore sequencing data with long read lengths and high error rates. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Schematic diagram of the workflow of base recognition in Example 1 of the present invention;
[0058] Figure 2 This is a diagram of the decoding process in Embodiment 1 of the present invention;
[0059] Figure 3 The prediction accuracy of the present invention under different restriction weights in Example 1 of the present invention;
[0060] Figure 4 The figure is a comparison of the prediction accuracy of different nanopore sequencing base recognition methods in Example 2 of the present invention. DETAILED DESCRIPTION
[0061] The technical scheme of the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the following embodiments are only exemplary descriptions and explanations of the present invention and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are included in the scope that the present invention is intended to protect.
[0062] Unless otherwise specified, the raw materials and reagents used in the following examples are commercially available or can be prepared by known methods.
[0063] Technical terms
[0064] The "self-attention mechanism" can simultaneously focus on all positions in the sequence, thereby capturing global dependencies. Moreover, this mechanism allows the model to automatically calculate the importance weights of each position according to different positions in the input sequence. This enables the model to focus on important information relevant to the current task and place more attention on meaningful positions. Compared with the traditional fixed-weight mechanism, the self-attention mechanism can more flexibly adapt to different tasks and inputs.
[0065] The "multi-head self-attention mechanism" allows the model to simultaneously learn multiple different attention representations. Each attention head can focus on different parts of the sequence, thereby providing multiple independent expressive capabilities. Multi-head attention can capture semantic information at different levels, enabling the model to better understand the input sequence and improve the representation and generalization capabilities.
[0066] "Autoregressive" means using the output of the decoder at the previous time step as the input to the decoder at the next time step for prediction. The first initial character received by the decoder is bos, and then autoregressive iteration begins until eos is generated and the decoder stops working.
[0067] Example 1 A method for nanopore sequencing base recognition based on the Transformer architecture
[0068] As Figure 1 shown, this method includes the following steps:
[0069] Step S1: The convolutional module receives the original current signal of nanopore sequencing, performs downsampling and feature information extraction, and performs positional encoding to obtain a fused vector;
[0070] Step S2: Input the fused vector from Step S1 into the encoding module to perform global context modeling and characterization on the current signal, generating a high-dimensional intermediate hidden vector;
[0071] Step S3: The decoding module simultaneously receives the hidden vector from Step S2 and the partial base sequence already generated by the decoding module, predicts the candidate base for the next time step, and when the decoder generates the end-of-decoding character (EOS), the base recognition ends, that is, the complete base sequence corresponding to the original current signal is generated.
[0072] Among them, in step S1, when using the convolutional module to downsample and extract features from the original current signal of nanopore sequencing, the convolutional module includes three groups of convolutional layers, and each group of convolutional layers is successively composed of a one-dimensional convolutional layer, a batch normalization layer, a swish activation function layer, and a scaling layer. When executing the steps, they are connected in series in sequence. Specifically, the one-dimensional convolutional layer downsamples and extracts features from the original current signal, and outputs a one-dimensional current signal feature map; the batch normalization layer converges the output one-dimensional current signal feature map to improve the generalization ability of the network; the swish activation function is introduced to increase the non-linear output; the scaling layer scales and displaces the non-linear output to prevent gradient explosion.
[0073] Furthermore, the specific calculation principle of the swish activation function is:
[0074] f(x) = x * sigmoid(x)
[0075]
[0076] where x is the input.
[0077] The output dimension of the one-dimensional convolutional layer of the first group is 4, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolutional layer of the second group is 16, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolutional layer of the last group is 384, the kernel size is 19, and the stride is 5. The final scaling layer of each group maps the feature values between -0.5 and 3.5.
[0078] Furthermore, a position encoding mechanism is introduced into the base sequence feature information, and the base order information in the base sequence is input into the model, enabling it to capture the relative position relationship of different nucleotides in the sequence, better handle the long-range dependence relationship in the sequence, and help understand the semantic differences at different positions in the sequence. Even further, the specific calculation mechanism of the position encoding (PE) is:
[0079]
[0080]
[0081] where pos is the position of the base in the sequence, i is the dimension in the encoding vector, and d model is the dimension of the model, which is 384 in this embodiment. The feature information and position information of the base sequence are weighted and fused to obtain a fused vector, which is used as the output vector of the encoder.
[0082] Further, it also includes the steps of preprocessing and segmenting the original current signal. Specifically, it includes: 1) preprocessing the original current signal, where the preprocessing includes identifying valid signals and noise signals, removing noise signals, and segmenting step signals to obtain a dataset with desired quality, and the dataset with desired quality includes multiple read electrical signal sequences; 2) segmenting each of the multiple read electrical signal sequences in the dataset of step 1), and for each read electrical signal sequence, N segmented subsequences are obtained. Among them, in the segmenting process, the length of each segment of the signal is not greater than 10,000, and the overlapping part at the head and tail of each segment of the signal is one-tenth of the length.
[0083] In the step where the encoding module in step S2 performs global context modeling and characterization on the sequencing electrical signal to generate a high-dimensional intermediate hidden vector, the encoder (encoding module) is composed of 8 groups of encoding layers connected in series. Each group of encoding layers sequentially includes a multi-head self-attention layer, a residual connection, a normalization layer, and a linear layer. When executing the steps, they are connected in series in sequence. The output vector of the encoder is input into the multi-head attention layer, and the multi-head attention layer performs context encoding on the input vector sequence. The residual connection is used to prevent gradient disappearance and explosion and avoid information loss; the normalization layer can accelerate convergence and improve robustness; the linear layer is used to increase the expression ability of the model.
[0084] Further, the calculation principle of the multi-head self-attention layer is as follows:
[0085] (Q, K, V) = X(W Q , W K , W V )
[0086]
[0087] Among them, X is the input matrix, and (W Q , W K , W V ) are three trainable parameter matrices. Q, K, and V are the query, key, and value matrices respectively, which are obtained by linearly transforming the input matrix, and d k is the dimension of the key.
[0088] NultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0089] Among them, W O is a trainable parameter matrix.
[0090] In step S3, the decoder (decoding module) simultaneously receives the hidden vector from step S2 and the partial base sequence (with the initial character [BOS]) already generated by the decoding module, and predicts the candidate bases for the next time step. When the decoder generates the end-of-decoding character (EOS), the base recognition ends, that is, the complete base sequence corresponding to the original waveform is generated. The decoder consists of 8 decoding layers connected in series. Each decoding layer sequentially includes a multi-head self-attention layer, a residual connection, a normalization layer, a cross-attention layer, a residual connection, a normalization layer, and a linear layer. The cross-attention layer is used to associate and interact the information between the encoder and the decoder. Except for the cross-attention layer, the other parts are the same as those of the encoder.
[0091] The specific algorithm of the decoder is to input the base sequence already generated by the decoder into the word embedding layer for word embedding representation and perform position encoding; together with the hidden vector obtained in step S2, it is input into the decoder (decoding module). The decoding module simultaneously receives the hidden variable generated by the encoder and the generated base sequence, and generates K candidate bases for the next time step according to the input vectors (hidden variable and generated base sequence) and the restricted weights (W) through the beam search strategy. This process is repeated until EOS is generated, and the sequence with the highest score is selected as the final base recognition result, that is, the complete base sequence corresponding to the original waveform is generated. If the Beam Size is set to 4, that is, at the inference time t, the model will retain the 4 sequences with the highest scores and input them into the prediction at time t+1, and select the 4 sequences with the highest scores among the 16 candidate sequences at time t+1, and so on until the decoding ends, and finally the sequence with the highest cumulative score is selected as the decoding result.
[0092] Furthermore, step S3 also includes merging and processing multiple decoding results. Except for retaining all bases in the first segment, the head of each subsequent segment of the base sequence removes 0.1275 * the length of the electrical signal overlapping part of bases. The decoding process is as Figure 2 shown.
[0093] This embodiment also includes using an optimizer for model training. Among them, the optimizer is AdamW, and the training process of the model is optimized by dynamically adjusting the learning rate. In the warm-up stage, the learning rate is gradually increased from a lower initial value to a higher value in a linear manner, and the calculation method is:
[0094] LR(t) = y0 + (y1 - y0) * t
[0095] where y0 is 0.0001, y1 is 0.001, and t is the time step.
[0096] After the warm-up stage ends, the learning rate is gradually decreased through cosine decay to better optimize the model, and its calculation principle is
[0097] LR(t) = [y1 + 0.5 * (y0 - y1) * (cos(t * π) + 1.0)] * LR(t)
[0098] Where y0 is 0.1, y1 is 0.01, and t is the time step.
[0099] The model is trained using the cross - entropy loss function, does not calculate the loss of padding labels, and performs smoothing on the label sequence with a smoothing coefficient of 0.1. The model is trained for a total of 16 epochs with a batch size of 2 and is trained in parallel on 8 Nvidia A100 40G, taking a total of 164.12 hours.
[0100] Furthermore, in step S3, the value range of w is [0, 1]. When w = 0, the generation of EOS by the model is not restricted. When w = 1, it means that the model does not consider the situation of generating EOS, and the model stops decoding at the maximum length. This embodiment further compares the prediction accuracy of this method under different w values on the same test set. The results are as Figure 3 shown. The results show that when the generation of EOS is not restricted (i.e., w = 0) or strictly restricted (w >= 0.8), the prediction accuracy will significantly decline. The model performance is stable and performs well in the range of 0.1 - 0.7.
[0101] Example 2 Comparison of the prediction accuracy of different models
[0102] The prediction accuracy of the method of the present invention is compared with Bonito CRF, Bonito CTC, and SACall on the same test set. The results are as Figure 3 shown.
[0103] After comparison, the results show that the method of the present invention performs better than the existing mainstream base recognition methods in terms of the three indicators of consistency, deletion rate, and insertion rate, and the mismatch rate is slightly inferior to the Bonito CRF model. The prediction accuracies of the other several base recognition software are Bonito CRF, SACall, and Bonito CTC in turn.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A base recognition method for nanopore sequencing based on the Transformer architecture, characterized in that, It includes the following steps: Step S1: The convolution module receives the original current signal of nanopore sequencing, performs downsampling and feature information extraction, and performs positional encoding to obtain a fused vector; Step S2: Input the fused vector of Step S1 into the encoding module to perform global context modeling and characterization on the current signal, and generate a high-dimensional intermediate hidden vector; Step S3: The decoding module simultaneously receives the hidden vector of Step S2 and the partial base sequence already generated by the decoding module, and predicts the candidate bases at the next time step. When the decoder generates the end decoding character, the base recognition ends, that is, the complete base sequence corresponding to the original current signal is generated; Among them, the decoding module includes several groups of decoding layers, and each decoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, a cross-attention layer, a residual connection, a normalization layer, and a linear layer; the method for the decoding module to predict the candidate bases at the next time step includes a method combining autoregression and beam search.
2. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, wherein, In Step S1, the original current signal goes through the steps of preprocessing and segmentation processing; The preprocessing includes: identifying the valid signal and noise signal of the original current signal, removing the noise signal, and segmenting the step signal; The segmentation processing includes: intercepting several current signal segments of a preset length from the current signal according to a preset overlap rate.
3. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, characterized in that, The length of each segment of the current signal is not greater than 10000, and the overlap rate at the head and tail of each segment is one-tenth of the signal length.
4. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, wherein, In Step S1, the original current signal is a long-read current signal obtained by nanopore sequencing.
5. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, wherein, The convolution module includes several groups of convolution layers, and each convolution layer includes a one-dimensional convolution layer, a batch normalization layer, a swish activation function, and a scaling layer.
6. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 5, wherein, The specific calculation principle of the swish activation function is: f(x) = x * sigmoid(x) where x is the input.
7. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 5, wherein, The convolution module includes three groups of convolution layers. The output dimension of the one-dimensional convolution layer in the first group is 4, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolution layer in the second group is 16, the kernel size is 5, and the stride is 1; the output dimension of the one-dimensional convolution layer in the third group is 384, the kernel size is 19, and the stride is 5.
8. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 5, characterized in that The scaling layer maps the eigenvalue to the range between -0.5 and 3.
5.
9. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, wherein, In Step S2, the encoding module includes several groups of encoding layers, and each encoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, and a linear layer.
10. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 9, wherein, The calculation principle of the multi-head self-attention layer is: (Q, K, V) = X(W Q , W K , W V ) Among them, X is the input matrix, and (W Q , W K , W V ) are three trainable parameter matrices; Q, K, and V are the query, key, and value matrices respectively, which are obtained by linearly transforming the input matrix, and d k is the dimension of the key; MultiHead(Q, K, V) = Concat(head1,..., head h )W O where W O is a trainable parameter matrix.
11. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, characterized in that, In Step S3, the method combining autoregression and beam search includes: the decoding module generates K candidate bases at the next time step through a beam search strategy according to the input vector composed of the hidden vector generated by the encoding module and the partial base sequence already generated by the decoding module and the constraint weight W.
12. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 11, wherein The method of inputting the partial base sequence already generated by the decoding module into the decoding module includes: inputting the partial base sequence already generated by the decoding module into the word embedding layer to obtain a word vector, performing positional encoding on the word vector, and then inputting it into the decoding module, where the initial character is [BOS].
13. A base recognition method for nanopore sequencing based on the Transformer architecture according to claim 11, wherein, The value range of the constraint weight W is 0 ≤ W ≤ 1.
14. A base recognition method for nanopore sequencing based on the Transformer architecture according to claim 13, wherein, The value range of the constraint weight W is 0.1 ≤ W ≤ 0.
7.
15. A base recognition method for nanopore sequencing based on the Transformer architecture according to claim 11, wherein The value range of the number of beam search candidates K is 1 ≤ K ≤ 4, and K is an integer.
16. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, wherein Step S3 further includes combining and processing multiple decoding results. Except for retaining all bases in the first segment, the head of each of the remaining base sequences is removed by 0.1275 * the length of the electrical signal overlapping part.
17. The base recognition method based on the Transformer architecture for nanopore sequencing according to claim 1, characterized in that, In steps S1 - S3, the optimizer used during model training is AdamW.
18. A base recognition system for nanopore sequencing based on the Transformer architecture, characterized in that It includes: a convolutional module, a high-dimensional intermediate hidden vector generation module, and a base sequence generation module; The convolutional module is used to downsample and extract features from the original current signal; The high-dimensional intermediate hidden vector generation module is used to perform global context modeling and characterization on the current signal to be sequenced to generate a high-dimensional intermediate hidden vector; The base sequence generation module is used to receive the hidden vector and the generated base sequence, predict the candidate bases at the next time step. When the decoding characters are all generated, the base recognition ends, that is, the base sequence corresponding to the original current signal is generated; The base sequence generation module includes several groups of decoding layers. The decoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, a cross-attention layer, a residual connection, a normalization layer, and a linear layer; the method for the decoding module to predict the candidate bases at the next time step includes a method that combines autoregression and beam search.
19. A base recognition system for nanopore sequencing based on the Transformer architecture according to claim 18, characterized in that, The convolutional module includes several groups of convolutional layers. The convolutional layer includes a one-dimensional convolutional layer, a batch normalization layer, a swish activation function, and a scaling layer; The high-dimensional intermediate hidden vector generation module includes several groups of encoding layers. The encoding layer includes a multi-head self-attention layer, a residual connection, a normalization layer, and a linear layer.
20. A base recognition system for nanopore sequencing based on the Transformer architecture according to claim 18, characterized in that, The base recognition system further includes a current signal preprocessing module and a current signal segmentation processing module; The current signal preprocessing module is used to identify the valid signal and noise signal of the original current signal, remove the noise signal, and segment the step signal; The current signal segmentation processing module is used to intercept several current signal segments with a preset length from the current signal according to a preset overlapping rate.
21. A base recognition device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the base recognition method according to any one of claims 1 - 17.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the base recognition method according to any one of claims 1 - 17.
Citation Information
Patent Citations
Nanopore sequence recognition network structure based on Transform
CN116312791A
Method for quickly identifying single-molecule nanopore sequencing bases based on deep network
CN112183486A
Encoder, decoder, coding and decoding system and method for nanopore signals
CN117113063A