Nanopore sequencing bar code demultiplexing method based on deep learning
By combining deep learning methods of one-dimensional convolutional neural network, Transformer and time convolutional network, the accuracy and efficiency limitations of barcode demultiplexing and signal processing in nanopore sequencing are solved, and more efficient signal decoding and feature extraction are achieved.
Patent Information
- Application Number
- CN202510113340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has limitations on accuracy and efficiency in barcode demultiplexing and signal processing in nanopore sequencing, especially the lack of effective molecular barcode protocols in RNA sequencing, and deep convolutional networks may lead to overfitting and reduce the ability to generalize to unseen data.
Deep learning-based method is adopted, combining one-dimensional convolutional neural network, Transformer and time convolution network to perform demultiplexing of nanopore sequencing signals. Specific steps include data preprocessing, preliminary feature extraction, global context modeling, time series dependency modeling and classification processing.
It significantly improves the accuracy and efficiency of demultiplexing of barcodes, effectively extracts local features of the signal, models global context information and captures long-range time dependence, reduces calculation errors and improves signal decoding accuracy.
Smart Images

Figure CN119993273A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the combination of bioinformatics and artificial intelligence, and relates to a nanopore sequencing barcode demultiplexing method based on deep learning. Background Art
[0002] The third generation sequencing technology has become an important tool for exploring complex genomes and transcriptomes. It can directly sequence DNA and RNA molecules and obtain single nucleotide resolution, which greatly reduces amplification bias and broadens its application range. In particular, ONT has made significant progress in genome assembly, transcriptome assembly, and mutation detection through its direct RNA sequencing technology. Although the third generation sequencing technology has great potential, it still faces many challenges in barcode demultiplexing and signal processing, especially in RNA sequencing, where there is a lack of effective molecular barcode protocols. Most of the existing barcode demultiplexing methods rely on traditional sequence-based algorithms, which face limitations in accuracy and efficiency when processing complex raw signals. In addition, although ONT provides a barcode toolkit for DNA sequencing, the barcode protocol for RNA sequencing is still immature. In this process, the conversion and decoding process of the original nanopore signal requires high-computation gene sequence base calling and there are a lot of errors, especially when the signal quality is poor, the calculation error will be more significant. In addition, due to the lack of effective feature extraction methods and long-range dependency modeling capabilities, existing demultiplexing methods perform poorly in terms of diversity and versatility, and cannot meet the needs of practical applications.
[0003] In recent years, deep learning-based technologies have been gradually applied to barcode demultiplexing, especially convolutional neural networks. However, although existing convolutional neural network-based methods can effectively extract local features, they often have limitations in dealing with long-range time series dependencies, especially in nanopore sequencing signals. The complexity of time series requires the model to be able to handle long-term dependencies. In addition, deep convolutional network models may lead to overfitting, reducing the model's generalization ability for unseen data. Summary of the invention
[0004] In view of the above problems, the purpose of the present invention is to propose a nanopore sequencing barcode demultiplexing method based on deep learning.
[0005] The technical solution of the present invention is: a nanopore sequencing barcode demultiplexing method based on deep learning according to the present invention comprises the following steps:
[0006] Step 1: extracting the barcode signal from the original nanopore sequencing signal and performing data preprocessing on the barcode signal;
[0007] Step 2: The barcode signal is subjected to preliminary feature extraction through a one-dimensional convolutional neural network to fuse information from different feature resolutions;
[0008] Step 3: Model the global context of the signal through the Transformer module to capture long-range dependencies;
[0009] Step 4: Use a temporal convolutional network to model the time series dependencies of the signal to preserve the temporal order of the signal;
[0010] Step 5: Use the classification layer to classify the extracted feature representation into barcodes to achieve barcode demultiplexing;
[0011] Step 6: Visualize the model attention map and signal temporal correlation map.
[0012] Furthermore, the specific steps of step 1 are as follows:
[0013] Step 1.1: Use the Z-Score normalization method to normalize the barcode raw signal. This process retains the dimensional shape of the signal and ensures the consistency of the analysis, as shown in the following formula:
[0014]
[0015] In the formula, x and x ′ Represent the input and output signals respectively, μ and σ are the mean and standard deviation of the signals;
[0016] Step 1.2: Use three data augmentation strategies to improve the robustness and generalization of the model.
[0017] The data enhancement strategies mainly include: baseline mutation: randomly scaling the signal values at different time positions and extracting the scaling factor from the normal distribution; duration mutation: adjusting the duration of each time point with a certain probability, effectively stretching or compressing part of the signal to simulate time changes; adding noise: introducing Gaussian distributed noise into the signal to simulate random fluctuations and environmental noise.
[0018] Furthermore, the specific steps of step 2 are as follows:
[0019] Step 2.1: Through the multi-layer convolution module, the module mainly includes: a one-dimensional convolution layer, which is used to extract local features in the time series signal; a batch normalization layer, which is used to stabilize the training process and improve the training efficiency of the model; an average pooling layer, which is used to downsample the signal;
[0020] The convolution module is as follows:
[0021] x=AvgPool(ELU(BatchNorm(Conv1d(X,W)))),
[0022] Where X and x represent the input signal and output signal respectively, W represents the learnable weight of the convolutional layer, Conv1d represents the one-dimensional convolution operation, BatchNorm represents batch normalization, ELU is the activation function, and AvgPool represents the average pooling operation;
[0023] Step 2.2: Residual downsampling is performed through multiple convolutional modules to retain low-level features and intermediate representations of the signal and capture the signal hierarchical pattern;
[0024] Step 2.3: Fuse the features of different levels of convolution modules and residual downsampling to obtain signal information of different feature resolutions. The fused features contain fine temporal features and high-level temporal features from different scales, which can better represent the multi-level structural information of the input data.
[0025] Furthermore, the specific steps of step 3 are as follows:
[0026] Step 3.1: Receive the output of the multi-layer feature fusion module, denoted as F fusion , for F fusion Add position encoding to obtain F0 to maintain the order information of the time series, as shown below:
[0027] F0=F fusion +P,
[0028] Where P is the position encoding matrix;
[0029] Step 3.2: Project F0 to generate query vector Q, key vector K and value vector V, as shown below:
[0030] Q=F0W Q ,K=F0W K ,V=F0W V ,
[0031] Where W Q ,W K ,W V is the weight matrix of linear projection;
[0032] The self-attention mechanism is used to calculate the weights and outputs as follows:
[0033]
[0034] Where, d k represents the dimension of the key vector K;
[0035] In order to improve computational efficiency and model expressiveness, multiple self-attention patterns are calculated in parallel through a multi-head attention mechanism. The output of each attention head is expressed as:
[0036]
[0037] Where i is the attention head number, is the weight matrix of the i-th attention head;
[0038] The outputs of all attention heads are concatenated and linearly transformed to obtain the final output, as shown below:
[0039] MHA(Q,K,V)=Concat(head1,…,head h )W o ,
[0040] Where h represents the number of attention heads, W o Represents the linear transformation matrix after splicing;
[0041] Step 3.3: Add the output of the multi-head attention mechanism to the input through residual connection to stabilize the model training, and apply layer normalization operation to normalize the feature distribution;
[0042] Step 3.4: The output of the attention mechanism is further processed using a feedforward neural network;
[0043] The feedforward network consists of two fully connected layers and activation functions to further enhance the feature expression capability;
[0044] Residual connections and layer normalization are also applied to the output of the feed-forward network to improve training stability.
[0045] Furthermore, the specific steps of step 4 are as follows:
[0046] Step 4.1: Perform causal convolution on the signal features processed in step 3 to ensure that the prediction at each moment depends only on the current and previous input data to avoid information leakage;
[0047] Step 4.2: Dilated convolution expands the receptive field by introducing gaps between consecutive filter positions in the convolution kernel, capturing long-distance temporal dependencies, allowing the model to capture dependencies over longer time spans without increasing computational costs;
[0048] The dilated convolution operation is as follows:
[0049]
[0050] In the formula, y t is the output at time step t, x t-r·i is the input of time step tr·i, r is the expansion rate, w i is the convolution weight, k is the convolution kernel size;
[0051] Step 4.3: The dilation rates of different layers in the temporal convolution module are set to exponential growth to achieve multi-level receptive field expansion.
[0052] Furthermore, the specific steps of step 5 are as follows: flatten the features extracted from step 2 to step 4, input them into a multi-layer fully connected network for classification, and apply an ELU activation function after each layer of linear transformation to introduce nonlinear features;
[0053] The output of the last layer is converted into category probability distribution through softmax, so as to achieve accurate classification of barcodes.
[0054] Furthermore, the specific steps of step 6 are as follows:
[0055] Step 6.1: Extract the attention weight matrix A∈R from the last encoder layer of Transformer T×T , where T is the number of time steps after the signal is fused with multiple layers of features, and each element A of the matrix ij represents the attention weight of the i-th time step to the j-th time step;
[0056] Calculate each column A j The total attention value, that is, the influence of each time step on the entire sequence, is obtained by summing each column of the attention weight matrix, as shown in the following formula:
[0057]
[0058] In the formula, Attention Value j represents the attention value at the jth time step;
[0059] Step 6.2: Attention Value of the signal time step j Mapped to the dimensional space T0 of the original signal by linear interpolation, as follows:
[0060]
[0061] In the formula, Represents the signal attention value after linear interpolation, LinearInterp represents the linear interpolation function, which maps the attention of each time step to the original signal dimension;
[0062] The weight vector after linear interpolation Visualization is used to analyze and interpret the model’s attention distribution. Visualization can intuitively show the impact of different time steps in the signal on the final barcode classification, helping researchers understand the degree to which the model pays attention to important features in the signal.
[0063] The beneficial effects of the present invention are as follows: the present invention significantly improves the accuracy and efficiency of barcode demultiplexing by combining convolutional neural networks, Transformer and time convolutional networks. This innovative method can effectively extract local features of signals, model global context information and capture long-range temporal dependencies, solving the shortcomings of traditional methods in modeling long-range dependencies. Compared with traditional tools based on gene sequence base calls, this method directly processes the original signal, reduces errors and improves signal decoding accuracy. In addition, the provided attention map and signal time correlation map visualization functions enhance the interpretability of the model and help researchers understand the basis for model decision-making. This method is efficient, scalable, and applicable to a variety of nanopore sequencing data. It is a reliable and accurate solution to the problem of barcode demultiplexing. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a specific method structure diagram of the present invention;
[0065] Figure 2 It is a schematic diagram of the division of the data set training set, validation set and test set used in the present invention;
[0066] Figure 3 Schematic diagram of 4 types of barcode sequencing signals in Dataset I of the present invention;
[0067] Figure 4 It is a schematic diagram of the comparison result of signal processing of the present invention;
[0068] Figure 5 It is a schematic diagram of the comparison results of the barcode demultiplexing model of the present invention on two data sets;
[0069] Figure 6 It is the attention map and signal time correlation map of the present invention. DETAILED DESCRIPTION
[0070] The specific technical scheme of the present invention is further described in detail below with reference to specific examples.
[0071] like Figure 1As shown, the example of the present invention discloses a nanopore sequencing barcode demultiplexing method based on deep learning. First, the barcode partial signal is extracted from the original nanopore sequencing signal, and the barcode signal is preprocessed; then the barcode signal is subjected to preliminary feature extraction through a one-dimensional convolutional neural network, and information from different feature resolutions is integrated; then the signal is globally context modeled through a Transformer module to capture long-range dependencies; then the time series dependencies of the signal are modeled using a temporal convolutional network to retain the time order of the signal; finally, the extracted feature representation is barcoded using a classification layer to achieve barcode demultiplexing; the model attention map and the signal time correlation map can be visualized to further analyze the barcode demultiplexing principle.
[0072] 1. Signal preprocessing:
[0073] The present invention uses two barcode original signal data sets for training and testing. The data volume and division of the two data sets are as follows: Figure 2 As shown in the figure, the visualization results of the original signals of different types of barcodes are as follows Figure 3 shown.
[0074] First, the Z-Score normalization method is used to normalize the barcode raw signal. This process retains the dimensional shape of the signal and ensures the consistency of the analysis. The calculation method is:
[0075] In the formula, x and x ′ Represent the input and output signals respectively, μ and σ are the mean and standard deviation of the signals.
[0076] Then, three data enhancement strategies were used to improve the robustness and generalization of the model. The data enhancement strategies mainly include: baseline mutation, duration mutation and noise addition; the original signal of nanopore sequencing and the signal visualization after data enhancement are shown in Figure 2. Figure 4 shown.
[0077] 2. Multi-layer feature fusion:
[0078] After preprocessing, the signal data first passes through a multi-layer convolution module, which mainly includes: a one-dimensional convolution layer, which is used to extract local features in the time series signal; a batch normalization layer, which is used to stabilize the training process and improve the training efficiency of the model; an average pooling layer, which is used to downsample the signal;
[0079] The convolution module is calculated as x = AvgPool (ELU (BatchNorm (Conv1d (X, W)))) ,
[0080] Where X and x represent the input signal and output signal respectively, W represents the learnable weight of the convolutional layer, Conv1d represents the one-dimensional convolution operation, BatchNorm represents batch normalization, ELU is the activation function, and AvgPool represents the average pooling operation;
[0081] Then, through residual downsampling of multiple convolution modules, the low-level features and intermediate representations of the signal are retained, and the hierarchical patterns of the signal are captured; the features of convolution modules at different levels and residual downsampling are fused to obtain signal information with different feature resolutions; the fused features contain fine temporal features and high-level temporal features from different scales, which can better characterize the multi-level structural information of the input data.
[0082] 3. Transformer captures long-term dependencies:
[0083] First, receive the output of the multi-layer feature fusion module, denoted as F fusion , for F fusion Add position encoding to obtain F0 to maintain the order information of the time series.
[0084] The calculation method is F0 = F fusion +P,
[0085] Where P is the position encoding matrix, which projects F0 to generate the query vector Q, key vector K and value vector V.
[0086] Q=F0W Q ,K=F0W K ,V=F0W V ,
[0087] Where W Q ,W K ,W V is the weight matrix of linear projection;
[0088] Use the self-attention mechanism to calculate weights and outputs,
[0089] The calculation method is:
[0090] Where, d k represents the dimension of the key vector K;
[0091] In order to improve computational efficiency and model expressiveness, multiple self-attention patterns are calculated in parallel through a multi-head attention mechanism.
[0092] The output of each attention head is expressed as
[0093] Where i is the attention head number, is the weight matrix of the i-th attention head;
[0094] The outputs of all attention heads are concatenated and linearly transformed to obtain the final output, which is calculated as MHA(Q,K,V)=Concat(head1,…,head h )W o ,
[0095] Where h represents the number of attention heads, W o Represents the linear transformation matrix after splicing;
[0096] The output of the multi-head attention mechanism is added to the input through residual connections to stabilize model training, and layer normalization operations are applied to normalize feature distribution;
[0097] The output of the attention mechanism is further processed using a feedforward neural network;
[0098] The feedforward network consists of two fully connected layers and activation functions to further enhance the feature expression capability;
[0099] Residual connections and layer normalization are also applied to the output of the feed-forward network to improve training stability.
[0100] 4. Temporal convolutional networks capture temporal features:
[0101] Perform causal convolution on the signal features after the Transformer module to ensure that the prediction at each moment depends only on the current and previous input data to avoid information leakage;
[0102] The dilated convolution expands the receptive field by introducing gaps between consecutive filter positions in the convolution kernel, capturing long-distance temporal dependencies, allowing the model to capture dependencies over longer time spans without increasing computational costs;
[0103] The calculation method of the dilated convolution operation is
[0104] In the formula, y t is the output at time step t, x t-r·i is the input of time step tr·i, r is the expansion rate, w i is the convolution weight, k is the convolution kernel size;
[0105] The expansion rates of different layers in the temporal convolution module are set to exponential growth to achieve multi-level receptive field expansion.
[0106] 5. Classification barcode signal:
[0107] The features finally extracted by the temporal convolutional network are flattened and input into a multi-layer fully connected network for classification. After each layer of linear transformation, the ELU activation function is applied to introduce nonlinear features. The output of the last layer is converted into a category probability distribution through softmax, thereby achieving accurate classification of the barcode. The barcode demultiplexing model of the present invention is compared with QuipuNet, DeepPlexiCon and Deepbinner on two data sets. The test results are as follows Figure 5 As shown; it can be seen from the figure that the accuracy, recall rate, F1 score and Kappa coefficient of the barcode demultiplexing model of the present invention on the two data sets are higher than those of the other three models, indicating that the present invention can demultiplex barcodes more accurately.
[0108] 6. Visualize the model attention map and signal time correlation map:
[0109] Extract the attention weight matrix A∈R from the last encoder layer of Transformer T×T , where T is the number of time steps after the signal is fused with multiple layers of features, and each element A of the matrix ij Represents the attention weight of the i-th time step to the j-th time step; calculate each column A j The total attention value, that is, the influence of each time step on the entire sequence, is obtained by summing each column of the attention weight matrix.
[0110] The calculation method is
[0111] In the formula, Attention Value j Represents the attention value of the jth time step; the attention value of the signal time step j Mapped to the dimensional space T0 of the original signal by linear interpolation,
[0112] The calculation method is
[0113]
[0114] In the formula, Represents the signal attention value after linear interpolation, LinearInterp represents the linear interpolation function, which maps the attention of each time step to the original signal dimension; the weight vector after linear interpolation Visualization is used to analyze and interpret the model’s attention distribution. Visualization can intuitively show the impact of different time steps in the signal on the final barcode classification, helping researchers understand the degree to which the model pays attention to important features in the signal.
[0115] The visualization results of the attention map and signal time correlation map of the original signals of the four types of barcodes in Dataset I are as follows: Figure 6 As shown in the figure, the high-attention time points identified by Transformer are related to key areas in the signal, such as a sudden drop or rise in current, and these features may come from specific nucleotide combinations, DNA or RNA base conversion, nucleic acid fragment binding or unbinding events, etc., while low-attention areas represent redundant or less effective signals, such as background noise or stable signals from repetitive molecular regions; in addition, the signal time correlation diagram shows the correlation between different time steps, which helps to distinguish noise and overlapping features, thereby accurately assigning molecular events.
Claims
1. A nanopore sequencing barcode demultiplexing method based on deep learning, characterized in that: Using a hybrid deep learning architecture, the raw nanopore sequencing signals are processed and classified. The operation steps are as follows: Step (1): extracting a barcode signal from the original nanopore sequencing signal and performing data preprocessing on the barcode signal; Step (2): The barcode signal is subjected to preliminary feature extraction through a one-dimensional convolutional neural network to fuse information from different feature resolutions; Step (3): Use the Transformer module to model the global context of the signal and capture long-range dependencies; Step (4): Use a temporal convolutional network to model the time series dependencies of the signal to preserve the temporal order of the signal; Step (5): Use the classification layer to classify the extracted feature representation into barcodes to achieve barcode demultiplexing; Step (6): Visualize the model attention map and signal time correlation map.
2. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The specific process of step (1) is as follows: Step (1.1): Use the Z-Score normalization method to normalize the barcode raw signal. This process retains the dimensional shape of the signal and ensures analysis consistency, as shown in the following formula: In the formula, x and x ′ Represent the input and output signals respectively, μ and σ are the mean and standard deviation of the signals; Step (1.2): Three data augmentation strategies are used to improve the robustness and generalization of the model.
3. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 2, characterized in that: The data enhancement strategy includes: Baseline mutation: Randomly scale the signal values at different time positions, extracting the scaling factor from a normal distribution; Duration mutation: adjust the duration of each time point with a certain probability, stretching or compressing parts of the signal to simulate time changes; Add Noise: Introduce Gaussian distributed noise into the signal to simulate random fluctuations and environmental noise.
4. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The specific process of step (2) is as follows: Step (2.1): Through a multi-layer convolution module; the convolution module is as follows: x=AvgPool(ELU(BatchNorm(Conv1d(X,W)))), Where X and x represent the input signal and output signal respectively, W represents the learnable weight of the convolutional layer, Conv1d represents the one-dimensional convolution operation, BatchNorm represents batch normalization, ELU is the activation function, and AvgPool represents the average pooling operation; Step (2.2): Retain low-level features and intermediate representations of the signal and capture the signal hierarchical pattern by using multiple convolutional modules through residual downsampling; Step (2.3): Fuse the features of convolution modules at different levels and residual downsampling to obtain signal information with different feature resolutions; the fused features contain fine temporal features and high-level temporal features from different scales, which can represent the multi-level structural information of the input data.
5. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: In step (2.1), the multi-layer convolution module includes: a one-dimensional convolution layer for extracting local features in the time series signal; a batch normalization layer for stabilizing the training process and improving the training efficiency of the model; and an average pooling layer for downsampling the signal.
6. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The specific process of step (3) is as follows: Step (3.1): Receive the output of the multi-layer feature fusion module, which represents the fusion , for F fusion Add position encoding to obtain F0 to maintain the order information of the time series, as shown below: F0=F fusion +P, Where P is the position encoding matrix; Step (3.2): Project F0 to generate query vector Q, key vector K and value vector V, as shown below: Q=F0W Q ,K=F0W K ,V=F0W V , Where W Q ,W K ,W V is the weight matrix of linear projection; The self-attention mechanism is used to calculate the weights and outputs as follows: Where, d k represents the dimension of the key vector K; Step (3.3): Add the output of the multi-head attention mechanism to the input through residual connection to stabilize the model training, and apply layer normalization operation to normalize the feature distribution; Step (3.4): Process the output of the attention mechanism using a feedforward neural network; The feedforward network consists of two fully connected layers and an activation function to enhance feature expression capabilities; Residual connections and layer normalization are also applied to the output of the feed-forward network to improve training stability.
7. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 6, characterized in that: In step (3.2), in order to improve computational efficiency and model expression ability, multiple self-attention patterns are calculated in parallel through the multi-head attention mechanism, and the output of each attention head is expressed as: Where i is the attention head number, is the weight matrix of the i-th attention head; The outputs of all attention heads are concatenated and linearly transformed to obtain the final output, as shown below: MHA(Q,K,V)=Concat(head1,…,head h )W o , Where h represents the number of attention heads, W o Represents the linear transformation matrix after splicing.
8. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The specific process of step (4) is as follows: Step (4.1): Perform causal convolution on the signal features processed by step (3) to ensure that the prediction at each moment depends only on the current and previous input data; Step (4.2): The dilated convolution expands the receptive field by introducing gaps between consecutive filter positions in the convolution kernel, capturing long-distance temporal dependencies, allowing the model to capture dependencies over longer time spans without increasing computational costs; The dilated convolution operation is as follows: In the formula, y t is the output at time step t, x t-r·i is the input of time step tr·i, r is the expansion rate, w i is the convolution weight, k is the convolution kernel size; Step (4.3): The expansion rates of different layers in the temporal convolution module are set according to exponential growth to achieve multi-level receptive field expansion.
9. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The step (5) specifically flattens the features extracted from the steps (2) to (4), inputs them into a multi-layer fully connected network for classification, and applies an ELU activation function after each layer of linear transformation to introduce nonlinear features; The output of the last layer is converted into category probability distribution through softmax to achieve accurate classification of barcodes.
10. The method for nanopore sequencing barcode demultiplexing based on deep learning according to claim 1, characterized in that: The specific process of step (6) is as follows: Step (6.1): Extract the attention weight matrix A∈R from the last encoder layer of Transformer T×T , where T is the number of time steps after the signal is fused with multiple layers of features, and each element A of the matrix ij represents the attention weight of the i-th time step to the j-th time step; Calculate each column A j The total attention value, that is, the influence of each time step on the entire sequence, is obtained by summing each column of the attention weight matrix, as shown in the following formula: In the formula, Attention Value j represents the attention value at the jth time step; Step (6.2): Attention Value of the signal time step j Mapped to the dimensional space T0 of the original signal by linear interpolation, as follows: In the formula, Represents the signal attention value after linear interpolation, LinearInterp represents the linear interpolation function, which maps the attention of each time step to the original signal dimension; The weight vector after linear interpolation Visualization is used to analyze and explain the model’s attention distribution. By visualizing the impact of different time steps in the signal on the final barcode classification, it helps researchers understand the degree to which the model pays attention to important features in the signal.
Citation Information
Patent Citations
Base sequence identification method and device and storage medium
CN111243674A
Nanopore sequence recognition network structure based on Transform
CN116312791A
Base identification method and device for nanopore sequencing based on Transform architecture and storage medium
CN118038972A
Nucleic acid constructs and related methods for nanopore readout and scalable DNA circuit reporting
US20220277814A1