A multi-sequence reconstruction method against intra-cluster noise in DNA storage
Patent Information
- Application Number
- CN202211190829.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-28
AI Technical Summary
严格的聚类算法也剔除了很多扰动序列,使得已知信息量减少,影响序列重建的效果
[0038]本申请具有鲁棒性。在有一定的比例的扰动序列的簇里,本申请提供的算法依然有着很高的序列重建成功率,增大扰动序列的比例并不会显著降低成功率,因此本申请提供的算法具有鲁棒性。
Smart Images

Figure CN115602251B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA storage technology, specifically to a method for multi-sequence reconstruction against intra-cluster noise in DNA storage. Background Technology
[0002] The exponential growth in data volume has posed enormous challenges to traditional data storage systems. Traditional storage systems such as portable hard drives, USB flash drives, and integrated circuits have gradually revealed shortcomings such as short effective storage time, data susceptibility to environmental factors, high energy consumption, and environmental pollution. Meanwhile, deoxyribonucleotides (DNA) molecules, as a medium for storing biological information, exhibit significant advantages in storage density, storage time, and energy consumption. Therefore, DNA molecules hold great promise as a highly promising storage medium for solving the problem of massive data storage in the information society.
[0003] The main process of DNA storage includes encoding information into the original DNA sequence (encoding), synthesizing the DNA sequence (synthesis), storing the DNA sequence (storage), reading the DNA sequence (sequencing), and decoding to recover the original information (decoding). Figure 1 It demonstrates the data stream stored in DNA. During synthesis and sequencing, errors can occur in the DNA sequence, such as the insertion, deletion, and substitution of bases.
[0004] Currently, the probability of an error per base in mainstream second-generation sequencing is 1%-2%. In third-generation sequencing, this error rate reaches 10%-25%. During the synthesis stage, some DNA sequences are copied hundreds or thousands of times, some may only be copied a few times, and some sequences may even be lost entirely. During the sequencing stage, because all sequences are stored disorderedly, the number of DNA sequences read from each sequence in the sequencing data is uneven. Furthermore, DNA sequence breaks and rearrangements frequently occur during the PCR-based synthesis and storage stages. Sequences with these errors are those with a large editing distance from the original DNA sequence (perturbed sequences). These errors occurring during DNA storage all lead to the failure of decoding and information recovery. Figure 2 It demonstrates the breaks and rearrangement errors in DNA molecules.
[0005] Taking next-generation sequencing as an example, after sequencing the DNA stored in the database, we read the noisy copies and perturbation sequences (DNA sequences) of each original sequence. The goal is to infer the original sequence from these reads and thus recover the information. Since all the noisy copies of the sequence are stored disorderedly in the sequencing file, it is necessary to first cluster the reads so that multiple noisy copies of a DNA strand are grouped into a cluster. The original information is then inferred from this cluster; this process is called sequence reconstruction. Currently, mainstream clustering algorithms are quite strict; if a sequence has a high error rate, it will be removed from the cluster. Perturbation sequences account for a portion of the sequencing file.
[0006] Current reconstruction algorithms are based on perfect clustering: all noisy copies from the same original sequence can be clustered relatively accurately into a single cluster, where each sequence contains only insertion, deletion, and substitution errors with a low error rate. However, in practice, perfect clustering is difficult to achieve due to perturbation sequences in sequencing files. During clustering, these perturbation sequences are often discarded due to algorithmic limitations, resulting in the loss of some original information during sequence reconstruction. Research has shown that these perturbation sequences still contain original information and can be helpful for sequence reconstruction.
[0007] The main challenges in sequence reconstruction during DNA storage currently lie in the following aspects:
[0008] (1) Current sequence reconstruction methods can only correct some insertion, deletion, and replacement errors and rely on strict clustering algorithms. Once there is one or more perturbation sequences in a cluster, it will affect the sequence reconstruction. Strict clustering algorithms also remove many perturbation sequences, which reduces the amount of known information and affects the effectiveness of sequence reconstruction.
[0009] (2) The current statistical inference method for sequence reconstruction algorithm mainly involves calculating the maximum a posteriori probability of the channel input symbol and comparing it. The computation is large, so there are problems such as slow computation speed and the number of sequences in the cluster need to be controlled to less than 10.
[0010] (3) Regarding the multiple comparison method for sequence reconstruction, since it is a position-by-position comparison, the reconstruction success rate is high on data with a low error rate and relatively neat sequences within the cluster, but the robustness is poor.
[0011] (4) Current deep learning algorithms for sequence reconstruction have multiple layers of encoders and decoders, resulting in too many layers and a large number of model parameters. Summary of the Invention
[0012] This application addresses the data reading stage of DNA storage, specifically solving the problem of inferring the original information-encoded DNA sequence from existing sequencing data. The DNA sequences in the sequencing data are the sequences of the original DNA sequence after passing through the DNA storage channel. They carry information from the original sequence, partly noise copies (sequences resulting from base insertion, deletion, and substitution errors) and partly perturbations. These sequences are numerous, unbalanced, and stored disorderly in a Fastq file.
[0013] Because the Fastq files generated from DNA sequencing contain a vast number of sequences, inferring the original information from these sequences requires clustering the file to group DNA sequences with similar characteristics into clusters. Inferring the original DNA sequence from these clusters is called sequence reconstruction. When the sequencing file contains perturbation sequences, due to limitations in clustering algorithms, these perturbation sequences may end up in different clusters. Our invention primarily addresses the problem of performing sequence reconstruction in the presence of perturbation sequences within clusters. Based on the above objectives, this application provides a method for multiple sequence reconstruction in DNA storage that combats intra-cluster noise.
[0014] To achieve the above objectives, the technical solution provided by the present invention is as follows:
[0015] A multi-sequence reconstruction method for DNA storage to combat intra-cluster noise employs the following multi-sequence reconstruction model: The multi-sequence reconstruction model is a deep learning model with an encoder-decoder architecture, capable of accepting clusters of arbitrary size. An attention module is added to the front end of the multi-sequence reconstruction model to score each sequence within the cluster, thereby increasing the weight of good sequences and suppressing the influence of bad sequences on the reconstruction result. A Conformer interactively combines local features obtained from a convolutional neural network and global features obtained from a weighting mechanism as the encoder. The decoder uses a single-layer long short-term memory network. The input of the multi-sequence reconstruction model is a cluster composed of multiple DNA sequences, and the output is the reconstructed sequence.
[0016] The attention module consists of a convolutional layer and an attention mechanism. The convolutional layer extracts one-dimensional features from each vector and inputs them into the attention mechanism. When the data is input into the multi-sequence reconstruction model, it first passes through the attention module. The attention mechanism scores sequences of different quality, thereby increasing the weight of good sequences and suppressing the influence of bad sequences. Through this module, the score of each sequence in the cluster is automatically learned, and the weighted data enters the next module.
[0017] For each input sequence x i After passing through the attention module, the output sequence is y. i :
[0018] y i =α i x i (1)
[0019] Weighted score α i The calculation formula is:
[0020] e i =v T f(Wx' i +b)+k (2)
[0021]
[0022] Where, x' i It is x i The sequence obtained after the convolutional layer transformation, f is a non-linear activation function, and the above formula performs vector mathematical operations; after the attention module, linear layers and convolutional upsampling are used to transform the sequence to the same length as the output.
[0023] The encoder consists of a single Conformer layer with four modules. The first and last modules are feedforward modules composed of two linear layers. The linear layer expands the feature space of the sequence to twice its original dimension, while the other layer restores the feature space to its original dimension to capture the complex relationships between each position in the sequence. In the second module, two depthwise separable convolutions are performed using a convolutional kernel of size 31 to capture local information between sequence positions. The inputs and outputs of each module are summed, and after layer normalization, the result is fed into the next module. The convolutional modules can correct insertion, deletion, and replacement errors by extracting rich semantic information, and the multi-head attention mechanism can extract multiple semantic information to capture dependencies of various ranges within the sequence.
[0024] For each input sequence x i After passing through the Conformer encoder module, the output sequence is y. i :
[0025]
[0026] x i =x i '+MHA(x i ')
[0027] x i ' = x i +Conv(x) i ”)
[0028]
[0029] Among them, the decoder of the single-layer long short-term memory network reduces the feature dimension of the sequence to 4 and finally outputs it through the Softmax function. The feature dimension of 4 represents the probability of the four bases at each position.
[0030] For each input sequence y i After passing through the LSTM decoder, the output sequence is z. i :
[0031] y i =LSTM(y) i )
[0032] y i =Linear(y') i )
[0033] z i =Softmax(y″) i )
[0034] This also includes data preprocessing, as detailed below:
[0035] (1.1) For n sequences x1, x2, ..., xg within a cluster, each consisting of four bases of ACGT... n The lengths of the n sequences are L1, L2, ..., L n Each sequence is one-hot encoded and becomes an L i A vector of size 4; let L' = max(L1, L2, ..., L n For vectors that are not long enough to be L', add zeros to the end to make them L'×4 vectors, so that each data input into the network is an n×L'×4 vector;
[0036] (1.2) Divide the input data into several batches for training. Sort the input data in descending or ascending order of cluster size so that clusters of the same size can be grouped into a batch. Therefore, the data input into the network to update the parameters once is a vector of b×n×L'×4, where B is the batch size.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] This application demonstrates robustness. Even in clusters with a certain proportion of perturbation sequences, the algorithm provided in this application still maintains a high sequence reconstruction success rate. Increasing the proportion of perturbation sequences does not significantly reduce the success rate; therefore, the algorithm provided in this application is robust.
[0039] The algorithm in this application achieves a high success rate in sequence reconstruction. The model achieves a reconstruction success rate of over 99.5% on the first two datasets; and a reconstruction success rate of over 96.4% on the third dataset, which is noisier and less regular, far exceeding the results of the comparison algorithm. Therefore, the algorithm has a high success rate in sequence reconstruction.
[0040] The algorithm in this application does not impose any restrictions on the data. Data of any encoding form can be reconstructed using the algorithm in this application. The algorithm makes no assumptions about the data, and clusters of any size and length can be input into the network. Therefore, the algorithm provided in this application does not impose any restrictions on the data.
[0041] The algorithm presented in this application is highly user-friendly. The encoder layer of the model consists of a single Conformer block, and the decoder layer consists of a single-layer LSTM, with only 2.5M parameters, far fewer than current deep learning-based sequence reconstruction algorithms that stack multiple layers, making it easy to train and test. Due to the small number of parameters, the requirements for computer hardware and systems are extremely low. The algorithm is based on the deep learning framework PyTorch; if speed is a priority, GPUs and parallel computing can be used for acceleration, thus making the algorithm highly user-friendly. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of DNA storage data flow in existing technologies;
[0043] Figure 2 This is a schematic diagram illustrating the high error rate of DNA strand reading during sequencing in existing technologies.
[0044] Figure 3 A schematic diagram of the sequence reconstruction model provided in this application;
[0045] Figure 4 The sequence reconstruction model architecture diagram provided in this application;
[0046] Figure 5 A schematic diagram of the attention mechanism provided in this application. Detailed Implementation
[0047] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0048] This application proposes a novel artificial neural network to address sequence reconstruction problems with numerous perturbations, and is the first sequence reconstruction algorithm to consider this scenario. It uses an attention mechanism to score each sequence within a cluster and assign weights to each sequence. The sequence reconstruction algorithm provided in this application has low dependence on the accuracy of the clustering algorithm; even with a coarse clustering algorithm that results in low accuracy, the algorithm provided in this application can still guarantee a high success rate. Furthermore, the multi-sequence reconstruction model provided in this application can accommodate clusters of different sizes, aiming to fully utilize known sequence information to improve the accuracy of sequence reconstruction. The model has a small number of parameters, facilitating training and testing. Details are as follows:
[0049] Sequence Reconstruction Problem Description
[0050] After X∈D passes through the DNA storage channel, it generates t noisy copies Y=(Y1,Y2,…,Y…). t )∈D t , Gathered together, here Z is a set of perturbation sequences, including noisy copies of other original sequences generated through the DNA storage channel, artificially added perturbation sequences, DNA strand breaks with high error rates mentioned above, and DNA rearrangement sequences. A common characteristic of perturbation sequences is their large edit distance from the original sequence. Sequence reconstruction involves reconstructing X from Y'+Z', which is equivalent to finding a mapping f', Y'+Z'∈D. p+q Under the influence of mapping f', it becomes Make X to The distance is the smallest. Figure 3 The sequence reconstruction model of this application is described.
[0051] Model Architecture
[0052] This application proposes a robust multi-sequence reconstruction model for situations with numerous perturbation sequences. The model is a deep learning model with an encoder-decoder architecture, capable of accepting clusters of arbitrary size. An attention mechanism is added to the front end of the model to score each sequence within a cluster, thereby increasing the weight of good sequences and suppressing the influence of bad sequences on the reconstruction results. The Conformer interactively combines local features obtained from a convolutional neural network and global features obtained from a weighted mechanism as the encoder. The decoder uses a single-layer Long Short-Term Memory (LSTM) network. Figure 4 The framework of the model is described, with arrows indicating the flow of data. The model's input is a cluster of multiple DNA sequences, and the output is the reconstructed sequence.
[0053] The model is described in detail below based on the flow of data.
[0054] Data preprocessing
[0055] 1. For n sequences x1, x2, ..., xg within a cluster, each consisting of the four bases ACGT... n Their lengths are L1, L2, ..., L n Each sequence is one-hot encoded and becomes an L i A vector of size 4. Let L' = max(L1, L2, ..., L...). n For vectors that are not long enough to be L', add zeros to the end to make them L'×4 vectors. In this way, each data input into the network is an n×L'×4 vector.
[0056] 2. The input data is divided into several batches for training, but the cluster sizes received by the model are different. Therefore, the input data is sorted according to the cluster size from largest to smallest or smallest to largest, so that clusters of the same size can be grouped into a batch. The model can receive clusters composed of sequences of different sizes and lengths. Therefore, the data input into the network to update the parameters once is a vector of size B×n×L'×4, where B is the batch size.
[0057] Attention module
[0058] The attention module consists of a convolutional layer and an attention mechanism. The convolutional layer extracts one-dimensional features from each vector and feeds them into the attention mechanism. Data input into the model first passes through the attention module. The attention mechanism scores sequences of different quality, thereby increasing the weight of good sequences and suppressing the influence of bad sequences. Through this module, the score of each sequence in the cluster is automatically learned. The weighted data then enters the next module.
[0059] For each input sequence x i After passing through the attention module, the output sequence is y. i :
[0060] y i =α i x i (1)
[0061] Weighted score α i The calculation formula is:
[0062] e i =v T f(Wx' i +b)+k (2)
[0063]
[0064] Where, x' i It is x i The sequence obtained after the convolutional layer transformation, f is a non-linear activation function, and the above formula performs vector mathematical operations. Figure 5 It demonstrates the attention mechanism.
[0065] Following the attention module, linear layers and convolutional upsampling are used to transform the sequence to the same length as the output. Since the existing feature space dimension is insufficient to characterize the sequence's features, the output of this part has a larger feature space. In the experiments, the feature dimension of the original sequence was expanded from 4 to 128.
[0066] Conformer Encoder Module
[0067] The encoder consists of a single Conformer layer with four modules. The first and last modules are feedforward modules composed of two linear layers. The linear layer expands the feature space of the sequence to twice its original dimension, while the second layer restores the feature space to its original dimension. This is to capture the complex relationships between each position in the sequence. In the second module, two depthwise separable convolutions are performed using a convolutional kernel of size 31 to capture local information between sequence positions. The input and output of each module are summed, and the normalized layers are fed into the next module. The convolutional modules can correct insertion, deletion, and replacement errors by extracting rich semantic information, and the multi-head attention mechanism can extract multiple semantic information to capture dependencies of various ranges within the sequence. The Conformer combines the local features obtained by the convolutional neural network with the global features obtained by the attention mechanism. This part combines... Figure 4 Let's take a look.
[0068] For each input sequence x i After passing through the Conformer encoder module, the output sequence is y. i :
[0069]
[0070] x i =x i '+MHA(x i ')
[0071] x i ' = x i +Conv(x) i ”)
[0072]
[0073] LSTM decoder module
[0074] Due to the powerful feature extraction capabilities of Conformer, complex decoders are not required. Furthermore, deep neural networks perform poorly when dealing with sequence problems and time-dependent issues. However, the bases in DNA sequences stored in DNA storage have contextual dependencies and can be viewed as sequences in the time dimension. LSTM, as a variant of recurrent neural networks, can effectively solve the problem of short-term dependencies in long sequences. Therefore, this application uses a single LSTM layer as the decoder. Finally, a linear layer reduces the feature dimension of the sequence to 4, and the output is finally passed through a Softmax function. A feature dimension of 4 represents the probability of the four bases at each position. This part combines... Figure 4 Let's take a look.
[0075] For each input sequence y iAfter passing through the LSTM decoder, the output sequence is z. i :
[0076] y i =LSTM(y) i )
[0077] y i =Linear(y') i )
[0078] z i =Softmax(y″) i ).
[0079] Experiments and Results
[0080] The model in this application requires an input-label pairing dataset. The dataset is divided into a training set and a test set. The training set is used to train the neural network, and the test set is used to evaluate the model by comparing the labels and the model's output.
[0081] The data formats for both the training and test sets are:
[0082] {Clusters of clustered DNA sequences (input), original DNA sequences (tags)}
[0083] This application uses three real-world next-generation sequencing datasets to simulate the performance of the proposed method. The data are from datasets published in papers by Erlich et al., Organic et al., and Chandak et al. Sequencing files and raw sequences of these datasets are stored separately in an unordered manner, and the alignment results of the Burrows-Wheeler alignment tool (BWA) are used as perfect clustering results.
[0084] Dataset 1: Erlich and Zielinski proposed fountain codes, which encode binary files into DNA sequences and store them within DNA. They synthesized 72,000 152-base sequences using Twist Bioscience technology and sequenced them using Illumina Miseq V4 technology. The reads used in the experiments were from the paired-end sequencing file ERR1816980. Paired-end sequences were assembled using the sequence assembly software Paired-EndreAdmergeR (PEAR) before sequence alignment.
[0085] Dataset 2: Organick et al. synthesized 35 independent data files into DNA sequences and stored them, carefully designing specific primers to tag the address of each file on the DNA sequence. They released two files: a Fastq format file containing sequencing data named id20 and 607,050 corresponding raw sequences, named id20.refs. These two datasets will be used for experiments.
[0086] Dataset 3: Chandak et al. proposed a practical and efficient scheme to balance the read and write costs in DNA storage. This scheme uses LDPC codes to encode binary information into DNA sequences for error correction and handling of missing sequences, and biological experiments were conducted. The first dataset used in the experiment was selected, which uses LDPC codes with 50% redundancy and contains 11,710 DNA sequences. The original sequence length is 150, and the read length is 151.
[0087] The success rate of sequence reconstruction is used to evaluate the model's performance. The calculation formula is as follows:
[0088]
[0089] The denominator is the total number of original sequences input into the model, and the numerator is the number of successfully predicted sequences. That is, if the output is identical to every position of the original sequence, then a sequence is successfully reconstructed.
[0090] In the experiment, the ratio of training to test sets was set to 1:1. It was also necessary to control the cluster size range for each dataset. Based on the BWA alignment results, it was observed that each original sequence in the sequencing files typically had more than 5 noisy copies; therefore, the minimum cluster size was set to 5. Larger clusters result in fewer clusters, and it was desirable for the cluster size to be average across all datasets; therefore, the maximum cluster size should not be too large. 30 was observed to be a suitable upper bound. Therefore, based on the BWA alignment results, 5-30 sequences were randomly selected from each cluster.
[0091] Perturbation sequences are injected into the sequencing file to simulate sequence reconstruction algorithms with a large number of perturbation sequences. They are added in a 1:1:1:1 ratio, and come from four sources:
[0092] 1. Sequences from other clusters. When perturbed sequences are present in the sequencing file, it is difficult to achieve perfect clustering. Inevitably, some sequences from other clusters will be clustered into this cluster.
[0093] 2. Reverse complementary strands in clusters. When double-stranded DNA sequences are sequenced, if errors introduced during sequencing are disregarded, the two strands are expected to be reverse complementary. However, in actual sequencing, the two strands are first separated, and typically only one is randomly selected for sequencing. Here, we consider the reverse complementary sequence as a perturbation sequence with a significant edit distance from the original sequence. Because reverse complementary sequences constitute a large proportion of the sequence, if we consider it a perturbation sequence, it will definitely be present in the resulting clusters.
[0094] 3. Randomly generated base sequences. One purpose of adding these perturbation sequences is to increase noise, because DNA sequences may be lost during storage. The loss of some sequences means a lower proportion of non-perturbation sequences in the sequencing file, while the proportion of perturbation sequences increases, thus increasing noise within the clusters. Another important source of these perturbation sequences is security considerations. From an information protection perspective, "false information" can be added to the original DNA sequence file to prevent attackers from successfully decoding it; from an attacker's perspective, "false information" can be added to the stored file, making it difficult to decode and recover the information. Randomly generated base sequences can simulate this "false information."
[0095] 4. Sequences composed of multiple DNA sequences spliced together within a cluster, as well as sequences composed of sequence fragments within a cluster and randomly generated sequences. These perturbation sequences simulate DNA breaks and rearrangement errors. Figure 2 This demonstrates the error.
[0096] Training and testing were performed on a single 2080ti GPU. The deep learning framework used was PyTorch, with a batch size of 64, an initial learning rate of 0.005, and cross-entropy as the loss function. The Adam optimizer was used to optimize the model during training. The L2 regularization coefficient was 1e-4 to prevent overfitting.
[0097] Table 1: Comparison of experimental results with different proportions of perturbation sequences
[0098]
[0099] Table 1 describes the success rate of the multi-read reconstruction algorithm on three datasets. It can be seen that the results on the third dataset are significantly lower than the first two datasets. This is due to the higher insertion, deletion, and replacement error rates on the third dataset, and the fact that the length of the input sequence is always one unit longer than the original sequence. The experiment also shows that gradually increasing the proportion of perturbation sequences leads to a gradual decrease in success rate, but the decrease is not significant. If the proportion of perturbation sequences continues to increase, the algorithm in this application should also perform well.
[0100] Table 2: Comparative Experiments
[0101]
[0102] Table 2 shows the comparison results of the algorithm of this application with four other algorithms: Iterative, BMA divider, Hybrid, and BMA Lookahead. When there is no perturbation sequence, the algorithm of this application performs slightly worse than Iterative and BMA Lookahead. However, as the proportion of perturbation sequence gradually increases, the advantage of the algorithm of this application becomes apparent.
[0103] Table 3: Ablation Experiment
[0104]
[0105] To verify the necessity and effectiveness of the "attention module" and "Conformer-encoder module" in the proposed network model, a series of ablation experiments were conducted. These included removing the attention mechanism before the encoder-decoder phase, removing the attention mechanism and normalizing all input sequences before summing them together, and replacing the Conformer with a Transformer (the Transformer has a similar structure to the Conformer, but with an additional convolutional module and a feedforward module). The reconstruction success rate of the model was evaluated under different datasets and different perturbation sequence ratios. Table 3 describes the experimental results. These results demonstrate that the attention mechanism in this application is very useful and can reduce the impact of perturbation sequences on the reconstruction results. The Conformer is also very useful and irreplaceable.
[0106] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not describe the various possible combinations separately.
Claims
1. A method for reconstructing multiple sequences against intra-cluster noise in DNA storage, characterized in that, The following multi-sequence reconstruction model is adopted. The multi-sequence reconstruction model is a deep learning model with an encoder-decoder architecture, which can accept clusters of arbitrary size. An attention module is added to the front end of the multi-sequence reconstruction model to score each sequence within the cluster, thereby increasing the weight of good sequences and suppressing the influence of bad sequences on the reconstruction results. The Conformer interactively combines the local features obtained by the convolutional neural network and the global features obtained by the weighting mechanism as the encoder. The decoder uses a single-layer long short-term memory network; the input of the multiple sequence reconstruction model is a cluster of multiple DNA sequences, and the output is the reconstructed sequence. The attention module consists of a convolutional layer and an attention mechanism. The convolutional layer extracts one-dimensional features from each vector and inputs them into the attention mechanism. When the data is input into the multi-sequence reconstruction model, it first passes through the attention module. The attention mechanism of the attention module scores sequences of different quality, thereby increasing the weight of good sequences and suppressing the influence of bad sequences. Through this module, the score of each sequence in the cluster is automatically learned, and the first sequence enters the next module. For each input sequence After passing through the attention module, the first output sequence is: : (1) Weighted score The calculation formula is: (2) (3) in, yes The sequence obtained after convolutional layer transformation f It is a non-linear activation function, and the above formula performs mathematical operations on vectors; after the attention module, linear layers and convolutional upsampling are used to transform the sequence to the same length as the output; The encoder consists of a single Conformer layer with four modules. The first and last modules are feedforward modules composed of two linear layers. The linear layer expands the feature space of the sequence to twice its original dimension, while the other layer restores the feature space to its original dimension to capture the complex relationships between each position in the sequence. In the second module, two depthwise separable convolutions are performed using a convolutional kernel of size 31 to capture local information between sequence positions. The inputs and outputs of each module are summed, and after layer normalization, the result is fed into the next module. The convolutional modules can correct insertion, deletion, and replacement errors by extracting rich semantic information. The multi-head attention mechanism inside the Conformer can extract multiple semantic information to capture dependencies of various ranges within the sequence. For each first sequence After passing through the Conformer encoder module, the second sequence is output as follows: : ; ; ; 。 2. The method for reconstructing multiple sequences against intra-cluster noise in DNA storage according to claim 1, characterized in that, The decoder of a single-layer long short-term memory network reduces the feature dimension of the sequence to 4 and finally outputs it through the Softmax function. The feature dimension of 4 represents the four-dimensional value output at each base position, which, after normalization, represents the probability of the four bases. For each second sequence After passing through the LSTM decoder, the third output sequence is: : ; ; 。 3. The method for reconstructing multiple sequences against intra-cluster noise in DNA storage according to claim 1, characterized in that, This also includes data preprocessing, as detailed below: (1.1) For n sequences within a cluster consisting of the four bases ACGT The lengths of the n sequences are respectively Each sequence is one-hot encoded and becomes a... The vector; let For length not enough The vector is padded with zeros to make it... A vector, such that each data input into the network is a... ; (1.2) Divide the input data into several batches for training. Sort the input data according to the cluster size from largest to smallest or smallest to largest, so that clusters of the same size are grouped into one batch. Therefore, the data input into the network to update the parameters once is... The vector, where, It refers to the batch size.
Citation Information
Patent Citations
Chinese text abstract generation method based on sequence-to-sequence model
CN111078866A
Depth decoupling time sequence prediction method
CN113177633A