Nanopore real-time high-throughput sequencing system and method based on DNA molecule passage

By using a nanopore real-time high-throughput sequencing system based on DNA molecular communication, combined with Watson-Crick base pairing and LT code encoding, and employing a deduplication algorithm, the redundant error problem caused by hopper bounce in nanopore sequencing was solved, achieving efficient and low-cost DNA sequencing error correction and communication.

CN115831231BActive Publication Date: 2026-04-24CHENGDU QUNZHI WEINA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU QUNZHI WEINA TECH CO LTD
Filing Date
2022-11-04
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing third-generation DNA sequencing technologies, nanopore sequencing methods suffer from redundant errors caused by molecular hopper bounce. Existing error correction methods are costly and inefficient, making it difficult to achieve high-throughput real-time error correction.

Method used

A real-time high-throughput sequencing system based on DNA molecular passage is adopted, which combines a DNA encoding module, a molecular hopper loading module, a nanopore sequencing module, and a DNA decoding module. Watson-Crick base pairing and LT code encoding are used, and redundant errors caused by hopper bounce are eliminated through a deduplication algorithm.

Benefits of technology

It achieves efficient and low-cost DNA sequencing error correction, reduces communication costs, improves the signal-to-noise ratio, and is suitable for DNA molecular communication applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831231B_ABST
    Figure CN115831231B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on DNA molecule traffic's nanopore real-time high-throughput sequencing system and method, including DNA encoding module, molecular hopper loads DNA sequence chain module, nanopore sequencing module and DNA decoding module;The transmission data required by user is stored in DNA sequence by base pairing;The DNA sequence obtained by encoding is linked with molecular hopper, is transported on the track of nanopore protein, and molecular motor of DNA chain transmission is transmitted on track;Real-time sequencing is completed by the different base pair blocking current different ability;By DNA decoding module, redundancy error caused by molecular hopper in the track of nanopore protein is eliminated.The system provided by the application solves the repeated data deletion error LT real-time error correction algorithm, to eliminate the redundant information generated by hopper bounce back;Eliminate the redundancy caused by hopper back, so as to reduce the high communication cost brought by this kind of error code, obtain greater signal-to-noise ratio, facilitate the application of DNA molecular communication to be popularized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of DNA molecular communication, DNA sequencing, and encoding / decoding, and specifically to a nanopore sequencing method based on a molecular hopper and its error correction algorithms for sequencing errors such as insertion, duplication, and deletion. Background Technology

[0002] Over the past three years since the COVID-19 pandemic spread globally, the demand for rapid pathogen detection and high-throughput DNA / RNA sequencing has grown exponentially. Simultaneously, in this era of information explosion, current storage technologies, due to physical limitations, cannot provide long-term storage, and improvements in storage performance come at the cost of high energy consumption. Due to the high stability and high information density of the DNA double helix structure, DNA storage can achieve high-information-capacity, ultra-long-term, and ultra-low-energy-consumption storage methods, surpassing current silicon-based magnetic storage methods and offering a revolutionary approach to bio-inspired computing and communication storage that is biocompatible and environmentally friendly. The rapid development of biotechnology such as PCR has enabled us to manipulate DNA molecules on a large scale in a more cost-effective manner, enabling large-scale DNA sequencing and DNA editing applications. Achieving high-throughput real-time DNA sequencing is a key challenge in these business-changing applications.

[0003] Currently, there are three main methods for DNA molecular sequencing:

[0004] First-generation sequencing methods, primarily Sanger sequencing, mainly employ the dideoxy chain termination method, based on the principle of DNA replication. Its core technology utilizes dideoxynucleotide triphosphates (ddNTPs). Lacking a 3'-OH group, these ddNTPs cannot form a phosphodiester bond with another deoxyribonucleoside triphosphate (dNTP). These dideoxynucleotide triphosphates (ddNTPs) can be used to terminate DNA chain elongation. Furthermore, these dideoxynucleotide triphosphates (ddNTPs) are attached with radioactive isotopes or fluorescent labeling groups, allowing them to be detected by automated instruments or gel imaging systems. This method has low sequencing throughput and is expensive, making it unsuitable for large-scale molecular communication.

[0005] Second-generation sequencing methods, primarily the Illumina sequencing platform, utilize cloning single-molecule array technology. First, the target DNA fragment is broken into 100-200 bp pieces and randomly ligated onto a solid matrix. After Bst polymerase extension and formate denaturation via bridged polymerase chain reaction (PCR) cycles, a large number of DNA clusters are generated. Subsequent reactions are similar to the Sanger method. Illumina's main drawback is its short sequencing length; error rates increase significantly beyond 100 bp. Furthermore, when performing short sequence reads and assembling genomes, large repetitive fragments become problematic.

[0006] Third-generation sequencing (NGS) technology refers to single-molecule sequencing, which eliminates the need for polymerase chain reaction (PCR) amplification during DNA sequencing, enabling the individual sequencing of each DNA molecule. The most popular NGS method is nanopore sequencing. It employs electrophoresis, using electrophoresis to drive individual molecules one by one through a nanopore for sequencing. Because the diameter of the nanopore is extremely small, only a single nucleic acid polymer can pass through. Since the four types of nucleotides have different spatial conformations, the current changes they cause as they pass through the nanopore differ. By detecting the change in the intensity of the current passing through the nanopore when a DNA or RNA chain composed of multiple nucleotides passes through it, the type of nucleotide that has passed through can be determined, allowing for real-time sequencing. Its main advantages are: single-molecule sequencing, long read lengths exceeding 150kb, high sequencing speed, real-time monitoring of sequencing data, and lightweight, portable sequencing equipment. NGS methods utilize a DNA molecular communication system architecture to achieve high-throughput, real-time error correction sequencing. A molecular hopper mechanism based on protein orbitals enables the DNA chain to move stably and rapidly along molecular orbitals, controlling the stable passage of the DNA chain through the nanopore to complete the sequencing process. The molecular hoppers commonly used in experiments contain kinesin and dynein. Through a series of thiol disulfide exchange reactions, the forces generated by the chemical reaction and the driving force of the external electric field jointly propel the DNA strand. The DNA sequence carrying information moves directionally along the internal tracks of the nanopore for hundreds of steps without detaching from the channel. The overall direction of movement of the DNA strand is determined by the strength of the externally applied electric field, while individual molecular hoppers move along tracks within the nanopore controlled by chemical ratchet mechanisms. However, even with a fixed potential applied, due to the conformational instability of the hoppers within the nanopore, a single base site may bounce back, causing a reread of the already sequenced base sequence. This results in similar repetitive segments in the externally measured ion current, leading to redundancy during decoding and incorrect information during decoding. Current methods for handling errors caused by molecular hopper bounce include: (1) multiple rereading; (2) multiple sequence alignment algorithm (MSA); (3) hybrid error correction; and (4) CD-Hit software clustering algorithm. The multiple rereading method involves changing the applied potential when the molecular hopper carrying the DNA sequence moves to the nanopore node, causing the molecular hopper to change its direction of movement and re-enter the protein orbital in the nanopore for ion current reading, thereby increasing the number of DNA sequence reads and reducing errors caused by molecular hopper bounce. However, the cost is too high and the sequencing time is too long, requiring a large number of reads to reduce the error. The multiple sequence alignment algorithm (MSA) compares the base sequences obtained from sequencing to place the same residue sites in the same column in order to find similar components between sequences that are not completely identical, forming as many columns as possible with the same characters. However, it cannot handle very similar base fragment sequences with different information. The algorithm will mistakenly identify them as the same base sequence fragment, thus causing new errors.Hybrid error correction methods receive two data sources at the receiver: long reads and short reads. If the detected errors are within the receiver's error correction capability, automatic error correction is performed. If the errors exceed the receiver's capability but can still be detected, a feedback channel requests the transmitter to retransmit the data to reduce errors. However, this method requires sufficient sequencing depth, making it too costly. CD-Hit software clustering algorithm: This is an incremental clustering method running on CD-Hit software. First, the input sequences are sorted according to their length, processed from longest to shortest. The longest sequence is automatically assigned to the first cluster and becomes its representative sequence. Then, the remaining sequences are compared with previously identified representative sequences, and based on sequence similarity, they are either assigned to one of the clusters or become the representative sequence of a new cluster. This process is repeated for all sequences to complete the clustering process. In the default mode, a sequence is only compared with the representative sequence (the longest sequence in that cluster) and not with other sequences in that cluster. In accurate mode, a sequence is compared with all sequences in each cluster and then a decision is made whether to become a new cluster or be assigned to one of the existing clusters. However, it involves a large workload and is costly. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a real-time high-throughput sequencing system and method based on DNA molecule passage through nanopores, which utilizes a deduplication algorithm to reduce redundant sequences in third-generation single-molecule sequencing technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] The present invention provides a real-time high-throughput sequencing system based on DNA molecular transport using nanopores, comprising a DNA encoding module, a molecular hopper module for loading DNA sequence strands, a nanopore sequencing module, and a DNA decoding module.

[0010] The DNA encoding module is used to store the data to be transmitted by the user in the DNA sequence through base pairing;

[0011] The molecular hopper is loaded with a DNA sequence chain module, which is used to link the encoded DNA sequence with the molecular hopper for transport on the nanoporous protein orbital, and the molecular motor loaded with the DNA chain is transported on the orbital.

[0012] The nanopore sequencing module achieves real-time sequencing by detecting different currents generated due to the different abilities of different base pairs to block current.

[0013] The DNA decoding module is used to eliminate redundant errors caused by the molecular hopper bouncing back on the protein orbitals in the nanopore.

[0014] Furthermore, the base pairing adopts Watson-Crick base pairing.

[0015] Furthermore, the DNA decoding module employs a deduplication algorithm.

[0016] Furthermore, the DNA encoding module employs LT codes, and the LT code encoding is performed according to the following steps:

[0017] First, the source information is divided into k input symbols of length L; the degree d is randomly generated from the degree distribution of each output symbol; the d input symbols are uniformly selected and XORed together to output coded symbols; the number of received coded symbols n should satisfy n≥k;

[0018] Subsequently, the encoded symbols are mapped to DNA bases, following...

[0019] Finally, the corresponding oligonucleotide sequences are output.

[0020] Furthermore, the deduplication algorithm in the DNA decoding module is performed according to the following steps:

[0021] The first step is data initialization, which involves inputting the oligonucleotide sequence X that has been read. n n is the oligonucleotide sequence X n The number of bases is defined as N, the maximum number of backsteps is defined as N, and the oligonucleotide sequence X is defined as... n The single base is x n The base sequence segment that causes a bit error due to a backstepping motion in the molecular hopper is defined as S. N+1 S N+1 ∈X n Define a redundant sequence Y;

[0022] The second step involves reversing backstepping during DNA sequence reading, with a maximum reversal count defined as N. If a reversal occurs in the molecular hopper, the already read oligonucleotide base x is... n Assigned to base sequence segment S N+1 In step (i), an error correction loop is performed. If an error correction sequence S is detected... N+1 The i-th base S in (i) N+1 (i) with the (i+2Nth)th base S N+1 If (i+2N) are equal, then the read oligonucleotide sequence X will be... n The value is assigned to the redundant Y, and the base sequence segment S that caused the bit error is also assigned. N+1 The (i+2)th base S in N+1 (i+2N) Read sequence X from the original DNA N Remove from the middle, and then read the DNA sequence X. nThe number of base sequences n in the sequence is updated, and the DNA read sequence X is updated simultaneously. n Update the number of backtracks N = N-1 until the maximum number of consecutive backtracks N = 0, then end the error correction process, thus achieving the goal of removing duplicate data.

[0023] The present invention provides a real-time high-throughput sequencing method based on nanopores for DNA molecule transport, comprising the following steps:

[0024] Base S was obtained through DNA nanopore sequencing. N+1 (i) Sequence;

[0025] Set the maximum number of consecutive backsteps N, and introduce a redundant sequence Y;

[0026] base S N+1 (i) Assign sequence value to X n ;

[0027] Determine the base S N+1 (i) Sequence and base S N+1 If the (i+2N) sequences are identical, then the base sequence is considered redundant due to the hopper bounce, and the redundant base is added to the redundant sequence Y; the base S is... N+1 (i+2N) from sequence X n Remove from the sequence to achieve deduplication, and update the sequence X after removing redundant bases. n If no, proceed to the next step; otherwise, proceed to the next step.

[0028] Update the number of backslides N = N-1 until the maximum number of consecutive backslides N = 0.

[0029] Furthermore, the method of obtaining base S through DNA nanopore sequencing N+1 (i) The sequence is performed according to the following steps:

[0030] The data that the user needs to transmit is stored in the DNA sequence through base pairing;

[0031] The encoded DNA sequence is linked to a molecular hopper for transport on nanoporous protein orbitals, and molecular motors carrying DNA strands are transported on these orbitals.

[0032] Different base pairs have different abilities to block current, resulting in different current detections, thus enabling real-time sequencing.

[0033] Furthermore, the base pairing adopts Watson-Crick base pairing.

[0034] Furthermore, the storage of the user's required transmission data in the DNA sequence via base pairing is achieved using LT codes, and the LT code encoding is performed according to the following steps:

[0035] First, the source information is divided into k input symbols of length L; the degree d is randomly generated from the degree distribution of each output symbol; the d input symbols are uniformly selected and XORed together to output coded symbols; the number of received coded symbols n should satisfy n≥k;

[0036] Subsequently, the encoded symbols are mapped to DNA bases, following...

[0037] Finally, the corresponding oligonucleotide sequences are output.

[0038] The beneficial effects of this invention are as follows:

[0039] This invention provides a nanopore real-time high-throughput sequencing system and method based on DNA molecular communication, targeting third-generation DNA sequencing methods and inspired by cutting-edge DNA molecular communication systems. The transmitter carries a DNA strand encoded according to the source, and a channel is constructed using a molecular hopper mechanism with orbitals to achieve stable and controllable high-speed transmission. The DNA sequence to be transmitted undergoes Luby Transform (LT) channel error correction encoding. The receiver uses a nanopore with potentials applied at both ends. As the DNA strand passes through the nanopore, current changes caused by varying degrees of current blocking at different bases are observed. These changes are converted into sequence information through real-time information processing, thereby achieving high-throughput DNA sequencing. Simultaneously, this invention provides an LT real-time error correction algorithm to address deduplication errors and eliminate redundant information caused by hopper bounce.

[0040] The beneficial effects of this invention are as follows:

[0041] This invention provides a real-time high-throughput sequencing system and method based on DNA molecular communication via nanopores. This system can be viewed as a molecular communication process. The starting point of the track carrying the encoded DNA strand is the transmitter, the information-carrying DNA strand is the messenger molecule, the molecular motor carrying the DNA strand on the track acts as a directional communication channel, and the terminal passing through the nanopore is considered the receiver. Real-time sequencing is achieved by detecting the current caused by the different base blocking current capabilities during the receiving process. This completes a mapping from DNA sequencing to molecular communication. This system provides an alternative paradigm to diffusion communication and offers high-speed, controllable communication. Furthermore, because the molecular motor may move back and forth on the track, duplicate and missed detections may occur during the nanopore perforation process, causing bit errors. A deduplication algorithm developed to reduce redundancy is in effect, eliminating redundancy caused by hopper retraction, thereby reducing the high communication costs caused by these bit errors, achieving a higher signal-to-noise ratio, and facilitating the widespread application of DNA molecular communication.

[0042] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0043] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:

[0044] Figure 1 This is a schematic diagram of a communication system for DNA nanopore sequencing. The DNA sequence strand passes through the nanopore under the influence of an electric potential, making single-step jumps on the protein orbitals.

[0045] Figure 2 The chemical structure of the molecular hopper.

[0046] Figure 3 A scheme for a molecular hopper communication system with a track.

[0047] Figure 4 This is a state model of molecular hoppers propagating along their orbits.

[0048] Figure 5 The repeated sequence was detected due to the back jump of the molecular hopper.

[0049] Figure 6 The pseudocode for the deduplication algorithm.

[0050] Figure 7 This is a flowchart of the deduplication algorithm.

[0051] Figure 8 This study aims to simulate and verify the proposed deduplication algorithm based on Matlab.

[0052] Figure 9 This section compares the best bit error rate performance based on Matlab and the best bit error rate performance based on LT coding.

[0053] Figure 10 For different redundancy levels and p based on Matlab j Comparison of bit error rates of encoding schemes.

[0054] In the diagram: 1. Cysteine, 2. Protein orbital, 3. Molecular hopper, 4. Disulfide bond, 5. Base sequence, α represents the rate at which the molecular hopper changes state to forward movement, β represents the rate at which the molecular hopper changes state to backward movement, and v represents the average movement rate of the molecular hopper on the protein orbital. Detailed Implementation

[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0056] Example 1

[0057] like Figure 1 As shown, the nanopore real-time high-throughput sequencing system and method based on DNA molecule passage provided in this embodiment aims to eliminate redundant information caused by the molecular hopper bouncing back on the protein orbitals of the nanopore. To achieve the above objective, this embodiment provides the following technical solution.

[0058] The nanopore real-time high-throughput sequencing system based on DNA molecular communication provided in this embodiment includes a DNA encoding module, a molecular hopper module for loading DNA sequence strands, a nanopore sequencing module, and a DNA decoding module.

[0059] The DNA encoding module is used to store the data to be transmitted by the user in the DNA sequence through Watson-Crick base pairing;

[0060] The molecular hopper is loaded with a DNA sequence chain module, which is used to link the encoded DNA sequence with the molecular hopper, so as to facilitate transport and sequencing on the nanoporous protein track. It can be regarded as the transmitter module in the communication system. The molecular motor loaded with DNA chain transmits on the track through a directional communication channel.

[0061] The nanopore sequencing module achieves real-time sequencing by detecting different currents generated due to the different abilities of different base pairs to block current.

[0062] The DNA decoding module is used to reduce redundant errors caused by the molecular hopper bouncing back on protein orbitals in the nanopore.

[0063] The DNA decoding module described in this embodiment employs a deduplication algorithm.

[0064] In this embodiment, DNA molecule-directed communication encoding uses LT codes, which include the following:

[0065] LT codes are a classic type of fountain code, characterized by flexible code rates and low encoding / decoding complexity, and were first proposed by Michael Luby. This work applies LT codes to orbital-based molecular communication (MC) systems because their encoding / decoding process does not require a strict code rate. The encoding process is described below:

[0066] 1) The source information is first divided into k input symbols of length L;

[0067] 2) The degree d is randomly generated from the degree distribution of each output symbol;

[0068] 3) d input symbols are uniformly selected and XORed together to output encoded symbols.

[0069] Additional redundancy is needed to decode the information in the LT code, and the number of received encoded symbols n should satisfy n≥k.

[0070] Subsequently, the encoded symbols are mapped into DNA bases, following... Finally, the corresponding oligonucleotide sequences are output.

[0071] In this embodiment, backstepping during sequencing can lead to unwanted repetitive sequences in the reads. In other words, interference in molecular hopper molecular communication (MMTMC) carrying orbitals does not affect the value of the information bits; its redundancy is increased through repetition of the information sequence. When consecutive backstepping occurs during sequencing, it forms a unique bisymmetric region. Furthermore, the repetition pattern is highly correlated with the source sequence itself and exhibits specific statistical characteristics.

[0072] The LT-encoding-based deduplication algorithm comprises two steps. The first step is data initialization, which involves inputting the read oligonucleotide sequence X. n Define variable S N+1 In the first stage, the read DNA sequence is added to a non-redundant possible sequence Y; in the second stage, if a backstep occurs during sequence reading, and the maximum backstep duration is defined as N, the read sequence X is... n In S N+1 The sequence is looped, and if the first i bases in the sequence are equal to the (i+2N)th base, then the read sequence X is... n Add it to sequence Y, and add S. N+1 The (i+2N)th base in sequence X N Remove from the middle, then X n The number of bases n in the sequence and the X sequence is updated, thus achieving the purpose of removing duplicate data.

[0073] Example 2

[0074] like Figure 1 As shown, Figure 1The schematic diagram of the communication system for DNA nanopore sequencing is as follows: A molecular hopper carrying a single-stranded DNA sequence passes through a nanopore under the influence of an electric potential, and performs single-step jumps along protein orbitals under the combined influence of chemical energy and electric potential. The continuous single-step jumps of the molecular hopper are accomplished through a series of thiol-disulfide bond exchange reactions at a transmembrane potential. The molecular hopper jumps along a designed orbital established within a protein nanopore, α-hemolysin (aHL), composed of an array of neatly arranged cysteine ​​residues. The thiol functional groups in the molecular hopper structure react with the sulfhydryl groups on the cysteine ​​residues at the anchorage to form disulfide bonds. Each single-step jump reaction exhibits strict regioselectivity, ensuring that the hopper remains bound to the orbital while maintaining high processability and directional movement.

[0075] like Figure 2 As shown, Figure 2 The structure is a carrier molecule hopper-disulfide conjugate, which is installed at the 5' end of the DNA strand, allowing it to move gradually through a nanopore.

[0076] like Figure 3 As shown, Figure 3 The molecular hopper (MMT) system for carrying protein orbitals consists of three parts: a transmitter, a channel, and a receiver.

[0077] The sending end consists of two steps: DNA encoding and loading the encoded DNA sequence into a molecular hopper. The encoding method uses the LT encoding scheme.

[0078] In the channel section, the molecular hopper loaded with DNA sequence jumps one step at a time along a track with potential at both ends. Different types of bases will cause corresponding changes in the potential difference at both ends of the track.

[0079] The receiving end consists of two parts: nanopore sequencing and information decoding. After sequencing, the DNA nanopore transmits the measured base sequence to the decoding end, which then performs information decoding.

[0080] like Figure 4 As shown, Figure 4 The following is a schematic diagram of the state model of the molecular hopper propagating along the track: the channel is considered a one-dimensional channel. During propagation, the hopper's motion randomly switches between two states at each node: forward movement and backward movement, which can be described as a Markov process. α is the probability that the hopper moves to the forward state (+), and β is the probability that it moves to the backward state (-). Its forward or backward velocity is V.

[0081] like Figure 5 As shown, Figure 5The diagram below illustrates the repetitive sequence detected due to the backflip of the molecular hopper: the top shows the original sequence, and the bottom shows the received sequence, with the red-marked bases indicating continuous backflipping. When continuous backflipping occurs during sequencing, the detected DNA sequence forms a unique bisymmetrical region. The blue letters with black underlines highlight the symmetrical region of this repetitive sequence.

[0082] like Figure 6 As shown, Figure 6 Here is the pseudocode for the deduplication algorithm. Figure 7 The deduplication algorithm flowchart is as follows:

[0083] The first step is data initialization, which involves inputting the oligonucleotide sequence X that has been read. n Define variable S N+1 The first step involves adding the read DNA sequence to the redundant sequence Y. The second step involves adding the read oligonucleotide base x to the redundant sequence Y. If a backstep occurs during DNA sequence reading, and the maximum number of backsteps is defined as N, and a backstep occurs in the molecular hopper, the read oligonucleotide base x is... n Assigned to base sequence segment S N+1 In step (i), an error correction loop is performed. If an error correction sequence S is detected... N+1 (i) The i-th base S N+1 (i) with the (i+2Nth)th base S N+1 If (i+2N) are equal, then the read oligonucleotide sequence X will be... n The value is assigned to the redundant Y, and the base sequence segment S that caused the bit error is also assigned. N+1 The (i+2N)th base S N+1 (i+2N) Read sequence X from the original DNA N Remove from the middle, and then read the DNA sequence X. n The number of base sequences n in the sequence is updated, and the DNA read sequence X is updated simultaneously. n Update the number of backtracks N = N-1 until the maximum number of consecutive backtracks N = 0, then end the error correction process, thus achieving the goal of removing duplicate data.

[0084] The following is combined with Figure 7 The flowchart in the document will be explained in detail:

[0085] The algorithm begins by sequencing the sequence X obtained from nanopore sequencing. n With the set deduplication base sequence S N+1 Input system, where n represents the number of bases read;

[0086] Two variables are introduced: redundant sequence Y, and maximum number of backtracking steps N during the sequencing process;

[0087] The already obtained read base sequence Xn Assigned to the set deduplication base sequence S N+1 ;

[0088] Determine S N+1 The i-th base S in the sequence N+1 (i) Sequence and the (i+2N)th base S N+1 If the (i+2N) sequences are identical, then the base sequence is considered redundant due to the hopper bounce, and the redundant base is added to the redundant sequence Y; the base S is... N+1 The (i+2N) sequence is derived from the already read sequence X. n Remove from the sequence to achieve deduplication, and update the sequence X after removing redundant bases. n If no, proceed to the next step; otherwise, proceed to the next step.

[0089] Update the back count, setting the back count minus 1 to N = N-1;

[0090] Check if the maximum number of consecutive backward steps N is 0. If yes, end the deduplication algorithm and output the deduplicated base sequence. If no, return to step four and run the code again.

[0091] In the flowchart of this embodiment, the number of backsteps N represents the maximum duration of the backstep.

[0092] In this embodiment, the base S is obtained through DNA nanopore sequencing. N+1 (i) The sequence was obtained using a nanopore real-time high-throughput sequencing system based on DNA molecule passage.

[0093] like Figure 8 As shown, Figure 8 To simulate and verify the proposed deduplication algorithm based on Matlab, this embodiment uses Matlab software to simulate the method. The average DNA sequence read length is 500 bp, the DNA sequence shift speed in the nanopore detection is 450 bp / s, and the molecular hopper bounce probability is 12%, as detailed below:

[0094] In Matlab, the horizontal axis is set to the detection sequence length, and the vertical axis is set to the percentage of successfully recovered original sequence. The four lines in the figure represent p. j =0.025, p j =0.05, p j =0.075, p j The simulation results when p = 0.1, where p j This refers to the probability of the molecular hopper retracting. It can be seen that the successful recovery curve decreases significantly with the increase of the detected sequence length. Furthermore, the retraction probability p... jThe smaller the DNA sequence, the higher the probability of successful recovery. Successful recovery refers to the probability of restoring the received sequence to the original sequence; this is the posterior probability that measures the performance of the deduplication algorithm.

[0095] like Figure 9 As shown, Figure 9 This section compares the best bit error rate performance based on Matlab and the best bit error rate performance based on LT coding.

[0096] In Matlab, the horizontal axis is set to the probability of backslip, and the vertical axis is set to the average bit error rate (BER). The red line in the graph represents the optimal BER performance, and the blue line represents the optimal BER performance of LT coding. In the optimal BER simulation, it is assumed that the transmission probability of each nucleotide is equal. As can be seen from the graph, the average BER increases with the backslip probability p. j The increase is due to the increase of p, which indicates that high p j This leads to more sequencing errors. The error rate curve of LT encoding is higher than the optimal error rate curve, and the deviation can be attributed to the randomness of LT encoding and decoding processes.

[0097] like Figure 10 As shown, Figure 10 For different redundancy levels and p based on Matlab j The bit error rates of the encoding schemes are compared. In Matlab, the horizontal axis is set to the proportion of redundant information required for successful decoding, and the vertical axis is set to the average bit error rate. The solid lines of different colors represent the backtracking probability p. j =0.025, p j =0.05, p j =0.075 and p j The bit error rate (BER) is 0.1 when the deduplication algorithm is used, while the dashed line, corresponding to the solid line, represents the BER when the deduplication algorithm is not used. The graph shows that as the backoff probability p... j With the increase of [symbols], the bit error rate curve generally tends to shift backward, indicating that more symbols are needed to derive the source information. Furthermore, the backward probability p [increases / increases]. j In specific cases, the proposed deduplication algorithm helps reduce sequence redundancy, accelerates the decoding process, and further reduces the bit error rate of the molecular hopper molecular communication (MMTMC) system carrying protein orbitals.

[0098] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.

Claims

1. A real-time high-throughput sequencing system based on nanopores for DNA molecule transport, characterized in that: It includes a DNA encoding module, a molecular hopper-loaded DNA sequence strand module, a nanopore sequencing module, and a DNA decoding module; The DNA encoding module is used to store the data to be transmitted by the user in the DNA sequence through base pairing; The molecular hopper is loaded with a DNA sequence chain module, which is used to link the encoded DNA sequence with the molecular hopper for transport on the nanoporous protein orbital, and the molecular motor loaded with the DNA chain is transported on the orbital. The nanopore sequencing module achieves real-time sequencing by detecting different currents generated due to the different abilities of different base pairs to block current. The DNA decoding module is used to eliminate redundant errors caused by the molecular hopper bouncing back on the protein orbitals in the nanopore; The deduplication algorithm in the DNA decoding module is performed according to the following steps: The first step is data initialization, which involves inputting the oligonucleotide sequence read. n is the oligonucleotide sequence The number of bases, defined as N, the maximum number of backsteps, and the definition of an oligonucleotide sequence. A single base in the middle is The base sequence segment that causes a bit error due to a backstepping motion in the molecular hopper is defined as... , ∈ Redundancy is defined as Y; The second step involves reversing the molecular hopper during DNA sequence reading, causing the already read oligonucleotide bases to... Assigned to base sequence segment During the error correction loop, if an error correction sequence is detected... The i-th base With the (i+2N)th base If they are equal, then the oligonucleotide sequence read will be... The value is assigned to the redundant Y, and the base sequence segment that caused the bit error is removed. The (i+2)th base in Sequence read from original DNA Remove from the middle, and then read the DNA sequence. The number of base sequences n in the sequence is updated, and the DNA read sequence is updated simultaneously. Update the number of backslides N=N-1 until the maximum number of consecutive backslides N=0, then end the error correction process.

2. The nanopore real-time high-throughput sequencing system based on DNA molecule passage as described in claim 1, characterized in that: The base pairing used is Watson-Crick base pairing.

3. The nanopore real-time high-throughput sequencing system based on DNA molecule passage as described in claim 1, characterized in that: The DNA decoding module employs a deduplication algorithm.

4. The nanopore real-time high-throughput sequencing system based on DNA molecule passage as described in claim 1, characterized in that: The DNA encoding module uses LT codes, and the LT code encoding is performed according to the following steps: First, the source information is divided into k input symbols of length L; the degree d is randomly generated from the degree distribution of each output symbol; the d input symbols are uniformly selected and XORed together to output coded symbols; the number of received coded symbols n should satisfy n ≥ k; The encoded symbols are then mapped to DNA bases, following the pattern {00, 01, 10, 11} ⇔ {A,C,T,G}. Finally, the corresponding oligonucleotide sequences are output.

5. A real-time high-throughput sequencing method based on nanopores for DNA molecule transport, characterized in that: Includes the following steps: Bases were obtained through DNA nanopore sequencing. sequence; Set the maximum number of consecutive backsteps N, and introduce a redundant sequence Y; bases Sequence value assignment ; Determine the base Sequence and bases If the sequences are identical, the base sequence is considered redundant due to the hopper bounce, and the redundant bases are added to the redundant sequence Y; the bases are then... Sequence from sequence Remove from the sequence to achieve deduplication, and update the sequence after removing redundant bases. If no, proceed to the next step; otherwise, proceed to the next step. Update the number of backslides N = N-1 until the maximum number of consecutive backslides N = 0.

6. The nanopore real-time high-throughput sequencing method based on DNA molecule passage as described in claim 5, characterized in that: The bases were obtained through DNA nanopore sequencing. The sequence is performed according to the following steps: The data that the user needs to transmit is stored in the DNA sequence through base pairing; The encoded DNA sequence is linked to a molecular hopper for transport on nanoporous protein orbitals, and molecular motors carrying DNA strands are transported on these orbitals. Different base pairs have different abilities to block current, resulting in different current detections, thus enabling real-time sequencing.

7. The nanopore real-time high-throughput sequencing method based on DNA molecule passage as described in claim 6, characterized in that: The base pairing used is Watson-Crick base pairing.

8. The nanopore real-time high-throughput sequencing method based on DNA molecule passage as described in claim 6, characterized in that: The storage of the user's required transmission data in the DNA sequence via base pairing is achieved using LT codes, and the LT code encoding is performed according to the following steps: First, the source information is divided into k input symbols of length L; the degree d is randomly generated from the degree distribution of each output symbol; the d input symbols are uniformly selected and XORed together to output coded symbols; the number of received coded symbols n should satisfy n ≥ k; The encoded symbols are then mapped to DNA bases, following the pattern {00, 01, 10, 11} ⇔ {A,C,T,G}. Finally, the corresponding oligonucleotide sequences are output.

Citation Information

Patent Citations

  • DNA sequencing method and system thereof

    CN104630358A

  • DNA-based data storage method, DNA-based data decoding method, DNA-based data storage system and DNA-based data decoding device

    CN111858507A

  • Information coding method and decoding method based on DNA and computer readable storage medium

    CN115206430A