Error correction method in DNA molecular communication
By using multi-strand characteristics to vote and compare and correct errors during the DNA decoding process, the problem of limited error correction capabilities in the existing technology is solved, efficient and accurate error correction effects are achieved, and the reliability of DNA communication and storage is improved.
Patent Information
- Application Number
- CN202510325928.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The existing DNA communication decoding technology cannot fully utilize the multi-strand characteristics of DNA, resulting in limited error correction capabilities and difficulty in efficiently and accurately decoding and correcting errors, becoming a bottleneck in the promotion of DNA communication and storage technology.
During the decoding process, the base value of the current position is calculated based on the voting mechanism of DNA multi-strands, and the results of the subsequent base information and other chains are combined to judge and correct the type of wrong bases to avoid adding redundant information.
It effectively improves the error correction performance of DNA communication, improves the accuracy and reliability of decoding, and is suitable for error correction in DNA communication and stored procedures, without reducing information density.
Smart Images

Figure CN120263346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular communication, and particularly relates to a DNA molecular communication error correction technology. Background Art
[0002] With the advent of the big data era, traditional electronic communication systems are encountering essential bottlenecks in coping with the rapidly increasing number of devices, improving energy efficiency, and adapting to micro-scale applications, and it is difficult to meet the growing data communication needs. Due to characteristics such as high storage density, long preservation time, and low energy consumption, DNA molecular communication has become an important research direction for big data communication technologies. The core processes of DNA communication technology include information encoding, DNA sequence synthesis, and decoding and recovery after sequencing. However, during DNA synthesis, sequencing, and storage, errors such as base substitution, insertion, or deletion will inevitably be introduced, which seriously affects the accuracy of decoding. Existing DNA communication decoding technologies mainly rely on hard decision error correction codes. By adding redundant information, such as Reed-Solomon codes (RS codes) and low-density parity-check codes (LDPC codes), a certain degree of error correction can be achieved. However, hard decision error correction technologies cannot fully utilize the multi-strand characteristics of DNA, resulting in limited error correction capabilities. In addition, the randomness and unpredictability of insertion and deletion errors in DNA communication also significantly increase the complexity of decoding. How to efficiently and accurately decode and correct errors has become a bottleneck problem restricting the popularization of DNA communication and storage technologies. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes an error correction method in DNA molecular communication. During the decoding process, an error correction method that utilizes the multi-strand characteristics of DNA is used, avoiding the redundant information brought by traditional error correction codes, effectively improving the performance of the error correction method, and being generally applicable to error correction problems in DNA communication and storage processes.
[0004] The technical solution adopted by the present invention is: an error correction method in DNA molecular communication. During the decoding process, starting from the first base position, a voting mechanism is adopted based on multiple DNA strands at the same position to calculate the base at the current position, and other different bases are regarded as errors occurring at the current position of the strand; combining the subsequent base information with the comparison results of other strands, the type of the error base is judged and corrected.
[0005] The process of calculating the base at the current position is as follows: a certain number of DNA sequences are extracted from the pool to form multiple DNA strands, and then starting from the first base position, a voting mechanism is adopted based on multiple DNA strands at the same position. Each time a base A, T, G, or C appears, one point is accumulated; finally, the base with the highest score is selected as the correct value at the current position, and other different bases are regarded as errors occurring at the current position of the strand.
[0006] Furthermore, when two bases with the highest scores appear, an additional DNA sequence is drawn from the pool for judgment.
[0007] The process of determining the type of incorrect base is as follows:
[0008] First, select the three consecutive bases following the incorrect base as a whole. If they correspond identically to the three consecutive bases following the correct base at the current position, it is considered that a substitution error has occurred at the current position.
[0009] Select the three consecutive bases following the incorrect base as a whole. If they correspond identically to the correct base at the current position and the two bases following it, it is considered that an insertion error has occurred at the current position.
[0010] Select the incorrect base and the two bases following it as a whole. If they are equal to the three consecutive bases following the correct base at the current position, it is considered that a deletion error has occurred at the current position.
[0011] The method for correcting errors is as follows: If a substitution error occurs at the current position, replace the incorrect base with the correct base at the current position; if an insertion error occurs at the current position, delete the incorrect base at the current position; if a deletion error occurs at the current position, insert the correct base in front of the base at the current position.
[0012] Advantages of the present invention: Compared with existing algorithms, the present invention uses the multi-strand characteristics of DNA for error correction without adding additional redundant information and without reducing the original information density. At the same time, this method effectively corrects errors, improves the reliability of DNA communication, and is generally applicable to base error correction in the processes of DNA communication and storage. Brief Description of the Drawings
[0013] Figure 1 It is a flowchart of the method for a specific embodiment of the present invention.
[0014] Figure 2 It is a schematic diagram of error correction provided by an embodiment of the present invention;
[0015] Among them, (a) is the original encoded sequence, (b) is a schematic diagram before error correction of the base position i = 0 of the 4 DNA sequences drawn, (c) is a schematic diagram after error correction of the base position i = 0 of the 4 DNA sequences drawn, (d) is a schematic diagram before error correction of the base position i = 2 of the 4 DNA sequences drawn, (e) is a schematic diagram after error correction of the base position i = 2 of the 4 DNA sequences drawn, (f) is a schematic diagram before error correction of the base position i = 4 of the 4 DNA sequences drawn, and (g) is the decoded sequence.
[0016] Figure 3 It is an error correction effect diagram provided by an embodiment of the present invention using the method of the present invention. Detailed Implementation Manner
[0017] To facilitate those skilled in the art to understand the technical content of the present invention, the content of the present invention will be further explained below with reference to the accompanying drawings.
[0018] The entire encoding and decoding process of DNA molecular communication is as follows: Information encoding is represented by DNA; then, multiple encoded DNAs are synthesized; next, the polymerase chain reaction technology is used to further amplify the number of synthesized DNAs; then, a certain amount (N pieces) of DNA is extracted from the synthesis pool for communication, and after reaching the receiving end, sequencing is performed on the N pieces of DNA; then, error correction is performed on these N pieces of DNA using an error correction method; finally, a correct DNA strand is obtained, and the information before encoding is restored from it.
[0019] During the decoding process, the biological characteristics of DNA are utilized. After it is encoded from the original information and needs to be synthesized, multiple DNA copies will definitely be generated in the synthesis pool. After sequencing, the property of multiple DNA copies is utilized to perform error correction on the method of the present invention.
[0020] An error correction method in DNA molecular communication of the present invention has the following principle: During the decoding process, starting from the first base position, based on the voting mechanism adopted by multiple DNA strands at the same position, the correct value at the current position is calculated, and other different bases are regarded as errors occurring at the current position of the strand. Subsequently, combined with the subsequent base information and the comparison results of other strands, the type of the error base is judged and corrected. After correcting the error, the second base position is judged. Finally, the correct bases at each position are added to the result list to complete the decoding process.
[0021] When the present invention is specifically operated, its process is as Figure 1 shown. The specific operation method of an error correction method in DNA molecular communication described above includes the following steps:
[0022] Step 1: Extract a certain amount of DNA copies from the pool to form multiple DNA strands, initialize the base position information (i = 0) and the result list, and calculate the length of the longest DNA strand in the DNA copies as Len;
[0023] Those skilled in the art should know that the pool mentioned in Step 1 refers to the DNA synthesis pool. Multiple DNA copies will be generated in the pool when the DNA sequence is synthesized and encoded. Subsequently, the polymerase chain reaction technology can be used to further multiply the number of specified DNA sequences. In this embodiment, the extraction quantity in Step 1 is 50 copies.
[0024] Step 2: Process the i-th base position and perform voting scoring based on the DNA multi-strands at the same position. The scoring process is as follows: Extract a certain number of DNA sequences from the pool to form DNA multi-strands. Then, starting from the first base position, use a voting mechanism based on the DNA multi-strands at the same position. Each occurrence of bases A, T, G, and C accumulates one point.
[0025] Step 3: Take the base X with the highest score at the current position as the correct value at the current position, and mark the different ones as incorrect base Y. If there are the same bases with the highest score, extract another DNA strand from the pool and add it to the operation, then repeat Step 2.
[0026] Step 4: Select the three consecutive bases after the incorrect base Y as a whole. If they correspond to the three consecutive bases after the correct base X at the current position, mark this position as a substitution error and execute Step 7. If not, execute Step 5.
[0027] Step 5: Select the three consecutive bases after the incorrect base Y as a whole. If they correspond to the correct base X at the current position and its two subsequent bases, mark this position as an insertion error and execute Step 7. If not, execute Step 6.
[0028] Step 6: Select the incorrect base Y and its two subsequent bases as a whole. If they correspond to the three consecutive bases after the correct base X at the current position, mark this position as a deletion error and execute Step 7. If not, delete the copy of the DNA strand with the incorrect base at the current position, process the next base position (i + 1), and execute Step 2.
[0029] Step 7: Perform error correction according to the error type at the current position. If it is a substitution error, replace it with the correct base X at the current position. If it is an insertion error, delete the base Y at the current position. If it is a deletion error, insert the correct base X in front of the base Y at the current position.
[0030] Step 8: Record the correct base information X at the current position in the result list. If i < Len - 3, process the next base position (i + 1) and execute Step 2. If i ≥ Len - 3, execute Step 9.
[0031] Step 9: Perform voting scoring based on the DNA multi-strands at positions Len - 3, Len - 2, and Len - 1 respectively. At this time, there may be empty information (i.e., there is no base information at the current position). In this case, there are five situations for each position: A, T, G, C, and empty. Select the one with the highest score as the result at the current position and fill it into the result list. Finally, output the result list as the error-corrected DNA sequence information, and the error correction process ends here.
[0032] As Figure 2(a) As shown, the encoded sequence is ATGTCAGGAA. In this embodiment, taking the example of extracting 4 DNA sequences from the pool to form a DNA multi-chain for illustration:
[0033] As Figure 2 (b) As shown, starting from the first base position i = 0, select the base at the i-th position of each sequence. The A with the highest score is used as the correct value at the current position, and there is no incorrect base at the current base position i = 0.
[0034] When judging the base position i = 1, as Figure 2 (b) shown, the correct base is T, and an error occurs at the base position i = 1 of sequence 4, and A is the incorrect base; then execute step 4. It is found that the three bases after the incorrect base are TGT, and there is no sequence among sequences 1 - 3 where the three consecutive bases after the correct base T are also TGT; then, execute step 5. It is found that the three bases TGT after the incorrect base A correspond to GT formed by the correct base T of sequence 1 and the two bases after it. Then, mark that an insertion error occurs at the base position i = 1 of sequence 4; execute step 7, correct the error in sequence 4, and delete the incorrect base at the base position i = 1, obtaining sequence 4 as shown in Figure 2 (c).
[0035] When judging the base position i = 2, as Figure 2 (d) shown, the correct base is G, and an error occurs at the base position i = 2 of sequence 3, and T is the incorrect base; then execute step 4. It is found that the three bases after the incorrect base are CAG, and there is no sequence among sequences 1, 2, and 4 where the three consecutive bases after the correct base G are also CAG; then, execute step 5. It is found that the three bases CAG after the incorrect base T do not match the correct base G and the two bases after it; execute step 6. It is found that the TCA formed by the incorrect base T and the two bases CA after it corresponds to the three bases TCA after the correct base G of sequence 2. Then, mark that a deletion error occurs at the current base position i = 2; execute step 7, correct the error in sequence 3, and insert the deleted base G in front of the base T, obtaining sequence 3 as shown in Figure 2 (e).
[0036] When judging the base position i = 3, as Figure 2 (f) shown, no error occurs, and the correct base at the current position is T.
[0037] When judging the base position i = 4, as Figure 2 (f) shown, the correct base is C, and an error occurs at the base position i = 4 of sequence 1, and T is the incorrect base; then execute step 4. It is found that the three bases after the incorrect base are AGG, which correspond to the three bases after the correct base C of the sequence. Then, mark that a substitution error occurs at the current position; execute step 7, correct the error in sequence 1, and replace the T at the current position with the correct base C. Then, it is found that there are no incorrect bases in the subsequent positions one by one until the last position, thus obtaining as shown in Figure 2The decoded sequence shown in (g).
[0038] As Figure 3 shown, when the DNA sequence is 1000 in length, the error correction effect of the method of the present invention with different copies is given as the error probability of DNA bases changes. It can be seen that by using the method of the present invention, generally 50 copies can correct base errors at an error rate of 3.5%.
[0039] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. An error correction method in DNA molecular communication, characterized in that, During the decoding process, starting from the first base position, a voting mechanism is adopted based on multiple DNA strands at the same position to calculate the correct base at the current position, and other different bases are regarded as errors occurring at the current position on that strand. Combining the subsequent base information with the alignment results of other strands, the type of the incorrect base is determined and corrected.
2. The error correction method in DNA molecular communication according to claim 1, characterized in that, The process of calculating the correct base at the current position is as follows: Extract several DNA sequences to form multiple DNA strands. Based on the voting mechanism of multiple DNA strands at the same position, each occurrence of bases A, T, G, and C accumulates one point. Finally, the base with the highest score is selected as the correct value at the current position, and other different bases are regarded as errors occurring at the current position on that strand.
3. The error correction method in DNA molecular communication according to claim 2, characterized in that, When there are two bases with the highest scores, another DNA sequence is extracted for judgment.
4. The error correction method in DNA molecular communication according to claim 3, wherein, The process of determining the type of the incorrect base is as follows: First, select the three consecutive bases after the incorrect base as a whole. If they correspond exactly to the three consecutive bases after the correct base at the current position of the other extracted DNA sequences, it is considered that a substitution error has occurred at the current position. Select the three consecutive bases after the incorrect base as a whole. If they correspond to the correct base at the current position of the other extracted DNA sequences and the two subsequent bases, it is considered that an insertion error has occurred at the current position. Select the incorrect base and the two subsequent bases as a whole. If they are equal to the three consecutive bases after the correct base at the current position of the other extracted DNA sequences, it is considered that a deletion error has occurred at the current position.
5. The error correction method in DNA molecular communication according to claim 4, wherein The method for correcting errors is as follows: If a substitution error has occurred at the current position, replace the incorrect base with the correct base at the current position; if an insertion error has occurred at the current position, delete the incorrect base at the current position. If a deletion error has occurred at the current position, insert the correct base in front of the base at the current position.