DNA storage random access address coding method and system
By constructing a reverse complementary pair resource pool and base conversion rules to generate DNA address sequences that satisfy mutual uncorrelation, GC balance, and error correction capabilities, the problems of secondary structure interference, weak mutual uncorrelation, and insufficient error correction capabilities in existing DNA storage are solved, thereby improving the robustness and data integrity of the DNA storage system.
Patent Information
- Application Number
- CN202511608111.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-06
AI Technical Summary
Existing DNA storage address encoding technologies suffer from problems such as secondary structure interference with PCR amplification, weak correlation between address sequences, large differences in melting temperature due to GC content fluctuations, and lack of error correction capabilities, which affect the accuracy and stability of random access.
By constructing a resource pool of reverse complementary pairs of a specific length, generating codes by combining prefix and suffix matching principles, applying base conversion rules and decoupled structure double binary mapping, generating address sequences that satisfy uncorrelatedness, GC balance and error correction capabilities, using Knuth's balancing technique to regulate 0/1 distribution, and adding LDPC check redundancy.
It significantly enhances the robustness and data integrity of DNA storage systems, improves PCR amplification efficiency, avoids cross-hybridization, ensures the distinguishability and thermodynamic stability of address sequences, and reduces the risk of data loss.
Smart Images

Figure CN121483335A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA storage technology, specifically to a method based on... A method and system for encoding random access addresses for unrelated DNA storage. Background Technology
[0002] With the exponential growth of global data volume, traditional storage media (such as hard drives and optical discs) are gradually becoming unable to meet the storage needs of massive amounts of data in terms of storage density, long-term stability, and energy consumption. DNA, as the main carrier of genetic information in living organisms, has advantages such as high storage density, strong stability, and low energy consumption, making it an important research direction for new data storage technologies.
[0003] In the 1960s, scientist Neiman first proposed the concept of using DNA for data storage. In the early 21st century, Professor Church's team at Harvard University and Professor Golaman's team at the European Institute for Bioinformatics (EIB) increased DNA storage capacity from less than 1KB to 739KB using techniques such as binary encoding and Huffman ternary conversion, respectively, propelling DNA storage into mainstream research. In DNA storage systems, random access capability is key to improving data management flexibility and access efficiency. By attaching address tag sequences to both ends of DNA fragments, PCR amplification technology can selectively amplify target fragments, enabling the rapid extraction of specific data.
[0004] However, existing DNA storage address encoding technologies have the following shortcomings: 1. Address sequences are prone to forming secondary structures such as stem-loops, which leads to reduced PCR amplification efficiency and may even cause erroneous amplification of non-target sequences; 2. The weak correlation between address sequences makes it difficult to effectively avoid cross-linking, affecting the accuracy of random access; 3. The thermodynamic stability of GC base pairs and AT base pairs differs significantly. If the GC content fluctuates too much, it will lead to significant differences in the melting temperature of different address sequences, further reducing amplification consistency. 4. Some random access address encoding schemes do not integrate error correction mechanisms, which cannot cope with base errors in the DNA synthesis and sequencing process, affecting data integrity.
[0005] To address this, the team led by Dai Junbiao at the Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, proposed the "Efficient DNA Storage (EDS)" method, which improves the random access efficiency of MRI images through regularized quaternary transcoding and simplified indexing techniques. However, this method still does not completely solve the problems of secondary structure suppression, GC balancing, and error correction capabilities. Summary of the Invention
[0006] The purpose of this invention is to propose a method and system for encoding random access addresses for DNA storage. The encoded address sequence not only satisfies the combined biological constraints but also integrates a certain error correction capability and exhibits excellent thermodynamic performance, which can significantly enhance the robustness of the DNA storage system.
[0007] According to a first aspect of the present disclosure, a method for encoding random access addresses for DNA storage is provided, comprising the following steps: Enumerate all DNA reverse complementary pairs of length m, screen for sequences whose sum of the number of "T" and "C" bases meets the preset conditions, and construct... Resource pool; From the above A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; According to the base conversion rules, the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Construct a set of uncorrelated codes of a predetermined length weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. For the weakly uncorrelated encoding set After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the decoupling structure is mapped using a bi-binary mapping to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and balancing GC. The bi-binary mapping rule for the decoupling structure is as follows: For the set of mutually uncorrelated codes By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words in the decoupling structure is mapped using a bi-binary mapping to generate an address sequence that is independent, avoids secondary structures, and has error correction capabilities. The bi-binary mapping rule of the decoupling structure is as follows: In one embodiment, the sum of the number of bases "T" and "C" satisfies a preset condition specifically: the sum of the number of bases "T" and "C" is greater than m / 2; the length m of the reverse complementary pair is in the range of 3-9, used to suppress the generation of DNA secondary structures with a stem length of 3-9.
[0008] In one embodiment, the prefix and suffix matching principle is: the prefix of the subsequent sequence... The bases and the following bases in the current sequence The bases are completely identical; the base cascade operation is: deleting the first base of the subsequent sequence. The target length is obtained by: [The text abruptly ends here, so the translation stops as well.] ,in For stem length, From The total number of times a sequence is selected from the resource pool.
[0009] In one embodiment, the base conversion rule is as follows: base "T" corresponds to binary "0", base "C" corresponds to binary "0", base "A" corresponds to binary "1", and base "G" corresponds to binary "1".
[0010] In one embodiment, the prefix constraint is: the prefix of each codeword The fixed bit constraint is: the first bit of each encoded word is "0"; Position, No. The bit cannot be "0"; the consecutive "0" constraint means that each encoded word starts from the first bit. The position reached the first The digits cannot be consecutive. The extended constraint is: in the first "0"; After that, cascade any number of bits. A binary sequence of bits.
[0011] In one embodiment, weakly uncorrelated codes are eliminated. From the middle Ranked first Bit exists A sequence of consecutive '1's is given, resulting in a code set denoted as . Then from that set Arbitrarily select weakly uncorrelated encoded words From the first The position reached the first By applying Knuth's balancing technique, a 0 / 1 balanced codeword is obtained. In addition, from Encoding set Choose any Encoded words and length Encoded words and The component encoding, used as a decoupling structure, undergoes a bi-binary mapping to obtain an address sequence that simultaneously satisfies the constraints of mutual independence, avoidance of secondary structures, and GC balance. .
[0012] In one embodiment, from an unrelated set of codes Selecting coded words After adding LDPC check redundancy, the encoded word is obtained. ,in It is a coded word The check digit; secondly, from Encoding set Select Encoded words and the length is Encoded words and The component encoding, used as a decoupling structure, undergoes a bi-binary mapping to obtain an address sequence that simultaneously satisfies the requirements of being uncorrelated, avoiding secondary structures, and possessing error correction capabilities. .
[0013] According to a second aspect of the present disclosure, a DNA storage random access address encoding system is provided, comprising: The resource pool construction module enumerates all DNA reverse complementary pairs of length m, selects sequences whose sum of the number of "T" and "C" bases meets preset conditions, and constructs the resource pool. Resource pool; The encoding generation module, from the A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; The encoding conversion module, according to the base conversion rules, converts the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Constraint module: Constructs a set of mutually uncorrelated codes of a preset length. weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. The balanced address sequence generation module processes the weakly uncorrelated encoding set. After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the middle is generated through a decoupled double binary mapping structure to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and GC balance; The error correction address sequence generation module performs the following steps on the uncorrelated encoding set: By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
[0014] According to a third aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the memory, wherein the processor executes the program to implement the aforementioned DNA storage random access address encoding method.
[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the aforementioned DNA storage random access address encoding method.
[0016] The advantages of the above technical solutions adopted in this invention compared with the prior art are as follows: 1. By constructing a resource pool based on inverse complementary pairs of specific lengths, combining prefix and suffix matching principles to achieve sequence cascading, and generating corresponding codes through base conversion rules, the formation pathways of DNA secondary structures within the stem length range of 3-9 are directly blocked at the binary level. This design not only solves the core problem of secondary structures interfering with PCR amplification in traditional coding, but also provides a feasible technical path for subsequent synergistic implementation of combined biological constraints (such as secondary structure avoidance and uncorrelated characteristics) through decoupling structures and unrelated codes. Simultaneously, it significantly enhances the length scalability of DNA sequences, flexibly adapting to the address length requirements of different storage scenarios.
[0017] 2. By introducing Knuth's balancing technique into uncorrelated codes, the 0 / 1 distribution of the coding words can be precisely controlled to achieve a 0 / 1 balance state. This balanced code, along with the code generated by base conversion, is then used as a component code of the decoupled structure, which is converted into a DNA address sequence via bibinary mapping of the decoupled structure. This process simultaneously meets three key requirements: first, the uncorrelated characteristic ensures the distinguishability between address sequences, avoiding non-target amplification caused by cross-hybridization; second, the secondary structure avoidance characteristic ensures PCR amplification efficiency; and third, the 0 / 1 balance characteristic is converted into GC content balance, reducing the thermodynamic performance differences between different address sequences and further improving the accuracy and stability of random access during DNA storage.
[0018] 3. LDPC encoding is performed on unrelated codewords of a specific length. By concatenating check bits at the end of the encoding to form an error-correcting codeword, it is then combined with a base conversion codeword of equal length to enter the decoupled structure double binary mapping process. The final generated address sequence, while retaining the two core biological constraints of unrelatedness and avoiding secondary structures, also possesses base error correction capabilities. This effectively addresses issues such as base mutations and deletions that may occur during DNA synthesis, sequencing, and storage, significantly reducing the risk of data loss and greatly improving the overall robustness and data integrity assurance capabilities of the DNA storage system. Attached Figure Description
[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.
[0020] Figure 1 This is a flowchart of a DNA storage random access address encoding method. Detailed Implementation
[0021] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0023] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0024] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems according to various embodiments of this disclosure. It should be noted that each block in a flowchart or block diagram may represent a module, segment, or portion of code, which may include one or more executable instructions for implementing the logical functions specified in the various embodiments. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or using a combination of dedicated hardware and computer instructions.
[0025] Example 1: like Figure 1 As shown, this embodiment provides a method for encoding random access addresses for DNA storage, including the following steps: Step 1: Enumerate all DNA reverse complementary pairs of length m, screen for sequences whose sum of the number of "T" and "C" bases meets the preset conditions, and construct... Resource pool; Specifically, based on stem length For example, enumerate 32 pairs of anticomplementary bases: {AAA,TTT}, {CCC,GGG}, ..., {TGA,TCA}. In each pair, the sum of the number of bases 'T' and 'C' is greater than... Selecting based on principles Sequence, construct Resource pool. Elements in this resource pool include: TTT, GTT, CTT, ATT, TGT, ACC, CGT, ACT, TCT, GCT, CCT, TAT, ATC, CAT, TTG, CAC, CTG, CCA, CCC, CCG, TCG, CGC, CTA, CTC, TTC, GTC, TGC, GCC, TCC, TAC, TTA, TCA.
[0026] Step 2: From the above A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; Specifically, with Taking the resource pool as an example, we select the element 'TTT' as the initial sequence. Due to the initial sequence The prefix is 'TT', therefore, the subsequent sequence The first two bases must also be 'TT'. Therefore, the next sequence can be 'TTC'. Delete the first two bases of the subsequent sequence 'TTC', and cascade the last base to the end of the initial sequence 'TTT'. The encoded word is 'TTTC'. Repeat this step until the sequence is complete. Reaching the target length, we obtain Encoding set.
[0027] Preferred, The target length is calculated as follows ,in For stem length, This represents the number of times a sequence is selected from the resource pool. Assume a target length. Selected The sequence is: TTT, TTC, TCG, CGC, GCT, CTA, which, after concatenation, yields a sequence of length 8. The encoded word is 'TTTCGCTA'. DNA sequences of arbitrary length constructed from this resource pool in this manner constitute... The set of encodings.
[0028] Step 3: According to the base conversion rules, the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Specifically, according to the base conversion rules The 'TTTCGCTA' encoded word was converted to The code is '00001001'.
[0029] Step 4: Construct a set of mutually uncorrelated codes of a preset length. weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. Specifically, the constructed codeword length is Uncorrelated sets of codes and the length of the encoded word is weakly uncorrelated coding set , where set The codewords in the set must satisfy conditions a, b, and c. The codewords in the code must meet conditions a, b, c, and d.
[0030] a. The beginning of each encoded word The bit must be "0"; b. The first of each encoded word Position, No. The bit cannot be "0"; c. Each encoded word starts from the first... The position reached the first The digits cannot be consecutive. A zero; d. in the After that, cascade any number of bits. A binary sequence of bits.
[0031] As ordered , , Therefore, according to the construction rule ac, the set In this code, the first three bits of each code are '000', the fourth and tenth bits are '1', and three consecutive '0's are not allowed from the fifth to the ninth bit. The sequence '00110' is one of many combinations that meet the conditions. Therefore, the constructed uncorrelated codes '0001001101' are a set. One of the elements. Finally, according to rule d, the uncorrelated code '0001001101' can be concatenated with '00', and the resulting sequence '000100110100' is a weakly uncorrelated code.
[0032] Step 5: Process the weakly uncorrelated encoding set After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the middle is generated through a decoupled double binary mapping structure to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and GC balance; Specifically, still make , , In order to maintain the construction rules after applying Knuth's balancing technique to weakly uncorrelated codes, the set is removed. The encoding set is obtained by finding three consecutive "0"s in the 5th to 10th positions. The encoded character '000100110100' is One of the elements in the code is used to apply Knuth's balancing technique starting from the 5th bit. When the 6th bit is determined, the encoded word becomes... And it has reached a 0 / 1 equilibrium state. Then, select... Encoding set Encoded words in ,Will , The component encoding, as a decoupling structure, yields an address sequence that simultaneously satisfies the requirements of non-correlation, avoidance of secondary structures, and GC balance. .
[0033] Step 6: For the set of mutually uncorrelated codes By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
[0034] Specifically, select mutually uncorrelated codes. After being fed into the LDPC system encoder, a sequence with added error correction redundancy is obtained. From the set Select the code words of the same length and will , Simultaneously, as a component encoding of the decoupling structure, an address sequence was obtained that simultaneously satisfies the requirements of mutual independence, avoidance of secondary structure constraints, and error correction capability. .
[0035] To test the performance of the encoded random access address sequence in practical applications, the encoding quality and GC content related to thermodynamic properties were verified and compared with other representative encoding algorithms. The experimental results are shown in Tables 1, 2 and 3. The results of this patent are significantly better than the experimental results of other algorithms.
[0036] Table 1 Comparison of thermodynamic properties of minimum free energy Table 2 Comparison of thermodynamic properties of depolymerization temperature Table 3 Comparison of GC content Example 2: This embodiment provides a DNA storage random access address encoding system, including: The resource pool construction module enumerates all DNA reverse complementary pairs of length m, selects sequences whose sum of the number of "T" and "C" bases meets preset conditions, and constructs the resource pool. Resource pool; The encoding generation module, from the A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; The encoding conversion module, according to the base conversion rules, converts the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Constraint module: Constructs a set of mutually uncorrelated codes of a preset length. weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. The balanced address sequence generation module processes the weakly uncorrelated encoding set. After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the middle is generated through a decoupled double binary mapping structure to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and GC balance; The error correction address sequence generation module performs the following steps on the uncorrelated encoding set: By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
[0037] Example 3: An electronic device includes a memory, a processor, and a computer program stored in the memory and running thereon. When the processor executes the program, it implements the aforementioned DNA storage random access address encoding method, comprising: Enumerate all DNA reverse complementary pairs of length m, screen for sequences whose sum of the number of "T" and "C" bases meets the preset conditions, and construct... Resource pool; From the above A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; According to the base conversion rules, the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Construct a set of uncorrelated codes of a predetermined length weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. For the weakly uncorrelated encoding set After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the middle is generated through a decoupled double binary mapping structure to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and GC balance; For the set of mutually uncorrelated codes By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
[0038] Example 4: A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned DNA storage random access address encoding method, comprising: Enumerate all DNA reverse complementary pairs of length m, screen for sequences whose sum of the number of "T" and "C" bases meets the preset conditions, and construct... Resource pool; From the above A DNA sequence is randomly selected from the resource pool as the initial sequence. Continue from the prefix and suffix matching principle Selecting subsequent sequences from the resource pool Generate the target length through base cascade operations Encoding, composition Encoding set; According to the base conversion rules, the... DNA sequences in the encoding set are converted into binary. Encoding, forming Encoding set ; Construct a set of uncorrelated codes of a predetermined length weakly uncorrelated coding set The encoded words in both types of encoding sets satisfy the prefix constraint, fixed bit constraint, and consecutive "0" constraint. The weakly uncorrelated encoding set also satisfies the extended constraint. For the weakly uncorrelated encoding set After screening and optimization, combined with balancing techniques, a 0 / 1 balanced codeword is obtained, which is then compared with the... Encoding set The corresponding length of the encoded words in the middle is generated through a decoupled double binary mapping structure to generate an address sequence that satisfies the requirements of being uncorrelated, avoiding secondary structures, and GC balance; For the set of mutually uncorrelated codes By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
[0039] Those skilled in the art will understand that the modules or steps described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, which can then be stored in a storage device for execution by a computer device. Alternatively, they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. This disclosure is not limited to any particular combination of hardware and software.
[0040] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0041] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A DNA storage random access address encoding method, characterized by, The method comprises the following steps: Enumerate all DNA reverse complementary pairs with length m, filter out sequences whose sum of the number of bases "T" and "C" meets the preset condition, and construct Resource pool; Randomly pick a DNA sequence from the resource pool as the initial sequence Randomly pick a DNA sequence from the resource pool as the initial sequence , continue to pick subsequent sequences from the resource pool according to the prefix-suffix matching principle , continue to pick subsequent sequences from the resource pool according to the prefix-suffix matching principle , generate the target length of the code through the base cascade operation , generate the target length of the code through the base cascade operation , generate the target length of the code through the base cascade operation According to the base conversion rule, the DNA sequences in the encoding set are converted to binary encoding, forming encoding set ; Constructing a set of mutually orthogonal codes of a predetermined length A set of weakly mutually orthogonal codes The code words in both sets of codes satisfy the prefix constraint, the fixed bit constraint and the consecutive "0" constraint, and the set of weakly mutually orthogonal codes also satisfies the extension constraint. to the set of weakly mutually independent codes Screening optimization, combined with balance technology to get 0 / 1 balanced code, with the code set The corresponding length of the code in the middle is generated by decoupling structure double two mapping, which generates address sequence that meets the mutually independent, two-level structure avoidance and GC balance. For the set of mutually uncorrelated codes By adding error-correcting redundancy to the encoded words in the code, an error-correcting encoded word is obtained, which is then compared with the code. Encoding set The corresponding length of the encoded words are generated through a decoupling structure double binary mapping to produce an address sequence that is mutually uncorrelated, avoids secondary structures, and has error correction capabilities.
2. The DNA storage random access address encoding method of claim 1, wherein, The sum of the number of bases "T" and "C" satisfies a preset condition, specifically, the sum of the number of bases "T" and "C" is greater than m / 2; the length m of the reverse complementary pair is in a range of 3-9, and is used for inhibiting the generation of a DNA secondary structure with a stem length of 3-9.
3. The method of claim 1, wherein the random access address is encoded by a DNA storage method, and The prefix and suffix matching principle is as follows: the prefix of the subsequent sequence... The bases and the following bases in the current sequence The bases are completely identical; the base cascade operation is: deleting the first base of the subsequent sequence. The target length is obtained by: [The text abruptly ends here, so the translation stops as well.] ,in For stem length, From The total number of times a sequence is selected from the resource pool.
4. The method of claim 1, wherein the random access address is encoded by a DNA storage method. The base conversion rule is that the base "T" corresponds to binary "0", the base "C" corresponds to binary "0", the base "A" corresponds to binary "1", and the base "G" corresponds to binary "1".
5. The method of claim 1, wherein the random access address is encoded by a DNA storage method. The prefix constraint is: the prefix of each encoded word... The fixed bit constraint is: the first bit of each encoded word is "0"; Position, No. The bit cannot be "0"; the consecutive "0" constraint means that each encoded word starts from the first bit. The position reached the first The digits cannot be consecutive. The extended constraint is: in the first "0"; After the position, cascade any A binary sequence of bits.
6. The method of claim 1, wherein the random access address is encoded by a DNA storage method. Pruning weakly mutually independent encoding set From the first bit to the first bit, there is a sequence of consecutive '1', get the encoding set denoted as , and then randomly select a weakly mutually independent encoding word from the set , and apply Knuth's balance technique to it from the first bit to the first bit, get the 0 / 1 balanced encoding word ; in addition, randomly select an encoding word from the encoding set , and decouple the encoding word with length and as the component encoding of the decoupling structure to get the address sequence that satisfies the mutually independent, two-level structure avoidance and GC balance constraints at the same time . 7. The method of claim 1, wherein the random access address is encoded by a DNA storage method. From uncorrelated sets of codes Selecting coded words After adding LDPC check redundancy, the encoded word is obtained. ,in It is a coded word The check digit; secondly, from Encoding set Select Encoded words and the length is Encoded words and The component encoding, used as a decoupling structure, undergoes a bi-binary mapping to obtain an address sequence that simultaneously satisfies the requirements of being uncorrelated, avoiding secondary structures, and possessing error correction capabilities. .
8. A DNA storage random access address encoding system, characterized by, The method comprises the following steps: The resource pool construction module enumerates all DNA reverse complementary pairs with length m, filters sequences whose sum of the number of bases "T" and "C" satisfies a preset condition, and constructs a resource pool; The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool The encoding generation module generates an initial sequence from the resource pool a code conversion module for converting the DNA sequence in the code set into binary according to a base conversion rule a code conversion module for converting the DNA sequence in the code set into binary according to a base conversion rule a code conversion module for converting the DNA sequence in the code set into binary according to a base conversion rule a code conversion module for converting the DNA sequence in the code set into binary according to a base conversion rule a code conversion module for converting the DNA sequence in the code set into binary according to a base conversion rule Constraint module: construct a set of mutually independent codes with a preset length and a set of weak mutually independent codes The code words in both sets of codes satisfy the prefix constraint, fixed bit constraint and continuous "0" constraint, and the set of weak mutually independent codes also satisfies the extension constraint a balance address sequence generation module, for generating the weakly mutually independent code set The 0 / 1 balanced code word is obtained by screening optimization combined with balance technology, and the code set The address sequence satisfying mutual independence, two-level structure avoidance and GC balance is generated by double two-mapping of the decoupling structure of the code word of the corresponding length in the code set. The error correction address sequence generation module appends error correction redundancy to the code words in the mutually independent code set to obtain code words with error correction function, and generates an address sequence satisfying mutual independence, two-level structure avoidance and error correction capability through a decoupling structure double two mapping of the code words of corresponding length in the code set 9. An electronic device comprising a memory, a processor, and a computer program stored on the memory to run on the processor, characterized in that, The processor implements the DNA storage random access address coding method according to any one of claims 1-7 when the program is executed.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the DNA storage random access address coding method according to any one of claims 1-7.