A DNA information storage method with preview function
By generating a preview file binary sequence and converting it into a DNA base sequence, combined with efficient data compression coding and error correction coding, the problem of low retrieval efficiency in existing DNA information storage is solved, the file preview function is realized, and the retrieval cost and resource waste are reduced.
Patent Information
- Application Number
- CN202211004824.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-08-22
AI Technical Summary
Existing DNA information storage methods require sequencing and decoding of all bases during retrieval, which increases time and cost, especially for large data files. The lack of a preview function leads to low retrieval efficiency and waste of resources.
By generating a preview file binary sequence and converting it into a DNA base sequence, combining efficient data compression coding and error correction coding, segmenting and storing it in independent DNA molecule pools, and using PCR technology and gel electrophoresis separation and purification technology to achieve file preview function, unnecessary sequencing and decoding operations are reduced.
The preview function of files in DNA storage is realized, which reduces unnecessary sequencing and decoding operations, reduces retrieval costs, and improves information retrieval efficiency.
Smart Images

Figure CN115292256B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of DNA information storage, in particular to a DNA information storage method. BACKGROUND
[0002] With the rapid development of artificial intelligence technology and Internet applications, the global information quantity presents an explosive growth trend. At present, more than 60% of the data quantity of massive archival data belongs to cold data, and the frequency of access is low 1% after 90-120 days of data generation, but these data may have high value, so it is necessary to use a more economical and effective way than the existing hard disk or magnetic tape and other storage media to store. DNA medium has the advantages of high storage density, low energy consumption and long storage time, and can be used to store massive archival data with low access frequency, effectively reducing the storage cost and maintenance cost of data, and meeting the demand for ultra-low energy storage. At present, DNA information storage has gradually become a global research hotspot in storage technology.
[0003] The architecture of DNA information storage determines the data retrieval method, and is also the key to further promote the actualization of DNA information storage. At present, DNA information storage has gone from theory to practical application, but there are still many limitations. The existing DNA data storage only converts the original file into base sequence by using compression coding and error correction coding to artificially synthesize DNA molecules, and when the target file needs to be accessed, the target file is extracted by adding specific primers, and then sequencing and decoding reconstruction are performed. The data stored by DNA medium has the characteristics of large number of files and large amount of data per file, and when the retrieval method does not have a preview function, DNA sequencing and decoding reconstruction of all bases of the original file are required each time the retrieval is performed, so that the original file content can be viewed, which inevitably causes a large waste of manpower and material resources, especially for image and video files with large data volume, the increase of time and decoding cost is more significant.
[0004] Storing only the base sequence of the original file cannot meet the demand of DNA information retrieval under different demands, for example, not every access to the file needs to completely restore all the information and details of the file, but only some details or quickly screen out some files. In this application environment, the traditional single file storage mode has the disadvantages of low retrieval efficiency and large redundancy of readout information entropy relative to the purpose information entropy. Therefore, it is necessary to design a suitable and efficient file preview method in the application of DNA data storage and file retrieval. SUMMARY
[0005] In view of the limitations of the current DNA information storage technology, the present application provides a DNA information storage method with a preview function, which realizes the file preview function by providing two types of DNA information sequence construction methods.
[0006] The application discloses a DNA information storage method with a preview function.
[0007] In step S1, a binary sequence of an original file to be stored and a binary sequence of a corresponding preview file are generated, wherein the length of the binary sequence of the preview file is set to be between 0.9 kbit and 1.8 kbit.
[0008] In step S2, the binary sequence of the original file and the binary sequence of the corresponding preview file are converted into a DNA base sequence of the original file and a DNA base sequence of the corresponding preview file by using a high-efficiency data compression coding mode and an error correction coding mode, so that the whole DNA base sequence meets a constraint of synthetic biology.
[0009] In step S3, it is determined whether the cost of synthesizing and assembling a plurality of 2kbp-length DNA base sequences to completely encode an original file information and a corresponding preview file information is lower than the cost of directly synthesizing and assembling a DNA base sequence capable of completely containing the DNA base sequence of the original file and the DNA base sequence of the corresponding preview file; if yes, a first type of DNA information sequence construction method is adopted, and the method proceeds to step S4; otherwise, a second type of DNA information sequence construction method is adopted, and the method proceeds to step S5.
[0010] In step S4, the first type of DNA information sequence construction method is executed, and the 16bp-length random seed base sequence used for generating a random sequence and the DNA base sequence of the original file are divided into a plurality of 1.99kbp-length subsequences; if the length of the last subsequence is less than 1.99kbp, a random DNA sequence meeting the constraint of synthetic biology is used to fill the last subsequence to make the length of the last subsequence still 1.99kbp; after the division, 10bp-length address information DNA sequences are added in front of each subsequence, and 4bp-length CRC-8 / ITU check bits are added behind each subsequence; the length of the DNA sequence of the preview file is 0.5kbp-1.0kbp; a pair of same upstream and downstream primer binding sites are added at both ends of each subsequence and the DNA sequence of the preview file, so that the added upstream and downstream primer binding sites meet the requirement of screening the original file DNA molecule or the preview file DNA molecule of the target file by using a PCR technology and a gel electrophoresis separation and purification technology, and the two artificially synthesized DNA molecules are respectively stored in independent DNA molecule pools.
[0011] S5. Execute the second DNA information sequence construction method. The random seed information used to generate the pseudorandom sequence is stored before the original file DNA base sequence obtained in step S2. A 4-bp CRC-8 / ITU check digit is added to the end of each sequence. Finally, a pair of identical upstream and downstream primer binding sites are added to both ends of the original file sequence and the preview file sequence, respectively, to form a single DNA sequence. The original file DNA sequence of each DNA molecule contains all valid encoding information for a file, with no length limit. The preview file DNA sequence length is 0.5 kbp to 1.0 kbp.
[0012] Different from the existing DNA information storage method, the DNA information storage method with a preview function of the present invention implements the file preview function on a traditional computer in the DNA information storage system, which greatly facilitates the file retrieval process of DNA storage and reduces unnecessary sequencing, decoding and reconstruction operations. When the preview file is the required target file and the information it contains cannot meet the requirements, the original file base sequence is sequenced, decoded and reconstructed, which can effectively reduce the cost of file retrieval in DNA storage and improve information retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is an overall flow chart of a DNA information storage method with a preview function according to the present invention.
[0014] Figure 2 This is an example diagram of the first type of DNA information sequence structure of the present invention.
[0015] Figure 3 This is an example diagram of the second type of DNA information sequence structure of the present invention.
[0016] Figure 4 This is a flow chart of the storage and retrieval of the first type of DNA information sequence of the present invention.
[0017] Figure 5 This is a flow chart of the storage and retrieval of the second type of DNA information sequence of the present invention. DETAILED DESCRIPTION
[0018] The technical solution of the present invention is described in further detail below with reference to the accompanying drawings.
[0019] like Figure 1 The figure shows a flow chart of a DNA information storage method with a preview function according to the present invention. The specific implementation steps are as follows:
[0020] Step S1, generate the original file binary sequence and the corresponding preview file binary sequence of all files to be stored, set the length of the preview file binary sequence between 0.9kbit-1.8kbit, ensure that the length is between 0.5kbp-1.0kbp after conversion to DNA base sequence, and the length of the original file binary sequence has no requirement; wherein, for video files, the video summary generation method is used to generate preview files; for image files, the wavelet transform downsampling method is used to generate preview files; for audio and text files, the preview file is generated by using the neural network based on Transformer;
[0021] Step S2, convert the original file binary sequence and the corresponding preview file binary sequence into original file DNA base sequence and corresponding preview file DNA base sequence by using high-efficiency data compression encoding method and error correction encoding method, and use the method of superimposing watermark on the pseudo-random sequence generated based on the random seed to make the whole DNA base sequence meet the constraints of synthetic biology, add random seed information at the front end and CRC-8 / ITU check bits at the end of the preview file DNA base sequence, and for the original file DNA base sequence, use the same method to add random seed information and check bit information;
[0022] Step S3, considering the cost constraint of DNA information synthesis, judge whether it meets the cost of synthesizing and assembling a plurality of 2kbp length original file DNA base sequences and a preview file DNA base sequence to make the complete coding of all effective information contained in a file lower than the cost of directly synthesizing and assembling a DNA base sequence that can completely contain the original file DNA base sequence and the corresponding preview file DNA base sequence; if so, use the first type of DNA information sequence construction method, go to step S4; otherwise, use the second type of DNA information sequence construction method, go to step S5; the DNA sequence length of the original file is very long, generally reaching Mbp or even Gbp order;
[0023] Step S4, a first type of DNA information sequence construction method is performed, the 16bp length random seed base sequence for generating a random sequence and the original file DNA base sequence are divided into several 1.99kbp length subsequences, if the length of the last subsequence is less than 1.99kbp, a random DNA sequence meeting the constraints of synthetic biology is used to fill to make the length still 1.99kbp, after the division, a 10bp length address information DNA sequence is added in front of each subsequence, and a 4bp length CRC-8 / ITU check bit is added behind each subsequence; the preview file DNA sequence length is 0.5kbp-1.0kbp; a same pair of upstream and downstream primer binding sites is added at both ends of each subsequence and the preview file sequence, so that the upstream and downstream primer binding sites meet the requirements of screening the original file DNA molecule or the preview file DNA molecule of the target file by using the PCR technology and the gel electrophoresis separation and purification technology, and the two artificially synthesized DNA molecules are stored in independent DNA molecule pools. When file retrieval is performed, PCR amplification is performed in the preview file DNA sequence pool, and the preview file of the target file is obtained through sequencing, decoding and reconstruction; the same primer is used for PCR amplification in the original file DNA sequence pool, and the original file of the target file is obtained through sequencing, decoding and reconstruction. As shown in Figure 2 Fig. 1 is a first type of DNA information sequence structure example diagram of the present application.
[0024] Step S5, a second type of DNA information sequence construction method is performed, the random seed information for generating a pseudo-random sequence is stored in front of the original file DNA base sequence obtained in step S2, a 4bp length CRC-8 / ITU check bit is added behind each sequence, and finally a same pair of upstream and downstream primer binding sites is added at both ends of the original file sequence and the preview file sequence, and is integrated into a DNA sequence. The original file DNA sequence of each DNA molecule contains all the effective encoding information of a file, and has no length limit, and the preview file DNA sequence length is 0.5kbp-1.0kbp; when file retrieval is performed, a specific primer is used for a PCR amplification reaction in the DNA molecule pool, and then a gel electrophoresis separation and purification technology is used to cut out a preview file short DNA molecule band length of 0.5kbp-1.0kbp as a target band, and the target preview file is obtained through sequencing, decoding and reconstruction; when the preview file is the target file And that needs to be retrieved, a long DNA molecule band length of 2kbp or above of the original file is cut out as a target band of the original file, and the target original file is obtained through sequencing, decoding and reconstruction. As shown in Figure 3 Fig. 2 is a second type of DNA information sequence structure example diagram of the present application.
[0025] The DNA information storage method with preview function of the present application, the first type of DNA information sequence construction method is that the method divides the original file DNA base sequence into a plurality of short sequences, and stores the divided original file DNA base sequence and the preview file DNA base sequence into two DNA pools respectively. The second type of DNA information sequence construction method is to integrate the original file DNA base sequence and the corresponding preview file DNA base sequence on the same DNA molecule.
[0026] The present application introduces the file preview function in the traditional computer into the DNA storage. For the original file, after compression encoding and error correction encoding, the DNA sequence composed of A, T, C and G four bases is generated, and the length of the sequence is generally in the order of Mbp or even Gbp; for the preview file, the DNA sequence with length of 0.5kbp-1kbp is generated through downsampling, feature extraction and other compression algorithms. The sequence length of the added upstream and downstream primer binding sites is 20bp, which is used to realize the random retrieval function of different files. After the original file sequence and the corresponding preview file sequence are generated, the subsequent operations such as storage, retrieval and preview are carried out in the form of actual DNA molecule according to the structure of two types of DNA information sequence.
[0027] The above examples are only for illustrating the technical concept and characteristics of the present application, the purpose is to enable the person skilled in the art to understand the content of the present application and to implement it, and cannot limit the protection scope of the present application. Any equivalent changes or modifications made according to the spirit and essence of the present application shall be covered within the protection scope of the present application.
Claims
1. A DNA information storage method with a preview function, characterized in that: The specific implementation steps of this method are as follows: Step S1: Generate a binary sequence of the original file to be stored and a corresponding binary sequence of the preview file, wherein the length of the binary sequence of the preview file is set between 0.9 kbit and 1.8 kbit; Step S2: using an efficient data compression coding method and an error correction coding method, and according to the constraints of synthetic biology, converting the original file binary sequence and the corresponding preview file binary sequence into the original file DNA base sequence and the corresponding preview file DNA base sequence respectively; Step S3: When synthesizing and assembling the original file DNA base sequence and the corresponding preview file DNA base sequence, determine whether the cost of using multiple 2 kbp-long original file DNA base sequences and one preview file DNA base sequence is lower than the cost of directly using one DNA base sequence that can completely contain the original file DNA base sequence and the corresponding preview file DNA base sequence; if so, use the first type of DNA information sequence construction method and go to step S4; otherwise, use the second type of DNA information sequence construction method and go to step S5; Step S4, executing the first type of DNA information sequence construction method, dividing the 16bp random seed base sequence used to generate the random sequence and the original file DNA base sequence into several subsequences of 1.99kbp in length. If the length of the last subsequence is less than 1.99kbp, it is padded with a random DNA sequence that meets the constraints of synthetic biology to make its length still 1.99kbp; after the segmentation, a 10bp address information DNA sequence is added before each subsequence and a 4bp CRC-8 / ITU check bit is added after each subsequence; the preview file DNA sequence length is 0.5kbp-1.0kbp; and the same pair of upstream and downstream primer binding sites are added at both ends of each subsequence and the preview file sequence, so that the added upstream and downstream primer binding sites meet the requirements of the utilization PCR technology and gel electrophoresis separation and purification technology are used to screen the original file DNA molecules or preview file DNA molecules of the target file, and the two artificially synthesized DNA molecules are stored in independent DNA molecule pools respectively. When performing file retrieval, a PCR amplification reaction is performed in the DNA molecule pool using specific primers, and then gel electrophoresis separation and purification technology is used to cut out a short DNA molecule band of the preview file with a length of 0.5kbp-1.0kbp as the target band, and then sequence, decode and reconstruct it to obtain the target preview file. When the preview file is the required target file and the effective information content or resolution of the preview file is low and cannot meet the retrieval requirements, a long DNA molecule band of the original file with a length of 2kbp or more is cut out as the target band of the original file, and then sequence, decode and reconstruct it to obtain the target original file. S5. Execute the second type of DNA information sequence construction method, save the random seed information used to generate the pseudo-random sequence before the original file DNA base sequence obtained in step S2, add a 4bp CRC-8 / ITU check bit after each sequence, and finally add a pair of identical upstream and downstream primer binding sites at both ends of the original file sequence and the preview file sequence, respectively, and integrate them into a DNA sequence; wherein, the original file DNA sequence of each DNA molecule contains all the valid coding information of a file and has no length limit. The preview file DNA sequence length is 0.5kbp-1.0kbp.
2. A DNA information storage method with a preview function as claimed in claim 1, characterized in that: The step S1 also includes: for video files, using a video summary generation method to generate a preview file; for image files, using a wavelet transform downsampling method to generate a preview file; for audio and text files, using a Transformer-based neural network to generate a preview file.
3. The DNA information storage method with preview function according to claim 1, characterized in that: The step S2 also includes: adding random seed information to the front end of the preview file DNA base sequence and adding CRC-8 / ITU check bit to the end. For the original file DNA base sequence, the same method is used to add random seed information and check bit information.
4. The DNA information storage method with preview function according to claim 1, characterized in that: The step S4 also includes: when performing file retrieval, performing PCR amplification in a DNA sequence pool storing preview files, and obtaining a preview file of the target file through sequencing, decoding, and reconstruction; performing PCR amplification using the same primers in a DNA sequence pool storing original files, and then obtaining the original file of the target file through sequencing, decoding, and reconstruction.
Citation Information
Patent Citations
Method for information storage with DNA (Deoxyribonucleic Acid)
CN106845158A