Novel block storage system and method based on DNA storage medium
By introducing two-layer translation tables and off-site update technologies in the DNA storage system, the problems of poor compatibility with existing DNA storage solutions and insufficient metadata storage capacity are solved, and efficient and low-cost DNA block storage and massive data management are achieved.
Patent Information
- Application Number
- CN202510172668.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing DNA storage solution is difficult to compatible with existing computer systems and cannot be applied in massive data storage scenarios, mainly due to its key-value pair-based storage interface design and metadata storage problems.
A new block storage system based on DNA storage media was designed, using two-layer translation tables and off-site update technology, providing block interfaces for compatibility with computer systems, and achieving efficient read and write operations through DNA sequencing and synthesis technology.
It realizes the connection between the DNA storage system and the computer system, improves the read and write performance and capacity utilization of DNA block equipment, reduces storage costs, and supports efficient garbage collection and low-cost DNA overwrite operations.
Smart Images

Figure CN119993239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a novel block storage system and method based on a DNA storage medium. Background Art
[0002] Current computer data is usually stored in traditional storage media, including solid-state drives, magnetic disks, and optical disks. However, due to the explosive growth of data worldwide, 64.2ZB of data was generated in 2020 alone. The storage density of traditional storage media is difficult to meet storage needs. The total capacity of storage devices produced in 2020 is only 6.7ZB, which has led to the loss of a lot of valuable data.
[0003] As a new type of storage medium, DNA storage has quickly attracted people's attention due to its ultra-high storage density and ultra-long storage life. Its storage density can reach 109GB / mm3, and its storage life can reach hundreds of years. With its advantages in density and life, DNA is expected to become a solution for global mass data storage.
[0004] However, existing DNA storage solutions have some limitations. First, existing DNA storage solutions are all based on key-value pair storage interface design, while the most widely used storage interface in computers is the block interface. This makes it difficult for existing DNA storage interfaces to be reused by existing computers, increasing the difficulty of accessing computer software stacks. Secondly, existing DNA storage solutions store metadata in traditional storage media, while DNA's EB-level massive storage capacity allows the metadata capacity to reach the PB level, which is difficult for traditional storage media to accommodate. This makes existing DNA storage solutions unsuitable for massive data storage scenarios. Summary of the invention
[0005] In view of the defects in the prior art, the object of the present invention is to provide a new block storage system and method based on DNA storage medium.
[0006] A novel block storage system based on a DNA storage medium provided by the present invention comprises: a DNA storage chip, a processor and a memory, and a storage device;
[0007] The DNA storage chip is used to store a first translation table and block data; wherein the first translation table is used to index the position of the block data in the DNA storage chip;
[0008] The processor and memory are used to process read and write requests and obtain the block number to be read or the data to be written;
[0009] The storage device is used to store the zeroth translation table; the zeroth translation table is used to index the position of the first translation table in the DNA storage chip.
[0010] Preferably, for the read operation, the block number in the user request is obtained based on the processor and the memory, the zeroth translation table in the solid-state drive is read according to the block number, the first translation table stored in the DNA storage chip is obtained according to the zeroth translation table, the location of the block is indexed according to the first translation table, the block data is read through the DNA sequencing technology based on the location of the block, and it is decoded into binary data and transmitted to the user end.
[0011] Preferably, in the reading operation, the revision record in the first translation table is obtained, and it is determined whether old data exists according to the revision record in the first translation table. If old data exists, an invalidation link is inserted at the position corresponding to the old data to invalidate it.
[0012] Preferably, for write operations, binary data is encoded into a DNA sequence, the DNA sequence is synthesized into corresponding DNA data through DNA synthesis technology and stored in an unused physical block address, and the first translation table and the zeroth translation table are updated based on the physical block address where the synthesized DNA data is located.
[0013] Preferably, the write operation includes:
[0014] Module M1: Encode binary data into DNA sequence, then synthesize the corresponding DNA data through DNA synthesis technology and store it in a test tube;
[0015] Module M2: Check whether there is a space that meets the preset requirements in the test tube corresponding to the first translation table. If so, module M6 is triggered; otherwise, module M3 is triggered;
[0016] Module M3: using DNA sequencing technology to read all valid translation entries corresponding to the first translation table;
[0017] Module M4: Select an idle DNA test tube and rewrite the read valid translation entries into the test tube using DNA synthesis technology;
[0018] Module M5: Update the position of the first translation table in the zeroth translation table;
[0019] Module M6: synthesize the DNA chain with the latest timestamp and insert it into the first translation table to update the corresponding entry.
[0020] Preferably, for the garbage collection operation, it is determined whether the free space in the DNA block device is greater than a preset value. If it is greater than the preset value, no garbage collection operation is required; if it is less than or equal to the preset value, a garbage collection operation is performed;
[0021] In the garbage collection operation, a test tube with the most garbage data is obtained, and DNA sequencing technology is used to read all data in the current test tube and metadata stored symbiotically with the data, and valid data therein is found through the symbiotic metadata; the valid data is written into a new free physical block, and the entry of the valid data in the first translation table is updated; and an erase process is performed for the test tube.
[0022] According to a novel block storage method based on a DNA storage medium provided by the present invention, the novel block storage system based on a DNA storage medium is used to perform the following steps:
[0023] Step S1: For the read operation, the logical block address provided by the user is translated into a physical block address, and the physical block address is DNA sequenced to read the block data;
[0024] Step S2: For a write operation, the block data is written into an unused physical block address by DNA synthesis, and the corresponding entry of the logical block in the translation table is modified to map it to the physical block address of the data storage.
[0025] Preferably, the step S1 comprises:
[0026] Step S1.1: read the zeroth translation table in the solid state drive according to the block number in the user request, and find the location of the first translation table in the DNA storage chip through the zeroth translation table;
[0027] Step S1.2: Read the first translation table in the DNA storage chip, and obtain the physical location of the data block according to the entries in the first translation table;
[0028] Step S1.3: judging whether there is old data according to the revision record in the first translation table, if yes, triggering step S1.4; if no, triggering step S1.5;
[0029] Step S1.4: For old data, insert an invalidation link at the position corresponding to the old data to invalidate it;
[0030] Step S1.5: Read the block data stored in the DNA storage chip through DNA sequencing technology and decode it into binary data.
[0031] Preferably, step S2 comprises:
[0032] Step S2.1: Encode the user's binary data into a DNA sequence, then synthesize the corresponding DNA data through DNA synthesis technology and store it in a test tube;
[0033] Step S2.2: Check whether there is sufficient space for writing in the test tube corresponding to the first translation table. If yes, step S2.6 is triggered; if not, step S2.3 is triggered;
[0034] Step S2.3: using DNA sequencing technology to read all valid translation entries corresponding to the first translation table;
[0035] Step S2.4: Select an idle DNA test tube and rewrite the read valid entries into the test tube using DNA synthesis technology;
[0036] Step S2.5: The position of the first translation table updated in the zeroth translation table;
[0037] Step S2.6: Synthesize the DNA chain with the latest timestamp and insert it into the first translation table to update the corresponding entry.
[0038] Preferably, the method further comprises:
[0039] Step S3: For garbage collection operation, determine whether the free space in the DNA block device is greater than a preset value. If it is greater than the preset value, no garbage collection operation is required; if it is less than or equal to the preset value, a garbage collection operation is performed;
[0040] The step S3 comprises:
[0041] Step S3.1: Determine whether there is sufficient free space in the DNA block device. If yes, terminate the operation directly; if no, trigger step S3.2;
[0042] Step S3.2: Find the test tube with the most junk data;
[0043] Step S3.3: Use DNA sequencing technology to read all the data in the test tube and the metadata stored symbiotically with the data, and find the valid data in it through the symbiotic metadata;
[0044] Step S3.4: Write valid data into a new free physical block;
[0045] Step S3.5: Update the entry of valid data in the first translation table;
[0046] Step S3.6: Perform an erase operation on this test tube.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] 1. The present invention provides a highly versatile block interface for DNA storage, allowing the DNA storage system to reuse existing computer software, thereby achieving the integration of DNA storage and computer systems;
[0049] 2. Through the technology of off-site update and double-layer translation table, high-performance and low-cost read and write operations of DNA block devices are achieved;
[0050] 3. The present invention achieves low overhead of garbage collection through the technology of symbiotic metadata;
[0051] 4. The present invention realizes low-cost DNA overwriting operation by delaying the invalidation of junk data. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0053] Figure 1 It is the reading and writing flow chart of the present invention.
[0054] Figure 2 Schematic diagram of the system architecture of the present invention.
[0055] Figure 3 This is a flow chart of the garbage collection operation of the present invention.
[0056] Figure 4 This is a schematic diagram of a novel block storage system based on DNA storage media according to the present invention. DETAILED DESCRIPTION
[0057] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0058] Example 1
[0059] According to a novel block storage system based on DNA storage medium provided by the present invention, Figure 4 As shown, it includes: DNA storage chip, processor and memory and storage device;
[0060] The DNA storage chip is used to store a first translation table and block data; wherein the first translation table is used to index the position of the block data in the DNA storage chip;
[0061] The processor and memory are used to process read and write requests and obtain the block number to be read or the data to be written;
[0062] The storage device is used to store the zeroth translation table; the zeroth translation table is used to index the position of the first translation table in the DNA storage chip.
[0063] In this embodiment, the storage device is a solid state drive.
[0064] A novel block storage system implementation method based on a DNA storage medium provided by the present invention includes:
[0065] like Figure 1 As shown, for the read operation, it includes:
[0066] Step 101, first read the zeroth translation table in the solid state drive through the block number in the user request, find the location of the first translation table in the DNA storage device through the zeroth translation table, and then execute step 102.
[0067] Step 102 , read the first translation table in the DNA, obtain the physical location of the data block according to the entries in the first translation table, and then execute step 103 .
[0068] Step 103, judging whether there is old data according to the revision record in the first translation table. If yes, step 104 is triggered; if not, step 105 is executed;
[0069] Step 104 , for the old data in step 103 , insert an invalidation link at the position corresponding to the old data to invalidate it, and then trigger the execution of step 105 .
[0070] Step 105, read the block data stored in the DNA storage chip through DNA sequencing technology, and decode it into binary data, and finally end the operation.
[0071] This embodiment invalidates the old data and further improves the writing performance through delayed invalidation technology. By postponing the invalidation of the old data during the overwriting process until the reading, the additional reading operation required during the writing is eliminated, the writing performance is greatly improved, and a low-cost DNA overwriting operation is achieved.
[0072] like Figure 2 As shown, for write operations, it includes:
[0073] Step 201, encode the user's binary data into a DNA sequence, then synthesize the corresponding DNA data through DNA synthesis technology and store it in a test tube, and then execute step 202.
[0074] Step 202, check whether there is sufficient space in the test tube corresponding to the first translation table for writing. If yes, execute step 206; if not, execute step 203.
[0075] Step 203 , using DNA sequencing technology to read all valid translation entries corresponding to the first translation table, and then executing step 204 .
[0076] Step 204 , select an idle DNA test tube, rewrite the read valid entries into the test tube using DNA synthesis technology, and execute step 205 .
[0077] Step 205 , update the position of the first translation table in the zeroth translation table, and then execute step 206 .
[0078] Step 206: synthesize a DNA chain with the latest timestamp and insert it into the first translation table to update the corresponding entry, and then end the operation.
[0079] like Figure 3 As shown, for garbage collection operations, including:
[0080] Step 301, determine whether there is sufficient free space in the DNA block device. If yes, directly terminate the operation; if not, execute step 302.
[0081] Step 302, find the test tube with the most junk data, and then execute step 303.
[0082] Step 303, using DNA sequencing technology to read all the data in the test tube and the metadata stored symbiotically with the data, find the valid data therein through the symbiotic metadata, and then execute step 304.
[0083] Step 304 , write the valid data into the new free physical block, and then execute step 305 .
[0084] Step 305 , updating the entry of valid data in the first translation table, and then executing step 306 .
[0085] Step 306, performing an erase operation on the test tube, and finally ending the operation.
[0086] This embodiment provides users with an efficient garbage collection mechanism through the technology of symbiotic metadata. The translation table entries are stored in the unused data space in the data block, eliminating the process of traversing the translation table during garbage collection, greatly improving the garbage collection performance and achieving low overhead for garbage collection.
[0087] Example 2
[0088] Embodiment 2 is a preferred embodiment of Embodiment 1
[0089] A novel block storage system based on a DNA storage medium provided by the present invention comprises: a DNA storage chip, a processor and a memory, and a solid state hard disk;
[0090] Among them, the DNA storage chip is used to actually store DNA data; the processor and memory are used to process read and write requests; and the solid-state hard disk is used to store metadata.
[0091] More specifically, the data block layer is several EB in size and is located in the DNA storage chip, which is responsible for storing the user's data blocks;
[0092] The first translation layer is several petabytes in size and is located in the DNA storage chip. It is a translation table that indexes the location of the blocks.
[0093] The zeroth translation layer is several GB in size and is located on a solid-state drive. It is used to index the location of the first translation layer.
[0094] In this embodiment, the novel block storage system based on DNA storage medium provides users with a read and write operation interface; the novel block storage system based on DNA storage medium abstracts data into fixed-size blocks, which are the minimum operation units for user read and write operations. Each block has its corresponding block number, and its size is usually 4,096 bytes. For read operations, the logical block address provided by the user is first translated into a physical block address using a translation table, and then the physical block address is DNA sequenced to read the block data and return it to the user. For write operations, the block data is first DNA synthesized and written to an unused physical block address, and then the corresponding entry of the logical block in the translation table is modified to map it to the physical block address of the data storage.
[0095] In this embodiment, the zeroth translation table has a relatively small capacity and is stored in a traditional storage medium, while the first translation table has a relatively large capacity and is stored in DNA. When reading the translation, the zeroth translation table in the traditional storage medium is first read to quickly find the location of the first translation table in the DNA, and then the first translation table is read through DNA sequencing to determine the physical block address corresponding to the logical block address; when writing and modifying the translation table entry, the first translation table stored in the DNA is updated by synthesizing a DNA chain with the latest timestamp.
[0096] Through the technology of off-site update, an efficient read and write operation interface is provided for users. Off-site update is realized by adding a double-layer translation table, in which the zeroth translation layer is stored in the traditional storage medium and the first translation layer is stored in DNA.
[0097] When the free space of the device is scarce, the device will first search for garbage data according to the translation table, then migrate the valid data in the same erase space as the garbage data through a read-and-rewrite operation, and finally erase all data in the erase space through an erase operation to free up the space.
[0098] The translation table is stored together with the data through symbiotic metadata to speed up garbage collection; the translation table entries are stored in the unused data space in the data block. During the process of garbage collection and migration of valid data, the corresponding translation table entry content can be obtained by directly reading the valid data without the need to read the entire translation table entry.
[0099] By delaying invalidation, the process of invalidating old data in the writing process is postponed to the reading process; when writing and updating the translation table entry, the first translation table is updated by synthesizing a DNA chain with the latest timestamp, without the need to read the translation table additionally to find the location of the old data. When reading the first translation table in the reading operation, the data block that needs to be outdated is found according to the timestamp, and a marker chain is inserted into the corresponding outdated data block to mark it as invalid.
[0100] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.
[0101] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A novel block storage system based on DNA storage medium, characterized in that: include: DNA memory chips, processors and memory, and storage devices; The DNA storage chip is used to store a first translation table and block data; wherein the first translation table is used to index the position of the block data in the DNA storage chip; The processor and memory are used to process read and write requests and obtain the block number to be read or the data to be written; The storage device is used to store the zeroth translation table; the zeroth translation table is used to index the position of the first translation table in the DNA storage chip.
2. The novel block storage system based on DNA storage medium according to claim 1 is characterized in that: For the read operation, the block number in the user request is obtained based on the processor and memory, the zeroth translation table in the solid-state drive is read according to the block number, the first translation table stored in the DNA storage chip is obtained according to the zeroth translation table, the location of the block is indexed according to the first translation table, the block data is read through DNA sequencing technology based on the location of the block, and it is decoded into binary data and transmitted to the user end.
3. The novel block storage system based on DNA storage medium according to claim 2 is characterized in that: In the reading operation, the revision record in the first translation table is obtained, and it is determined whether old data exists according to the revision record in the first translation table. If old data exists, an invalidation link is inserted at the position corresponding to the old data to invalidate it.
4. The novel block storage system based on DNA storage medium according to claim 1 is characterized in that: For write operations, the binary data is encoded into a DNA sequence, and the DNA sequence is synthesized into corresponding DNA data through DNA synthesis technology and stored in an unused physical block address, and the first translation table and the zeroth translation table are updated based on the physical block address where the synthesized DNA data is located.
5. The novel block storage system based on DNA storage medium according to claim 4 is characterized in that: The write operation includes: Module M1: Encode binary data into DNA sequence, then synthesize the corresponding DNA data through DNA synthesis technology and store it in a test tube; Module M2: Check whether there is a space that meets the preset requirements in the test tube corresponding to the first translation table. If so, module M6 is triggered; otherwise, module M3 is triggered; Module M3: using DNA sequencing technology to read all valid translation entries corresponding to the first translation table; Module M4: Select an idle DNA test tube and rewrite the read valid translation entries into the test tube using DNA synthesis technology; Module M5: Update the position of the first translation table in the zeroth translation table; Module M6: synthesize the DNA chain with the latest timestamp and insert it into the first translation table to update the corresponding entry.
6. The novel block storage system based on DNA storage medium according to claim 1 is characterized in that: For garbage collection operation, it is determined whether the free space in the DNA block device is greater than a preset value. If it is greater than the preset value, no garbage collection operation is required; if it is less than or equal to the preset value, a garbage collection operation is performed; In the garbage collection operation, a test tube with the most garbage data is obtained, and DNA sequencing technology is used to read all data in the current test tube and metadata stored symbiotically with the data, and valid data therein is found through the symbiotic metadata; the valid data is written into a new free physical block, and the entry of the valid data in the first translation table is updated; and an erase process is performed for the test tube.
7. A novel block storage method based on DNA storage medium, characterized in that: The novel block storage system based on DNA storage medium according to claim 1 is used to perform the following steps: Step S1: For the read operation, the logical block address provided by the user is translated into a physical block address, and the physical block address is DNA sequenced to read the block data; Step S2: For a write operation, the block data is written into an unused physical block address by DNA synthesis, and the corresponding entry of the logical block in the translation table is modified to map it to the physical block address of the data storage.
8. The novel block storage method based on DNA storage medium according to claim 7 is characterized in that: The step S1 comprises: Step S1.1: read the zeroth translation table in the solid state drive through the block number in the user request, and find the location of the first translation table in the DNA storage chip through the zeroth translation table; Step S1.2: Read the first translation table in the DNA storage chip, and obtain the physical location of the data block according to the entries in the first translation table; Step S1.3: judging whether there is old data according to the revision record in the first translation table, if yes, triggering step S1.4; if no, triggering step S1.5; Step S1.4: For old data, insert an invalidation link at the position corresponding to the old data to invalidate it; Step S1.5: Read the block data stored in the DNA storage chip through DNA sequencing technology and decode it into binary data.
9. The novel block storage method based on DNA storage medium according to claim 7, characterized in that: The step S2 comprises: Step S2.1: Encode the user's binary data into a DNA sequence, then synthesize the corresponding DNA data through DNA synthesis technology and store it in a test tube; Step S2.2: Check whether there is sufficient space for writing in the test tube corresponding to the first translation table. If yes, step S2.6 is triggered; if not, step S2.3 is triggered; Step S2.3: using DNA sequencing technology to read all valid translation entries corresponding to the first translation table; Step S2.4: Select an idle DNA test tube and rewrite the read valid entries into the test tube using DNA synthesis technology; Step S2.5: The position of the first translation table updated in the zeroth translation table; Step S2.6: Synthesize the DNA chain with the latest timestamp and insert it into the first translation table to update the corresponding entry.
10. The novel block storage method based on DNA storage medium according to claim 7, characterized in that: The method further comprises: Step S3: For garbage collection operation, determine whether the free space in the DNA block device is greater than a preset value. If it is greater than the preset value, no garbage collection operation is required; if it is less than or equal to the preset value, a garbage collection operation is performed; The step S3 comprises: Step S3.1: Determine whether there is sufficient free space in the DNA block device. If yes, terminate the operation directly; if no, trigger step S3.2; Step S3.2: Find the test tube with the most junk data; Step S3.3: Use DNA sequencing technology to read all the data in the test tube and the metadata stored symbiotically with the data, and find the valid data in it through the symbiotic metadata; Step S3.4: Write valid data into a new free physical block; Step S3.5: Update the entry of valid data in the first translation table; Step S3.6: Perform an erase operation on this test tube.