Reading method and system for DNA data storage

CN120660141APending Publication Date: 2025-09-16SHENZHEN HUADA GENE INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202380092857.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The information stored in DNA data cannot be read instantly, and the reading speed of a single file is much lower than that of existing storage media, resulting in storage efficiency problems.

Method used

The specific sequence is used to locate the bases to be sequenced, combine the single-round sequencing of high-throughput sequencing output data, and decode part of the data after each round of sequencing. By capturing the DNA sequence library designed by the localization area and the data area, the sequencing chip and Magnetic beads with specific marks enable fast data reading.

Benefits of technology

Realize instant decoding of DNA data storage, improve information reading efficiency, and improve the possibility of large-scale application of DNA data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660141A_ABST
    Figure CN120660141A_ABST
Patent Text Reader

Abstract

The invention discloses a DNA sequence library for continuous data storage, a continuous data storage method and a corresponding reading method and system. The reading method comprises the following steps: 1) capturing a DNA sequence of a DNA sequence library by using a capturing and positioning region, wherein the capturing and positioning region corresponds to spatial information of continuous data; and 2) sequencing the data area of the DNA sequence, the data area corresponding to the continuous data content, and restoring the continuous data content along with the sequencing. According to the method, the reading mode of DNA data storage is innovated to realize instant decoding, and the problem of information reading efficiency in DNA data storage is solved, so that the large-scale application possibility of DNA data storage is improved.
Need to check novelty before this filing date? Find Prior Art

Description

DNA data storage reading method and system Technical Field

[0001] The present invention relates to the fields of biotechnology and information, and in particular to a method and system for reading DNA data storage. Background Art

[0002] With the development of modern technology, especially the Internet, the amount of data in the world is increasing exponentially. The ever-increasing amount of data places higher and higher demands on storage technology. Traditional storage technologies, such as magnetic tape and optical disc storage, are increasingly unable to meet current data needs due to limited storage density and time. In recent years, the development of DNA storage technology has provided a new way to solve these problems. Compared with traditional storage media, DNA as a medium for information storage has the advantages of long storage time (up to thousands of years, more than a hundred times that of existing magnetic tape and optical disc media), high storage density (reaching ~10 9 Gb / mm 3 , which is more than 10 million times that of existing tape and CD media) and has good storage security.

[0003] DNA data storage typically involves the following steps: 1) Encoding: Converting the binary 0 / 1 code of computer information into DNA sequence information of A / T / C / G; 2) Synthesis: Using DNA synthesis technology to synthesize the corresponding DNA sequence, and storing the obtained chemical DNA molecules in an in vitro medium or living cells; 3) Sequencing: Using sequencing technology to read the DNA sequence of the stored DNA molecules; 4) Decoding: Using the method corresponding to the encoding process in step 1, the sequenced DNA sequence is converted into a binary 0 / 1 code, and further converted into computer information.

[0004] One of the current challenges with DNA storage is that information cannot be accessed instantly; the data must be fully sequenced before it can be decoded. While sequencing throughput is very high, the absolute read speed for a single file is far lower than that of existing storage media. Therefore, to address the information access efficiency issues associated with DNA storage, a method for "instant" access, similar to the buffered loading of web page content, is needed.

[0005] Summary of the Invention

[0006] The present invention aims to provide a method for reading DNA data storage, using specific sequences to locate the bases to be sequenced. This method combines the output data of a single round of high-throughput sequencing with an encoding scheme to obtain partial raw data after each round of sequencing and perform corresponding decoding. This method allows for relatively rapid acquisition of partial data during the sequencing process.

[0007] In order to achieve the above objectives, the following technical solutions are adopted:

[0008] A DNA sequence library for continuous data storage, wherein the DNA sequence comprises a capture location region and a data region, wherein the capture location region corresponds to the spatial information of the continuous data and the data region corresponds to the content of the continuous data.

[0009] Furthermore, both ends of the DNA sequence also include primer regions.

[0010] A method for continuous data storage, the method comprising:

[0011] 1) generating a DNA sequence library for continuous data, wherein the DNA sequence includes a capture location region and a data region, wherein the capture location region corresponds to spatial information of the continuous data, and the data region corresponds to content of the continuous data;

[0012] 2) Synthesizing the DNA sequence library.

[0013] A method for reading DNA data storage, the method comprising:

[0014] 1) capturing DNA sequences of a DNA sequence library using a capture location region, wherein the capture location region corresponds to spatial information of continuous data;

[0015] 2) Sequencing a data region of the DNA sequence, the data region corresponding to continuous data content, and restoring the continuous data content as the sequencing proceeds.

[0016] Furthermore, the capture of DNA sequences is achieved through a sequencing chip.

[0017] Furthermore, the capture of DNA sequences is achieved by magnetic beads with specific labels.

[0018] Furthermore, each microarray on the sequencing chip is connected to a specific DNA sequence capture probe via a solid phase carrier and is matched with a coordinate positioning sequence.

[0019] Furthermore, the capture positioning region is a specific sequence that matches the coordinate positioning sequence and can be captured by the capture probe.

[0020] A DNA data storage reading system, the system comprising:

[0021] a capture unit configured to capture a DNA sequence storing continuous data;

[0022] a sequencing unit configured to sequence the captured DNA sequence;

[0023] A decoding unit configured to decode DNA sequence sequencing data;

[0024] A storage unit is configured to store the sequencing data and the decoding data.

[0025] Furthermore, the DNA sequence includes a capture location region and a data region, the capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data.

[0026] By adopting the above scheme, the beneficial effects of the present invention are: innovating the reading method of DNA data storage to achieve instant decoding, solving the problem of information reading efficiency in DNA data storage, and thus improving the possibility of large-scale application of DNA data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] FIG1 is a schematic diagram of a capture chip according to an embodiment of the present invention.

[0028] FIG2 is a schematic diagram of a DNA sequence for storing data according to one embodiment of the present invention.

[0029] FIG3 is a schematic diagram of a flow chart of encoding video data according to an embodiment of the present invention.

[0030] FIG4 is a schematic diagram of a process for sequencing, reading and decoding video data stored in a DNA sequence according to an embodiment of the present invention.

[0031] FIG5 is a schematic diagram of data sequencing and decoding in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0033] The present invention provides a DNA sequence library for continuous data storage. The DNA sequence includes a capture location region and a data region. The capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data.

[0034] The DNA sequence also includes primer regions at both ends.

[0035] The present invention also provides a method for continuous data storage, comprising:

[0036] 1) Generate a DNA sequence library for continuous data, the DNA sequence includes a capture location region and a data region, the capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data;

[0037] 2) Synthesize DNA sequence library.

[0038] The present invention also provides a method for reading DNA data storage, comprising:

[0039] 1) using a capture location region to capture the DNA sequence of the DNA sequence library, where the capture location region corresponds to the spatial information of the continuous data;

[0040] 2) Sequencing the data region of the DNA sequence, where the data region corresponds to continuous data content, and restoring the continuous data content as sequencing proceeds.

[0041] In this context, continuous data refers to a sequential, continuous set of data, which can be either temporal or spatial. For example, continuous data can include images, videos, or the location of cells in tissues. Specifically, if an image is divided into different parts, the different parts of the image are spatially continuous data; a video is essentially composed of different frames, and the multiple frames of a video are temporally continuous data; and cells in a specific spatial location within an organism collaborate with their microenvironment to exert their unique biological functions, so the location of cells in a tissue over a period of time is continuous data.

[0042] In the present invention, DNA sequence capture can be achieved through a sequencing chip. The specific sequencing chip is shown in Figure 1. The original sequencing chip is based on the BGI spatiotemporal sequencing chip. Each microarray is connected to a specific DNA sequence capture probe through a solid phase carrier and matched with a coordinate positioning sequence (Figure 1).

[0043] Capturing DNA sequences can also be achieved using magnetic beads with specific labels.

[0044] During the DNA sequence design and synthesis process, each sequence ( FIG. 2 ) needs to carry a portion of a specific sequence that matches the coordinate positioning sequence and can be captured by the capture probe.

[0045] Taking video storage as an example, matrix encoding is combined with coordinate positioning during encoding (Figure 3). Each information matrix encodes one frame of the picture, forming a base matrix. Multiple base matrices are connected to form a DNA sequence library. When reading data, as shown in Figure 4, the synthesized DNA library is first amplified and poured onto the chip surface and fully reacted. At this time, the specific capture probe will capture the DNA sequence with the corresponding positioning sequence. After that, library construction and sequencing are carried out. Each round of sequencing obtains the base information of a matrix. Each base matrix can obtain partial data (such as the Nth frame image) by decoding. This is repeated until all information is fully restored after sequencing is completed.

[0046] If single-molecule sequencing is used, the base data read out per unit time can be used as data blocks for semi-instant decoding.

[0047] Example:

[0048] This embodiment takes the BGI logo image as an example to illustrate the database construction method, data storage method, and data reading method disclosed in the present invention.

[0049] 1. Encoding: During the buffered read encoding process, we segment the bit sequence of the file to be stored (the BGI logo image) and do not directly encode it. Instead, we multiply the dimension of each bit sequence to form a two-dimensional bit matrix. Adjacent matrices are superimposed, ultimately generating continuous data with a depth equal to the total number of segmented sequences. This encoding then forms a DNA sequence library. The length of the continuous data multiplied by the width is the total number of DNA molecules, and the depth is the length of the DNA molecules. Any missing length is padded with random sequences and marked with the file size, making it possible to distinguish which bases are redundant during decoding. The necessary parameters corresponding to the original file type are also returned, facilitating the determination of the recovered file type corresponding to the read sequence during the immediate read phase.

[0050] 2. Synthesis and Storage: Synthesize the DNA sequence encoded in step 1 using methods including but not limited to column synthesis, electrochemical synthesis, chip-based inkjet printing synthesis, sorting-based chip synthesis, photodeprotection synthesis, etc. The synthesis length is consistent with the sequence length, and the synthetic product is stored as an in vitro oligonucleotide / gene fragment or in vivo. The storage environment includes but is not limited to sealed capsules, inorganic packaging, living cells, etc.

[0051] 3. Sequencing: The stored DNA sequence is amplified (optional) and then a library is constructed using a selected sequencing technology, including but not limited to high-throughput sequencing and single-molecule sequencing. During the sequencing process, the sequencing results are output based on the base matrix produced in each sequencing round or per unit time.

[0052] 4. Decoding: Sequence decoding is performed in rounds or per unit time according to the output sequencing results, gradually completing the decoding of the source file, as shown in Figure 5.

[0053] The present invention also provides a DNA data storage reading system, the system comprising

[0054] a capture unit configured to capture a DNA sequence storing continuous data;

[0055] a sequencing unit configured to sequence the captured DNA sequence;

[0056] A decoding unit configured to decode DNA sequence sequencing data;

[0057] The storage unit is configured to store sequencing data and decoding data.

[0058] The DNA sequence in the reading system includes a capture positioning area and a data area. The capture positioning area corresponds to the spatial information of continuous data, and the data area corresponds to the continuous data content.

[0059] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A DNA sequence library for continuous data storage, characterized in that: The DNA sequence includes a capture location region and a data region, wherein the capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data.

2. The DNA sequence library according to claim 1, characterized in that The two ends of the DNA sequence also include primer regions.

3. A method for continuous data storage, characterized in that: The method comprises: 1) generating a DNA sequence library for continuous data, wherein the DNA sequence includes a capture location region and a data region, wherein the capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data; 2) Synthesizing the DNA sequence library.

4. A method for reading DNA data storage, characterized in that: The method comprises: 1) capturing a DNA sequence of a DNA sequence library using a capture location region, wherein the capture location region corresponds to spatial information of continuous data; 2) Sequencing a data region of the DNA sequence, the data region corresponding to continuous data content, and restoring the continuous data content as the sequencing proceeds.

5. The method according to claim 4, characterized in that The capture of DNA sequence is achieved through a sequencing chip.

6. The method according to claim 4, characterized in that The capture of DNA sequences is achieved by magnetic beads with specific labels.

7. The method according to claim 5, characterized in that Each microarray on the sequencing chip is connected to a specific DNA sequence capture probe via a solid phase carrier and matched with a coordinate positioning sequence.

8. The method according to claim 7, characterized in that The capture positioning region is a specific sequence that matches the coordinate positioning sequence and can be captured by the capture probe.

9. A DNA data storage reading system, characterized in that: The system comprises: a capture unit configured to capture a DNA sequence storing continuous data; A sequencing unit configured to sequence the captured DNA sequence; A decoding unit, configured to decode DNA sequence sequencing data; A storage unit is configured to store the sequencing data and the decoding data.

10. The system according to claim 9, characterized in that The DNA sequence includes a capture location region and a data region, wherein the capture location region corresponds to the spatial information of the continuous data, and the data region corresponds to the content of the continuous data.

Citation Information

Patent Citations

  • Construction method and application of single-stranded sequencing library

    CN109536579A

  • DNA information storage and reading method and system

    CN114743602A

  • Spatial barcoding

    CN115461469A

  • Method and System for Decoding Information Stored on a Polymer Sequence

    US20220364991A1

  • Method and apparatus for performing data storage by using DNA, and storage device

    WO2022266802A1