Nucleic acid storage medium for storing index through structure and preparation and reading method thereof
Through the nucleic acid composite structure and DNA-PAINT method, the problems of decreased DNA storage density and PCR dependence were solved, efficient and accurate data access and index information reading were achieved, and the performance of the DNA storage system was improved.
Patent Information
- Application Number
- CN202410307922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-19
AI Technical Summary
In existing DNA storage technology, the macroscopic physical scaffold leads to a decrease in storage density, and the reliance on PCR technology for reading index information leads to a high error rate, which limits the storage scale and data access efficiency of a single data pool.
A nucleic acid composite structure is adopted, including a backbone chain, a staple chain and an index chain. The nucleic acid composite structure is formed through self-assembly. The index information storage segment is exposed in the structure. Huffman coding is used to compress the index information, and optical reading is performed through the DNA-PAINT method, reducing dependence on PCR technology.
It improves the theoretical scale and storage density of DNA storage, reduces the error rate, and enables fast and accurate data access and index information reading.
Smart Images

Figure CN120673859A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of DNA storage technology, and more particularly to a nucleic acid storage medium and a preparation and application method thereof. Background Art
[0002] In recent years, the amount of information generated by humans has grown exponentially. However, the means by which we store information face fundamental material, energy, environmental, and spatial limitations. DNA storage, as a cutting-edge storage method, has shown significant potential as a data storage medium due to its extremely high density (storage of 1 exabyte of data per cubic millimeter), durability (retention time of up to several centuries), and efficient resource conservation. Figure 1 As shown in Figure 2, advances in DNA synthesis and sequencing technologies have enabled the development of DNA-based data storage systems with capacities up to 1GB. However, in addition to continuing to reduce the costs of DNA synthesis and sequencing, we must also address the challenges of effectively managing, accessing, and searching the data stored in DNA. To take the next step toward industrial applications of DNA data storage, it is essential to develop a random-access data management system specifically suited for DNA data storage.
[0003] Currently, DNA storage data management systems face some challenges in data access. Unlike silicon-based storage, DNA data storage units do not have spatial addressability because all synthesized DNA exists in a random and chaotic solution space. Since there are a large number of different and disordered DNA molecules densely adjacent and freely floating in the storage system, an addressing system that can function in complex and information-intensive molecular mixtures is required. In 2018, Organick et al. proposed embedding index sequences in DNA chains to randomly access any DNA sequence in the data pool by using PCR-specific amplification, such as Figure 2 As shown in (a) in the figure. Although this embedded scheme helps to speed up the search process, the qualified rate of primers decreases rapidly as the scale of data storage increases. This is because when multiple candidate files contain very similar information, multiple erroneous amplifications are prone to occur, which limits the storage scale of a single data pool to the TB level. In addition, this indexing method will fundamentally undermine the integrity of the original data because the process involves DNA enrichment technology (PCR). During the enrichment process, the DNA molecules corresponding to the original data are mixed with the reactants, resulting in irregular enrichment or deviation of different DNA molecules. After multiple accesses, that is, multiple PCR, PCR deviations will randomly accumulate, resulting in the absence of specific subsets in the sequencing results, and the error rate will inevitably increase, which can only be compensated by increasing the sequencing coverage.
[0004] On the other hand, in order to overcome the difficulties in PCR design of a single data pool, Newman et al. in 2019 used a physical scaffold to arrange the DNA sub-data pools, such as Figure 2 This approach can, to some extent, address the PCR-based addressing challenge, as the spatial physical isolation is equivalent to addressing data on a traditional tape drive. Since the DNA data pools are independent of each other, the same addressing scheme can be used. However, this approach sacrifices the density advantage of DNA storage, as the physical scaffold itself takes up a lot of space. In short, the PCR-based DNA data storage indexing system limits the theoretical storage scale of a single data pool.
[0005] In addition to storing data in DNA sequences, DNA nanotechnology can store digital information in structures by constructing complex one-dimensional, two-dimensional, and three-dimensional structures. These DNA nanostructures or systems may contain one or more structural chains that are designed to utilize Watson-Crick pairing of nucleotides, allowing DNA fragments to self-assemble into final predictable structures. Various design structures and self-assembly methods can be used, among which DNA origami is the most commonly used method to construct DNA-based data storage structures at the nanoscale. Importantly, all of these bottom-up methods can produce asymmetric patterns, which is a key criterion for data storage applications: data can be stored in the three-dimensional shape of these assemblies rather than directly encoded in the base sequence.
[0006] Compared to encoding data directly in nucleotide sequences, data storage based on DNA nanotechnology has a major drawback: low data density. By using various ensemble-level or single-molecule characterization methods, the size and morphology of the resulting DNA structures can be assessed, enabling access to the information stored in the shape and structure of these nanoscale assemblies. While this structural information-based storage approach is unlikely to replace the current DNA storage paradigm, it could serve as a complementary approach.
[0007] Therefore, our invention aims to solve two problems:
[0008] First, solve the problem of a sharp drop in the overall storage density of DNA data storage caused by macroscopic physical scaffolds; second, solve the problem of DNA file index information reading relying on PCR technology. Summary of the Invention
[0009] The purpose of the present invention is to provide a nucleic acid storage medium that stores index information through a structure and a preparation and reading method thereof.
[0010] In a first aspect of the present invention, a nucleic acid storage medium is provided, comprising:
[0011] (a) a nucleic acid composite structure configured to store predetermined data;
[0012] Wherein, the nucleic acid composite structure comprises:
[0013] (i) one or more backbone chains;
[0014] (ii) multiple staple chains;
[0015] (iii) one or more index chains, wherein the index chains include (M1) staple segments and (M2) index information storage segments;
[0016] wherein the backbone chain, the staple chain, and the staple segments of the index chain self-assemble to form the nucleic acid composite structure;
[0017] Furthermore, the index information storage segment of the index chain is exposed to the nucleic acid composite structure;
[0018] The index information of the data is stored in the index chain.
[0019] In another preferred embodiment, the index information storage segments of the index chain are spatially resolvable.
[0020] In another preferred embodiment, the index chain is a single chain structure.
[0021] In another preferred example, the index information storage segment of the index chain is a single chain.
[0022] In another preferred embodiment, the nucleic acid composite structure is a sheet-like structure, and the index information storage segment of the index chain is exposed on one side surface of the sheet-like structure.
[0023] In another preferred embodiment, the nucleic acid storage medium includes: (b) a nucleic acid composite structure and a storage solution.
[0024] In a preferred embodiment, the invention comprises:
[0025] The content information of the data is stored in the skeleton chain, or the staple chain, or a combination thereof;
[0026] The nucleic acid composite structure includes nucleic acid nanostructure, DNA composite structure, DNA nanostructure and DNA origami structure.
[0027] In another preferred embodiment, the content information of the data is stored in the skeleton chain.
[0028] In another preferred embodiment, the address information of the data is stored in the skeleton chain.
[0029] In another preferred embodiment, the nucleic acid composite structure is a nucleic acid nanostructure.
[0030] In another preferred embodiment, the nucleic acid complex structure includes a DNA complex structure.
[0031] In another preferred embodiment, the nucleic acid composite structure includes a DNA nanostructure.
[0032] In another preferred embodiment, the nucleic acid composite structure is a DNA origami structure.
[0033] In a preferred embodiment, the invention comprises:
[0034] (a) an index base sequence, the index base sequence comprising:
[0035] (i) an index sequence of digital bits, wherein each digital bit index sequence identifies a digit of an index digital sequence, wherein the index digital sequence is formed by encoding index information of the data;
[0036] (ii) an index sequence of a direction position, wherein each index sequence of a direction position identifies a site on the nucleic acid composite structure and is used to confirm the direction of the index digital sequence during reading.
[0037] The index base sequence identifying each digital bit or direction bit forms an index information storage segment of one or more index chains.
[0038] In another preferred embodiment, the index sequence identifying the digital bits forms an index information storage segment of an index chain.
[0039] In another preferred embodiment, the index sequence identifying the direction bit forms index information storage segments of multiple index chains.
[0040] In another preferred embodiment, the index digital sequence includes an array of multi-base integers, which is used to identify the file to which the nucleic acid complex structure belongs.
[0041] In another preferred embodiment, the nucleic acid storage medium comprises:
[0042] (b) a content base sequence, the content base sequence comprising:
[0043] (i) a content sequence of a sub-file, wherein the content sequence of the sub-file is formed by encoding the content information of the sub-file;
[0044] (ii) an address sequence of a sub-file, wherein the address sequence of the sub-file is used to mark the sequence number of the sub-file in the file.
[0045] In a preferred example, a key-value architecture is used to associate the index information of the data with the content information of the data.
[0046] In a preferred embodiment, the invention comprises:
[0047] The index digital sequence comprises a quinary matrix array of three rows and four columns;
[0048] The spatial distance between the index chains identifying the digital bits ranges from less than 30 nm;
[0049] The index chains are evenly arranged on the surface of the nucleic acid composite structure in the order of the matrix array;
[0050] The spatial position of the index strand that identifies the orientation on the nucleic acid complex structure includes:
[0051] (i) the index chain identifying the first direction bit is located between the index chains identifying the four digits in the upper left corner of the matrix array;
[0052] (ii) the index chain identifying the second direction bit is located between the index chains identifying the four digits in the lower left corner of the matrix array;
[0053] (iii) The index strand identifying the third orientation is located at the upper right corner of the nucleic acid composite structure.
[0054] In another preferred embodiment, the spatial distance between the index chains identifying the digital bits ranges from 20 nm.
[0055] In another preferred embodiment, the index chains are evenly arranged on one side surface of the sheet structure in the order of the matrix array.
[0056] In a preferred embodiment, the invention comprises:
[0057] The index information storage segment of the index chain that identifies the digital bit has a variety of protrusion shapes when combined with the same index probe.
[0058] When the numerical values of the digits are the same, the index sequences of the digits are the same, and the shapes of the protrusions on the index chains identifying the digits are the same.
[0059] In a second aspect of the present invention, a method for preparing the nucleic acid storage medium is provided, comprising:
[0060] (a) encoding index information of predetermined stored data into an index digital sequence;
[0061] (b) encoding the content information of the predetermined stored data into a content base sequence, designing a nucleic acid composite structure based on the content base sequence and the index digital sequence, and generating a backbone chain, a staple chain, and an index chain respectively;
[0062] (c) annealing and assembling the backbone strand, the staple strand, and the index strand to form the nucleic acid composite structure.
[0063] In another preferred embodiment, Huffman coding is used to compress the index information of the data to form a digital index sequence and determine the index base sequence.
[0064] In another preferred embodiment, the payload and the error correction code are encoded and converted into a sequence consisting of four ATCG nucleotide bases to form a content sequence of the sub-file.
[0065] In another preferred embodiment, the content sequence of the sub-file is appended with the address sequence to generate the backbone chain sequence of the nucleic acid composite structure 1 .
[0066] In another preferred embodiment, an index chain site is designed according to the index information of the data and a corresponding nucleic acid complex structure is synthesized.
[0067] In another preferred embodiment, the staple chain and index chain are designed based on the desired structural shape and the skeleton chain.
[0068] In another preferred embodiment, the synthesized nucleic acid complex structure is stored in a nucleic acid storage database for archiving.
[0069] In a third aspect of the present invention, a method for reading the nucleic acid storage medium is provided, comprising:
[0070] (a) Generate an index probe based on the index information of the data to be searched, and insert the index probe into the base sequence on the index strand as described above for complementary pairing;
[0071] (b) obtaining metadata of a plurality of nucleic acid complex structures based on optical interpretation of the pairing reaction of the index strands;
[0072] (c) selecting a nucleic acid complex structure whose metadata meets the search requirements, decoding and sequencing the nucleic acid complex structure, and obtaining data that meets the search requirements.
[0073] In a preferred embodiment, the invention comprises:
[0074] The index probe is of a single type, and the time it stays when binding and dissociating with index chains with different identifier values is different;
[0075] The DNA-PAINT method is used to collect spatial and temporal information of the index data domain on a single nucleic acid complex structure.
[0076] In another preferred embodiment, an index probe is generated according to index information of the file to be searched.
[0077] In another preferred embodiment, the index information is used as input to obtain a corresponding index chain site design scheme and generate an index probe of the corresponding sequence.
[0078] In another preferred embodiment, the index probe carries a fluorescent group, the index chain carries a docking chain, and the fluorescent group is complementary to the base sequence of the docking chain.
[0079] In another preferred embodiment, the fluorescent group and the docking chain repeatedly bind and dissociate to generate multiple fluorescent positioning points.
[0080] In another preferred embodiment, a sample is physically extracted from a nucleic acid storage database containing stored data, wherein the sample contains a large number of nucleic acid complex structures.
[0081] In another preferred embodiment, all nucleic acid complex structures are deposited on a glass slide, and a fluorescent group-labeled index probe is added to the solution for DNA-PAINT imaging to collect the index chain sites.
[0082] In another preferred embodiment, the pairing reaction imaging process of the index chain is recorded at multiple fluorescence positioning points, image post-processing is performed, metadata is obtained, and the sample is read multiple times for error correction.
[0083] In another preferred embodiment, among the templates of index digital sequences of multiple files, the template closest to the metadata is selected as the index digital sequence of the nucleic acid complex structure.
[0084] In another preferred embodiment, image post-processing is performed using an image averaging method to enhance the signal-to-noise ratio.
[0085] In another preferred embodiment, metadata is obtained by reading the bit depth through dynamic characteristics.
[0086] In another preferred embodiment, an error correction algorithm based on physical copy redundancy is established to ensure error-free data recovery.
[0087] In another preferred embodiment, among the templates of index digital sequences of multiple files, a clustering method based on similarity calculation is adopted to select the template closest to the metadata as the index digital sequence of the nucleic acid complex structure.
[0088] In another preferred embodiment, the DNA-PAINT method is used to collect spatial information of the index data domain on a single nucleic acid complex structure.
[0089] In another preferred embodiment, the DNA-PAINT method is used to collect the time information of the index data field on a single nucleic acid complex structure.
[0090] In another preferred embodiment, DNA-PAINT images of similar structures are aligned based on templates to preliminarily obtain sites encoding 0 and non-0.
[0091] In another preferred embodiment, the residence time of the index chain site is analyzed to obtain metadata to restore the index information of the data.
[0092] In another preferred embodiment, a plurality of nucleic acid complex structures whose index digital sequences meet the search requirements are separated from the nucleic acid storage database.
[0093] In another preferred embodiment, multiple nucleic acid complex structures whose stored index information meets the search requirements are separated from the nucleic acid storage database by visual labeling and magnetic bead separation methods.
[0094] In another preferred embodiment, a plurality of nucleic acid composite structures are sequenced, content sequences of a plurality of sub-files are synthesized according to the address sequences of the sub-files, and files meeting the search requirements are obtained by decoding.
[0095] In another preferred embodiment, the index probes are of multiple types, and different index probes bind and dissociate with index chains with different marker values to produce different imaging effects.
[0096] In a fourth aspect of the present invention, the present application further discloses a nucleic acid storage database, comprising:
[0097] The nucleic acid storage database comprises N nucleic acid storage media as described above, wherein N is a positive integer ≥ 2;
[0098] There is no macroscopic physical isolation between the N nucleic acid storage media.
[0099] In another preferred example, the nucleic acid storage database stores data of M files, where M is a positive integer ≥1.
[0100] In another preferred embodiment, the content information of the file is composed of content information of one or more sub-files.
[0101] In another preferred embodiment, the data includes the file and the sub-file.
[0102] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 It is a schematic diagram of the DNA information storage process in one embodiment of the prior art;
[0104] Figure 2 (a) and (b) are schematic diagrams of random access strategies for DNA data storage based on PCR addressing and spatial physical isolation, respectively, in the prior art;
[0105] Figure 3is a schematic diagram of a nucleic acid complex structure according to one embodiment of the present application;
[0106] Figure 4 is a schematic diagram of a nucleic acid storage database according to one embodiment of the present application;
[0107] Figure 5 is a schematic diagram of encoding data information into a base sequence according to one embodiment of the present application;
[0108] Figure 6 (a) and (b) are schematic diagrams of the design of digital bits and directional bits on a nucleic acid composite structure according to one embodiment of the present application, and a DNA-PAINT image of a single structural digital bit, respectively;
[0109] Figure 7 (a) and (b) are schematic diagrams of the combination of an index chain and an index probe, and the design of protrusions of different sizes on the index chain according to an embodiment of the present application;
[0110] Figure 8 (a) and (b) are schematic diagrams of the preparation process and reading process of the nucleic acid storage medium according to one embodiment of the present application;
[0111] Figure 9 is a schematic diagram of the DNA-PAINT imaging principle according to one embodiment of the present application;
[0112] Figure 10 (a) and (b) are schematic diagrams of determining the order of digital bits according to the direction bit and selecting the site according to the center circle of the digital bits of the template according to one embodiment of the present application;
[0113] Figure 11 is a schematic diagram of a pattern alignment process according to one embodiment of the present application;
[0114] Figure 12 is a schematic diagram of the difference in signal-to-noise ratio between a single structure and an average structure according to an embodiment of the present application;
[0115] Figure 13 This is a schematic diagram of redundancy error correction for physical copies according to one embodiment of the present application;
[0116] Figure 14 is a schematic diagram of clustering nucleic acid complex structures according to index readout results and reaching sequencing consensus within the cluster according to one embodiment of the present application;
[0117] Figure 15 is a schematic diagram of Huffman coding according to one embodiment of the present application;
[0118] Figure 161 is a schematic diagram of the bit length for retrieving index information at a 20 nm interval according to one embodiment of the present application;
[0119] Figure 17 1 is a schematic diagram of the distribution of the number of digital bits in the detection structure and the detection efficiency of each digital bit according to one embodiment of the present application;
[0120] Figure 18 is a schematic diagram of a single-molecule determination method for studying the robustness of bit depth to index information according to one embodiment of the present application;
[0121] Figure 19 This is a schematic diagram of error correction based on replica redundancy using value replacement according to an embodiment of the present application;
[0122] Figure 20 1 is a schematic diagram of the residence time distribution of each digit according to an embodiment of the present application;
[0123] Figure 21 is a schematic diagram of a statistical significance test using a two-tailed Student's t test according to one embodiment of the present application;
[0124] Figure 22 is a schematic diagram of the photobleaching results according to one embodiment of the present application;
[0125] Figure 23 Schematic diagram of index information for recovering a file from a single read according to one embodiment of the present application;
[0126] Figure 24 This is a schematic diagram of the residence time of each point in a mixed data set used for testing according to one embodiment of the present application;
[0127] Figure 25 This is a schematic diagram of calculating the similarity between each unknown structure and a known template based on a distance function in one embodiment of the present application.
[0128] In the accompanying drawings, the following are marked:
[0129] 1-nucleic acid complex structure;
[0130] 101-skeleton chain;
[0131] 102-staple chain;
[0132] 103-index chain;
[0133] 103.1-Stapling Section;
[0134] 103.2- Index information storage segment;
[0135] 103.3-protrusion;
[0136] 2- Nucleic acid storage database;
[0137] 3-Stored data information;
[0138] 301-Data;
[0139] 301.1-Information on the content of data;
[0140] 301.2-Data index information;
[0141] 302-File;
[0142] 302.1-Content information of the file;
[0143] 302.2-File index information;
[0144] 303-subfile;
[0145] 303.1-Content information of sub-files;
[0146] 303.2-Address information of sub-files;
[0147] 303.3-Index information of sub-files;
[0148] 4- index number sequence;
[0149] 5-base sequence;
[0150] 501-content base sequence;
[0151] 501.1-Content sequence of sub-files;
[0152] 501.2-Sequence of addresses of subfiles;
[0153] 502-index base sequence;
[0154] 502.1 - Index sequence of digital bits;
[0155] 502.2 - Index sequence of direction bits;
[0156] 6-index chain site;
[0157] 601-digital bit position;
[0158] 602-direction position;
[0159] 7-index probe;
[0160] 8-Metadata;
[0161] 801-stay time;
[0162] 802 - index number read;
[0163] 9-Template;
[0164] 901-Customized template for nucleic acid complex structure;
[0165] 902-Template with known index information. DETAILED DESCRIPTION
[0166] After extensive and in-depth research, the inventors have provided for the first time a nucleic acid storage medium and a method for preparing and applying the same. Specifically, the nucleic acid storage medium includes a nucleic acid composite structure, which is configured to store predetermined data; the nucleic acid composite structure includes one or more backbone chains, multiple staple chains, and one or more index chains; the backbone chain, the staple chain, and the staple segments of the index chain self-assemble to form a nucleic acid composite structure; and the index information storage segment of the index chain is spatially distinguishable and exposed to the nucleic acid composite structure; the index information of the data is stored in the index chain. The nucleic acid storage database contains N nucleic acid storage media, and there is no macroscopic physical isolation between the N nucleic acid storage media. The present application stores the index information of the data in the DNA structure information, which reduces the dependence on PCR technology, reduces the decrease in the overall storage density of DNA data storage due to the macroscopic physical scaffold, and greatly improves the theoretical scale of DNA storage. The present invention was completed on this basis.
[0167] Nucleic acid storage medium of the present invention
[0168] This application provides a nucleic acid storage medium, such as Figure 3 As shown, the nucleic acid storage medium includes:
[0169] A nucleic acid composite structure 1 is configured to store predetermined data 301; the nucleic acid composite structure 1 includes one or more backbone chains 101, multiple staple chains 102, and one or more index chains 103. The index chain 103 includes a staple segment 103.1 and an index information storage segment 103.2. The index chain 103 is a special staple chain 102 that serves both the function of a staple chain 102 and the function of recording index information 301.2 of the data. Generally, the nucleic acid composite structure 1 includes a backbone chain 101, multiple staple chains 102, and multiple index chains 103.
[0170] The backbone chain 101, the staple chain 102, and the staple segments 103.1 of the index chain 103 self-assemble to form the nucleic acid composite structure 1. That is, the nucleic acid composite structure 1 is assembled using an origami method, wherein the relatively long backbone chain 101 is folded into a pre-designed shape by interacting with the relatively short staple chain 102.
[0171] Furthermore, the index information storage segment 103 . 2 of the index chain 103 is exposed to the nucleic acid composite structure 1 ; the index information 301 . 2 of the data is stored in the index chain 103 .
[0172] In one embodiment, the index information storage segment 103.2 is spatially distinguishable to facilitate subsequent optical imaging to collect index chain site information; the index chain 103 is a single-stranded structure to facilitate subsequent binding and dissociation with the index probe 7; the nucleic acid composite structure 1 is a sheet-like structure, and the index information storage segment 103.2 of the index chain 103 is exposed on one side surface of the sheet-like structure.
[0173] In one embodiment, the data content information 301 . 1 is stored in the skeleton chain 101 , or the staple chain 102 , or a combination thereof; generally, the data content information 301 . 1 is stored in the skeleton chain 101 .
[0174] The nucleic acid composite structure 1 includes a nucleic acid nanostructure, a DNA composite structure, a DNA nanostructure, and a DNA origami structure; the nucleic acid composite structure 1 mentioned below mostly takes a DNA nanostructure or an origami structure as an example.
[0175] 1. Base sequence of nucleic acid storage medium
[0176] In one embodiment, it is characterized in that Figure 5 As shown, the base sequence 5 in the nucleic acid storage medium includes a content base sequence 501 and an index base sequence 502 .
[0177] Index base sequence 502 forms index information storage segment 103.2 of index chain 103. Index base sequence 502 includes a digital index sequence 502.1 and a direction index sequence 502.2. Each digital index sequence 502.1 identifies a digit of index digital sequence 4, which is encoded from index information 301.2 of the data. Each direction index sequence 502.2 identifies a site on nucleic acid composite structure 1 and is used to confirm the direction of index digital sequence 4 during reading.
[0178] The index base sequence 502 identifying each digit or direction bit forms an index information storage segment 103.2 for one or more index chains 103. In one embodiment, the digit index sequence 502.1 forms an index information storage segment 103.2 for one index chain 103, and the direction index sequence 502.2 forms an index information storage segment 103.2 for multiple index chains 103.
[0179] In one embodiment, the content base sequence 501 includes a sub-file content sequence 501.1 and a sub-file address sequence 501.2, wherein the sub-file content sequence 501.1 is formed by encoding the sub-file content information 303.1, and the sub-file address sequence 501.2 is used to mark the sequence number of the sub-file 303 in the file 302.
[0180] Because the size of the nucleic acid composite structure 1 is limited, a file 302 is divided into multiple sub-files 303 and encoded into multiple nucleic acid composite structures 1. Here, the data 301 contained in each nucleic acid composite structure 1 is an independent sub-file 303. To improve encoding efficiency, the content information 303.1 of the sub-file is generally encoded and written into the backbone chain 101 of a nucleic acid composite structure 1.
[0181] That is, the skeleton chain 101 has two structural fields. The first is a sequence recording the content information 303.1 of the sub-files, and the second is a sequence recording the address information 303.2 of the sub-files. A corresponding number of address information sequences 501.2 need to be generated according to the number of times the original file 302 is split, and this is used to determine the order between the file subsets.
[0182] Correspondingly, the area where the index information storage segment 103.2 of the index chain 103 is located is the index data domain, and the amount of information data that the stored index information 301.2 can contain is the index space. The index information 301.2 is stored in the form of digital information, preferably in the form of an array of digital information, represented as an index digital sequence 4 on the nucleic acid composite structure 1. Since the index information 301.2 can be stored in the form of digital bits on the nucleic acid composite structure 1, any type of message or data can be saved, including but not limited to binary, octal, decimal, hexadecimal, text or graphics. The message or data can also be encrypted and / or compressed before being stored in the nucleic acid composite structure 1. The message or data can be converted or encoded, for example, converting a text message or image data into binary, or encoding it using a code. For the evaluation of the index space, the number of digital bit sites 601 is used as the bit length (L) of the index information 301.2, and the number of selectable sequences at the site 601 is used as the bit depth (D), then the theoretical index space is D L , that is, you can manage / index D L 303 sub-files.
[0183] 2. Site arrangement of nucleic acid storage media
[0184] In one embodiment, the index chain site 6 of the nucleic acid composite structure 1 is used to encode the index information 301.2 of the data. In the design scheme of the index chain site 6, the index chain site 6 is subdivided into a digital site 601 and a direction site 602 according to different functions, such as Figure 6As shown in (a) in FIG. 1 , a nucleic acid complex structure 1 encoded in each different file 301 is assigned a number position 601 and a direction position 602 .
[0185] The index sequence 4 is unique for each file 302 and comprises an array of multi-base integers, used to identify the file 302 to which the nucleic acid complex structure 1 written into the sub-file belongs. All nucleic acid complex structures 1 within a file 302 use the same index chain site 6 design. Digital bit sites 601 are added to the nucleic acid complex structures 1 within its multiple sub-files 303. This allows identification of differently encoded files 302 without distinguishing between nucleic acid complex structures 1 within the same file 302. Specifically, within the same file 302, the sub-file index information 303.3 and the file index information 302.2 are identical, while the sub-file content information 303.1 and the sub-file address information 303.2 differ. The state of the digital bit is determined by the index sequence 4, whose value is based on the unique scintillation dynamics during optical reading and can be 0, 1, 2, or more.
[0186] A direction bit 602 is added to all nucleic acid composite structures 1 to identify the direction of the index digital sequence 4 during the decoding process. While any direction bit system can be used, such as pairing certain direction bits with digital bits, as a general approach, the direction bits are the same on all nucleic acid composite structures 1.
[0187] In one embodiment, the index number sequence 4 includes a quinary matrix array with three rows and four columns; the index chains 103 are evenly arranged on the surface of the nucleic acid complex structure 1 according to the order of the matrix array.
[0188] In one embodiment, the spatial distance between the index chains 103 identifying the digital digits is less than 30 nm. Theoretically, the smaller the distance, the larger the theoretical storage capacity that can be achieved. In one embodiment, the spatial positions of the index chains 103 identifying the directional digits on the nucleic acid composite structure 1 include: the index chain 103 identifying the first directional digit is located between the index chains 103 identifying the four digits in the upper left corner of the matrix array; the index chain 103 identifying the second directional digit is located between the index chains 103 identifying the four digits in the lower left corner of the matrix array; and the index chain 103 identifying the third directional digit is located at the corner position of the upper right corner of the nucleic acid composite structure 1.
[0189] Nucleic acid storage database of the present invention
[0190] This application provides a nucleic acid storage database 2 composed of nucleic acid storage media, such as Figure 4 As shown, the nucleic acid storage database 2 includes N nucleic acid storage media, where N is a positive integer ≥ 2; there is no macroscopic physical isolation between the N nucleic acid storage media;
[0191] In one embodiment, the nucleic acid storage database 2 stores data of M files 302 , where M is a positive integer ≥ 1; the content information 302 . 1 of a file is composed of the content information 303 . 1 of one or more sub-files, and the data 301 includes the file 302 and the sub-file 303 .
[0192] Compared with the DNA data storage method based on spatial physical isolation in the prior art, the storage method in this application allows the nucleic acid storage media of different files 302 to be stored in the same DNA data pool, effectively improving the storage density.
[0193] Preparation method of the present invention
[0194] (a) Encoding the index information 301.2 of the data to be stored as an index number sequence 4;
[0195] (b) encoding the content information of the predetermined stored data into a content base sequence 501, designing the nucleic acid composite structure 1 based on the content base sequence 501 and the index digital sequence 4, and generating a backbone chain 101, a staple chain 102, and an index chain 103;
[0196] (c) Annealing and assembling the backbone chain 101 , the staple chain 102 , and the index chain 103 to form the nucleic acid composite structure 1 .
[0197] 1. Index information Huffman coding method
[0198] In the era of big data, efficient data storage and fast retrieval are crucial. Index information 301.2, a key component for accelerating data retrieval, has a significant impact on system performance due to its size. We propose a technical solution for losslessly compressing index information 301.2 using Huffman coding, converting the information into a digital array to obtain the index digital sequence 4. The specific steps are as follows:
[0199] First, a Huffman tree is constructed. The index information 301.2 is treated as a character set, and a Huffman tree is constructed based on the frequency of each character. Through the Huffman tree, characters with high frequency are represented using short codes, thereby achieving efficient information representation.
[0200] Then, a Huffman coding table is generated. The constructed Huffman tree is traversed to generate a corresponding Huffman code for each character. This coding table is used for the subsequent lossless compression of the index information 301.2.
[0201] Next, the index information 301.2 is encoded. The original index information 301.2 is encoded according to the generated Huffman coding table. By replacing each character with its corresponding Huffman code, the index information 301.2 is compressed.
[0202] Finally, a digital array is generated. The resulting binary code is converted into a digital array. Each Huffman code is treated as a binary number, and then the binary code is converted into a specific digital array to adapt to the most appropriate base to ensure optimal data representation.
[0203] Compared with the traditional index information 301.2 storage method, this technical solution has the following advantages:
[0204] (i) Efficient compression: Huffman coding is used to achieve efficient compression of index information and reduce storage requirements.
[0205] (ii) Fast retrieval: The losslessly compressed index information 301.2 still maintains the accuracy of the original information, ensuring that data can be quickly and accurately located during retrieval.
[0206] (iii) Versatility: The index information 301.2 is applicable to various data types and has strong versatility and applicability.
[0207] 2. Design of Nucleic Acid Composite Structure Index Strand
[0208] The payload and error correction code (ie, the data content information 301.1) are encoded and converted into a sequence consisting of four ATCG nucleotide bases, forming the sub-file content sequence 501.1.
[0209] The content sequence 501 . 1 of the sub-file is appended with the address sequence 501 . 2 to generate the backbone chain sequence 101 of the nucleic acid composite structure 1 .
[0210] Since the staple chain sequence 102 and index chain sequence 103 of the folded backbone chain sequence 101 are determined by the specific backbone chain 101 and structure design, the user can input the desired structural shape (e.g., as designed as the nucleic acid complex structure 1 digital position 601) and the backbone chain 101 sequence into the software. Once determined, the software will provide the staple chain sequence 102 and index chain sequence 103 for creating the desired structure. Such complex structures can usually be designed using software such as caDNAno, minimizing errors and design time.
[0211] In one embodiment, Figure 7 As shown, the index information storage segment 103.2 of the index chain 103 for identifying a digital digit has various shapes of protrusions 103.3 when combined with the index probe 7.
[0212] When the digit values are the same, the digit index sequences 502.1 are the same, and the protrusions 103.3 on the index chains 103 identifying the digits are the same in shape. Conversely, when the digit values are different, the digit index sequences 502.1 are different, and the protrusions 103.3 on the index chains 103 identifying the digits are different in shape. Protrusions of varying sizes result in varying degrees of duplex affinity, facilitating subsequent bit depth reading.
[0213] 3. Key-value architecture design
[0214] In one embodiment, a key-value architecture is used to associate the index information 301.2 of the data with the content information 301.1 of the data.
[0215] All storage systems need a way to assign identifying tags to data objects so that they can be retrieved later. We chose a simple key-value architecture, where the put(key, value) operation associates a key with a value, and the get(key) operation retrieves the value assigned to a key.
[0216] In order to implement a key-value interface in a DNA storage system, we need: a function that maps a key to the DNA pool where the nucleic acid composite structure 1 containing the data is located; and a mechanism to selectively retrieve the required part of the pool (i.e., random access).
[0217] Reading method of the present invention
[0218] (a) generating an index probe based on the index information of the data to be searched, and inserting the index probe into the nucleic acid storage database according to claim 2, wherein the nucleotide sequence on the index probe is complementary to the base sequence on the index strand;
[0219] (b) obtaining metadata of a plurality of nucleic acid complex structures based on optical interpretation of the pairing reaction of the index strands;
[0220] (c) selecting a nucleic acid complex structure whose metadata meets the search requirements, decoding and sequencing the nucleic acid complex structure, and obtaining data that meets the search requirements.
[0221] In one embodiment, the index probe is of a single species, and the time for the index probe to bind and dissociate with index chains with different identifier values is different;
[0222] The DNA-PAINT method is used to collect spatial and temporal information of the index data domain on a single nucleic acid complex structure.
[0223] 1. DNA-PAINT Super-Resolution Optical Imaging Method
[0224] The key (i.e., index information 301.2) is used as input to obtain the corresponding index chain site 6 design scheme, and generate the corresponding sequence index probe 7, the sequence of the probe is reverse complementary to the index sequence 502.1 of the digital bit on the nucleic acid complex structure 1.
[0225] The storage system then physically extracts a sample from the DNA pool containing the stored data, along with a large number of unrelated nucleic acid complex structures1.
[0226] All nucleic acid complex structures 1 are deposited on a glass slide, and a fluorescent group (such as Cy3B) is added to the solution to label the index probe 7 for DNA-PAINT imaging. Figure 9 As shown, the index chain site 6 information of the complementary sequence is collected.
[0227] DNA-PAINT is a super-resolution optical imaging method based on single-molecule localization, enabling optical imaging of origami structures within the diffraction limit of light. In the DNA-PAINT method, a short oligonucleotide chain (docking chain) is labeled on the target, and a complementary sequence (index probe) is labeled with a fluorophore as an affinity probe. When a pair of fluorophore-labeled index probes bind to the docking chain, their combined fluorescence, which generates an ON signal, can be detected as a single spot (typically within 1 × 1 μm² based on the point spread function) after evanescent field illumination using a total internal reflection microscope. When the fluorophore-labeled index probes dissociate from the docking chain, they diffuse rapidly in the solution, resulting in no detectable fluorescent spot, resulting in a shutdown of the unbinding signal. Through repeated binding and dissociation, a large number of localization points can be generated at the target, thus achieving super-resolution imaging.
[0228] Rather than using AFM or electron microscopy, widely used characterization techniques in DNA nanotechnology, to read information about nucleic acid complex structures, DNA-PAINT super-resolution imaging was employed because it relies on positioning data. On the one hand, the positioning points contain time-scale information, which can extend the bit depth, rather than simply binary digits 0 (absence) and 1 (presence). This work provides index information 301.2 bit-depth recording in the time dimension by programming the dwell time 801 at the site. Furthermore, positioning points are highly compatible with many clustering algorithms, such as supervised machine learning, Bayesian methods, and unsupervised machine learning for clustering, which greatly facilitates highly automated information reading.
[0229] 2. Image averaging to improve signal-to-noise ratio
[0230] The DNA-PAINT readout results of the nucleic acid complex structure 1 digit are prone to false negatives, that is, the detection efficiency cannot reach 100%, such as Figure 6As shown in (b), there are missing sites in the image. This is because the nucleic acid complex structure 1 undergoes a series of mechanical disturbances such as PEG purification, freeze-thaw cycles, pipetting, and sample preparation mixing, as well as local Joule heating caused by the high laser power density used in DNA-PAINT imaging, which reduces the incorporation efficiency of the staple chain 102. To further improve the detection efficiency of the index data domain, the incorporation efficiency of the staple chain 102 can be improved by extending the binding region between the staple chain 102 and the backbone chain 101 and increasing the melting temperature between the two in terms of sequence design details. Here, we specify a more general image processing method that improves the signal-to-noise ratio of the final readout information by aligning and averaging multiple images of the same structure.
[0231] We first used a personalized template matching strategy. Since the nucleic acid complex structure 1 has a certain degree of flexibility, it will produce a certain degree of deformation in the solution, resulting in slight differences between the structures. The symmetry of the rectangular grid is broken by marking the direction of each structure, such as Figure 10 As shown, a customized template 901 is generated for each nucleic acid complex structure 1, and index strand sites 6 are obtained based on the sites in template 901. Template 901 is generated using a semi-automated labeling method, requiring manual labeling of three digital sites 601 (asymmetric triangles), and an algorithm automatically generates all remaining digital sites 601. Of course, machine learning algorithms, including but not limited to supervised learning, unsupervised learning, or reinforcement learning algorithms, can also be used to achieve full automation.
[0232] Then, the template 901 of each structure is used for translation and rotation alignment, and the obtained translation vector and rotation matrix are applied to the corresponding positioning point of the structure, that is, the average alignment of multiple structures is achieved, such as Figure 11 We quantitatively compared the spatial resolution of a single structure and the averaged structure. After averaging, the image signal-to-noise ratio was significantly improved, the spatial resolution remained basically unchanged, and false negatives were eliminated, ensuring the integrity of the structure. Figure 12 shown.
[0233] 3. Digital multi-base encoding method for time dimension
[0234] For bit-depth encoding, we rely on precisely tunable kinetic features (the dwell time of index probe 7 at the digital position is 801) to provide a unique kinetic barcode for the multi-bit number. This partially complementary sequence design allows for temporal discrimination of different sequences, further increasing data density. This allows for the synergistic effect of sequence-specific spatial and temporal separation, allowing for long spatial readouts and deep temporal readouts, significantly increasing the data density of structural information.
[0235] Specifically, we adjust the flashing duration of a given site by changing the size of the protrusion 103.3 structure formed by the digital bit index sequence 502.1 and the index probe 7 (for example, a protrusion 103.3 sequence site with 2 Ts results in a longer binding event, while a protrusion 103.3 sequence site with 6 Ts results in a shorter binding event). Since the digital bit index sequence 502.1 is exactly the same except for the protrusion 103.3 part, only one index probe 7 needs to be used to simultaneously read multiple digital bits, such as Figure 7 As shown, we can identify the shape of the protrusion 103.3 by recording the dwell time 801, thereby obtaining the preliminary index number 802 of the digital position 601, and obtain the metadata 8 after interpreting the numbers of all positions.
[0236] 4. Error correction method based on physical copy redundancy
[0237] Due to the limitations of DNA molecular hybridization reactions, incorrect information reading is inevitable in our system. Image averaging can effectively overcome false negatives in index information reading. Apart from this, the main error is the numerical deviation of the digital bits (bit depth reading error), specifically the measurement deviation of the dwell time 801. Since the binding of the index probe 7 and the digital bits of the nucleic acid complex structure 1 is a thermodynamically driven interaction, these interactions are not completely all-or-nothing (two distinct states) like traditional electronic storage addresses. The binding and dissociation of the double strands is a Poisson process, and the measurement of the dwell time 801, that is, the time interval between consecutive events, satisfies the exponential distribution. This also causes a certain degree of fluctuation in the dwell time 801 of the index data domain, and there is a probability that it will be incorrectly decoded to a nearby number.
[0238] Here, an error correction algorithm based on physical copy redundancy is established to ensure error-free data recovery. Specifically, by increasing the sampling volume of the same nucleic acid complex structure 1, the mean of the residence time 801 of each digital bit is calculated. According to the Central Limit Theorem (CLT), as the sample size increases, regardless of the shape of the original population distribution, the distribution of the sample mean tends to be normal, as shown in the following example: Figure 10As shown. At the same time, according to the Law of Large Numbers (LLN), as the number of independent and identically distributed samples increases, the sample mean tends to the population mean in probability. In scenarios where accurate estimation of the population mean is required, allocating resources to obtain a larger sample size is a wise choice. However, a balance must be struck between the benefits of increased accuracy and the time and resource costs required for data collection. Currently, our error correction algorithm tends to copy redundancy, which has the advantage of lower index encoding difficulty, but requires physical copy redundancy. We believe that when the encoding space of the index information is large enough, logical redundancy error correction schemes (parity check, fountain code and other encoding schemes) can be considered, but smaller data storage spaces make it difficult to take advantage of logical redundancy.
[0239] 5. Clustering method based on similarity calculation
[0240] In the information reading stage, we are concerned about whether the nucleic acid complex structure 1 encoded by different index information 301.2 can be correctly identified at the single structure level. We adopt a clustering method based on similarity calculation, by calculating the similarity between the single structure index information 301.2 and the known index information (template 902) respectively, and taking the one with the highest similarity as the final predicted index information of the structure, such as Figure 14 As shown in Figure 9 , our index data is high-dimensional, L-dimensional (L represents bit length). It is independent and uncorrelated, and contains missing values (not a number, NaN). Therefore, the correlation coefficient is not suitable for calculating correlation. We use a distance function (Euclidean distance) to calculate the correlation between the test subset and template 902.
[0241] Definition of distance function: Assume that we have a template 902 set T (the real index number sequence of each file is known 4), which contains m N-dimensional arrays, that is, T = t1, t2, ..., t m For each test data x (index number sequence 4 of a nucleic acid complex structure 1), we use a distance function D to measure its similarity with all templates 902. This distance function can be the average distance square, defined as:
[0242]
[0243] Where Neffective is the effective dimension, which indicates the number of dimensions actually involved in the calculation when calculating the distance, because a single structure can easily produce false negative results due to missing digital bits (NaN in the array)
[0244] Calculate distance: For each test data x, calculate its distance with all templates 902
[0245] distances={D(t1,x),D(t2,x),...,D(t m ,x)}
[0246] Select the minimum distance: Find the minimum value in the distance array distance, that is, find the best matching template 902:
[0247] min_index = argmin(distances)
[0248] Classification: Assign the test data x to the template with the smallest distance 902
[0249] prediction(x)=min_index
[0250] The mathematical expression of the whole process is:
[0251] prediction(x)=argmin i (D(t i ,x))
[0252] This expression indicates that we select the index number sequence 4 of template 902 that minimizes the distance function as our prediction result. We also ensure that dimensions containing NaNs are skipped when calculating the distance, thus handling the missing values in the test set. The advantage of this method is that it can be applied to various distance functions and different datasets, making it flexible in many pattern recognition and classification problems.
[0253] The main advantages of the present invention include:
[0254] (1) The problem of DNA storage index information 301.2 reading relying on PCR technology is solved.
[0255] (2) It reduces the decline in the overall storage density of DNA data storage caused by macroscopic physical scaffolds and greatly improves the theoretical scale of storage.
[0256] (3) It is expandable. By designing the nucleic acid composite structure 1, the bit length and bit depth of the data index information 301.2 can be increased.
[0257] (4) When the structural information is sufficiently encoded, it is possible to realize a preview function of the data 301 , and preview a low-resolution version of the file 302 by optical methods without having to fully access or download the file 302 .
[0258] The present invention will be further described below with reference to specific examples. It should be understood that these examples are only intended to illustrate the present invention and are not intended to limit the scope of the present invention.
[0259] Example 1: Huffman Coding Example of Index Information
[0260] In one embodiment, the skeleton chain 101 is of type p7249 (M13mp18), and the number of staple chains 102 is 184.
[0261] We then demonstrated how to use Huffman coding to losslessly compress the index information 301.2 “dnaMetadata” and generate the index number sequence 4.
[0262] like Figure 15 As shown, the string is traversed, counting the number of occurrences of each character. A Huffman tree is constructed based on character frequency, with each character becoming a leaf node in the tree, and characters with higher frequencies are closer to the root of the tree. A Huffman encoding table is then generated. Starting from the root node, each character along the tree path is assigned a code. The rule is: left branches are 0, right branches are 1. The length of the code is related to the depth of the character in the tree.
[0263] Next, using the generated Huffman encoding table, each character in the original index information 301.2 "dnaMetadata" is replaced with the corresponding Huffman code. The resulting combination produces a 27-bit binary code number, but using ASCII code would require 88 bits of binary code.
[0264] Finally, the binary number is converted into a quinary number (because the first bit of the binary code is 0, the first bit of the quinary number is regarded as a leading 0), and the index number sequence 4 "042322130241" is obtained.
[0265] Example 2: Spatial Information Processing of DNA-PAINT Imaging Data
[0266] For details about the DNA-PAINT method, see above. For index base sequence 502, when index probe 7 is a single species, we performed NUPACK simulations to calculate the Gibbs free energy and dissociation constant for uniform imaging with different protrusion 103.3 sequences (see Table 1 for detailed sequence information). Aside from the portion that forms the protrusion, the multi-digit index sequence 502.1 is identical. Each digit of index sequence 502.1 corresponds to a unique dwell time 801 of index probe 7. By mapping the multi-base digits to the sequence (reflected in the imaging process as the dwell time 801 at that digit), deep multi-base digital encoding is achieved.
[0267] Table 1 Index sequence 502.1 candidate, index probe 7 sequence is TCTCCTTCCTCT
[0268]
[0269] (ΔG calculation refers to FluorogenicDNA-PAINT forfaster, low-backgroundsuper-resolutionimaging|Nature Methods; Kd calculation refers to Non-complementary strandcommutation as a fundamental alternative for informationprocessingbyDNA andgene regulation|Nature Chemistry)
[0270] We demonstrated accurate readout of the index information in the nucleic acid complex 1 at a spatial distance of 20 nm, relying on the "image averaging method to enhance signal-to-noise ratio." Due to the precise positioning of the nucleic acid complex 1, index strands 103 can be placed at any distance from each other. However, the distance between two index strand sites 6 may depend on the resolution capability of the microscope.
[0271] We designed four square structures carrying 12 digital bits with a spacing of 20 nm, only changing the index sequence 502.1 containing the digital bits (see Table 1 for detailed sequence information), and the sequence of the index data domain of each nucleic acid complex structure 1 was unified, and used index probe 7 to perform DNA-PAINT imaging on them, as shown in Figure 1. Figure 16 shown.
[0272] Figure 16 (a) shows a schematic diagram of nucleic acid composite structure 1, illustrating the arrangement of digital site 601 and directional site 602 on the structure's surface. The imaging tag for engineered protrusion 103.3 is located at digital site 601 of nucleic acid composite structure 1. DNA-PAINT super-resolution imaging is achieved through transient reversible hybridization of index probe 7. (b) shows a DNA-PAINT averaged image of four engineered protrusions 103.3 on nucleic acid composite structure 1 (left) and a cross-sectional histogram analysis of a selected region (right).
[0273] Super-resolution reconstructions of the four square structures revealed a well-resolved 20nm pattern, derived from images averaged from over 200 different structures, demonstrating the structural integrity of the DNA origami. Cross-sectional histogram analysis of the DNA-PAINT averaged images confirmed that the average distances along horizontal lines (i) and vertical lines (ii) were 20.6±0.6nm and 21.9±0.4nm, respectively, closely matching the distances designed by computer software. This relatively high spatial resolution and precise site arrangement are crucial for the readout of index information 301.2, ensuring accurate identification of each site, a prerequisite for accurate decoding of bit length and bit depth.
[0274] Based on the spatial information of the imaging, false negative errors (bit length readout errors) in the reading of index information 301.2 can be quantitatively analyzed. By setting a threshold for the number of locations at index chain position 6, and considering those below the threshold as undetectable, the detection efficiency of digital position 6 can be quantitatively evaluated.
[0275] Nucleic acid complex structures 1 targeting four different protrusion 103.3 sequences, Figure 17 The distribution of the number of index strands 103 detected is shown, as well as the detection efficiency for all index strand sites 6 in the right panel.
[0276] In the designed digital index sequence 502.1, the detection probability of T2, T4, and T6 was similar, averaging 85%, and the average number of sites detected ranged from 10.03 to 10.45. Using the ~7% offset established by Strauss et al. to convert detection efficiency into incorporation efficiency, an average incorporation efficiency of 92% was obtained.
[0277] However, the detection efficiency of A2 is significantly lower than that of T2, T4, and T6, with an average detection efficiency of only 64%. Since the A2 sequence is completely complementary to index probe 7, the dwell time 801 is significantly longer than that of the other protrusion 103.3 sequences. The long dwell time 801 will increase the probability of "double blinking". This is not conducive to single-molecule localization, and will lead to an increase in the proportion of unqualified localization points, which will be filtered out in the filter step, resulting in a small number of localization points at the A2 site and a significant decrease in detection efficiency. We can improve its detection efficiency by reducing the proportion of the A2 sequence in the index data domain. Otherwise, due to the low detection efficiency of individual index chain site 6, the copy redundancy level must be greatly increased to ensure sufficient sampling (similar to the bucket effect).
[0278] Example 3: Example of temporal information processing of DNA-PAINT imaging data
[0279] We demonstrate a "time-dimensional digital multi-base encoding method" by designing protrusions 103.3 of different sizes to control the time dimension of digital bits, i.e., the dwell time 801. The specific digital bit index sequence 502.1 is shown in Table 1.
[0280] We used a self-written python script to analyze several positioning points at each site and counted the dwell times 801 of all binding and dissociation events, as shown in Figure 18 shown. Figure 18(a) shows, in the left panel, representative fluorescence intensity traces of four engineered protrusions 103.3 on the nucleic acid complex structure 1. The middle panel shows a representative histogram (cumulative distribution function, CDF) of blinking event lengths and their corresponding exponential fits, used to quantitatively estimate the dwell time 801. The right panel shows the distribution of the average dwell time 801 for each site and its corresponding Gaussian fit. (b)-(c) show the quantitative calculation results of the dwell time 801 and dissociation time of the four protrusions 103.3. (d) shows the Gibbs free energy difference between A2 and the three protrusion 103.3 sequences.
[0281] The characteristic residence times 801 of the four sequences were obtained by exponential fitting, which showed large differences and were consistent with the expectation that the larger the size of the protrusion 103.3, the shorter the residence time 801. In addition, we used the fully complementary A2 as a kinetic reference to compare the subtle differences in binding strength of the three protrusion 103.3 sizes. It is worth noting that when quantitatively calculating the residence time 801 at a site, it is usually not possible to obtain a sufficient number of residence times 801 for exponential fitting. Therefore, the average residence time 801 at the site is calculated as an indicator of sequence recognition. From the Gaussian distribution results of the four sequences, it can be seen that the degree of separation is not large. This is an inherent thermodynamic barrier to the double-strand dissociation process, which can be overcome later by copy redundancy, such as Figure 19 shown. Figure 19 The figure shows the change of the average dwell time 801 (solid line) over all protrusions 103.3 sequences with increasing sampling volume. The data set contains 200 repeated samples for each sampling volume to calculate the average dwell time 801. The shaded area surrounds all data points from the set, i.e., from minimum to maximum.
[0282] In addition, we also studied the position dependence of the dwell time 801 of the four protrusion 103.3 sequences, e.g. Figure 20 For each homogeneous structure of the index data domain, the two-tailed Student's test was used to compare the residence time 801 distribution of each position to obtain statistical significance, as shown in Figure 21 This test calculates the probability that the dwell time 801 distributions at two arbitrary spatial locations are identical. Within the explored spatial range (12 digits), there are generally slight differences, but the dwell time 801 length distributions remain essentially consistent, with negligible variations. Figure 21 The dwell time 801 length distributions for each digit were compared. P values less than 0.05 indicated statistical significance.
[0283] When designing the index data domain and the index probe 7 sequence, it is important to pay attention to the upper limit of the residence time 801, because photobleaching during imaging may cause an underestimation of the actual residence time 801 (especially for longer binding times). If the photobleaching rate of the Cy3B fluorescent group is greater than the slowest dissociation rate of the index probe 7 in the index data domain (the slowest dissociation here is the A2 sequence, koff = 0.6 sec-1), the scheme of kinetically encoding the index information 301.2 bits deep will fail because the actual site residence time 801 cannot be observed. We conducted photobleaching experiments, such as Figure 22 As shown, the photobleaching rate of the Cy3B fluorophore (kphotobleaching = 0.3 sec-1) was significantly lower than the slowest dissociation rate in the index data domain. This indicates that the dwell time 801 measurement of index strand position 6 under our index information 301.2 reading system is not affected by the photophysical properties of the fluorophore.
[0284] Example 4: Example 301.2 of recovering index information from a single read
[0285] In one embodiment, based on the replica redundancy settings used in the experiment, we found that only approximately 70 nucleic acid complex structures 1 were required to recover index information 301.2 with a probability close to 100%. As with the error analysis of index information 301.2, this number is largely driven by the presence of structures in our sample, which are prone to site deletions and, due to read time limitations, the limited number of binding and dissociation events obtained (used to calculate dwell time 801), resulting in a certain degree of fluctuation in the average site time.
[0286] In practical applications, we randomly access the target nucleic acid complex structure 1 from the mixed structure. We generate a mixed test set containing 4 types of index information 301.2 (i.e., 4 index number sequences 4), such as Figure 24 As shown. The distance function in the “clustering method based on similarity calculation” is used to calculate the similarity between the unknown nucleic acid complex structure 1 index digital sequence 4 and the known 4 templates 902, and finally the predicted clustering result is obtained. We found that the smaller mean square distance (<R2_score> ) can improve clustering accuracy but it can only improve by 2%, such as Figure 25 As shown in (c), however, due to overly stringent screening criteria, half of the nucleic acid complex structures 1 cannot participate in cluster recognition.
[0287] Figure 25(a) shows the clustering accuracy of individual structures at different R2_scores; (b) shows the proportion of correct cluster structures at different R2_scores. A higher R2_score indicates a higher proportion of correct cluster structures in the dataset; (c) shows that as the sample size of nucleic acid complex 1 increases, the probability of reaching consensus in sequencing results increases.
[0288] Therefore, we do not filter the mean square distance results and use the<R2_score> The results showed that a single nucleic acid complex structure 1 could achieve 79% clustering accuracy. Subsequently, through replica redundancy, more than 25 similar nucleic acid complex structures 1 were sampled, and consensus was reached in the final sequencing results, accurately recovering data 301.
[0289] All documents mentioned herein are incorporated herein by reference as if each document were individually incorporated by reference. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the appended claims.
Claims
1. A nucleic acid storage medium, characterized in that: The nucleic acid storage medium comprises: (a) a nucleic acid composite structure configured to store predetermined data; Wherein, the nucleic acid composite structure comprises: (i) one or more backbone chains; (ii) multiple staple chains; (iii) one or more index chains, wherein the index chains include (M1) staple segments and (M2) index information storage segments; wherein the backbone chain, the staple chain, and the staple segments of the index chain self-assemble to form the nucleic acid composite structure; Furthermore, the index information storage segment of the index chain is exposed to the nucleic acid composite structure; The index information of the data is stored in the index chain.
2. The nucleic acid storage medium according to claim 1, wherein include: The content information of the data is stored in the skeleton chain, or the staple chain, or a combination thereof; The nucleic acid composite structure includes nucleic acid nanostructure, DNA composite structure, DNA nanostructure and DNA origami structure.
3. The nucleic acid storage medium according to claim 1, wherein A key-value architecture is used when associating the index information of the data with the content information of the data.
4. The nucleic acid storage medium according to claim 1, wherein (a) an index base sequence, the index base sequence comprising: (i) an index sequence of digital bits, wherein each digital bit index sequence identifies a digit of an index digital sequence, wherein the index digital sequence is formed by encoding index information of the data; (ii) an index sequence of a direction position, wherein each index sequence of a direction position identifies a site on the nucleic acid composite structure and is used to confirm the direction of the index digital sequence during reading. The index base sequence identifying each digital bit or direction bit forms an index information storage segment of one or more index chains.
5. The nucleic acid storage medium according to claim 4, characterized in that include: The index digital sequence comprises a quinary matrix array of three rows and four columns; The spatial distance between the index chains identifying the digital bits ranges from less than 30 nm; The index chains are evenly arranged on the surface of the nucleic acid composite structure in the order of the matrix array; The spatial position of the index strand that identifies the orientation on the nucleic acid complex structure includes: (i) the index chain identifying the first direction bit is located between the index chains identifying the four digits in the upper left corner of the matrix array; (ii) the index chain identifying the second direction bit is located between the index chains identifying the four digits in the lower left corner of the matrix array; (iii) The index strand identifying the third orientation is located at the upper right corner of the nucleic acid composite structure.
6. The nucleic acid storage medium according to claim 4, characterized in that include: The index information storage segment of the index chain that identifies the digital bit has a variety of protrusion shapes when combined with the same index probe; When the numerical values of the digits are the same, the index sequences of the digits are the same, and the shapes of the protrusions on the index chains identifying the digits are the same.
7. A method for preparing a nucleic acid storage medium according to claim 1, characterized in that: include: (a) encoding index information of predetermined stored data into an index digital sequence; (b) encoding the content information of the predetermined stored data into a content base sequence, designing a nucleic acid composite structure based on the content base sequence and the index digital sequence, and generating a backbone chain, a staple chain, and an index chain respectively; (c) annealing and assembling the backbone strand, the staple strand, and the index strand to form the nucleic acid composite structure.
8. A method for reading a nucleic acid storage medium according to claim 1, characterized in that: include: (a) generating an index probe based on the index information of the data to be searched, and inserting the index probe into a nucleic acid storage database, wherein the nucleotide sequence on the index probe is complementary to the base sequence on the index strand; (b) obtaining metadata of a plurality of nucleic acid complex structures based on optical interpretation of the pairing reaction of the index strands; (c) selecting a nucleic acid complex structure whose metadata meets the search requirements, decoding and sequencing the nucleic acid complex structure, and obtaining data that meets the search requirements.
9. The method for reading a nucleic acid storage medium according to claim 7, wherein: include: The index probe is of a single type, and the time it stays when binding and dissociating with index chains with different identifier values is different; The DNA-PAINT method is used to collect spatial and temporal information of the index data domain on a single nucleic acid complex structure.
10. A nucleic acid storage database, characterized by: The nucleic acid storage database comprises N nucleic acid storage media according to claim 1, wherein N is a positive integer ≥ 2; There is no macroscopic physical isolation between the N nucleic acid storage media.
Citation Information
Patent Citations
Sequencing of nucleic acids by emergence
CN111566211A
Nucleic acid memory (NAM) / digital nucleic acid memory (DNAM)
US20220025428A1
Methods and compositions for molecular authentication
WO2019183359A1