Data storage medium and use thereof

By combining nucleic acid molecules and vectors on a DNA origami substrate and utilizing the barrier strand, activation strand, and erasure strand, addressable writing, modification, and reading of DNA storage systems have been achieved. This addresses the shortcomings of existing DNA storage systems, improves storage density and reading accuracy, and supports multiple operations.

CN119604935BActive Publication Date: 2026-02-03SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380056815.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-01
Filing Date
2023-07-31
Publication Date
2026-02-03
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

Existing DNA storage systems cannot achieve addressable writing, addressable modification, and addressable reading. In particular, they lack the ability to address and modify arbitrary data stored in DNA storage systems, and errors are easily introduced after multiple reads.

Method used

By combining nucleic acid molecules containing data information with carriers containing addressable information, and by designing barrier strands, activation strands, and erasure strands, addressable reading, modification, and multiple erasure and writing of data can be achieved. The nanostructures and complementary sequences on the DNA origami substrate are used to programmably combine and separate information.

Benefits of technology

It achieves addressability of DNA storage system, improves areal density, supports multiple in-situ erasures and selective reads, does not destroy original data during the read process, has high address information accuracy, and is suitable for data operation under conditions from room temperature to 42 degrees Celsius.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119604935B_ABST
    Figure CN119604935B_ABST
Patent Text Reader

Abstract

The application provides a data storage medium and its application, and also provides a nucleic acid molecule, which can be combined with a carrier with addressable information, and data information contained in the nucleic acid molecule can be randomly read and erased in situ on the carrier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data storage, specifically to a data storage medium and its application. Background Technology

[0002] With the rise of big data and artificial intelligence technologies, the demand for storing massive amounts of data has exploded, and mainstream storage media are increasingly unable to meet this rapidly growing demand. Deoxyribonucleic acid (DNA), as the storage medium for the carbon-based genetic code of life, selected through millions of years of natural evolution, possesses extremely high storage density and robustness. DNA's inherent coding and efficient replication capabilities offer the potential to provide a novel strategy for high-density data storage and high-performance computing. DNA storage boasts high physical stability, unlike electronic media which degrade with repeated reads, providing a fundamental solution for long-term data storage. Furthermore, DNA combines information processing and computing capabilities, offering new ideas for developing novel storage-computing architectures and systems.

[0003] The DNA data storage process mainly includes six technical steps: information encoding, writing, storage, retrieval, reading, and decoding. Existing DNA storage systems are mainly based on artificially synthesized short-chain or long-chain DNA. Digital information is encoded into DNA sequences and corresponding DNA strands are synthesized, stored in cells or in vitro, and then read out using sequencing technology.

[0004] Compared to existing mature data storage systems, current DNA storage systems mainly include two storage modes. One mode is a non-random access storage architecture, where the data to be stored is encoded and written as a whole. Therefore, reading the data requires sequencing all sequences in the system and decoding the overall sequencing results to obtain the data information. Non-random access storage architectures cannot search for partial information within a file, and therefore cannot modify the written information in an addressable manner. The other mode is a DNA storage architecture with a random access mode. In this type of system, the data to be written is first segmented into data fragments, which are then encoded and indexed. Therefore, this type of DNA storage architecture has the ability to selectively read information fragments. However, selective reading in this type of system mainly relies on PCR technology, requiring the direct extraction of a portion of the original data or the extraction of the data to be read using magnetic beads followed by PCR amplification and then sequencing. However, this method disrupts the original data composition; after multiple reads and PCR cycles, many errors are introduced into the data, affecting the recovery of the original data. Similar to the in-situ read method of traditional semiconductor storage media, a data reading method that enables selective data reading without destroying the original data is still lacking. Furthermore, current DNA storage systems primarily offer single-write archiving capabilities, while the ability to modify data within the storage system after writing remains significantly insufficient, particularly the addressable modification of arbitrary data stored in the DNA storage system.

[0005] Therefore, there is an urgent need in the field for an addressable write, addressable modify and / or addressable read DNA storage method. Summary of the Invention

[0006] This application provides a DNA storage system with complete data manipulation capabilities, enabling addressable functions such as data writing, deletion, modification, and reading. In particular, it addresses the functional limitations of existing DNA storage systems by allowing for multiple write / erase operations, data modification, and multiple readouts. For example, this application enables programmable binding and separation of storage addresses and data on DNA molecules.

[0007] On the one hand, this application provides a system comprising 1) a nucleic acid molecule containing data information and 2) a carrier having addressable information, wherein the nucleic acid molecule can bind to the carrier and the data information contained in the nucleic acid molecule can be stored (also known as written), read in situ and / or edited in situ (including erasing and / or writing new data information) in the carrier.

[0008] On the other hand, this application provides a data storage method, the method comprising providing 1) the nucleic acid molecule containing data information, and 2) the carrier having addressable information.

[0009] On the other hand, this application provides a method for reading data, the method comprising providing 1) the nucleic acid molecule containing data information, and 2) the carrier having addressable information.

[0010] On the other hand, this application provides a data editing method, the method comprising providing 1) the nucleic acid molecule containing data information, and 2) the carrier having addressable information, and replacing the data information in the nucleic acid molecule.

[0011] In some embodiments, different physical locations of the carrier have different sequence of address sequences.

[0012] In some embodiments, the carrier comprises a DNA origami substrate, the DNA origami substrate comprising a staple chain, the staple chain comprising an address sequence having addressable information.

[0013] In some embodiments, the address sequences at the ends of the staple chains are arranged in a matrix within the carrier. The matrix is ​​at least a 2×2 matrix. In some embodiments, the spacing between two adjacent staple chains is approximately 6-24 nm. For example, the spacing between two adjacent staple chains is approximately 6 nm. For example, the spacing between two adjacent staple chains is approximately 12 nm. For example, the spacing between two adjacent staple chains is approximately 18 nm. For example, the spacing between two adjacent staple chains is approximately 24 nm.

[0014] In some embodiments, the nucleic acid molecules are bound to the staple chain and arranged in a matrix. The matrix is ​​at least a 2×2 matrix. In some embodiments, the spacing between two adjacent nucleic acid molecules is approximately 6-24 nm. For example, the spacing between two adjacent nucleic acid molecules is approximately 6 nm. For example, the spacing between two adjacent nucleic acid molecules is approximately 12 nm. For example, the spacing between two adjacent nucleic acid molecules is approximately 18 nm. For example, the spacing between two adjacent nucleic acid molecules is approximately 24 nm.

[0015] In some embodiments, the nucleic acid molecule includes a complementary sequence that is complementary to an address sequence on the vector. Through the complementarity between the complementary sequence and the address sequence, the nucleic acid molecule binds to a specific physical location on the vector.

[0016] In some implementations, the address complementary sequence is about 15 or more nucleotides in length.

[0017] In some embodiments, the nucleic acid molecule contains a data sequence, and the data information in the data sequence can be read while the nucleic acid molecule is substantially inseparable from the vector. For example, a readable data strand can be synthesized by activating a reaction, thereby allowing the data information in the data sequence to be read.

[0018] In some embodiments, the data sequence is about one or more nucleotides in length. The data sequence may be single-stranded.

[0019] In some embodiments, the nucleic acid molecule includes a read-initiating sequence that can induce the data sequence to be synthesized or transcribed into a sequencing strand, and the sequencing strand is complementary to the data sequence.

[0020] In some implementations, the read-initiating sequence includes a promoter. For example, the read-initiating sequence is the T7 promoter. The read-initiating sequence enables the data sequence to be transcribed. By sequencing and decoding the transcription product, the data information contained in the data sequence can be read.

[0021] In some embodiments, the system and / or method further includes 3) a barrier chain, which prevents the data information from being read. In this application, the barrier chain may also be referred to as a closing chain.

[0022] In some embodiments, the blocking chain includes a blocking sequence that is complementary to the read-initiating sequence, thereby preventing the data from being read. In one specific embodiment, the blocking sequence includes a promoter complement sequence that binds to the promoter, which is the read-initiating sequence, preventing transcription from starting.

[0023] In some embodiments, the blocking strand includes a blocking extension sequence located upstream and / or downstream of the blocking sequence, and the blocking extension sequence is substantially non-complementary to the nucleic acid molecule. In this application, the blocking extension sequence may also be referred to as a teohold. A teohold can provide the energy drive and specificity for transcriptional activation. In one specific embodiment, the blocking extension sequence is about 4-10 nucleotides long, for example, 6 nucleotides.

[0024] In this application, for the data information at each address, in the storage state or when it is not needed to be read, the barrier strand and the reading initiation sequence on the nucleic acid molecule are complementary to bind, and transcription cannot be initiated.

[0025] In some embodiments, the system and / or method further includes 4) an activation chain, which can be added, for example, when data reading and writing are required. The activation chain prevents the blocking chain from binding to the nucleic acid molecule, and is complementary to both the blocking sequence and the blocking extension sequence. The activation chain corresponds to a specific physical address. When data information at a specific address needs to be read, an activation chain corresponding to that address can be added. This activation chain binds to the blocking chain, causing the blocking chain to detach from the read-initiating sequence of the nucleic acid molecule, thereby initiating reading, for example, activating the transcription function at that address. In one specific embodiment, adding an activation chain corresponding to a specific address activates the transcription function at that address. Under the action of T7 RNA polymerase, the data at that address is transcribed into an RNA chain. Unactivated addresses do not have transcriptional ability. The obtained RNA chains are collected for subsequent sequencing to decode the nucleic acid sequence into data information, thereby achieving addressable data reading.

[0026] In some implementations, the data reading process can be performed at room temperature. At room temperature, the barrier strand to which a nucleic acid molecule at a specific physical location is bound can bind to the activation strand.

[0027] In some embodiments, the nucleic acid molecule includes an erasure sequence located upstream and / or downstream of the address complementary sequence, and the erasure sequence is substantially non-complementary to the address sequence on the vector.

[0028] In some embodiments, the system and / or method further includes 5) an erasure strand, for example, which can be added when data erasure is required. The erasure strand prevents the nucleic acid molecule from binding to the vector, and the erasure strand is complementary to both the address complementary sequence and the erase / write functional sequence. In some embodiments, the erasure strand can bind complementary to the erase / write functional sequence of the nucleic acid molecule. The erasure strand corresponds to a specific physical address. When data information at a specific address needs to be erased, an erasure strand corresponding to that specific address can be added, and then, for example, through a DNA strand substitution reaction, the data sequence at that address is bound, thereby removing the data sequence from the vector and restoring the address to an unwritten state.

[0029] In some embodiments, the erasure process can be performed at room temperature. At room temperature, after the erasure strand is added, the nucleic acid molecule at a specific physical location can bind to the erasure strand, and the nucleic acid molecule does not substantially bind to the carrier.

[0030] In some implementations, when the original data information is erased—that is, when the original data sequence is erased by the erasure chain—the address can be rewritten with a new data sequence, for example, by adding a new nucleic acid molecule containing the data information. This enables in-situ modification and / or editing of the data.

[0031] In some embodiments, the molar ratio of the nucleic acid molecule to the address sequence in the vector is about 1:1 or higher. In some embodiments, the molar ratio of the nucleic acid molecule to the address sequence in the vector is about 2:1 or higher. In some embodiments, the molar ratio of the nucleic acid molecule to the address sequence in the vector is about 3:1 or higher. In some embodiments, the molar ratio of the nucleic acid molecule to the address sequence in the vector is about 4:1 or higher. In some embodiments, the molar ratio of the nucleic acid molecule to the address sequence in the vector is about 5:1 or higher.

[0032] On the other hand, this application provides a nucleic acid molecule in the system and / or method, the nucleic acid molecule comprising the data sequence, address complement sequence, and / or read-initiating sequence. In one embodiment, upstream and / or downstream of the data sequence, the nucleic acid molecule further comprises primers for verifying the feasibility of erasing and reading. In other specific embodiments, the primer portion may be replaced by a data sequence containing data information, further increasing storage capacity. Figure 2 The image shown is an example of a nucleic acid molecule described in this application.

[0033] On the other hand, this application provides a carrier for the said system and / or method.

[0034] On the other hand, this application provides a storage medium that includes the method of this application.

[0035] On the other hand, this application provides an apparatus comprising a storage medium of this application and a processor coupled to the storage medium, the processor being configured to execute, based on a program stored in the storage medium, the method of this application.

[0036] In one specific embodiment, the inventors of this application constructed a 6nm DNA chip to achieve multiple in-situ erasure and / or multiple in-situ selective reads (WMRM). In a more specific embodiment, utilizing the addressability and programmability of DNA nanostructures, data sequences storing information are arranged at 6nm intervals on DNA origami, each data sequence is assigned a physical address at a 6nm interval, and by designing erasure strands, the information processing and computing capabilities provided by DNA strand substitution reactions are utilized to achieve repeated erasure and writing of DNA data sequences on the origami, and modification of arbitrary data. Furthermore, promoter switches (e.g., T7 promoters) are set on the DNA strands as read initiation sequences. By designing blocking strands and / or activating strands, the transcriptional activity of the DNA strands is regulated by DNA strand substitution reactions, selectively transcribing DNA information for RNA sequencing (e.g., using Illumina technology), achieving selective and multiple reads of data information.

[0037] Compared with existing technologies, this application has at least the following characteristics: Based on DNA assembly technology, it possesses nanoscale addressability, which is not available in current DNA storage systems; utilizing this nanoscale addressability, the areal density of DNA storage can be increased to 1 bit / square nanometer, surpassing existing inorganic and DNA storage architectures; based on the controllable dynamic assembly of DNA origami surfaces, a fully functional storage system is achieved, including data reading, writing, and modification operations; the readout process employs in-situ transcription into RNA molecules, followed by sequencing using the transcription products, without altering the original storage system, achieving non-destructive DNA data readout. For example, the method and data carrier of this application can achieve multiple read / write and multiple read operations. For example, after multiple reads, the structure of the data system of this application remains essentially unchanged. For example, the address information of the data carrier of this application has high accuracy and specificity, enabling a higher level of discrimination. For example, the method and data system of this application can achieve data reading at room temperature to approximately 42 degrees Celsius.

[0038] Other aspects and advantages of this application will readily be apparent to those skilled in the art from the detailed description below. Only exemplary embodiments of this application are shown and described in the following detailed description. As will be appreciated by those skilled in the art, the content of this application enables them to make modifications to the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application pertains. Accordingly, the descriptions in the accompanying drawings and specification of this application are merely exemplary and not restrictive. Attached Figure Description

[0039] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:

[0040] Figure 1 The illustration shows an exemplary operation flow of the data storage medium of this application. For example, the stored data can be used in all data formats, including but not limited to Chinese characters; the storage capacity can be infinitely expanded based on the length of the data sequence, including but not limited to 16 bits / site or 120 bits / site; the data can be read using sequencing technologies known in the art, including but not limited to high-throughput in situ sequencing, transcription-RNA sequencing, and transcription-reverse transcription-amplification-DNA sequencing.

[0041] Figure 2 The diagram shows an exemplary structural composition of the nucleic acid molecule of this application, namely, a data chain containing a data sequence, wherein the design of primers 1 and 2 is only for the purpose of more easily verifying the feasibility of selective erasure and read. This part can be replaced with actual stored information to further increase the storage capacity.

[0042] Figure 3 This illustrates an exemplary process for selective reading and reversible erasure / writing of the data storage medium of this application. Transcription activation process: The keychain is the activation chain; when the nucleic acid molecule described in this application is blocked by the blocking chain, transcription cannot start. When the keychain is added, the blocking chain binds to the keychain instead of binding to the promoter. After the complete T7 promoter is added, RNA polymerase binds to the promoter, initiating transcription. Reversible erasure / writing: When the erasure chain (i.e., the erase chain) is added, the data chain binds to the erasure chain and detaches from the vector. After a data chain with a complementary sequence at the same address is re-added, the data chain is rewritten onto the vector.

[0043] Figure 4 This displays the structural characterization results after data was written onto the DNA origami surface. The upper part shows a schematic diagram of the addresses where the written data was located on the DNA origami surface. The lower part shows the AFM characterization results after data writing.

[0044] Figure 5 The display shows a heatmap obtained from the statistical analysis of PAGE data showing the orthogonality of readout activations. Data can only be read efficiently when the activation chain matches the data chain.

[0045] Figure 6 The display shows the next-generation sequencing results after reading data from seven addresses. The sequence information obtained from sequencing at each address is completely consistent with the information written.

[0046] Figure 7The image displayed shows the super-resolution fluorescence microscopy imaging results of the addressing data erasure and writing process. The upper part shows the data point erasure and writing process, along with a schematic diagram of the corresponding data addresses. The lower part shows the fluorescence imaging structural characterization results for each step of this process.

[0047] Figure 8 The results displayed show the fluorescence test results demonstrating the feasibility of repeatedly modifying the data.

[0048] Figure 9 The image displayed is a single-point TIRF (total internal reflection fluorescence microscopy) image of seven address data that has been repeatedly erased and rewritten 10 times.

[0049] Figure 10 The image displayed is a TIRF (Total Internal Reflection Fluorescence Microscopy) image and colocalization statistics of seven address data repeated 10 times. Red and green fluorescence colocalization represents the presence of data on the origami.

[0050] Figure 11 This shows a qPCR quantification graph of the transcribed RNA during repeated reads of the data chain.

[0051] Figure 12 The display shows the data chain write results under four different matrix spacings.

[0052] Figure 13 The display shows the data chain readout results under four different matrix spacings. Detailed Implementation

[0053] The following specific embodiments illustrate the implementation of the invention. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification.

[0054] Terminology Definition

[0055] In this application, the term "addressability" generally refers to associating specific information with a location on a storage medium. For example, in order to selectively access, read, and / or modify information at a specific location, the data carrier needs to be addressable. For example, when recording data, the information carrier can simultaneously record substantially unique index information (addresses) corresponding to different pieces of data. For example, the address information of specific data can be recorded through the physical location, spatial location, etc. of the data carrier. For example, when the data carrier is addressable, the information that is desired to be accessed, read, and / or modified can be accessed selectively, without having to access each piece of information one by one.

[0056] In this application, the terms "nucleic acid molecule," "nucleic acid sequence," and "nucleic acid fragment" are used interchangeably and generally refer to deoxyribonucleotides or ribonucleotides of various lengths, or analogues thereof. Exemplary nucleotides include deoxyribonucleotides (DNA) or ribonucleotides (RNA), or non-standard nucleotides, nucleotide analogues, and / or modified nucleotides.

[0057] In this application, the term "carrier" generally refers to a substance capable of loading nucleic acid molecules. For example, the carrier may comprise a nucleic acid nanostructure, which (also called a nanostructure) can be a two-dimensional or three-dimensional nanostructure made of nucleic acids (e.g., DNA, RNA, locked nucleic acids (LNA), peptide nucleic acids (PNA), or any combination thereof). For example, single-stranded or double-stranded nucleic acids (e.g., having only a helical structure) may not be considered "nanostructures." In some embodiments, the nucleic acid nanostructure acts as a scaffold for forming more complex structures such as molecular complexes. In some embodiments, the nucleic acid nanostructure is a DNA origami structure assembled using DNA origami methods. For example, a nucleic acid origami nanostructure can refer to a nucleic acid nanostructure formed by assembling two or more "staple chains" with one or more "scaffold" chains into a prescribed shape. Staple chains are typically short (e.g., 50 nucleotides or less) nucleic acid chains (single-stranded nucleic acids); scaffold chains are typically longer (e.g., longer than 200 nucleotides) nucleic acid chains (single-stranded nucleic acids). The nucleic acid origami nanostructure can be a DNA origami nanostructure.

[0058] DNA origami nanostructures can be folded (e.g., through self-assembly) into discrete and unique geometric patterns, such as two-dimensional (2D) and three-dimensional (3D) shapes, which can be further self-assembled to create larger nanostructures or microstructures containing two or more discrete origami nanostructures. In some embodiments, the scaffold chain has a sequence derived from M13 phage. Other scaffold chains can be used. In some embodiments, the staple chain is a fluorophore-labeled staple chain. In some embodiments, the staple chain is 4 to 30 nucleotides in length (e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides). In some embodiments, for example at room temperature, the staple chain binds stably to the scaffold chain (for longer than 10 seconds). In some embodiments, for example at room temperature, the staple chain binds stably to the scaffold chain (for longer than one week). In some implementations, the staple chain is longer than 30 nucleotides.

[0059] In this application, the term "in situ" generally refers to an operation performed in its original location. For example, in situ reading refers to reading data from the nucleic acid molecule being recorded at its original location on the vector, without needing to release the nucleic acid molecule into solution before reading the data. The term "amplification" generally refers to the production of copies of nucleic acid molecules via repeated cycles of initiated enzymatic synthesis. The reading step in this application may include a polymerization step, which includes, but is not limited to, polymerase chain reaction (PCR), transcription to RNA, etc., and may also include any other nucleic acid amplification and / or transcription techniques known to those skilled in the art.

[0060] In this application, the terms "complementary" or "complementarity" generally refer to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. For example, the sequence AGT is complementary to the sequence TCA. The complementarity percentage indicates the percentage of residues in the nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with the second nucleic acid sequence (e.g., 5%, 6%, 7%, 8%, 9%, and 10% complementarity, respectively). For example, "complete complementarity" means that all consecutive residues in the nucleic acid sequence will hydrogen bond with the same number of consecutive residues in the second nucleic acid sequence. For example, "substantially complementary" refers to a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% between two nucleic acids that hybridize under strict conditions (i.e., strict hybridization conditions). For example, "binding" generally refers to, for example, a class of things that binds uniquely to a particular class of things in a sequence-specific manner.

[0061] In this application, the term "comprising" generally means including the explicitly specified features, but does not exclude other elements.

[0062] In this application, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value. Invention Details

[0064] On the one hand, this application provides a nucleic acid molecule that can bind to a carrier with addressable information, and the data information contained in the nucleic acid molecule can be randomly read and erased in situ on the carrier. For example, the nucleic acid molecule provided in this application can serve as a data chain that binds to a specific location on the carrier based on the principle of base complementarity. Combined with dynamic DNA assembly technology, programmable binding and separation of storage addresses and data can be achieved. For example, when reading the data information, the original storage system can be maintained, achieving non-destructive data reading. For example, when reading the data information recorded in the nucleic acid molecule, the nucleic acid molecule does not need to be released from the carrier first, achieving in-situ random reading of the required information.

[0065] For example, address sequences at different physical locations on the vector have different sequences, and the nucleic acid molecule may contain address-complementary sequences, which are complementary to the address sequences on the vector. For example, the address-complementary sequence may be an address recognition sequence that can recognize and specifically bind to the address sequences on the vector. For example, address information can be recorded using address sequences at different physical locations on the vector. For example, the vector may have two or more coordinate points. For example, different coordinate points represent different physical locations. For example, a specific coordinate point may be located by the end of a staple chain at a specific location in DNA origami. For example, different coordinate points extend into address sequences with unique sequence compositions. For example, the address sequences of this application may be linear single-stranded or branched strand structures. For example, address sequences at specific physical locations, due to their unique sequences, can be used to record index information, and the data information can be bound to the address information when a substantially uniquely complementary address sequence (containing the address-complementary sequence in a specific data strand) binds to the address sequence. In some embodiments, for example at room temperature, the data strand and the address strand bind stably (e.g., for longer than 10 seconds). In some implementations, such as at low temperatures (approximately 4 degrees Celsius), the data link and address link are stably bound (e.g., for more than 10 seconds). In some implementations, such as when stored in a dry powder state on an antibacterial, antioxidant, and heat-resistant plate, the data link and address link are stably bound (e.g., for more than 10 seconds).

[0066] For example, the length of the address complementary sequence is about 15 or more nucleotides. For example, the address complementary sequence of this application can be a linear single-stranded or branched strand structure. For example, the length of the address complementary sequence can be about 10 or more, about 11 or more, about 12 or more, about 13 or more, about 14 or more, about 15 or more, about 16 or more, about 17 or more, about 18 or more, about 19 or more, about 20 or more, about 25 or more, about 30 or more, about 40 or more, about 50 or more, or about 100 or more nucleotides. For example, the nucleotides can comprise natural nucleotides and / or nucleotides with artificial modifications, such as including, but not limited to, methyl, amino, fluorinated modifications, etc.

[0067] For example, the nucleic acid molecule may further include an erasure / write functional sequence located upstream and / or downstream of the address complementary sequence, and the erasure / write functional sequence is substantially non-complementary to the address sequence on the vector. For example, the erasure / write functional sequence may also be referred to as an erasure / write functional region. For example, the erasure / write functional sequence is located upstream and / or downstream of the address complementary sequence. For example, the sequence formed by the erasure / write functional sequence and the address complementary sequence may not be completely complementary to the address sequence on the vector. For example, an exemplary address complementary sequence may be 20 nucleotides, the erasure / write functional sequence may be 10 nucleotides, and the address sequence on the vector may only be complementary to the 20 nucleotides of the address complementary sequence, and may not be completely complementary to the 30 nucleotides formed by the erasure / write functional sequence and the address complementary sequence. For example, a strand with higher binding capacity (e.g., the erasure strand of this application) can therefore be introduced so that the nucleic acid sequence of the data strand can be removed (erased) from a specific location on the vector.

[0068] For example, when the eraser strand is present, the nucleic acid molecule can avoid binding to the vector, and the eraser strand is complementary to both the address strand and the erase / write functional sequence. For example, at room temperature, the nucleic acid molecule (data strand) of this application can have a stronger binding affinity to the eraser strand. For example, compared to binding to the address sequence on the vector, the data strand can have more and / or stronger binding bases to the eraser strand.

[0069] For example, to achieve addressable erasure and writing, the complementary sequence of the nucleic acid molecule (data chain), the address sequence corresponding to a specific position in the vector, and the sequence of the corresponding eraser strand can be designed with a specific base sequence, possessing a unique complementary matching mode. For instance, the address sequence at a specific position in the vector is uniquely complementary to the complementary sequence of the specific nucleic acid molecule (data chain) to achieve addressable writing. Similarly, the complementary sequence of the specific nucleic acid molecule (data chain) bound to a specific position in the vector is uniquely complementary to the sequence of the corresponding eraser strand to achieve addressable erasure.

[0070] For example, the nucleic acid molecule may contain a data sequence, and the data information in the data sequence can be read while the nucleic acid molecule and the vector are substantially inseparable. For example, the nucleic acid molecule (data strand) can be used to read data information in its in-situ storage location. For example, reading methods include, but are not limited to, in-situ sequencing, in-situ transcription into other nucleic acid molecules for information reading, in-situ reading into other nucleic acid molecules for information reading, etc. For example, exporting the information stored in the nucleic acid molecule (data strand) in this application can be considered as reading. For example, sequencing the exported (e.g., transcribed / amplified) other nucleic acid molecules can be considered as a subsequent optional additional reading step.

[0071] For example, the length of the data sequence is about one or more nucleotides. For example, the data sequence of this application can be a linear single-stranded or branched strand structure. For example, the length of the data sequence can be about one or more, about two or more, about three or more, about four or more, about five or more, about six or more, about seven or more, about eight or more, about nine or more, about ten or more, about fifteen or more, about twenty or more, about thirty or more, about forty or more, about fifty or more, about one hundred or more, about one hundred and twenty or more, about one hundred and twenty or more, about one hundred and fifty or more, about two hundred and twenty or more, about five hundred or more, about seven hundred or more, or about one hundred and ten thousand or more. For example, the storage method of this application can be compatible with data sequences of any length.

[0072] For example, the nucleic acid molecule may contain a read-initiating sequence that triggers the reading of the data sequence into a sequencing strand, and the sequencing strand is complementary to the data sequence. For example, the nucleic acid molecule may contain a read-initiating sequence that triggers the reading of the data sequence. For example, the read-initiating sequence of this application may be a promoter, such as, but not limited to, the T7 promoter. For example, the storage method of this application may be compatible with read-initiating sequences of any length and type, and the specific sequence of the read-initiating sequence can be used for the binding and polymerization initiation of polymerases known in the art.

[0073] For example, the nucleic acid molecule (data chain) of this application can have selective reading capabilities. For example, selective reading can be achieved by introducing a blocking strand, the blocking sequence of which is complementary to the read-initiating sequence. For example, when the blocking strand is present, the data sequence can be prevented from being read. The blocking strand can contain a blocking sequence, and the blocking sequence of the blocking strand is partially or completely complementary to the read-initiating sequence. For example, the blocking sequence of the blocking strand is complementary to the read-initiating sequence and a region approximately 5 nucleotides upstream / downstream of the read-initiating sequence. For example, the blocking sequence is 22 nucleotides long, of which 17 nucleotides are complementary to the read-initiating sequence, and the other 5 nucleotides are complementary to a region approximately 5 nucleotides upstream / downstream of the read-initiating sequence. For example, the blocking strand has a higher binding affinity to the data chain compared to the T7 promoter binding. For example, the blocking strand of this application can bind to the read-initiating sequence of the nucleic acid molecule (data chain), the data chain, or any location capable of preventing the export of information stored in the nucleic acid molecule (data chain).

[0074] For example, the blocking strand may further include a blocking extension sequence located upstream and / or downstream of the blocking sequence, the blocking extension sequence being substantially non-complementary to the nucleic acid molecule. For example, the blocking strand has a blocking sequence and a blocking extension sequence, where, when the blocking strand binds to the nucleic acid molecule (data strand), the blocking extension sequence is substantially non-binding to the nucleic acid molecule (data strand), for example, forming a dangling structure. For example, the blocking extension sequence may be approximately 8 nucleotides in length. For example, the blocking extension sequence can act as a lever to enhance the binding affinity between the activating strand and the blocking strand, causing the activating strand to remove the blocking strand from the data strand.

[0075] For example, when the activating strand is present, the blocking strand can not bind to the nucleic acid molecule, and the activating strand is complementary to both the blocking sequence and the blocking extension sequence. For example, during annealing, the activating strand of this application can have a stronger binding affinity to the blocking strand. For example, compared to sequences on the nucleic acid molecule (data strand) that bind to the blocking strand, the activating strand can have more and / or stronger binding bases to the blocking strand.

[0076] For example, to achieve addressable reading, the sequence used for blocking reading, the corresponding blocking sequence of the specific blocking strand, and the corresponding activating strand of the nucleic acid molecule (data chain) can be designed with specific base sequences to have unique complementary matching methods. For example, the sequence used for blocking reading of the nucleic acid molecule (data chain) is uniquely complementary to the blocking sequence of the specific blocking strand to achieve addressable locked reading. For example, the sequence of the specific activating strand is uniquely complementary to the blocking sequence of the specific blocking strand to achieve addressable unlocked (activated) reading. For example, when a specific nucleic acid molecule (data chain) is in a locked reading state, the data sequence of the specific nucleic acid molecule (data chain) can be substantially not read, transcribed, and / or amplified. For example, when a specific nucleic acid molecule (data chain) is in an unlocked (activated) reading state, the data sequence of the specific nucleic acid molecule (data chain) can be read, transcribed, and / or amplified.

[0077] For example, the carrier may include a DNA origami substrate, and the staple chains of the DNA origami may contain address sequences with addressable information. For example, the carrier may have two or more coordinate points. For example, different coordinate points represent different physical locations. For example, a specific coordinate point may be located by the end of a staple chain at a specific location in the DNA origami. For example, different coordinate points extend into address sequences with unique sequences. For example, the address sequence at a specific physical location, due to its unique sequence, can be used to record index information, and the data information can be combined with the address information when a substantially unique complementary address sequence (containing the complementary address sequence in a specific data chain) is combined with the address sequence. For example, the spacing between two or more of the staple chains is about 6 nanometers or greater. For example, the spacing between adjacent staple chains is about 6 nanometers or greater, about 7 nanometers or greater, about 8 nanometers or greater, about 9 nanometers or greater, about 10 nanometers or greater, about 15 nanometers or greater, about 20 nanometers or greater, about 25 nanometers or greater, or about 30 nanometers or greater.

[0078] This application provides a system that may include the nucleic acid molecule of this application and a vector. For example, the system may also include an erasure strand, a barrier strand, and / or an activation strand of this application.

[0079] This application provides a data storage method, which may include providing the nucleic acid molecule of this application and / or the system of this application. A data editing and / or data reading method, wherein the data editing method may include replacing data information stored in the nucleic acid molecule of this application. A data reading method, wherein the data reading method may include determining the data information stored in the nucleic acid molecule of this application.

[0080] For example, the method may further include providing a carrier, which may comprise a DNA origami substrate, the staples of which may comprise an address sequence having addressable information, and the method may include providing the nucleic acid molecule and the address sequence in a molar ratio of about 2:1 or higher. For example, the method may include providing the nucleic acid molecule and the address sequence in a molar ratio of about 2:1 or higher, about 2.1:1 or higher, about 2.2:1 or higher, about 2.3:1 or higher, about 2.4:1 or higher, about 2.5:1 or higher, about 3:1 or higher, about 4:1 or higher, about 5:1 or higher, about 10:1 or higher, about 20:1 or higher, about 50:1 or higher, or about 100:1 or higher. For example, selecting a suitable molar ratio can improve the storage success rate of the data link. For example, selecting a suitable molar ratio can improve the storage cost-effectiveness of the data link.

[0081] For example, the method may also include providing an erasure strand of the present application, wherein the nucleic acid molecule at a specific physical location binds to the erasure strand at room temperature, and the nucleic acid molecule substantially does not bind to the carrier. The terms "room temperature" and "ambient temperature" generally refer to a temperature between about 16 degrees Celsius and about 40 degrees Celsius. For example, a temperature between about 16 degrees Celsius and about 25 degrees Celsius. For example, about 25 degrees Celsius.

[0082] For example, the method may further include providing a barrier strand of this application, which binds to the nucleic acid molecule, and the data information of the nucleic acid molecule is substantially unreadable. For example, the method may further include providing an activation strand of this application, wherein during heating and cooling, the barrier strand bound to the nucleic acid molecule at a specific physical location binds to the activation strand, and the barrier strand is substantially not bound to the nucleic acid molecule. For example, the heating and cooling process of this application includes a process of annealing the nucleic acid. For example, the process of separating part of the double strand of the nucleic acid molecule of this application (e.g., heating) and then restoring it to a partially double-stranded structure (e.g., cooling) can be considered a nucleic acid annealing process. For example, heating at approximately 95 degrees Celsius for approximately 3 minutes, and then cooling to room temperature at a rate of approximately 1.2 degrees Celsius per minute can be considered a nucleic acid annealing process.

[0083] On the other hand, this application provides a storage medium that records a program capable of running the methods of this application.

[0084] On the other hand, this application provides an apparatus that may include the storage medium of this application. On the other hand, this application provides a non-volatile computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement any one or more methods described in this application. For example, the non-volatile computer-readable storage medium may include floppy disks, flexible disks, hard disks, solid-state storage (SSS) (e.g., solid-state drives (SSDs)), solid-state cards (SSCs), solid-state modules (SSMs)), enterprise-grade flash drives, magnetic tape, or any other non-transitory magnetic media. The non-volatile computer-readable storage medium may also include punched cards, paper tape, cursor sheets (or any other physical medium with perforated patterns or other optically identifiable markings), compact disc read-only memory (CD-ROM), rewritable optical disc (CD-RW), digital versatile optical disc (DVD), Blu-ray disc (BD), and / or any other non-transitory optical media.

[0085] For example, the device of this application may also include a processor coupled to the storage medium, the processor being configured to execute the methods of this application based on a program stored in the storage medium. For example, the device may implement various mechanisms to ensure that the methods described in this application, executed on a database system, produce correct results. In this application, the device may use a disk as a persistent data storage device. In this application, the device may provide database storage and processing services to multiple database clients. The device may store database data across multiple shared storage devices and / or may utilize one or more execution platforms with multiple execution nodes. The device can be organized such that storage and computing resources can be effectively and infinitely expanded.

[0086] This application provides the following implementation methods:

[0087] The fabrication of a nanoscale addressable, fully functional DNA storage system includes the following steps:

[0088] (1) Design and assembly of a nanoscale addressable DNA platform. The addresses on the surface of the DNA origami are represented in the form of DNA sequences. An orthogonal sequence library with a length of, for example, 15 to 18 bases or longer is designed using a sequence screening algorithm to ensure that the data strands can be stored at specific addresses in a stable hybridization manner. Using a single-layer rectangular DNA origami as a template, the address sequences are extended from the corresponding sites on the origami surface by designing its staple chains. The staple chains with address design are combined with the backbone chain and assembled through annealing to form an addressable blank DNA storage platform.

[0089] The DNA origami assembly process can refer to known DNA origami assembly techniques in this field, such as Rothemund, PWK, Folding DNA to Create Nanoscale Shapes and Patterns. Nature 2006, 440(7082), 297–302.

[0090] An exemplary overall storage process can be as follows: Figure 1 As shown, the exemplary data chain (nucleic acid molecule described herein) structure can be as follows: Figure 2 As shown, selective reading and reversible erasure / writing of data can be performed as follows: Figure 3 As shown

[0091] (2) Encode the data corresponding to the address. The data to be written is divided into segments that match the capacity of a single address, and the T7 promoter sequence and address sequence are added to form the data chain sequence to be written, which is then synthesized to obtain the data chain. The method of synthesizing nucleic acid chains is well known in the art, for example, by chemical synthesis of data chains.

[0092] A data operation system for an addressable DNA storage system, the system comprising:

[0093] (1) A data writing (or storage) step, which may include using the DNA complementary pairing principle to add a data strand containing the data sequence to bind it at a specific address.

[0094] (2) Data erasure step, which may include adding an erasure operation DNA strand (also known as an Erase strand, i.e., an erasure strand) to a specific address, and through a DNA strand substitution reaction, binding the data sequence at that address to restore the address to an unwritten state.

[0095] (3) Data modification (or editing) steps, which may include erasing the original data and then writing the new data.

[0096] (4) Data readout (or reading) step, which may include adding an activation strand (also called an input strand) corresponding to a specific address to activate the transcription function of that address. Under the action of T7 RNA polymerase, the data at that address is transcribed into an RNA strand. Unactivated addresses do not have transcription ability. The obtained RNA strands are collected for subsequent sequencing, thereby achieving addressable data reading.

[0097] The embodiments described below are not intended to be limited by any theory, but are merely for illustrating the products, preparation methods and uses of this application, and are not intended to limit the scope of the invention.

[0098] Example

[0099] Example 1 Preparation of a Nano-Addressable All-Functional DNA Storage System

[0100] The design process of the nano-addressable DNA platform is as follows:

[0101] (1) The DNA origami scaffold strand uses the M13mp18 single strand, and the staple strand combination is obtained according to the target template shape (two-dimensional rectangular structure). An exemplary method can be, optionally, marking the correspondence between the staple strand and the address number according to the shape information of the seven-segment addressing structure. Randomly generate and screen out seven orthogonal address sequences with a length of 20 bases each, and extend the staple strand sequences corresponding to addresses 1-7. The extended sequences are the designed address sequences, thereby obtaining the staple strand combination for assembling the addressable DNA platform. As Figure 1 , the information that can be stored in the 7 data strands corresponding to addresses 1-7 is the seven Chinese characters "When heaven and earth change, all things become connected".

[0102] (2) The switchable data strand adopts a partially complementary paired double-stranded structure. A single Chinese character is converted into binary coding and then encoded into a 10-base sequence using a cyclic coding algorithm. The complete data strand is composed of an address complementary sequence, a rewrite functional region, a T7 promoter sequence, and a data sequence (including a pair of primers) connected together ( Figure 2 ). The closed strand contains a 17-base complementary sequence of the T7 promoter and 5 downstream bases, as well as an 8-base hanging sequence corresponding to the address (referred to as toehold). The data strand and the closed strand form a double-stranded structure through 22-base partial complementary pairing.

[0103] Based on the above design, the structure of the nano-addressable platform is assembled, and the assembly process is as follows:

[0104] (1) The M13 single strand and the addressable staple strands are mixed in 1×TAE-Mg 2+ solution, with a final concentration of the scaffold strand of 10 nM and the staple strand of 50 nM.

[0105] (2) Annealing assembly is carried out in a PCR instrument. The annealing program is: keep at 95 degrees for 3 minutes, then cool down to 25 degrees at a rate of 1 degree per minute, and finally maintain at 4 degrees.

[0106] (3) The obtained assembly product is purified by PEG precipitation to obtain a pure assembly structure of about 20 nM as the substrate for information storage. Only for the visualization display effect, the obtained DNA structure can be characterized by an atomic force microscope (AFM).

[0107] (4) The data strand and the corresponding closed strand are in 1×TAE-Mg 2+Mix in solution to a final concentration of 10 μM. Anneal and hybridize in a PCR instrument. Characterize the hybridization products using 10% polyacrylamide gel electrophoresis (PAGE).

[0108] Example 2: Data Manipulation Method of a Nanoscale Addressable Fully Functional DNA Storage System

[0109] Data writing and reading:

[0110] (1) 100 nM of 1-7 data strands were added to 10 nM DNA origami and hybridized at room temperature for 1 hour to perform fully addressable information writing. At this time, the origami surface exhibited an "8" shape, composed of 7*7 (49 in total) 1-7 data strands. AFM was used to characterize the data writing success rate. Some combinations of 1-7 data strands were added, forming "0-9" shapes on the origami surface. For example, the number "1" consisted of two data strands, the number "2" consisted of five data strands, the number "3" consisted of five data strands, and so on. The robustness of the data writing platform was visualized and verified by AFM morphology characterization. The results showed ( Figure 4 The data chain at the target write location can be visualized and verified in the AFM characterization results.

[0111] (2) Add the activation strands of data 1-7 to the solutions of data 1-7 respectively, react at room temperature for 1 hour, then add the complete T7 promoter strand and react at room temperature for 1 hour. Add T7 transcription mixture and react in a 42℃ water bath for 1 hour. PAGE characterizes the feasibility of transcription after activation and the orthogonality of addressable readout. The results obtained by statistical analysis of PAGE data on the orthogonality of activation readout are as follows ( Figure 5 The data can only be exported if the sequence of the activation chain matches that of the data chain.

[0112] (3) The RNA molecules obtained by transcription were reverse transcribed using a reverse transcription kit to obtain the corresponding DNA strands. The obtained DNA strands were quantified using real-time PCR to confirm the data export.

[0113] (4) The exported data was sequenced using a next-generation sequencer, and the sequencing results were decoded. Next-generation sequencing of the sequences read from the seven addresses showed that the sequence information obtained at each address was completely consistent with the written information. Figure 6 ).

[0114] Addressable erasure and modification of data:

[0115] (1) After data chains 1-7 are fully written, eraser chains 5 and 6, which are complementary to data chains 5 and 6, are added. The mixture is incubated at room temperature for 1 hour to erase the corresponding data. Super-resolution fluorescence microscopy confirms that the target data chains 5 and 6 have been erased, resulting in the number "3". Further addition of eraser chains 1, 4, and 7, which are complementary to data chains 1, 4, and 7, is performed, and the reaction is characterized. The target data chains 1, 4, and 7 are then erased, resulting in the number "1". This process can be found in [reference needed]. Figure 7 In the diagram, the number "8" is transformed into the number "3" and then the number "1".

[0116] (2) Add data link 1, incubate at room temperature for 1 hour, and rewrite data at address 1 to obtain the result of adding data link 1, which is represented by the number "7". Super-resolution fluorescence microscopy imaging results show that the fluorescence imaging structure characterization results of each step in the addressable data erasure and writing process correspond to the target erasure and writing positions. This process can be found in [reference needed]. Figure 7 In the text, the process of the number "1" transforming into the number "7" is shown.

[0117] (3) Taking address 1 as an example to test the ability to repeatedly modify data, DNA origami was fixed on the surface of magnetic beads, and two types of data (two data strands) to be written to address 1 were labeled with Alexa488 and Cy5 fluorescence, respectively. The concentration of the residual data strand in the supernatant after writing was tested using a fluorescence spectrophotometer after adding the Alexa488-labeled data strand, and the erasure operation was performed to test the erasure efficiency. The writing efficiency was tested after writing the Cy5-labeled data strand, and the erasure efficiency was tested after erasure. The specific experiment was as follows: First, 647 data was written, at which point the writing efficiency of the first 488 data was basically 0. Then, 647 data was erased, at which point the erasure efficiency of the 488 data was basically 0. Next, 488 data was written, at which point the writing efficiency was about 54%. Then, 488 data was erased, at which point the erasure efficiency of the 488 data was about 110%. This cycle was repeated more than 5 times to determine the addressable rewrite capability of the system. The results showed that after repeated erasure and writing, the obtained fluorescence test results corresponded to the expected erasure and writing results, indicating that repeated data modification is feasible. Figure 8 ), Figure 8 The vertical axis represents the erasure efficiency of 488 data points above and the write efficiency of 488 data points below. The horizontal axis represents the number of erasures or writes.

[0118] Multiple data erase and write operations:

[0119] The 20 pM biotin-linked origami assembled in Example 1 was fixed onto a glass slide using the PEG-biotin-Avidin method. After cleaning, 1 μM of Cy5-labeled data chain 1 was added, and after incubation for 30 min, it was cleaned again, and fluorescence detection was performed under TIRF. Then, 1 μM of erase chain 1 was added, and after incubation for 1 h, it was cleaned with buffer and fluorescence imaging was performed again. The erase-write process was repeated 10 times, and the fluorescence changes were recorded by TIRF (total internal reflection fluorescence microscopy). Figure 9 It is a single-point TIRF imaging image of seven address data repeatedly erased and rewritten 10 times. Figure 9 The results show that the single fluorescent dot is stable after 10 erases and rewrites. Figure 10 It is a TIRF (Total Internal Reflection Fluorescence Microscopy) image and colocalization statistics of seven address data repeated 10 times. Red and green fluorescence colocalization represents the presence of data on the origami. Figure 10 The data shows that after 10 write / erase cycles, the co-location ratio stabilizes at around 85% after writing and around 10% after erasing.

[0120] Multiple data reads:

[0121] Assemble the data chain onto the origami, add the activation chain for activation, then attach the 2nM origami to the mica sheet, using 1×TAE Mg 2+ After washing, the transcription system was added to mica, and the mica was incubated in a 42°C oven for 2 hours. The liquid on the mica was then collected and stored. The above transcription process was repeated 11 times, and the RNA transcribed each time was collected. Finally, the RNA transcribed from the 11 times was quantified by reverse transcription-PCR. Figure 11 The data link is a repeatedly read qPCR quantification graph, indicating that the storage medium of this application can be repeatedly read at least 10 times.

[0122] Example 3: Reading and Detection of Origami Platforms with Different Matrix Spacing

[0123] DNA origami was designed with different matrix spacings, including 6nm, 12nm, 18nm, and 24nm. Storage information strands were arranged on the DNA origami with different spacings, and each strand was assigned a physical address with a different spacing. The writing and reading of data strands under different matrix spacings were detected according to the method in Example 2. The results are as follows... Figure 12 and 13 As shown, when the matrix spacing is 6nm, 12nm, 18nm, and 24nm, there is no significant difference in the write efficiency and read efficiency of the data link.

[0124] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Various variations of the embodiments listed herein will be apparent to those skilled in the art and are reserved within the scope of the appended claims and their equivalents.

Claims

1. A system comprising: 1) Nucleic acid molecules containing data information, 2) A carrier with addressable information, Furthermore, the data information contained in the nucleic acid molecule can be read in situ from the vector. in, The carrier comprises a DNA origami substrate, the DNA origami substrate comprises a staple chain, and the staple chain comprises an address sequence with addressable information. The nucleic acid molecule contains a complementary sequence, which is complementary to the address sequence on the vector. The nucleic acid molecule contains a data sequence, and the data information in the data sequence can be read while the nucleic acid molecule is not separated from the vector. The nucleic acid molecule contains a read-initiating sequence that can initiate the synthesis of the data sequence into a sequencing strand, and the sequencing strand is complementary to the data sequence. The system further includes 3) a blocking chain, which comprises a blocking sequence that is complementary to the read trigger sequence, thereby preventing the data information from being read. The blocking strand includes a blocking extension sequence, which is located upstream and / or downstream of the blocking sequence, and the blocking extension sequence is not complementary to the nucleic acid molecule. The system further includes 4) an activation strand, which prevents the blocking strand from binding to the nucleic acid molecule, and the activation strand is complementary to both the blocking sequence and the blocking extension sequence. The nucleic acid molecule contains an erasing / writing functional sequence, which is located upstream and / or downstream of the address complementary sequence, and the erasing / writing functional sequence is not complementary to the address sequence on the vector. The system further includes 5) an erasure strand, which enables the nucleic acid molecule to prevent binding to the vector, and the erasure strand is simultaneously complementary to the address complementary sequence and the erase / write functional sequence.

2. The system of claim 1, wherein the different physical locations of the carrier have different sequence of address sequences.

3. The system of claim 1, wherein in the carrier, the address sequence of the staple chain is arranged in a matrix.

4. The system of claim 1, wherein the spacing between two adjacent staple chains is 6-24 nm.

5. The system of claim 1, wherein the length of the address complementary sequence is 15 or more nucleotides.

6. The system of claim 1, wherein the length of the data sequence is more than one nucleotide.

7. The system of claim 1, wherein the read trigger sequence includes a promoter.

8. The system of claim 1, wherein the erasure strand is capable of binding to the nucleic acid molecule in place of the carrier.

9. The system of claim 1, wherein the molar ratio of the nucleic acid molecule to the address sequence in the vector is 2:1 or higher.

10. The nucleic acid molecule in the system according to any one of claims 1-9.

11. The carrier in the system according to any one of claims 1-9.

12. A method for data storage, the method comprising providing the system according to any one of claims 1-9.

13. A method for data editing, the method comprising replacing data information in a nucleic acid molecule in the system of any one of claims 1-9.

14. The method of claim 13, wherein the method comprises providing the erasure strand, wherein at room temperature, the nucleic acid molecule at a predetermined physical location binds to the erasure strand, and the nucleic acid molecule does not bind to the carrier.

15. A method for reading data, the method comprising determining data information in nucleic acid molecules in the system of any one of claims 1-9.

16. The method of claim 15, further comprising providing the activation strand, wherein at room temperature, the barrier strand to which a nucleic acid molecule at a specific physical location is bound binds to the activation strand, and the barrier strand does not bind to the nucleic acid molecule.

17. A storage medium comprising the method of any one of claims 12-16.

18. An apparatus comprising the storage medium of claim 17, and a processor coupled to the storage medium, the processor being configured to execute, based on a program stored in the storage medium, the method of any one of claims 12-16.

Citation Information

Patent Citations

  • Nucleic acid-based electrically readable, read-only memory

    CN112585152A