Compositions, systems and methods for nucleic acid data storage

JP2024530614A5Pending Publication Date: 2025-08-05NAIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505225
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-14
Filing Date
2022-07-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Current techniques for encoding data into nucleic acid molecules are inefficient, limited by yield, chain length, time, and cost, resulting in the production of relatively short strands that are not ideal for single molecule sequencing and require significant reagent consumption.

Method used

Development of polymers with convertible residues covalently linked to the backbone, allowing for conversion between different states using light, electrical voltage, enzymatic agents, or redox agents, enabling efficient and high-density data encoding and storage in nucleic acids.

Benefits of technology

The solution enables the creation of long nucleic acid strands that can be efficiently encoded and read, significantly increasing data storage density and reducing costs while maintaining stability over long durations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023009674000001
    Figure 2023009674000001
  • Figure 2023009674000002
    Figure 2023009674000002
Patent Text Reader

Abstract

Presented herein are writeable polymers (e.g., writeable nucleic acid polymers) and related methods for data storage. In general, writeable polymers (e.g., nucleic acid polymers) contain one or more convertible residues (e.g., convertible nucleic acid bases) that can be converted from a first state to a second state, and the first state and the second state are different. Various methods, such as polymerase extension by rolling circle reaction or chemical synthesis and ligation, can be used to generate writeable nucleic acid polymers. Also presented herein are various methods for writing or encoding into writeable nucleic acid polymers by selectively converting nucleic acid bases to a second state. Also presented herein are various methods for reading or decoding data from encoded nucleic acid polymers.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 226,720, filed July 28, 2021, and U.S. Provisional Patent Application No. 63 / 269,324, filed March 14, 2022, the contents of each of which are incorporated by reference in their entirety.

[0002] Technical Field The present disclosure is generally directed to compositions, systems, and methods for storing data in nucleic acid molecules. [Background technology]

[0003] background As the amount of digital data increases, the vexing problem of preserving digital data for the long term is becoming a rapidly growing problem. Digital data archived electronically or magnetically can be easily manipulated, distorted, and / or lost during storage. Efficient solid-state electronic methods for archival data storage exist, but are not stable for many years, and data must be periodically rewritten or migrated to new devices before it is lost. Similarly, magnetic tape is commonly used for data archiving, but it also degrades over time. Thus, methods for efficiently encoding and preserving data, especially for long periods of time, are being very actively sought.

[0004] Nucleic acid molecules, particularly DNA, offer a potential solution to overcome the problems associated with data storage. Nucleic acid polymers are biochemical molecules that have sequences of repeating bases and are essentially digital information that can be stably stored at high densities and for very long durations. Natural DNA contains digital information encoded in the four bases: A, C, T, and G, and can be used to encode binary data in the sequence of the synthesized strands. A single polymer of DNA can be very long (e.g., for a chromosome), with millions of bits of data encoded. One cubic inch of DNA can contain 10 18 It has been estimated that DNA can encode up to 10 ...

[0005] Furthermore, to facilitate access to data stored in nucleic acid molecules, the stored data can be read quickly and inexpensively by high-throughput sequencing techniques. Advances in sequencing technology have significantly reduced costs and increased the speed of sequencing, making it possible to efficiently read data in DNA. New long-read single-molecule techniques allow the bases of a single DNA molecule that is tens of thousands of bases long to be read quickly. New nanopore technologies allow the sequence of a single molecule of DNA to be read in seconds to minutes (see N Kono and K. Arakawa, Dev Growth Differ. 2019; 61: 316-326; and Q Chen and Z. Liu, Sensors (Basel). 2019; 19: 1886, the disclosures of each of which are incorporated herein by reference), allowing the sequence of a strand that is tens of thousands of base pairs long or longer to be read.

[0006] Although nucleic acids are an important potential source of data storage, the process of synthesizing nucleic acids, particularly sequences that define data, is inefficient, and thus the process of encoding into nucleic acids is a substantial barrier to using nucleic acids as data storage. Current methods for storing data in DNA involve chemically or enzymatically synthesizing strands of arbitrary sequences that are encoded with digital information (see GM Church, Y. Gao, and S. Kosuri Science. 2012; 337: 1628; X. Chengtao, et al., Nucleic Acids Res. 2021; 49: 5451-5469; and E. Yoo, et al., Comput Struct Biotechnol J. 2021; 19: 2468-2476, the disclosures of each of which are incorporated herein by reference). Oligonucleotide synthesizers can create DNA lengths of approximately 100-200 nucleotides. Specialized synthesizers can generate hundreds or thousands of oligonucleotides in one go, allowing for higher throughput data writing. In addition to chemical DNA synthesis, enzymatic methods involving polymerases or other enzymes are also being explored for creating DNA with arbitrary data encoded in its sequence. These enzymatic methods involve the addition of specialized nucleotides one at a time, or the stepwise addition of short segments of DNA. Methods for encoding data into DNA during synthesis are limited by yield, length of strand, time, and cost. Current efficient DNA synthesizers produce strands of up to approximately 200 nucleotides, thus encoding relatively small amounts of information. To compensate for the shortness of sequences, many different oligonucleotides must be synthesized. Oligonucleotide synthesis requires excess reagents to achieve stepwise high yields, necessitating costly consumption of reagents and solvents. Oligonucleotide synthesis also requires time to achieve these high yields for each nucleotide addition (typically 1-5 minutes for each step), which means that longer times are required to encode larger amounts of data. Common enzymatic methods under development also add nucleotides or groups of nucleotides in a stepwise fashion, and have yet to make significant improvements in their ability to make very long strands and encode large amounts of data. Because enzymatic synthesis methods are also stepwise, they are similarly limited in the speed of data encoding. Furthermore, both the chemical and enzymatic strategies described above generally produce relatively short strands that may not be ideal for single molecule sequencing and may instead rely on sequencing methods that require larger amounts of individual written DNA. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] N Kono and K. Arakawa, Dev Growth Differ. 2019; 61: 316-326 [Non-Patent Document 2] Q Chen and Z. Liu, Sensors (Basel). 2019; 19: 1886: 1628 [Non-Patent Document 3] GM Church, Y. Gao, and S. Kosuri Science. 2012; 337 [Non-Patent Document 4] X. Chengtao, et al., Nucleic Acids Res. 2021; 49: 5451-5469 [Non-Patent Document 5] E. Yoo, et al., Comput Struct Biotechnol J. 2021; 19: 2468-2476 Summary of the Invention [Means for solving the problem]

[0008] Summary of the Disclosure In one aspect, there is provided a polymer for encoding data, comprising: comprising a plurality of convertible residues covalently linked to the backbone of the polymer at repetitive intervals along the backbone of the polymer; each of the plurality of convertible residues has a first state and is convertible from the first state to a second state, the first state and the second state being distinct, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state being readable by a polymerase enzyme; Provided herein are polymers having a plurality of transformable residues covalently linked to the polymer in a first state and a second state.

[0009] In certain embodiments, the polymer is a nucleic acid polymer and the plurality of convertible residues are convertible nucleobases.

[0010] In certain embodiments, the nucleic acid polymer is a single-stranded nucleic acid polymer.

[0011] In certain embodiments, the nucleic acid polymer is a double-stranded nucleic acid polymer.

[0012] In certain embodiments, the nucleic acid polymer comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), phosphorothioate DNA, glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or a combination thereof.

[0013] In certain embodiments, the nucleic acid polymer comprises more than 10 alterable residues.

[0014] In certain embodiments, the ratio of the total number of nucleotides to convertible residues in a nucleic acid polymer is between 2 and 100.

[0015] In certain embodiments, the plurality of convertible nucleobases are non-naturally occurring nucleobases.

[0016] In certain embodiments, the plurality of convertible nucleobases are modified naturally occurring nucleobases or derivatives of naturally occurring nucleobases.

[0017] In certain embodiments, each of the plurality of convertible nucleobases comprises a chemically modifiable moiety.

[0018] In certain embodiments, the chemically modifiable moiety of each of the plurality of convertible nucleobases is directly attached to the base of the convertible nucleobase.

[0019] In certain embodiments, the chemically modifiable moiety of each of the plurality of convertible nucleobases is attached to the base without a linker or side chain.

[0020] In certain embodiments, the plurality of convertible nucleobases are covalently linked to the backbone of the nucleic acid via the sugar.

[0021] In certain embodiments, the chemically modifiable moiety is activatable by light, voltage, an enzymatic agent, a chemical reagent, or a redox agent, thereby converting the first state to the second state.

[0022] In certain embodiments, the chemically modifiable moiety is activatable by light, thereby converting the first state to the second state.

[0023] In certain embodiments, the conversion from the first state to the second state occurs via an irreversible reaction.

[0024] In certain embodiments, the convertible nucleobase becomes a naturally occurring nucleobase after conversion to the second state.

[0025] In certain embodiments, the convertible nucleobase becomes guanine, adenine, thymine, uracil or cytosine after conversion to the second state.

[0026] In certain embodiments, the backbone of the polymer (eg, the phosphates and sugars of a nucleic acid polymer) remains unchanged during conversion from a first state to a second state.

[0027] In certain embodiments, the polymer comprises two or more distinct sets of convertible residues, each set of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being distinct.

[0028] In certain embodiments, each of the plurality of transformable residues comprises a chemically modifiable moiety that can be activated by light.

[0029] In certain embodiments, two or more distinct sets of convertible residues are activatable by light of different wavelengths.

[0030] In certain embodiments, a first set of convertible residues is activatable by a first wavelength of light and a second set of convertible residues is activatable by a second wavelength of light, wherein the first wavelength and the second wavelength are different.

[0031] In certain embodiments, the chemically modifiable moiety comprises one or more photoremovable groups.

[0032] In certain embodiments, the chemically modifiable moiety is a leaving group.

[0033] In certain embodiments, the one or more photoremovable groups are [ka] wherein X represents NR2, NHR, OR, or SR, and R is a nucleobase having a photoremovable group attached thereto. It is.

[0034] In certain embodiments, the plurality of convertible nucleobases are convertible by light of wavelengths of 325 nm, 360 nm, or 400 nm.

[0035] In certain embodiments, the plurality of convertible nucleobases are convertible by light of a wavelength between 400 nm and 850 nm.

[0036] In certain embodiments, each of the plurality of convertible nucleobases comprises a chemically modifiable moiety that is activatable by oxidation-reduction.

[0037] In certain embodiments, the chemically modifiable moiety is capable of being activated by localized oxidation.

[0038] In certain embodiments, the chemically modifiable moiety can be activated by oxidation using an electrode.

[0039] In certain embodiments, the nucleotide comprising the convertible nucleobase is [ka] is selected from the group consisting of:

[0040] In certain embodiments, the convertible nucleobase is selected from the group consisting of O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0041] In certain embodiments, the first state and the second state of the plurality of convertible nucleobases are readable by a sequencing method capable of detecting and distinguishing non-naturally occurring and / or modified nucleobases.

[0042] In certain embodiments, the first state and the second state of the plurality of convertible nucleobases are readable by nanopore sequencing.

[0043] In certain embodiments, the first state and the second state of the plurality of convertible nucleobases are readable by sequencing by synthesis.

[0044] In certain embodiments, the plurality of convertible nucleobases, when converted to the second state, have altered properties of the plurality of convertible nucleobases compared to the first state (e.g., reduced size, altered shape, altered H-bonding, and / or altered polymerase substrate ability).

[0045] In certain embodiments, one or more of the plurality of convertible nucleobases are capable of converting from a second state to a third state, and one or more of the plurality of convertible nucleobases are covalently attached to the nucleic acid polymer in the third state.

[0046] In certain embodiments, each of a plurality of convertible residues can be independently and selectively converted.

[0047] In certain embodiments, the polymers presented herein further comprise a plurality of spacer residues linked via the backbone of the polymer, where each of the plurality of convertible residues is separated by one or more spacer residues of the plurality of spacer residues.

[0048] In certain embodiments, the repeating spacing between the multiple convertible residues is compatible with the resolution of a writing mechanism for encoding data onto the polymer.

[0049] In certain embodiments, the repeat spacing between two adjacent convertible residues is equal to or greater than the resolution of the data encoding mechanism for encoding the data in the polymer.

[0050] In one particular embodiment, the resolution of the writing mechanism is at least 1 nm.

[0051] In certain embodiments, the spacer residues do not interfere with the reading of the convertible residues.

[0052] In certain embodiments, multiple spacer residues in a polymer are the same spacer residue.

[0053] In certain embodiments, the plurality of spacer residues comprises two or more different spacer residues (eg, different nucleobases, eg, different naturally occurring nucleobases).

[0054] In certain embodiments, the polymer consists essentially of spacer moieties.

[0055] In certain embodiments, each of the plurality of convertible nucleobases is separated by 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 spacer residues.

[0056] In certain embodiments, each of the plurality of convertible nucleobases is separated by six spacer residues.

[0057] In certain embodiments, the plurality of spacer residues are naturally occurring nucleobases, non-naturally occurring nucleobases, tetrahydrofuran abasic residues, or ethylene glycol residues.

[0058] In certain embodiments, the spacer residues are naturally occurring nucleobases.

[0059] In certain embodiments, the polymers presented herein further comprise one or more delimiters linked to the backbone of the polymer.

[0060] In certain embodiments, each of the one or more delimiters comprises one or more naturally occurring or non-naturally occurring nucleobases.

[0061] In certain embodiments, one or more of the delimiters comprises a naturally occurring nucleobase.

[0062] In certain embodiments, one or more delimiters separate two or more adjacent data fields within a polymer.

[0063] In certain embodiments, the polymers presented herein further comprise one or more data tags.

[0064] In certain embodiments, one or more data tags comprise one or more naturally occurring or non-naturally occurring nucleobases.

[0065] In certain embodiments, the polymer is a nucleic acid polymer and the one or more data tags are present at the 5' or 3' end of the nucleic acid polymer.

[0066] In certain embodiments, the one or more data tags are incorporated into the nucleic acid polymer by ligation during synthesis of the nucleic acid polymer, during conversion of the plurality of convertible nucleobases to the second state, or after conversion of the plurality of convertible nucleobases to the second state.

[0067] In certain embodiments, the polymer can be stored under standard nucleic acid storage protocols.

[0068] In certain embodiments, the polymer is a nucleic acid polymer that can be stored at room temperature or at low temperature (e.g., −20° C.) in an appropriate nuclease-free solution.

[0069] In certain embodiments, the polymer can be stored at room temperature without stabilizers.

[0070] In another aspect, a system for writing data is provided, comprising: a writeable polymer comprising a plurality of convertible residues covalently linked to the backbone of the polymer at repetitive intervals along the backbone of the polymer, each of the plurality of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being distinct, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state being readable by a polymerase enzyme, and the plurality of convertible residues are covalently linked to the polymer in the first state and in the second state; A data writing device for writing data onto a writable polymer. Also presented herein is a system that includes:

[0071] In certain embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleobases.

[0072] In one particular embodiment, the data writing device comprises a nanopore.

[0073] In one particular embodiment, the data writing device includes a microscope with a light source.

[0074] In certain embodiments, the data writing device converts the plurality of convertible nucleobases to the second state by a light pulse, a voltage pulse, an enzymatic agent, or a redox agent.

[0075] In certain embodiments, the data writing device converts a plurality of convertible nucleobases to the second state by a light pulse.

[0076] In one particular embodiment, the data writing device includes a light illuminating device.

[0077] In another aspect, there is provided a method for generating a writeable nucleic acid polymer, comprising: providing a circular single stranded oligonucleotide template complementary to the repeats of a data field comprising convertible nucleobases; incubating the circular single-stranded oligonucleotide template in the presence of a nucleic acid primer, a polymerase, and a nucleotide triphosphate, the nucleotide triphosphate comprising a convertible nucleobase in a first state and capable of being converted from the first state to a second state, the first state and the second state being different; Also presented herein are methods, including:

[0078] In certain embodiments, the circular single-stranded oligonucleotide template comprises nucleobases complementary to the convertible nucleobases, with the complementary nucleobases repeatedly spaced apart, such that incubation of the template with a nucleic acid primer, a polymerase, and nucleotide triphosphates results in a nucleic acid polymer comprising a plurality of convertible nucleobases covalently linked through the backbone of the nucleic acid polymer and repeatedly spaced apart along the backbone of the nucleic acid polymer, the plurality of convertible nucleobases being covalently linked to the nucleic acid polymer in the first state and in the second state.

[0079] In certain embodiments, the repeats of the data field further comprise a spacer nucleobase and the nucleotide triphosphate further comprises a spacer nucleotide triphosphate.

[0080] In yet another aspect, there is provided a method for generating a writeable nucleic acid polymer, comprising: chemically synthesizing a plurality of oligomers, each oligomer comprising a plurality of convertible nucleobases linked via a nucleic acid polymer backbone at repetitively spaced intervals along the nucleic acid polymer backbone, each of the plurality of convertible nucleobases having a first state and capable of being converted from the first state to a second state, the plurality of convertible nucleobases being covalently attached to the nucleic acid polymer in the first state and in the second state, the first state and the second state being distinct; ligating a plurality of oligomers to form a writeable nucleic acid polymer; A method is presented herein, comprising:

[0081] In certain embodiments, each of the plurality of oligomers comprises a plurality of spacer residues linked via the backbone of the nucleic acid polymer, and each of the plurality of convertible nucleobases is separated by one or more spacer residues of the plurality of spacer residues.

[0082] In certain embodiments, the ligating step is by chemical ligation.

[0083] In certain embodiments, the ligating step is by enzymatic ligation.

[0084] In certain embodiments, the ligating step uses a complementary DNA splint.

[0085] In certain embodiments, the method further comprises annealing the plurality of complements to the oligomer prior to the ligating step.

[0086] In yet another aspect, a method for writing data onto a writeable polymer comprises the steps of: providing a writeable polymer including a plurality of convertible residues covalently linked through the backbone of the polymer at repetitive intervals along the backbone of the polymer, each of the convertible residues of the plurality of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being distinct, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state being readable by a polymerase enzyme; utilizing a data writing device to selectively convert one or more of the plurality of convertible residues to a second state, thereby generating a data encoded polymer; A method is presented herein, comprising:

[0087] In certain embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleobases.

[0088] In certain embodiments, the data writing device includes a nanopore, and the method further includes passing the writeable polymer through a nanopore of the writing device, where the nanopore converts one or more of the plurality of convertible residues to a second state.

[0089] In certain embodiments, the nanopore is a plasmonic nanopore that provides a pulse of light or redox energy to selectively convert a convertible nucleobase from a first state to a second state.

[0090] In certain embodiments, the data writing device includes a plasmonic well or channel, and the method further includes transferring the writeable polymer to the plasmonic well or channel of the data encoding device, and providing a light pulse or redox energy by the plasmonic well or channel to selectively convert the convertible nucleobases from a first state to a second state.

[0091] In certain embodiments, the data writing device selectively converts the convertible moiety to the second state by a light pulse, a voltage pulse, an enzymatic agent, or a redox agent.

[0092] In certain embodiments, the data writing device selectively converts the convertible moieties to the second state by a light pulse.

[0093] In certain embodiments, the convertible residue becomes a naturally occurring nucleobase after conversion to the second state.

[0094] In certain embodiments, the plurality of convertible residues comprises two or more types of convertible residues, where a first type of convertible residue is activatable by a first wavelength of light and a second type of convertible residue is activatable by a second wavelength of light.

[0095] In certain embodiments, the repeating spacing between the multiple convertible residues is compatible with the resolution of a data writing device for selectively converting the convertible residues.

[0096] In certain embodiments, the selectively converting step does not require specific positioning of the writeable polymer.

[0097] In certain embodiments, the conversion of the convertible residues to the second state is not uniform across the data encoded polymer.

[0098] In certain embodiments, conversion of a convertible residue to the second state is not limited to a particular location on the data encoded polymer.

[0099] In certain embodiments, the method further comprises the step of stretching or combing the writeable polymer (eg, writeable DNA) onto the solid support.

[0100] In certain embodiments, the method further comprises visualizing the location of the convertible residues using a dye.

[0101] In certain embodiments, the method further comprises the step of locally illuminating or locally exciting the writeable polymer.

[0102] In certain embodiments, the locally illuminating or locally exciting step uses a stimulated emission depletion (STED) laser.

[0103] In certain embodiments, the method further includes the step of bonding two or more data fields from two or more writable polymers end-to-end, thereby resulting in a bonded polymer that includes two or more data fields.

[0104] In certain embodiments, the method further comprises controlling the rate of passage of the writeable polymer through the nanopore of the writing device.

[0105] In certain embodiments, multiple writable polymers are passed through a data writing device to write the same data (eg, creating data redundancy).

[0106] In yet another aspect, there is provided a method for reading data from a data encoded polymer, comprising the steps of: providing a data-encoded polymer comprising convertible residues covalently linked via the polymer backbone at repetitive intervals along the polymer backbone, a first subset of the convertible residues being in a first state and a second subset of the convertible residues being in a second state, the first state and the second state being distinct, a plurality of the convertible residues in the first state and a plurality of the convertible residues in the second state being readable by a polymerase enzyme; passing the writable data encoded polymer through a data reading device to read the encoded data on the data encoded polymer; Also presented herein are methods, including:

[0107] In certain embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleobases.

[0108] In certain embodiments, a convertible residue in a first state can be converted to a second state by light.

[0109] In certain embodiments, the data reading device comprises a nanopore.

[0110] In certain embodiments, the data reading device is a sequencing device.

[0111] In certain embodiments, the sequencing device is a sequencing-by-synthesis device.

[0112] In certain embodiments, the method further comprises measuring the current flow in the electrolyte while the writeable polymer passes through.

[0113] In certain embodiments, the method further includes determining whether each of the plurality of convertible residues is in a first state or a second state based on a measured current flow in the electrolyte during passage of the writeable polymer.

[0114] In certain embodiments, the method further includes passing the data encoded polymer again through the data reading device to again read the encoded data on the data encoded polymer.

[0115] In certain embodiments, the method further includes verifying and correcting the encoded data on the data-encoded polymer by comparing the encoded data on multiple copies of the data-encoded polymer.

[0116] In yet another aspect, there is provided a method for reading or decoding data from a data encoded nucleic acid polymer, comprising: a plurality of converted nucleobases, each converted nucleobase comprising a first nucleobase structure, the first converted nucleobase being converted from a first state to a second state, the first state and the second state being different; a plurality of convertible nucleobases, each convertible nucleobase comprising a second nucleobase structure and a directly linked leaving group, the convertible nucleobase being provided in a first state and capable of being converted from the first state to a second state by releasing the second leaving group from the second nucleobase structure, the first state and the second state being different; providing a plurality of overlapping copies of a data-encoded nucleic acid polymer, comprising: the converted nucleobase and the convertible nucleobase are linked via a nucleic acid polymer backbone; determining the sequence of each of the multiple overlapping copies of the nucleic acid polymer; Also presented herein are methods, including: In certain embodiments, the method further includes detecting the plurality of converted nucleobases and the plurality of convertible nucleobases, and decoding the data based on the detected plurality of converted nucleobases.

[0117] In certain embodiments, the plurality of converted nucleobases in the first state and the plurality of converted nucleobases in the second state are readable by a polymerase enzyme.

[0118] In certain embodiments, the plurality of convertible nucleobases in the first state and the plurality of convertible nucleobases in the second state are readable by a polymerase enzyme.

[0119] In certain embodiments, the plurality of converted nucleobases and the plurality of convertible nucleobases are detected based on sequencing results of overlapping copies of the data-encoded nucleic acid polymer.

[0120] The description and claims may be more fully understood with reference to the following drawings and data graphs, which are presented as exemplary embodiments and should not be construed as a complete description of the scope of the disclosure. [Brief description of the drawings]

[0121] [Figure 1] 1A and 1B present schematic diagrams of a writeable nucleic acid polymer according to various embodiments.

[0122] [Diagram 2] 2A and 2B present schematic diagrams of data-encodeable nucleic acid polymers according to various embodiments.

[0123] [Figure 3-1] 3A-3G show structures of various examples of convertible nucleobases for use in writeable nucleic acid polymers. [Figure 3-2] Same as above. [Figure 3-3] Same as above.

[0124] [Figure 4] FIG. 4 provides an example of a convertible nucleobase O6-nitrobenzyl-guanine according to various embodiments.

[0125] [Figure 5-1]5A and 5B show structures of various examples of nucleotides containing convertible nucleobases for use in writeable polymers, according to various embodiments. [Figure 5-2] Same as above.

[0126] [Figure 6] FIG. 6 presents molecular structure diagrams of various removable groups (eg, leaving groups) on convertible nucleobases for use in writeable polymers, according to various embodiments.

[0127] [Figure 7] FIG. 7 presents a schematic diagram of the generation of writeable nucleic acid polymers utilizing polymerase extension via a rolling circle reaction, according to various embodiments.

[0128] [Figure 8] FIG. 8 presents a schematic of the generation of writable nucleic acid polymers using chemical synthesis and ligation, according to various embodiments.

[0129] [Figure 9] 9A-9C present schematic diagrams for encoding data in a writeable nucleic acid polymer using nanopores and optical energy, according to various embodiments.

[0130] [Figure 10] 10A-10C present schematic diagrams for encoding data in a data-encodeable nucleic acid polymer comprising transducible nucleobase pairs using a nanopore and optical energy, according to various embodiments.

[0131] [Figure 11]11A-11C illustrate the encoding of data in a writeable nucleic acid polymer containing convertible nucleobases using a nanopore and light energy according to various embodiments. Figure 11A: A writeable nucleic acid polymer containing convertible nucleobases Ca and Cb; Figure 11B: A writeable nucleic acid polymer passes through a nanopore and certain convertible nucleobases (e.g., Ca at the 3' end) are converted by light energy to the written converted nucleobases (e.g., Ca'); and Figure 11C: Certain convertible nucleobases Ca and Cb are selectively converted to converted nucleobases Ca' and Cb', respectively, resulting in a data-encoded nucleic acid polymer containing stochastically or irregularly spaced converted nucleobases Ca' and Cb'.

[0132] [Figure 12] 12A-12C present schematic diagrams for encoding data in a writeable nucleic acid polymer containing duads using nanopores and optical energy, according to various embodiments.

[0133] [Figure 13] 13A-13C present molecular structure diagrams of dual-bit convertible nucleobases for use in writeable nucleic acid polymers, according to various embodiments.

[0134] [Figure 14] 14A and 14B present data decoding strategies using nanopore current-based sequencing (FIG. 14A) and sequencing-by-synthesis (FIG. 14B) according to various embodiments.

[0135] [Figure 15] FIG. 15 illustrates a particular embodiment of a data-encodeable nucleic acid polymer that includes convertible nucleobases. [ka] to T, and [ka] 1 illustrates an example of encoding with binary data 1010010 by selectively converting T and G, respectively. Certain convertible nucleobases within a data-encodeable nucleic acid polymer are skipped during the data encoding process, and the resulting data-encoded nucleic acid polymer includes stochastically and / or irregularly spaced converted nucleobases (e.g., T and G). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0136] Detailed Description Provided herein are compositions of data-encodeable polymers (e.g., nucleic acid polymers) and methods and systems thereof for data encoding / decoding (writing / reading) and data storage. Also provided herein are methods of making the polymers (e.g., nucleic acid polymers) described herein.

[0137] Referring now to the figures and data, compositions and systems for nucleic acid data storage, methods of use and synthesis are disclosed according to various embodiments. In some embodiments, the data storage system comprises a writeable (i.e., data-encodeable) nucleic acid polymer having one or more convertible nucleobases. Thus, the writeable nucleic acid polymer is analogous to an encodable blank tape, and encoding in the writeable nucleic acid polymer is achieved by converting one or more nucleobases of the writeable nucleic acid polymer. The conversion of the nucleobases can be considered as a binary code, with each convertible nucleobase being analogous to a "bit", with an unconverted nucleobase being analogous to a "0" and a converted nucleobase being analogous to a "1". However, it should be understood that not only binary codes are possible, but codes can also be written in ternary, quaternary or other numeric code schemes, which can be done by utilizing multiple types of convertible bases or by performing multiple writes to further change the state of the convertible bases. In some embodiments, the conversion of the convertible nucleobases is stable or permanent, allowing for long-term archiving. In some embodiments, the combination of two convertible nucleotides comprises a "bit".

[0138] In some embodiments, a convertible residue (e.g., a convertible nucleobase) is referred to as a writeable "bit" and a converted residue (e.g., a converted nucleobase, e.g., a native nucleobase) is referred to as a written "bit."

[0139] In some embodiments, the terms "writable" and "data encoding" are used interchangeably herein. In some embodiments, the terms "write" and "data encoding" are used interchangeably herein.

[0140] In some embodiments, the terms "leaving group" and "removable group" are used interchangeably herein. In some embodiments, when referring to convertible nucleobases, the terms "pair" and "duad" are used interchangeably herein. "Pair," as used herein, refers to a pair of different convertible nucleobases (e.g., writable bits) that are located close enough to each other in a polymer (e.g., a nucleic acid polymer) described herein so that both are exposed to a single writing action or event (e.g., the same pulse of light or the same voltage pulse). Thus, the convertible nucleotides that comprise the pair are closer than the resolution of the writing action or event.

[0141] In other embodiments of the systems provided herein, the systems include two or more sets of convertible nucleobases (e.g., nucleobases having different structures, e.g., those having different chemically modifiable moieties), where conversion of the nucleobases (e.g., removal of a cage group of the nucleobases) can be thought of as a binary code, with each convertible nucleobase (or two or more convertible base sets) analogous to a writeable "bit" of data, and each converted nucleobase (or two more converted nucleobase sets) analogous to a written "bit" of data. In some embodiments, the convertible nucleobases are utilized to encode data bits, where conversion of a first nucleobase structure (i.e., the first set of convertible nucleobases) is analogous to a "0" and conversion of a paired second nucleobase structure (i.e., the second set of convertible nucleobases) is analogous to a "1," and data can be encoded by selective conversion of nucleobases along a polymer (e.g., a nucleic acid polymer). In some embodiments, pairs of convertible nucleobases are utilized to encode data into writable bits, where conversion of one nucleobase of the pair is similar to a "0" and conversion of both nucleobases of the pair is similar to a "1", and data can be encoded by nucleobase pair conversion along the polymer. However, it should be understood that not only binary codes are possible, but codes can also be written in ternary, quaternary, or other numeric code schemes, and this can be done by utilizing multiple types of convertible bases or by performing multiple writes to further change the state of the convertible bases. In some embodiments, the conversion of the convertible nucleobases is stable or permanent for long periods of time, allowing for long-term archiving.

[0142] In some embodiments, the nucleic acid polymer is a single-stranded nucleic acid polymer or a double-stranded nucleic acid polymer. In some embodiments, the nucleic acid polymer is a single-stranded nucleic acid polymer. In some embodiments, the nucleic acid polymer is a double-stranded nucleic acid polymer.

[0143] Some embodiments are directed to compositions of writeable nucleic acid polymers. Any suitable nucleic acid polymer may be utilized, including, but not limited to, DNA, RNA, phosphorothioate DNA, glycerol nucleic acid (GNA), threose nucleic acid (TNA). Furthermore, the nucleic acid polymer may be single-stranded or double-stranded. In some embodiments, the writeable nucleic acid polymer comprises a plurality of convertible nucleic acid bases linked by a polymer backbone. In certain embodiments, the convertible nucleic acid bases are spaced apart to provide spatial resolution such that each of the nucleic acid bases can be independently and selectively converted according to encoding. In some embodiments, spacer residues linked through the polymer backbone are utilized to provide spacing between the convertible nucleic acid bases. In some embodiments, the spacer residues are non-reactive to the writing mechanism. In various embodiments, the writeable nucleic acid polymer may further comprise delimiters and / or data tags for labeling data, each of which may be provided by a specific sequence of nucleic acid bases.

[0144] In some embodiments, any suitable nucleic acid polymer may be utilized, including, but not limited to, DNA, RNA, phosphorothioate DNA, glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), and combinations thereof.

[0145] In some embodiments, multiple convertible nucleotides are capable of being incorporated into a nucleic acid polymer by one or more polymerase enzymes.

[0146] In some embodiments, the plurality of convertible nucleobases are non-naturally occurring nucleobases. In some embodiments, the plurality of convertible nucleobases are modified naturally occurring nucleobases or derivatives of naturally occurring nucleobases.

[0147] In some embodiments, each of the plurality of convertible nucleobases comprises a chemically modifiable moiety. In some embodiments, the chemically modifiable moiety of each of the plurality of convertible nucleobases is directly attached to the base of the convertible nucleobase. In some embodiments, the chemically modifiable moiety of each of the plurality of convertible nucleobases is attached to the base without a linker or side chain. In some embodiments, the plurality of convertible nucleobases is covalently linked to the backbone of the nucleic acid via the sugar of the backbone of the nucleic acid. In some embodiments, the removable group of the plurality of convertible nucleobases is covalently linked to the backbone of the nucleic acid via the nucleobase.

[0148] In some embodiments, the convertible nucleobase is linked to the backbone of a nucleic acid polymer in the same manner that the nucleobase of a native nucleotide is linked to the backbone of a nucleic acid polymer (via the sugar of the nucleotide), either without an intervening linker or as a side chain.

[0149] In some embodiments, conversion of the nucleobase (i.e., from a first state to a second state) is achieved by removing one or more removing groups from the nucleobase. In some embodiments, the removable groups are caging groups.

[0150] In one embodiment, the chemically modifiable moiety is activatable by light, thereby converting the first state to the second state. In some embodiments, the conversion from the first state to the second state occurs by an irreversible reaction. In some embodiments, the convertible nucleobase becomes a naturally occurring nucleobase after conversion to the second state. In some embodiments, the convertible nucleobase becomes a native nucleobase after conversion to the second state. In one embodiment, the convertible nucleobase becomes guanine, adenine, thymine, uracil, or cytosine after conversion to the second state. In some embodiments, the backbone of the polymer (e.g., the phosphate and sugar of the nucleic acid polymer) remains unchanged during the conversion from the first state to the second state. In some embodiments, the chemically modifiable moiety is activatable by light, voltage, an enzymatic agent, a chemical reagent, or a redox agent or electrode, thereby converting the first state to the second state. In some embodiments, the chemically modifiable moiety comprises one or more photoremovable groups.

[0151] In some embodiments, the one or more photoremovable groups are [ka] wherein X represents NR2, NHR, OR, or SR, and R is a nucleobase having a photoremovable group attached thereto. It is.

[0152] In some embodiments, the plurality of convertible nucleobases are convertible by light of wavelengths of 325 nm, 360 nm, or 400 nm.

[0153] In some embodiments, the plurality of convertible nucleobases are convertible by light of a wavelength between 400 nm and 850 nm.

[0154] In some embodiments, each of the plurality of convertible nucleobases comprises a chemically modifiable moiety that is activatable or removable by oxidation-reduction. In some embodiments, the chemically modifiable moiety is capable of being activated by localized oxidation. In some embodiments, the chemically modifiable moiety is capable of being activated by oxidation or reduction using one or more electrodes.

[0155] In some embodiments, the nucleotide comprising the convertible nucleobase is [ka] is selected from the group consisting of:

[0156] In some embodiments, the convertible nucleobase (having a specific substitution position of the removable group) is selected from the group consisting of O6-guanine, O6-thioguanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, O4-uracil, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0157] In some embodiments, the first and second states of the plurality of convertible nucleobases are readable by a sequencing method capable of detecting and distinguishing between non-naturally occurring and / or modified nucleobases. In some embodiments, the first and second states of the plurality of convertible nucleobases are readable by nanopore sequencing. In some embodiments, the first and second states of the plurality of convertible nucleobases are readable by sequencing by synthesis. In some embodiments, the plurality of convertible nucleobases, when converted to the second state, have altered properties of the plurality of convertible nucleobases compared to the first state (e.g., reduced size, altered shape, altered H-bonding, and / or altered polymerase substrate capability and / or polymerase coding). In some embodiments, one or more of the plurality of convertible nucleobases are capable of converting from the second state to a third state, and one or more of the plurality of convertible nucleobases are covalently attached to the nucleic acid polymer in the third state. In some embodiments, each of the plurality of convertible residues is capable of being independently and selectively converted.

[0158] In some embodiments, a polymer (e.g., a nucleic acid polymer) described herein comprises two or more distinct sets of convertible residues, each set of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being different. In some embodiments, each of the plurality of convertible residues comprises a chemically modifiable moiety that can be activated and / or removed by light, the two or more distinct sets of convertible residues being activatable and / or removable by light of different wavelengths. In some embodiments, the first set of convertible residues is activatable by a first wavelength of light and the second set of convertible residues is activatable by a second wavelength of light, the first wavelength and the second wavelength being different.

[0159] In certain embodiments, the convertible nucleobases (or pairs of convertible bases) in the writeable nucleic acid polymer described herein are spaced repeatedly to provide spatial resolution such that each nucleobase (or each set or pair) can be independently and selectively converted according to encoding. In certain embodiments, the convertible nucleobases are regularly or irregularly spaced, but certain nucleobases are identified and selectively converted to encode data, resulting in a data-encoded nucleic acid polymer. In some embodiments, the data encoding mechanism may skip any convertible nucleobases as needed until the correct convertible nucleobase is reached according to the code.

[0160] In some preferred embodiments, convertible nucleobases are regularly spaced (e.g., by spacers), but certain nucleobases are identified and selectively converted to encode data, resulting in a data-encoded nucleic acid polymer that includes stochastically spaced converted nucleobases (i.e., written bits).One advantage of the writeable nucleic acid polymer presented herein is that it is not necessary to control the position or passage speed of the writeable nucleic acid polymer.Certain convertible nucleobases can be skipped.

[0161] In some embodiments, a writing procedure is utilized to encode data into the writeable nucleic acid. Data encoding can be performed by selectively converting convertible nucleobases of the nucleic acid molecule, such that the written nucleic acid molecule contains a sequence of unconverted and converted nucleobases, similar to the binary code of "0" and "1". Any suitable mechanism for chemically converting the nucleobases to a second structure can be utilized. According to various embodiments, the nucleobases are modified by light, voltage, enzymatic agents, chemical reagents, and / or redox agents.

[0162] In some embodiments, the data-written (data-encoded) nucleic acid molecule contains a sequence of converted nucleobases comprising a converted first set of nucleobases and a converted second set of nucleobases, resembling a binary code of "0" and "1."

[0163] In some embodiments, the nucleic acid polymer with written data (encoded) is stored according to standard nucleic acid storage protocols. For example, the nucleic acid polymer with written data can be dried and stored as a precipitate or in a suitable nuclease-free solution at room temperature or at a lower temperature (e.g., -20°C). Stabilizers such as (e.g.) alcohol, chelating agents and nuclease inhibitors can be included with the stored nucleic acid. To read the data on the written nucleic acid polymer, any suitable sequencer capable of reading non-natural and / or modified nucleic acid bases can be utilized, such as Oxford Nanopore Technologies' PromethION, MinION, and GridION sequencing platforms (Oxford, UK) or Pacific Bioscience's Single Molecule, Real-Time (SMRT) sequencing platform (Menlo Park, CA). Alternatively, a nanopore device for reading data can be fabricated or manufactured. The nanopore can be composed of solid state material or can contain one or more proteins.

[0164] In some embodiments, it is also contemplated to use solid support, such as polymer beads, glass beads, or inorganic solids, to isolate and stabilize nucleic acid.In some embodiments, the data written (encoded) on the nucleic acid polymer is decoded or read by sequencing by synthesis (SBS).In addition, in some embodiments, data can be decoded or read by utilizing a sequencer that can read modified and / or unmodified nucleic acid bases, such as Oxford Nanopore Technologies' PromethION, MinION, and GridION sequencing platform (Oxford, UK) or Pacific Bioscience's Single Molecule, Real-Time (SMRT) sequencing platform (Menlo Park, CA).

[0165] The present disclosure overcomes many of the limitations associated with traditional nucleic acid data storage by separating synthesis and data encoding into separate steps. The present disclosure provides a molecular strategy for making long strands of writeable nucleic acid that do not themselves encode data, but instead provide a template that can undergo writing. Writable nucleic acid polymers can be made in bulk prior to data encoding. The present disclosure further provides compositions and systems that include convertible nucleobases (and pairs of convertible nucleobases) that can be switched from a first state to a second state, thus acting as writeable "bits" of data that define "0" and "1" in binary code. The present disclosure further provides methods for writing data into the writeable nucleic acid polymers presented herein at the single molecule level, thus consuming negligible amounts of material. Data writing can be accomplished chemically or physically, utilizing (for example) light or voltage pulses. Finally, the written nucleic acid polymers are long, allowing more data to be encoded per molecule than with short DNA, and can be efficiently and quickly read by a variety of sequencers currently on the market. The compositions, systems, and methods described herein significantly increase the speed and density of nucleic acid data encoding while reducing the cost. Writable polymers for encoding data

[0166] In one aspect, provided herein is a polymer for encoding data, comprising a plurality of convertible residues covalently linked to the backbone of the polymer at repetitive intervals along the backbone of the polymer, each of the plurality of convertible residues having a first state and capable of converting from the first state to a second state, the plurality of convertible residues being covalently linked to the polymer in the first state and the second state. In some embodiments, the first state and the second state are different (e.g., the convertible residues have different structures in the first state and in the second state). In some embodiments, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state are readable by a polymerase enzyme. In some embodiments, the plurality of convertible residues are present at repetitive intervals along the backbone of the polymer.

[0167] In some embodiments, the polymer described herein is a nucleic acid polymer and the plurality of convertible residues are convertible nucleobases.

[0168] In certain embodiments, the convertible residues are repeatedly spaced apart to provide a spatial resolution that allows each residue to be converted independently. In some embodiments, any suitable spacer (e.g., non-writable, i.e., non-reactive to the data writing mechanism) is present between the convertible residues. In some embodiments, residues linked by a polymer backbone can be utilized as spacers. In some embodiments, the convertible residues are spaced apart by spacers depending on the spatial resolution of the writing mechanism and / or writing device. In some embodiments, the spacers can be residues and non-reactive to the writing mechanism. In some embodiments, these spacers are unmodified DNA nucleotides. In various embodiments, the polymer further comprises a delimiter and / or a data tag for labeling data.

[0169] In some embodiments, the polymer (e.g., nucleic acid polymer) described herein further comprises a plurality of spacer residues linked through the backbone of the polymer, where each of the plurality of convertible residues is separated by one or more spacer residues of the plurality of spacer residues. In some embodiments, the repeating spacing between the plurality of convertible residues matches the resolution of a writing mechanism for encoding data on the polymer. In some embodiments, the repeating spacing between two adjacent convertible residues is equal to or greater than the resolution of a data encoding mechanism for encoding data in the polymer. In some embodiments, the resolution of the writing mechanism is at least 1 nm. In some embodiments, the plurality of spacer residues does not interfere with the reading of the convertible residue. In some embodiments, the plurality of spacer residues in the polymer are the same spacer residue. In some embodiments, the plurality of spacer residues comprises two or more different spacer residues (e.g., different nucleobases, e.g., different naturally occurring nucleobases).

[0170] In some embodiments, the polymer described herein is a blank tape. In some embodiments, the polymer described herein is a blank tape of DNA. Blank tape, as used herein, refers to a writable nucleic acid polymer that includes convertible nucleobases that are repeatedly spaced along the writable nucleic acid polymer, and thus the conversion of the convertible nucleobases from a first state to a second state results in the encoding of data. The blank tape itself does not contain data, but by using a suitable writing system (e.g., by light), it is possible to encode data by converting the convertible nucleobases. In some embodiments, the blank tape can be written continuously from one end to the other end to encode data.

[0171] In some embodiments, the blank tape is writable over its entire length, hi some embodiments, each convertible nucleobase in the blank tape is independently and individually writable.

[0172] In some embodiments, a polymer (eg, a nucleic acid polymer) described herein consists essentially of spacer residues.

[0173] In some embodiments, the polymers described herein (eg, nucleic acid polymers) do not include delimiters or data tags.

[0174] In some embodiments, a polymer (eg, a nucleic acid polymer) described herein is comprised of a spacer residue and a convertible residue (eg, a convertible nucleobase).

[0175] In some embodiments, each of the plurality of convertible nucleobases is separated by 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 spacer residues.In some embodiments, each of the plurality of convertible nucleobases is separated by 6 spacer residues.In some embodiments, the plurality of spacer residues is a naturally occurring nucleobase, a non-naturally occurring nucleobase, a tetrahydrofuran abasic residue, or an ethylene glycol residue.The plurality of spacer residues is a naturally occurring nucleobase.

[0176] In some embodiments, the polymers (e.g., nucleic acid polymers) described herein further comprise one or more delimiters linked to the backbone of the polymer. In some embodiments, each of the one or more delimiters comprises one or more naturally occurring or non-naturally occurring nucleobases. In some embodiments, the one or more delimiters comprise naturally occurring nucleobases. In some embodiments, the one or more delimiters separate two or more adjacent data fields in the polymer.

[0177] In some embodiments, the polymers (e.g., nucleic acid polymers) described herein further comprise one or more data tags. In some embodiments, the one or more data tags comprise one or more naturally occurring or non-naturally occurring nucleobases. In some embodiments, the polymer is a nucleic acid polymer and the one or more data tags are present at the 5' or 3' end of the nucleic acid polymer. In some embodiments, the one or more data tags are incorporated into the nucleic acid polymer by ligation during synthesis of the nucleic acid polymer, during conversion of the plurality of convertible nucleobases to the second state, or after conversion of the plurality of convertible nucleobases to the second state.

[0178] In some embodiments, a polymer can have any number or length of monomer units, for example, as few as 10 monomer units to more than 100,000 monomer units, in various embodiments, the polymer has more than 500 monomer units, more than 1,000 monomer units, more than 5000 monomer units, more than 10,000 monomer units, more than 50,000 monomer units, or more than 100,000 monomer units.

[0179] In some embodiments, the nucleic acid polymer comprises more than 10 convertible residues. In some embodiments, the nucleic acid polymer comprises more than 100 convertible residues. In some embodiments, the nucleic acid polymer comprises more than 500 convertible residues. In some preferred embodiments, the nucleic acid polymer comprises more than 1,000 convertible residues. In some embodiments, the nucleic acid polymer comprises more than 10,000 convertible residues. In some embodiments, the nucleic acid polymer comprises more than 100,000 convertible residues.

[0180] In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 2 and 500. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 2 and 200. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 2 and 100. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 2 and 10. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 10 and 50.

[0181] In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 10 and 100. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 20 and 100. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is between 20 and 50. In some embodiments, the ratio of the total number of monomeric units (e.g., nucleotides) to convertible residues (e.g., convertible nucleobases) in a polymer (e.g., a nucleic acid polymer) is greater than 100. Writable Nucleic Acid Polymers

[0182] In certain embodiments, the polymers described herein (e.g., writeable polymers) are nucleic acid polymers, and the plurality of convertible residues are convertible nucleobases. In certain embodiments, the polymers described herein are nucleic acid polymers, and include a plurality of convertible nucleobases covalently linked to the backbone of the nucleic acid polymer at repetitive intervals along the backbone of the nucleic acid polymer, each of the plurality of convertible nucleobases having a first state (e.g., having a structure in a first state) and capable of converting from the first state to a second state (e.g., having a structure in a second state), and the plurality of convertible nucleobases are covalently linked to the nucleic acid polymer in the first state and the second state. In some embodiments, the first state and the second state are different and both are readable by a polymerase enzyme. In some embodiments, the nucleobase in the second state is a natural nucleobase. In some embodiments, the nucleobase in the second state is traceless (i.e., guanine, adenine, thymine, thiothymine, thioguanine, or 5-methylcytosine, or the native form of the nucleobase, such as cytosine).

[0183] In some embodiments, the unwritten state is also referred to as the untransformed state, and the written state is also referred to as the transformed state.

[0184] Compounds according to embodiments of the present disclosure are based on nucleic acids with multiple convertible nucleobases, which are analogous to writeable data bits. Each convertible nucleobase can exist in two or more states, an unwritten state (e.g., a first state) analogous to a "0", and at least a first written state (e.g., a second state of the nucleobase) analogous to a written bit representing a "1", and in some embodiments, a second written state (e.g., a third state of the nucleobase), and / or additional written states (i.e., the written bit is further writable). In some embodiments, a writeable nucleic acid polymer is synthesized with multiple convertible nucleobases in an "unwritten" state that can be converted to a "written" state. In some embodiments, two different convertible nucleobases are used as a pair to encode a single bit, where one conversion encodes a "0" and the other conversion encodes a "1". These writeable nucleic acids can be created in long lengths (e.g., 5-50 kb, or longer) and can be made in bulk prior to data writing.

[0185] In some embodiments, a single convertible nucleobase is utilized to encode the bit data. In some embodiments, a set of two or more convertible nucleobases is utilized to enable encoding of the bit data. In some embodiments, a pair of two different convertible nucleobases is used as a pair to enable encoding of a single bit. In some embodiments, a pair of two different convertible nucleobases is utilized, where conversion of the first nucleobase is encoded to a "0" and conversion of the other nucleobase is encoded to a "1". In some embodiments, a pair of two different convertible nucleobases is utilized, where conversion of one nucleobase is encoded to a "0" and conversion of both nucleobases is encoded to a "1".

[0186] In some embodiments, the writeable nucleic acid polymer comprises a plurality of convertible nucleobases linked to a polymer backbone. In certain embodiments, the convertible nucleobases are repeatedly spaced apart to provide a spatial resolution that allows each nucleobase to be converted independently. In some embodiments, the spatial resolution depends, at least in part, on the writing mechanism. For example, when modifying nucleobases using an optical light source and device with a resolution of 1 nm, each convertible base needs to be separated by at least 1 nm. Any suitable spacer can be utilized between the convertible nucleobases. In some embodiments, residues linked by the polymer backbone can be utilized as spacers. Since the distance between nucleobases in a double-stranded DNA polymer is about 0.34 nm, according to many embodiments, three spacers are utilized per nanometer of spatial resolution of the source inducing the modification. In some embodiments, the spacer is a nucleobase that can be non-reactive to the writing mechanism. In various embodiments, the writeable nucleic acid polymer can further comprise delimiters and / or data tags for labeling data, each of which can be provided by a specific sequence of residues.

[0187] In some embodiments, the data-encodeable nucleic acid polymer comprises a plurality of convertible nucleobases linked by a polymer backbone. In certain embodiments, the convertible nucleobases are regularly or irregularly spaced, and the data is encoded by identifying and selectively converting the nucleobases to obtain an encoded polymer. Some embodiments utilize regularly or irregularly spaced convertible nucleobases, and the data encoding mechanism can optionally skip any convertible nucleobase until it reaches the correct convertible nucleobase according to the code, resulting in a data-encoded nucleic acid polymer that includes stochastically and / or regularly spaced converted nucleobases. In certain embodiments, the convertible nucleobases (or sets of nucleobases) are repeatedly spaced apart to provide a spatial resolution that allows each nucleobase (or each set of nucleobases) to be independently converted. The spatial resolution depends, at least in part, on the writing mechanism. For example, when modifying nucleobases using an optical light source and device with a resolution of 1 nm, each convertible base (or each set of nucleobases) should be separated by at least 1 nm. Any suitable spacer can be utilized between the convertible nucleobases (or sets of nucleobases). In some embodiments, residues linked by a polymer backbone can be utilized as spacers. Since the distance between nucleobases in a double-stranded DNA polymer is about 0.34 nm, in accordance with many embodiments, three spacers are utilized per nanometer of spatial resolution of the source inducing the modification. In some embodiments, the spacer is a nucleobase that can be non-reactive to the writing mechanism. In various embodiments, the data-encoding nucleic acid polymer can further include delimiters and / or data tags for labeling data, each of which can be provided by a specific sequence of residues.

[0188] In some embodiments, the writeable nucleic acid polymers provided herein are capable of being written (e.g., selectively and sequentially converting convertible nucleobases to converted (e.g., naturally occurring or native nucleobases)) in either direction (e.g., in either the 5' to 3' direction or the 3' to 5' direction).

[0189] FIG. 1A illustrates an example of a writeable nucleic acid polymer having multiple writeable nucleobases. The writeable nucleic acid polymer includes a repeat strand sequence that can exist as a single-stranded or double-stranded molecule. The repeat unit includes a convertible nucleobase, which can be natural or unnatural, and can undergo a chemical change from a first structural state to a second structural state, which is similar to switching from a "0" state to a "1" state. Each of these convertible bases is similar to a "bit" for data encoding. It is understood that the definitions of "1" and "0" are arbitrary and merely indicate binary code. Prior to any data writing, the convertible nucleobases are initially provided in an unconverted state. In some embodiments, the repeat unit of the writeable nucleic acid polymer includes a data field that includes multiple convertible nucleobases and may also contain spacers or sequences that separate or separate the bits. FIG. 1B presents another example of a data field sequence with multiple convertible nucleobases separated by spacers. For example, as shown, three spacers are used between each convertible nucleic acid base, which results in a spatial resolution of 1 nm. It is understood that longer spacer sequences can be used for low bit writing resolution. In some embodiments, the writeable nucleic acid polymer comprises one or more unique data tag sequences that indicate documentation such as data type, date, or other information. The unique data tag sequence can be written during the synthesis of the writeable DNA, can be written during the data writing process, or can be added to the end via a primer, or can be added to the data strand via ligation after data writing.

[0190] FIG. 2A illustrates yet another example of a data-encodeable nucleic acid polymer having a plurality of convertible nucleobases, where each bit is a pair of convertible nucleobases that are repeated iteratively along the polymer. The data-encodeable nucleic acid polymer may exist as a single-stranded or double-stranded molecule. Each convertible nucleobase contains a removable group, such that the nucleobases can be converted from one structural state to a second structural state by removing the removable group with light or redox energy. With reference to FIG. 2A, in some embodiments, "C a " Convert the nucleic acid base to "C b " is not converted and is kept unchanged to obtain a "0" bit, and "C b " Convert the nucleic acid base to "C a " is obtained by keeping "C a " Convert the nucleic acid base to "C b " is not converted and is kept unchanged to obtain a "0" bit, and "C a " and "C b Both conversions of the nucleobases yield a "1" bit. It is understood that the definitions of "0" and "1" are arbitrary and merely indicate the binary code.

[0191] FIG. 2B illustrates a further example of a data-encodeable nucleic acid polymer having a plurality of convertible nucleobases, where each bit is a convertible nucleobase spaced along the nucleic acid polymer. The data-encodeable nucleic acid polymer may exist as a single-stranded or double-stranded molecule. Each convertible nucleobase contains a removal group, such that the nucleobase can be converted from one structural state to a second structural state by removing the removable group with light or redox energy. As shown in FIG. 2B, in some embodiments, "C a "The conversion of the nucleic acid base results in a "0" bit, and "C bConversion of a "1" nucleobase results in a "1" bit. In these embodiments, there are convertible nucleobases that can remain unconverted and therefore do not contribute to the code of the data.

[0192] In some embodiments, the data-encodable nucleic acid polymer includes one or more unique data tag sequences that indicate documentation such as the type of data, date, or other information. The unique data tag sequences can be incorporated during synthesis of the encodable polymer, can be added to the ends via primers, or can be added to the data strands via ligation after data encoding.

[0193] In various embodiments, the writeable nucleic acid polymer can be any length, for example, from as short as 15 nucleotides to more than 100 kilobases in length. In various embodiments, the writeable nucleic acid polymer is more than 500 nucleotides in length, more than 1000 nucleotides, more than 5000 nucleotides, more than 10,000 nucleotides, more than 50,000 nucleotides, or more than 100,000 nucleotides. The maximum length is only limited by the stability of DNA, by the method used to create them, and by the method used to read the written data. In some embodiments, longer chains have the advantage that more data is contained per molecule. Notably, current sequencing technologies are capable of handling nucleic acid strands that are tens to hundreds of thousands of bases long (see N Kono and K. Arakawa, Dev Growth Differ. 2019; 61: 316-326; and Q Chen and Z. Liu, Sensors (Basel). 2019; 19: 1886, the disclosures of each of which are incorporated herein by reference).

[0194] Some embodiments are directed to convertible nucleobases that can be incorporated into a writeable nucleic acid polymer. A convertible nucleobase according to various embodiments is a nucleobase that can be converted from a first chemical state to a second chemical state by controlled reaction chemistry. Any suitable mechanism for converting a nucleobase from a first state to a second state can be utilized, including but not limited to light pulses, voltage pulses, enzymatic agents, chemical reagents, and / or redox agents. It is understood that "nucleobase" is not limited to naturally occurring structures, but may also embody non-natural nucleobases, such as designer nucleobases.

[0195] In some embodiments, a convertible nucleobase is a nucleobase that can be converted from a first structural state to a second structural state by controlled reaction chemistry. In some embodiments, a convertible nucleobase comprises a removable group that can be removed (e.g., as a leaving group) to cause a structural change. Any suitable mechanism for converting a nucleobase from a first state to a second state can be utilized, including, but not limited to, light pulses, voltage pulses, enzymatic agents, chemical reagents, and / or redox agents. It is understood that "nucleobase" is not limited to naturally occurring structures, but may also embody non-natural nucleobases, such as designer nucleobases.

[0196] In some embodiments, the structural change results in the conversion of an unnatural nucleobase (e.g., a nucleobase in a first structural state) to a natural or native nucleobase (e.g., a nucleobase in a second structural state). A natural or native nucleobase in this definition can be identified by standard sequencing methods. In some embodiments, the nucleobase in the second state is a natural nucleobase. In some embodiments, the nucleobase in the second state is unmarked. In some embodiments, the nucleobase in the first state comprises a chemically modifiable moiety. In some embodiments, the nucleobase in the first state comprises no linker (or linker moiety) or side chain between the base of the nucleobase and the chemically modifiable moiety. In some embodiments, when the nucleobase in the first state is converted to the second state, the chemically modifiable moiety is removed, thereby leaving the nucleobase in the second state as a natural or native nucleobase. In some embodiments, the nucleobase in the first state and the nucleobase in the second state are readable or recognizable by a polymerase. In some embodiments, the written nucleic acid polymer is readable by various sequencing methods, for example, sequencing-by-synthesis (SBS).

[0197] In some embodiments, "trace" as used herein refers to a group that is not normally found in naturally occurring DNA (e.g., a part of a linker or side chain, etc.) that remains after a covalent bond is broken.Trace is frequently observed in some DNA sequencing techniques when a label is released by breaking a linker during the sequencing step.

[0198] 3A-3G provide examples of unconverted and converted states of convertible nucleobases. In some embodiments, the convertible nucleobases can be encoded into "bit" data, which allows conversion from a first structural state to a second structural state, analogous to the digital bit designations "0" or "1". In some embodiments, each state of the nucleobase should be readable by a sequencing method capable of detecting and distinguishing unnatural and / or modified bases, such as (for example) sequencing by synthesis or nanopore sequencing. FIGs. 3A-3G provide examples of convertible nucleobases designed to be converted from a first state to a second state by a localized pulse of light that removes the base's caging group, reduces its size, or alters its shape or H-bonds. A variety of photoremovable groups can be incorporated into the photoconvertible nucleobase (see, for example, DD Young and A. Deiters, Org Biomol Chem. 2007; 5: 999-1005; and Y. Wu, Z. Yang, and Y. Lu, Curr Opin Chem Biol. 2020; 57: 95-104, the disclosures of each of which are incorporated herein by reference). Although a few examples are provided, it is understood that any suitable photoremovable group and other nucleobases can be used in accordance with various embodiments. Figure 3E presents a convertible nucleobase that can be converted by localized enzymatic activity that removes a group, resulting in changes in size, shape, and H-bonds (see AE Pegg and TL Byers, FASEB J 1992; 6: 2302-10). Figure 3F presents a convertible nucleobase that is converted by localized oxidation resulting in a change in shape and polymerase substrate ability (K. Kino, et al., Genes Environ. 2017; 39: 21). Figure 3G presents a convertible nucleobase that is converted with a redox-removable group that similarly results in a change in size, shape, and / or polymerase substrate ability.With reference to Figures 3A-3G, both the unconverted and converted forms of these nucleobases are uniquely identifiable by current sequencing methods.

[0199] Figure 4 illustrates the conversion of the convertible nucleobase O6-nitrobenzyl-guanine to guanine by cleaving the bond to the nitrobenzyl group using light energy. This conversion can represent a bit of data or can be used in combination with one or more other convertible nucleobases to represent a writable bit of data. When decoding data by sequencing by synthesis, the unconverted O6-nitrobenzyl-guanine is read as a mixture of A and G, and after conversion, the resulting guanine is read as >99% G.

[0200] 5A-5B show further examples of convertible nucleobases that can be converted from a first state to a second state by a localized pulse of light that removes the caging group, thereby resulting in a natural nucleobase structure. Each of the exemplary convertible nucleobases includes a caging or removable group, which is shown in the structural diagram as "CG." Although several examples are provided, it is understood that any suitable convertible nucleobase structure that includes a photoremovable caging group can be used in accordance with various embodiments. With reference to FIG. 5A-5B, both the unconverted and converted states of these nucleobase structures are uniquely identifiable by current sequencing methods.

[0201] FIG. 6 provides further examples of photoremovable caging groups that can be utilized with nucleobase structures to provide convertible nucleobases that can be converted from a first state to a second state by a localized pulse of light. In various embodiments, any one of the photoremovable caging groups of FIG. 6 can be combined with the nucleobase structures of FIGS. 4 and 5A-5B. The photoremovable caging group includes a linker shown as "X" connected to the nucleobase structure shown as R. In addition to the examples provided, various other photoremovable caging groups can be incorporated into the photoremovable nucleobase (see, for example, D.D. Young and A. Deiters, Org Biomol Chem. 2007; 5: 999-1005; and Y. Wu, Z. Yang, and Y. Lu, Curr Opin Chem Biol. 2020; 57: 95-104, the disclosures of each of which are incorporated herein by reference).

[0202] Many embodiments are also directed to writeable nucleic acid polymers further incorporating one or more of spacers, delimiters, and data tags. According to various embodiments, the spacer is a molecular residue incorporated into the writeable nucleic acid polymer that provides the necessary spacing between convertible nucleic acid bases depending on the spatial resolution of the data writing mechanism. In many embodiments, the spacer is distinguishable from the convertible nucleic acid bases, and therefore, when reading the data with a sequencer, the spacer does not interfere with the ability to read the convertible nucleic acid bases. In some embodiments, the spacer is non-reactive with the data writing mechanism. In some embodiments, the writeable nucleic acid polymer utilizes repeats of the same residue in every spacer. In some embodiments, however, the writeable nucleic acid polymer utilizes two or more different residues as spacers. Any suitable residue that is distinguishable from the convertible nucleic acid bases can be utilized as a spacer, including naturally occurring nucleic acid bases, non-natural nucleic acid bases, tetrahydrofuran abasic residues, and / or ethylene glycol residues.

[0203] In some embodiments, the spacer is distinguishable from the convertible nucleobase and / or the converted nucleobase, and thus, when reading the data with a sequencer, the spacer does not interfere with the ability to encode the data and decode / read the encoded data. In some embodiments, the spacer is non-reactive with the data encoding mechanism.

[0204] According to various embodiments, the delimiter is a residue that indicates a boundary. In some embodiments, the delimiter is used to separate two adjacent data fields. Any suitable residue that can be distinguished from a convertible nucleobase can be used as the delimiter, including naturally occurring nucleobases, non-natural nucleobases, tetrahydrofuran abasic residues, and / or ethylene glycol residues.

[0205] In some embodiments, a data tag is a series of residues (typically four or more residues) that indicates a particular data. For example, a data tag can indicate the type of data, the date, the data source, or any other information. Any suitable residue that can be distinguished from a convertible nucleobase can be utilized as a data tag residue, including naturally occurring nucleobases, non-natural nucleobases, tetrahydrofuran abasic residues, and / or ethylene glycol residues.

[0206] In another aspect, also provided herein is a method for generating a writeable nucleic acid polymer comprising providing a circular single-stranded oligonucleotide template complementary to a repeat of a data field comprising convertible nucleobases; and incubating the circular single-stranded oligonucleotide template in the presence of a nucleic acid primer, a polymerase, and a nucleotide triphosphate, wherein the nucleotide triphosphate comprises a convertible nucleobase in a first state and is capable of being converted from the first state to a second state, wherein the first state and the second state are different.

[0207] In some embodiments, the circular single-stranded oligonucleotide template comprises nucleobases complementary to the convertible nucleobases, with the complementary nucleobases repeatedly spaced apart, such that incubating the template with a nucleic acid primer, a polymerase, and nucleotide triphosphates results in a nucleic acid polymer comprising a plurality of convertible nucleobases covalently linked through the backbone of the nucleic acid polymer and repeatedly spaced apart along the backbone of the nucleic acid polymer, the plurality of convertible nucleobases being covalently linked to the nucleic acid polymer in the first state and in the second state.

[0208] In some embodiments, the repeats of the data field further comprise a spacer nucleobase and the triphosphate nucleotide further comprises a triphosphate spacer nucleotide.

[0209] In another aspect, also presented herein is a method for generating a writeable nucleic acid polymer comprising: chemically synthesizing a plurality of oligomers, each oligomer comprising a plurality of convertible nucleobases linked via a nucleic acid polymer backbone at repetitively spaced intervals along the nucleic acid polymer backbone, each of the plurality of convertible nucleobases having a first state and capable of being converted from the first state to a second state, the plurality of convertible nucleobases being covalently attached to the nucleic acid polymer in the first state and the second state, the first state and the second state being distinct; and ligating the plurality of oligomers to form a writeable nucleic acid polymer.

[0210] In some embodiments, each of the plurality of oligomers comprises a plurality of spacer residues linked through the backbone of the nucleic acid polymer, and each of the plurality of convertible nucleic acid bases is separated by one or more spacer residues of the plurality of spacer residues. In some embodiments, the ligation step is by chemical ligation. In some embodiments, the ligation step is by enzymatic ligation. In some embodiments, the ligation step uses a complementary DNA splint.

[0211] In some embodiments, the multiple oligomers have the same sequence. In some embodiments, the multiple oligomers are multiple copies of the same sequence. In some embodiments, the multiple oligomers have different sequences.

[0212] In some embodiments, the method further comprises annealing the plurality of complements to the oligomer prior to the ligating step.

[0213] Writable nucleic acids can be generated by any suitable method for generating long nucleic acid polymers. Generally, according to various embodiments, polymerase extension or chemical synthesis is used to generate the writeable nucleic acid polymer. When polymerase extension is used, suitable convertible nucleic acid bases and residues that can be polymerized by polymerase are used. When chemical synthesis is used, a wider range of convertible nucleic acid bases and residues are used, but generally the nucleic acid strands that are generated by synthesis are short (e.g., between 10 and 200 residues) that can be ligated together to generate longer nucleic acid polymers. It is understood that both polymerase and ligation methods can build the writeable polymer repeats in either single-stranded or double-stranded states.

[0214] FIG. 7 illustrates an example of the generation of writable nucleic acids using polymerase extension, specifically, FIG. 7 illustrates an enzymatic rolling circle reaction method. In certain embodiments, a circular single-stranded DNA oligonucleotide is utilized as a template (MG Mohsen and ET Kool, Acc Chem Res. 2016; 49: 2540-2550, the disclosure of which is incorporated herein by reference). The circular single-stranded DNA oligonucleotide is complementary to the repeats of a data field that includes convertible nucleobases. In various embodiments, the circular single-stranded DNA oligonucleotide further comprises a spacer, a delimiter, and / or a data tag. In various embodiments, the circular DNA size is 2-2000 nucleotides long, preferably 2-200 nucleotides long, and more preferably 45-95 nucleotides long.

[0215] Once the nucleic acid circular template that encodes the repeat of the data field is constructed, the template is incubated with a nucleic acid primer, a polymerase, a suitable buffer to support the polymerase activity, and a suitable nucleoside triphosphate to generate a writable nucleic acid. The primer binds to the circle, and then the polymerase produces a long repeat complement of the circle. Rolling circle nucleic acid synthesis has been demonstrated to proceed for thousands of nucleotides, thereby generating long DNA repeats (see MM Ali, et al., Chem Soc Rev. 2014; 43: 3324-41; and MG Mohsen and ET Kool, Acc Chem Res. 2016 Nov 15; 49 (11): 2540-2550, the disclosures of which are incorporated herein by reference). In some embodiments, a data tag is utilized, which can be included at the distant 5' end of the primer and left non-complementary to the DNA circle. In this case, rolling circle DNA synthesis results in repeats of the writeable nucleic acid with a data tag attached to the 5' end. If it is desired that the writeable nucleic acid polymer is double stranded, a primer complementary to the repeat of the data field can be used together with a polymerase and nucleotides complementary to the first polymer to generate a complementary strand.

[0216] FIG. 8 illustrates a chemical synthesis and ligation method for generating writeable nucleic acids. In some cases, the nucleotides for incorporation into the writeable nucleic acid are not efficient polymerase substrates, particularly many unnatural nucleobases, which impede the ability to effectively use the polymerase to generate long nucleic acid polymers. Chemical synthesis and ligation techniques involve constructing short writeable nucleic acid polymers on a DNA synthesizer, which can be done using phosphoramidite synthesis protocols, typically resulting in polymer lengths of 10-200 nucleotides. To aid in ligation, in some embodiments, the short polymers synthesized further include a 5'-phosphate group and a native, unmodified 3'-hydroxyl group. In the presence of ATP, the short polymers are joined together to generate long repeat polymers by a DNA ligase enzyme (e.g., T4 DNA ligase). In some embodiments, ligation is aided by utilizing a complementary "sprint" nucleic acid oligonucleotide that can hybridize to the reactive ends.

[0217] In some embodiments, to generate a double-stranded writeable nucleic acid, a nucleic acid complement is synthesized that includes a 5'-phosphate group. The complementary strand is hybridized with the writeable nucleic acid prior to ligation. In some embodiments, hybridization of the complementary strand results in a duplex with sticky ends that can be efficiently ligated into a double-stranded writeable nucleic acid polymer using a ligase enzyme.

[0218] The polymer molecules extracted by ligation can result in a range of polymer lengths. In some embodiments, a mixture of polymers of different lengths is used for data encoding. In some embodiments, specific lengths are enriched and / or isolated (e.g., by electrophoresis) and then used for data encoding.

[0219] Some embodiments are directed to polymerase amplification of writeable nucleic acid polymers by repeat extension using a thermostable polymerase (e.g., DNA polymerase from Thermococcus litoralis). For further details regarding polymerase extension of repeat regions, see JS Hartig and ET Kool, Nucleic Acids Res. 2005; 33: 4922-7, the disclosure of which is incorporated herein by reference.

[0220] If the ends of the data field DNA to be ligated are inefficient ligase enzyme substrates due to poor hybridization or non-native structures that interfere with the enzyme, then according to various embodiments, natural nucleic acid bases can be added to the ligation sites to ensure good hybridization / ligation. Some embodiments utilize chemical ligation to generate the writable nucleic acid polymer. Chemical ligation can be achieved using cyanogen bromide, using carbodiimide reagents, or by the nucleophilic reaction of a phosphorothioate group at one nucleic acid polymer strand end with a leaving group such as iodide (for example) at the other nucleic acid polymer strand end. Chemical ligation involves joining phosphate and hydroxyl ends, but the reaction can be performed using 5'-phosphate and 3'-hydroxyl, or 3'-phosphate and 5'-hydroxyl. Such chemical ligation methods have been described (see E. T. Kool, Acc Chem Res. 1998; 31: 502-510; C. Obianyor, et al., Chembiochem. 2020; 21: 3359-3370; and Y. Xu and E. T. Kool, Nucleic Acids Res. 1999; 27: 875-81, the disclosures of each of which are incorporated herein by reference). Method and system for writing and reading data - Patents.com

[0221] In another aspect, provided herein are systems and methods for writing to or reading written polymers (e.g., nucleic acid polymers) as provided herein. system

[0222] In another aspect, presented herein is a system for writing data, the system including: a writeable polymer comprising a plurality of convertible residues covalently linked to a backbone of the polymer at repetitively spaced intervals along the backbone of the polymer, each of the plurality of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being distinct, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state being readable by a polymerase enzyme, the plurality of convertible residues being covalently linked to the polymer in the first state and the second state; and a data writing device for writing data onto the writeable polymer.

[0223] In some embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleic acid bases. In some embodiments, the data writing device comprises a nanopore. In some embodiments, the data writing device converts the plurality of convertible nucleic acid bases to the second state by a light pulse, a voltage pulse, an enzymatic agent, or a redox agent. In some embodiments, the data writing device converts the plurality of convertible nucleic acid bases to the second state by a light pulse. In some embodiments, the data writing device comprises a light irradiation device. Methods for writing / encoding on writable polymers

[0224] In yet another aspect, presented herein is a method for writing data onto a writeable polymer, the method comprising: providing a writeable polymer comprising a plurality of convertible residues covalently linked via the backbone of the writeable polymer at repetitively spaced intervals along the backbone of the writeable polymer, each of the convertible residues of the plurality of convertible residues having a first state and capable of being converted from the first state to a second state, the first state and the second state being distinct, the plurality of convertible residues in the first state and the plurality of convertible residues in the second state being readable by a polymerase enzyme; and utilizing a data writing device to selectively convert one or more of the plurality of convertible residues to the second state, thereby generating a data encoded polymer.

[0225] Some embodiments are directed to writing and reading data on a nucleic acid polymer. In many embodiments, a writeable nucleic acid polymer is provided having convertible nucleic acid bases that are repeatedly spaced along the writeable polymer. The provided writeable nucleic acid polymer may also have spacers, delimiters, and data tags as described herein. To write data onto a nucleic acid polymer, according to various embodiments, individual strands are passed through a device having a nanopore. The device having a nanopore further provides a means for selectively converting the convertible nucleic acid bases from a first state to a second state. A number of means can be utilized to convert the convertible nucleic acid bases, including, but not limited to, light pulses, voltage pulses, enzymatic agents, chemical reagents, and / or redox agents. Examples of nanopore devices for passing DNA and encoding with localized light pulses are described in the examples provided in the exemplary embodiments.

[0226] In some embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleic acid bases. In some embodiments, the data writing device comprises a nanopore and the method further comprises passing the writeable polymer through a nanopore of the writing device, where the nanopore converts one or more of the plurality of convertible residues to a second state.

[0227] In some embodiments, the nanopore is a plasmonic nanopore that provides localized excitation energy for selectively converting the convertible nucleobase from a first state to a second state. In some embodiments, the data writing device includes a plasmonic well or channel, and the method further includes transferring the writeable polymer to a plasmonic well or channel of the data encoding device, the plasmonic well or channel providing localized excitation from a light pulse to selectively convert the convertible nucleobase from a first state to a second state. In some embodiments, the data writing device selectively converts the convertible residue to the second state by a light pulse, a voltage pulse, an enzymatic agent, or a redox agent. In some embodiments, the data writing device selectively converts the convertible residue to the second state by a light pulse.

[0228] In some embodiments, the convertible residue becomes a naturally occurring nucleobase after conversion to the second state.

[0229] In some embodiments, the start and / or end positions for writing on a writeable polymer can be any position (i.e., any convertible residue, e.g., a convertible nucleic acid base) within the writeable polymer (e.g., a writeable nucleic acid polymer), and no specific start and / or end positions are required.

[0230] In some embodiments, the selectively converting step is initiated at either end of the writeable polymer (e.g., the 5' or 3' end of the nucleic acid polymer). In some embodiments, the selectively converting step is initiated at the 5' or 3' end of the nucleic acid polymer. In some embodiments, the selectively converting step selectively converts convertible residues (e.g., convertible nucleobases) in either direction of the writeable polymer. In some embodiments, the selectively converting step selectively converts convertible nucleobases (e.g., writeable bits) in either the 5' to 3' direction or the 3' to 5' direction. In some embodiments, the selectively converting step is initiated at the 5' end of the nucleic acid polymer. In some embodiments, the selectively converting step is initiated at the 3' end of the nucleic acid polymer.

[0231] In some embodiments, writing begins at any position on the writeable polymer (e.g., any convertible residue, e.g., a convertible nucleobase). In some embodiments, writing ends at any position on the writeable polymer (e.g., any convertible residue, e.g., a convertible nucleobase). In some embodiments, writing begins and ends at any position on the writeable polymer (e.g., any convertible residue, e.g., a convertible nucleobase).

[0232] In some embodiments, a writable polymer is writable along its entire length, with writing beginning at a beginning position (e.g., the 3' end of the nucleic acid polymer) and ending at an ending position (e.g., the 5' end of the nucleic acid polymer).

[0233] In some embodiments, the plurality of convertible residues includes two or more types of convertible residues, a first type of convertible residue being activatable by a first wavelength of light and a second type of convertible residue being activatable by a second wavelength of light. In some embodiments, the repetitive spacing between the plurality of convertible residues is compatible with the resolution of a data writing device for selectively converting the convertible residues. In some embodiments, the selectively converting step does not require specific positioning of the writeable polymer. In some embodiments, the conversion of the convertible residues to the second state is not uniform on the data encoded polymer. In some embodiments, the conversion of the convertible residues to the second state is not limited to certain locations on the data encoded polymer.

[0234] In some embodiments, the writeable polymer comprises a plurality of convertible residues regularly spaced along the writeable polymer, hi some embodiments, the data encoded polymer after the data has been written comprises converted nucleobases stochastically or irregularly spaced.

[0235] In some embodiments, the plurality of convertible nucleobases are convertible by light of wavelengths of 325 nm, 360 nm, or 400 nm.

[0236] In some embodiments, the plurality of convertible nucleobases are convertible by light of a wavelength between 400 nm and 850 nm.

[0237] In some embodiments, the method further comprises stretching or combing the writeable polymer (eg, writeable DNA) onto a solid support.

[0238] In some embodiments, the method further comprises visualizing the location of the convertible residues using a dye.

[0239] In some embodiments, the method further comprises locally illuminating or locally exciting the writeable polymer, hi some embodiments, the locally illuminating or locally exciting uses a stimulated emission depletion (STED) laser.

[0240] In some embodiments, the method further comprises joining two or more data fields from two or more writable polymers end-to-end, thereby resulting in a joined polymer comprising two or more data fields.

[0241] In some embodiments, the method further comprises controlling the rate of passage of the writeable polymer through the nanopore of the writing device.

[0242] In some embodiments, multiple writable polymers are passed in parallel through a data writing device or devices to write the same data (eg, data redundancy occurs).

[0243] In some embodiments, data-encoded polymers produced by selectively converting convertible nucleobases include different polymer molecules encoded with the same data, hi some embodiments, the data-encoded nucleic acid polymers include converted nucleobases at different positions (e.g., differently and optionally irregularly spaced) along the nucleic acid polymer, but with the same data encoded therein (e.g., the order of written data bits is the same among the different encoded polymer molecules).

[0244] In some embodiments, to encode data onto the writeable nucleic acid polymers presented herein, according to various embodiments, individual polymers have light energy or redox energy that repeatedly strikes the polymer, thus allowing for controllably and selectively converting convertible nucleic acid bases to encode a data code (e.g., a binary data code).

[0245] Although devices with nanopores are described, any device can controllably and selectively convert convertible nucleobases according to a data code, hi some embodiments, the device utilizes plasmonic channels or wells to controllably and selectively convert convertible nucleobases.

[0246] In some embodiments, the device selectively provides a means for converting the convertible nucleobase as the writeable nucleic acid polymer passes through the nanopore. For example, if the nucleobase should be converted to a second state by a light pulse, the device provides light as the nucleic acid polymer passes through the nanopore, so that the light can contact the convertible nucleobase and convert the convertible nucleobase to the second state. If the nucleobase should remain in the first state, no light is provided by the device, so that the convertible nucleobase passes through the nanopore unconverted. In many embodiments, the convertible nucleobase can be sandwiched between spacers that depend on the writing resolution of the device to ensure that only a single nucleobase is converted by the device. For example, when modifying nucleobases using an optical light source and device with a resolution of 1 nm, each convertible base must be separated by at least 1 nm.

[0247] In certain embodiments, if the nucleobases are to be converted to a second state by a light pulse, the device provides light as the nucleic acid polymer passes through the nanopore, and thus the light can contact only the set of convertible nucleobases to be converted. If the nucleobases are to remain in their original state, no light is provided by the device, and thus the convertible nucleobases pass through the nanopore unconverted. In many embodiments, to ensure that only the set of nucleobases is converted by the device, the set of convertible nucleobases can be sandwiched between spacers that correspond to the writing resolution of the device.

[0248] In some embodiments, to ensure that only a single nucleobase (or set of nucleobases) is converted by the device, the device utilizes two or more means for converting nucleobases; the first means can convert the first nucleobase structure but cannot convert the second nucleobase structure, and the second means can convert the second nucleobase structure but cannot convert the first nucleobase structure.For example, the device can utilize two wavelengths of light to provide energy, so that the first wavelength can convert the first nucleobase structure but cannot convert the second nucleobase structure, and the second wavelength can convert the second nucleobase structure but cannot convert the first nucleobase structure.

[0249] In some embodiments, to ensure that only a single nucleobase (or set of nucleobases) is converted by the device, the device utilizes two or more means for converting nucleobases; a first means is capable of converting a first nucleobase structure but not a second nucleobase structure, and a second means is capable of converting both the first nucleobase structure and the second nucleobase structure simultaneously as a pair. For example, the device can utilize two wavelengths of light to provide energy, such that the first wavelength can convert the first nucleobase structure but not the second nucleobase structure, and the second wavelength can convert both the first nucleobase structure and the second nucleobase structure simultaneously as a pair.

[0250] In many embodiments, the writing device is provided with a code for writing data into the nucleic acid polymer. Thus, the writing device selectively converts various nucleobases of the polymer that resemble the binary code "1", while selectively allowing the nucleobases of the polymer that resemble "0" to pass through the pore unconverted. After the data code is written into the nucleic acid polymer, the nucleic acid polymer with the written data code can be stored by any suitable means for storing nucleic acid molecules. For example, the nucleic acid polymer with the written data can be dried and stored as a precipitate or in a suitable nuclease-free solution at room temperature or at a lower temperature (e.g., -20°C). Stabilizers such as (for example) alcohol, chelating agents and nuclease inhibitors can be included with the nucleic acid to be stored.

[0251] In some embodiments, the polymers (e.g., nucleic acid polymers) presented herein can be stored under standard nucleic acid storage protocols. In some embodiments, the polymers are nucleic acid polymers that can be stored in a suitable nuclease-free solution at room temperature or at low temperature (e.g., -20°C). In some embodiments, the polymers can be stored at room temperature without the use of stabilizers.

[0252] In many embodiments, the data encoding device is provided with a code for writing data into a nucleic acid polymer. Thus, in some embodiments, the encoding device selectively converts various nucleobases of the polymer according to the code. In some embodiments using single nucleobases as bits, some of the nucleobases are selectively converted and others are selectively not converted, resulting in a binary code of converted and unconverted nucleobases, thereby encoding data. In some embodiments using single nucleobases as bits, some of the nucleobases are selectively converted to a first converted structure and others are selectively converted to a second converted structure, resulting in a binary code of converted nucleobases, thereby encoding data. In this case, any unconverted nucleobases remain uncoded and are not used to decode the data code.

[0253] In some embodiments utilizing sets of nucleobases to encode bits, each set includes at least two convertible nucleobases, and the encoding device selectively converts a first nucleobase of a portion of the set to the converted structure and a second nucleobase of the other set to the converted structure, resulting in a binary code. In some embodiments utilizing sets of nucleobases to encode bits, each set includes at least two convertible nucleobases, and the encoding device selectively converts a first nucleobase of a portion of the set to the converted structure and both nucleobases of the other set to the converted structure, resulting in a binary code.

[0254] In some embodiments, nucleic acid polymers are the ones that store data most efficiently at the single molecule level, thereby providing the highest potential density of information. However, in some embodiments, if data redundancy is required for better data storage accuracy, multiple nucleic acid polymers can be used to write the same data redundantly to each polymer of the multiple polymers. Error correction algorithms for digital data storage are already well developed, and some of these algorithms can be applied to the present approach (see J. Li, et al., IEEE Transactions on Emerging Topics in Computing. 2021; 9: 651-663, the disclosure of which is incorporated herein by reference).

[0255] In various embodiments of decoding encoded data by sequencing by synthesis (SBS), it may be desirable to have redundancy in the data, and therefore have the same data on each polymer of multiple polymers.For example, when using a nucleobase structure such as O6-nitrobenzyl-guanine, the structure is read as a mixture of A and G using SBS, and therefore redundancy in the reading of the structure is required to interpret whether the structure is O6-nitrobenzyl-guanine, guanine, or adenine.In some methods of SBS, the redundancy is inherent to each single sequence that is read. Methods for reading / decoding writable polymers

[0256] In another aspect, also presented herein is a method for reading data from a data encoded polymer, the method comprising: providing a data encoded polymer comprising convertible residues covalently linked through the backbone of the polymer at repetitively spaced intervals along the backbone of the polymer, a first subset of the convertible residues being in a first state and a second subset of the convertible residues being in a second state, the first state and the second state being distinct, a plurality of the convertible residues in the first state and a plurality of the convertible residues in the second state being readable by a polymerase enzyme; and passing the writable data encoded polymer through a data reading device to read the encoded data on the data encoded polymer.

[0257] In some embodiments, the writeable polymer is a writeable nucleic acid polymer and the plurality of convertible residues are convertible nucleic acid bases. In some embodiments, the convertible residues in a first state can be converted to a second state by light. In some embodiments, the data reading device comprises a nanopore. In some embodiments, the data reading device is a sequencing device. In some embodiments, the sequencing device is a sequencing by synthesis device.

[0258] In some embodiments, the method further comprises measuring the current flow in the electrolyte while the writeable polymer passes through.

[0259] In some embodiments, the method further includes determining whether each of the plurality of convertible residues is in a first state or a second state based on a measured current flow in the electrolyte during passage of the writeable polymer.

[0260] In some embodiments, the method further includes passing the data encoded polymer again through the data reading device to again read the encoded data on the data encoded polymer.

[0261] In some embodiments, the method further includes verifying and correcting the encoded data on the data encoded polymer by comparing the encoded data on multiple copies of the data encoded polymer.

[0262] In another aspect, there is provided a method for reading or decoding data from a data encoded nucleic acid polymer, comprising: a plurality of converted nucleobases, each converted nucleobase comprising a first nucleobase structure, the first converted nucleobase being converted from a first state to a second state, the first state and the second state being different; a plurality of convertible nucleobases, each of which comprises a second nucleobase structure and a directly linked removable group, the convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by releasing the second removable group from the second nucleobase structure, the first state and the second state being different; providing a plurality of overlapping copies of a data-encoded nucleic acid polymer, wherein the converted nucleobases and the convertible nucleobases are linked via a nucleic acid polymer backbone; determining the sequence of each of the multiple overlapping copies of the nucleic acid polymer; Also presented herein are methods, including:

[0263] In some embodiments, the method further comprises detecting the plurality of converted nucleobases and the plurality of convertible nucleobases, and decoding the data based on the detected plurality of converted nucleobases.

[0264] In some embodiments, the plurality of converted nucleobases in the first state and the plurality of converted nucleobases in the second state are readable by a polymerase enzyme. In some embodiments, the plurality of convertible nucleobases in the first state and the plurality of convertible nucleobases in the second state are readable by a polymerase enzyme. In some embodiments, the plurality of converted nucleobases and the plurality of convertible nucleobases are detected based on sequencing results of overlapping copies of the data-encoded nucleic acid polymer.

[0265] In some embodiments, the sequence determining step is initiated at either end of the writeable polymer (e.g., the 5' or 3' end of the nucleic acid polymer). In some embodiments, the sequence determining step is initiated at the 5' or 3' end of the nucleic acid polymer. In some embodiments, the sequence determining step is initiated at the 5' end of the nucleic acid polymer. In some embodiments, the sequence determining step is initiated at the 3' end of the nucleic acid polymer.

[0266] 9A-9C illustrate an example of using a device 501 with a nanopore to write data into a writeable nucleic acid polymer 503. The device includes a substrate 505 including a plasmonic nanostructure 507 for providing localized optical energy to the writeable polymer 503. The writeable polymer 503 is controllably passed through the nanopore 501 at a constant rate. The nanopore may be composed of a protein, an in silico engineered pore, or an artificial one such as another inorganic solid (see N Kono and K. Arakawa, Dev Growth Differ. 2019; 61: 316-326; and Q Chen and Z. Liu, Sensors (Basel). 2019; 19: 1886, the disclosures of each of which are incorporated herein by reference). Methods for constructing nanopores and controlling the rate of passage have been described previously (see Y. Zhishan, et al., Nanoscale Res Lett. 2020; 15: 80, the disclosure of which is incorporated herein by reference). As the writable nucleic acid polymer 503 passes through the nanopore 501 at a controlled rate, the device selectively converts each convertible nucleobase as encoded as the convertible nucleobase passes through the pore. As shown in FIG. 9B, a pulse of light 509 can be applied to the convertible nucleobase via the plasmonic nanostructure 507 locally at the same time as the convertible nucleobase passes through the pore, which can be timed appropriately to control the rate of passage through the pore. As a result of selective nucleobase conversion, binary digital data is encoded within the polymer (FIG. 9C).

[0267] 10A-10C illustrate another example utilizing a device 701 with a nanopore for encoding data into an encodable nucleic acid polymer 703 that includes multiple sets of convertible nucleobases that are repeated iteratively along the polymer. The device includes a substrate 705 that includes a plasmonic nanostructure 707 for providing localized optical energy of multiple wavelengths to the data encodable polymer 703. The polymer 703 is controllably passed through the nanopore 701 at a constant rate. As the data encodable nucleic acid polymer 703 passes through the nanopore 701 at a controlled rate, the device selectively converts one or both of the convertible nucleobases of each set as the set passes through the pore, as defined by the data code. In this example, the encoded data code is 1001, where 1 is a C a ', and 0 is C a 'C b As shown in FIG. 10A , a pulse of light 709 of a first wavelength (e.g., 400 nm) can be applied via the plasmonic nanostructure 707 to the set locally as it passes through the pore, resulting in conversion of a single convertible base (as shown, base C a C a As shown in FIG. 10B, a pulse 711 of light at a second wavelength (e.g., 365 nm) can be applied via the plasmonic nanostructure 707 locally to the set as it passes through the pore, resulting in conversion of both of the convertible bases (as shown, base C a and C b C a ' and C b As a result of the selective nucleobase conversions, binary digital data is encoded within polymer 703, which is encoded by sets with single nucleobase conversions 713 and sets with dual nucleobase conversions 715 (FIG. 10C).

[0268] 11A-11C illustrate yet another example utilizing a device 801 with a nanopore for encoding data into an encodable nucleic acid polymer 803 that includes a plurality of two convertible nucleobase structures repeated stochastically or randomly along the polymer. The device includes a substrate 805 including a plasmonic nanostructure 807 for providing localized optical energy of one or more wavelengths to the data encodable polymer 803. The polymer 803 is controllably passed through the nanopore 801 at a constant rate. As the data encodable nucleic acid polymer 803 passes through the nanopore 801 at a controlled rate, the device selectively converts one convertible nucleobase structure at a time as prescribed by the data code. In this example, the encoded data code is 10110, where 1 is C a ', and 0 is C b As shown in FIG. 11A, a pulse of light 809 can be applied via the plasmonic nanostructure 807 to the first nucleobase structure locally as it passes through the pore, resulting in conversion of the nucleobase (as shown, base C a C a As shown in FIG. 11B, a pulse of light 809 can be applied via the plasmonic nanostructure 807 to the second nucleobase structure locally as it passes through the pore, resulting in conversion of the nucleobase (as shown, base C b C b '). Additionally, as shown in Figures 11B and 11C, convertible bases 813, 815, and 817 are skipped in accordance with the code. As a result of the selective nucleobase conversion, binary digital data is encoded in polymer 803, which is encoded by the converted nucleobases Ca'Cb'Ca'Ca'Cb' in accordance with the data code, with any convertible bases skipped.

[0269] Highly localized optical excitation can be achieved by subwavelength focusing strategies with specialized microscopes such as STEDX, or by using nanoplasmonic structures such as bowties, or by using zero-mode waveguides (see Y. Fang and M Sun, Light Sci Appl. 2015; 4:e294; and X. Shi, et al. Small. 2018; 14:e1703307, the disclosures of each of which are incorporated herein by reference). When using redox for the conversion of nucleobases, applied potentials of electrodes near or within the nanopore or nanochannel can be used. With a constant passage rate, electronic pulses of timed voltage potentials can provide appropriate intervals for the conversion of nucleobases. For the conversion of nucleobases by enzymes, a writable nucleic acid polymer can be passed through two adjacent nanopores at a controlled rate. Once the convertible nucleobase enters the volume between the two pores, an enzyme contacts the chain at a localized portion / base / bit (e.g., by microfluidics). The timing of the microfluidic flow and the controlled passage of the writeable polymer can be coordinated with appropriate intervals so that data is encoded with fidelity.

[0270] Some embodiments are also directed to positive bit writing using dual bits. Thus, in certain embodiments, the writable nucleic acid polymer comprises one or more repeated pairs of convertible nucleobases, where each convertible base of the pair is within the same field of resolution of the writing mechanism. In some embodiments, each convertible nucleobase of the pair is adjacent to the other nucleobase of the pair. In some embodiments, each convertible nucleobase of the pair is sufficiently close so that the other nucleobase of the pair is addressed by the same conversion signal. In some embodiments, one of the convertible nucleobases of the pair has a different reaction condition for converting the nucleobase than the other nucleobase of the pair. For example, in some embodiments, the first convertible nucleobase of the pair is converted by light of a first wavelength, and the second convertible nucleobase of the pair is converted by light of a second wavelength. Thus, in certain embodiments encoding a writeable nucleic acid polymer containing one or more pairs, as each pair enters the nanopore, specific reaction conditions are provided to convert the first convertible nucleobase, or the second convertible nucleobase, or both the first convertible nucleic acid residue and the second convertible nucleobase according to the code.

[0271] 12A-12C illustrate an example of using a nanopore-equipped device 601 to write data into a writeable nucleic acid polymer 603 that includes multiple pairs. The device includes a substrate 605 that includes a plasmonic nanostructure 607 for providing localized optical energy of multiple wavelengths to the writeable polymer 603. The writeable polymer 603 is controllably passed through the nanopore 601 at a constant rate. As the writeable nucleic acid polymer 603 passes through the nanopore 601 at a controlled rate, the device selectively converts each convertible nucleic acid base of the pair as coded upon the pair passing through the pore. As shown in FIG. 12A, a pulse 609 of light at a first wavelength (e.g., 400 nm) can be applied via the plasmonic nanostructure 607 to the pair locally upon its passage through the pore, resulting in the conversion of a single convertible base (base W as shown). a Wa As shown in FIG. 12B, once the pair has passed through the pore, a pulse of light 611 of a second wavelength (e.g., 325 nm) can be applied locally to the pair via the plasmonic nanostructure 607, resulting in conversion of both of the convertible bases (as shown, base W a and W b W a ' and W b '). As a result of selective nucleobase conversion, binary digital data is encoded within the polymer 603, which is encoded by single nucleobase converted pairs 613 and dual nucleobase converted pairs 615 (FIG. 12C). Examples of convertible nucleobases that convert at specific wavelengths are presented in FIGS. 13A-13C.

[0272] In many embodiments, any suitable sequencer capable of reading non-natural and / or modified nucleobases can be utilized to read data on the written nucleic acid polymer. In certain embodiments, the device is capable of writing to and reading from the nucleic acid polymer. In certain embodiments, the nanopore has dual functionality for both writing to and reading from the nucleic acid polymer, although some devices may include separate nanopores for writing and reading. Examples of commercially available nanopore sequencers include Oxford Nanopore Technologies' PromethION, MinION, and GridION sequencing platforms (Oxford, UK) and Pacific Bioscience's Single Molecule, Real-Time (SMRT) sequencing platform (Menlo Park, CA). Alternatively, nanopore devices for writing and / or reading data can be fabricated or manufactured. The nanopore can be composed of solid state material or can contain one or more proteins.

[0273] In many embodiments, any suitable sequencer capable of reading non-natural and / or modified nucleic acid bases can be utilized to decode the data on the encoded nucleic acid polymer. Examples of sequencing techniques used to decode DNA include, but are not limited to, shotgun sequencing, long-read sequencing, nanopore sequencing, and sequencing by synthesis.

[0274] In various embodiments of decoding encoded data by sequencing by synthesis (SBS), it may be desirable to provide redundancy in the data, so that each polymer of a plurality of polymers has the same data.For example, when using a nucleobase structure such as O6-nitrobenzyl-guanine, the structure is read as a mixture of A and G using SBS, and therefore redundancy in the reading of the structure is required to interpret whether the structure is O6-nitrobenzyl-guanine, guanine, or adenine.

[0275] FIG. 14A provides an example of using a nanopore to read the nucleobase sequence of a convertible nucleobase and a converted nucleobase. In this example, O4-nitrobenzylthymine (T-4-ONB) is used as the convertible base, and the nucleobase is converted to thymine by removal of the nitrobenzyl group. The weak current of T-4-ONB is low, while the current of thymine is larger, making these two structures distinguishable from the resulting current readout. Although T-4-ONB is used in this example, any convertible nucleobase that can recognize changes in structure size and / or charge can be used, including (but not limited to) the structures presented in FIGS. 4 and 5A-5B.

[0276] In certain embodiments, sequencing by synthesis (SBS) is performed to decode data in nucleic acid polymers. SBS can be useful for decoding between certain bases that have been converted and / or remain unconverted. Standard SBS utilizes polymerase a to read a strand of DNA sequence and create a complementary copy of that strand. The converted nucleic acid base has the ability to serve as a polymerase substrate, resulting in a predictable sequence outcome, which should allow the polymerase to incorporate a base on the opposite side and continue synthesis. For example, O6-nitrobenzylguanine (O6NBG) is intended as a convertible base, which is a suitable substrate for DNA polymerase enzymes and can therefore be read by SBS. Sequencing of an O6NBG nucleobase results in a read with a mix of A and G nucleobases encoded at that position (see, e.g., AM Kietrys, WA Velema, and ET Kool, J Am Chem Soc. 2017; 139: 17074-17081, the disclosure of which is incorporated herein by reference). However, once the nitrobenzyl group is removed and converted to a guanine structure, the sequencing read has an unambiguous G signal. When utilizing SBS, sequencing multiple copies of the encoded nucleic acid can help distinguish whether the nucleobase at a given position is a converted structure (e.g., guanine) or an unconverted structure (e.g., O6-nitrobenzylguanine), and thus indicate the presence of whether data is encoded at that position. Notably, sequencing multiple copies of the encoded nucleic acid can help distinguish several convertible / converted nucleobase structures, such as those presented in Figures 4 and 5A-5B.

[0277] FIG. 14B provides an example of using SBS to read the nucleobase sequence of a convertible nucleobase and a converted nucleobase. In this example, O4-nitrobenzylthymine (T-4-ONB) is used as the convertible base, and removal of the nitrobenzyl group converts the nucleobase to thymine. SBS of T-4-ONB results in a mixed base read, while removal of the nitrobenzyl group results in a specific read of thymine (see, e.g., AM Kietrys, WA Velema, and ET Kool, J Am Chem Soc. 2017; 139: 17074-17081, the disclosure of which is incorporated herein by reference). While this example uses T-4-ONB, any convertible nucleobase that changes the sequencing read as a result of conversion can be used, including (but not limited to) the structures provided in FIGS. 4 and 5A-5B. Certain embodiments

[0278] Embodiment 1. A nucleic acid polymer for encoding data, comprising: a plurality of pairs of convertible nucleobases, the pairs being repeatedly spaced along the nucleic acid polymer, each of the convertible nucleobases being linked via a nucleic acid polymer backbone; A nucleic acid polymer, wherein each of the convertible nucleobases of each pair comprises a nucleobase structure and a leaving group, the leaving group being linked to the nucleobase structure via a linker, and each of the convertible nucleobases of each pair is provided in a first state and can be converted from the first state to a second state by light energy or redox energy that releases the leaving group from the nucleobase structure.

[0279] Embodiment 2. The nucleic acid polymer of embodiment 1, further comprising a first plurality of sets of spacer residues, each spacer residue being linked via the nucleic acid polymer backbone, each set of the first plurality of sets comprising two or more spacer residues, each set of the first plurality of sets being interposed between each pair of the plurality of pairs of convertible nucleobases to provide repeating spacing between the plurality of pairs of convertible nucleobases.

[0280] Embodiment 3. The nucleic acid polymer of embodiment 2, further comprising a second plurality of sets of spacer residues, each spacer residue being linked via the nucleic acid polymer backbone, each set of the second plurality of sets comprising one or more spacer residues, each set of the second plurality of sets being interposed between a respective convertible nucleobase of a pair of nucleobases, and the number of spacer residues in each set of the second plurality of sets being less than the number of spacer residues in each set of the first plurality of sets.

[0281] Embodiment 4. The nucleic acid polymer of embodiment 1 or 2, wherein the repeating spacing between pairs of convertible nucleobases is equal to or greater than the resolution of the data encoding mechanism for encoding data in the nucleic acid polymer.

[0282] Embodiment 5. The nucleic acid polymer of any one of embodiments 1 to 4, wherein each convertible nucleobase comprises one of the following nucleobase structures: O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0283] Embodiment 6. The leaving group is: [ka] where X is a linker to the nucleobase structure, the linker being one of NR2, NHR, OR, or SR, and R is a nucleobase structure. 6. The nucleic acid polymer of any one of embodiments 1 to 5, comprising one of:

[0284] Embodiment 7. The nucleic acid polymer of embodiment 1, wherein light energy is used to release each leaving group, wherein a first wavelength of light provides energy capable of converting a first convertible nucleobase of each pair to its second state, and a second wavelength of light provides energy capable of converting a second convertible base of each pair to its second state.

[0285] Embodiment 8. The nucleic acid polymer of embodiment 7, wherein the second wavelength of light provides energy further capable of converting the first convertible nucleobase of each pair to its second state.

[0286] Embodiment 9. A nucleic acid polymer for encoding data, comprising: a first plurality of convertible nucleobases linked via a nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the first plurality of convertible nucleobases comprising a first nucleobase structure and a first leaving group, the first leaving group being linked to the first nucleobase structure via a first linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the first leaving group from the first nucleobase structure; a second plurality of convertible nucleobases linked via the nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the second plurality of convertible nucleobases comprising a second nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a second linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the second leaving group from the second nucleobase structure; A nucleic acid polymer comprising:

[0287] Embodiment 10. The nucleic acid polymer of embodiment 9, further comprising a plurality of spacer residues linked via the nucleic acid polymer backbone, the spacer residues being stochastically or randomly spaced between the convertible nucleobases.

[0288] Embodiment 11. The nucleic acid polymer of embodiment 9 or 10, wherein each convertible nucleobase comprises one of the following nucleobase structures: O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0289] Embodiment 12. The leaving group is: [ka] where X is a linker to the nucleobase structure, the linker being one of NR2, NHR, OR, or SR, and R is a nucleobase structure. 12. The nucleic acid polymer of any one of embodiments 9 to 11, including one of:

[0290] Embodiment 13. A convertible nucleobase for use in a data-encodeable polymer, comprising a nucleobase structure and a leaving group, wherein the leaving group is linked to the nucleobase structure via a linker, and wherein the leaving group is capable of being removed from the nucleobase structure by light energy or redox energy.

[0291] Embodiment 14. The convertible nucleobase of embodiment 13, wherein the nucleobase structure comprises O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0292] Embodiment 15. The leaving group is: [ka] where X is a linker to the nucleobase structure, the linker being one of NR2, NHR, OR, or SR, and R is a nucleobase structure. 14. The convertible nucleobase of embodiment 13, comprising:

[0293] Embodiment 16. The convertible nucleobase of embodiment 15, wherein the linker comprises NR2, NHR, OR, or SR, and R is a nucleobase structure.

[0294] Embodiment 17. A data-encoded nucleic acid polymer comprising a plurality of pairs of nucleobases, each nucleobase pair comprises at least a first converted nucleobase, the first converted nucleobase comprising a first nucleobase structure, the first converted nucleobase being converted from a first state to a second state by light energy or redox energy that releases a first leaving group from the first nucleobase structure; Each pair of nucleobases is a convertible nucleobase comprising a nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a linker, the convertible nucleobase being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the second leaving group from the second nucleobase structure; or a second converted nucleobase, the second converted nucleobase comprising a second nucleobase structure and converted from a first state to a second state by light energy or redox energy that releases a second leaving group from the second nucleobase structure. and further comprising at least one of the nucleobase pairs are repeatedly spaced along the nucleic acid polymer, and the nucleobases are linked via the nucleic acid polymer backbone; Nucleic acid polymer.

[0295] Embodiment 18. The nucleic acid polymer of embodiment 17, further comprising a first plurality of sets of spacer residues, each spacer residue being linked via the nucleic acid polymer backbone, each set of the first plurality of sets comprising two or more spacer residues, each set of the first plurality of sets being interposed between each pair of the plurality of pairs of nucleobases to provide repeating spacing between the plurality of pairs of nucleobases.

[0296] Embodiment 19. The nucleic acid polymer of embodiment 18, further comprising a second plurality of sets of spacer residues, each spacer residue being linked via the nucleic acid polymer backbone, each set of the second plurality of sets comprising one or more spacer residues, each set of the second plurality of sets being interposed between a respective convertible nucleobase of a pair of nucleobases, and the number of spacer residues in each set of the second plurality of sets being less than the number of spacer residues in each set of the first plurality of sets.

[0297] Embodiment 20. The nucleic acid polymer of embodiment 17 or 18, wherein the repeating spacing between pairs of nucleobases is equal to or greater than the resolution of the data encoding mechanism used to encode the data in the data-encoded nucleic acid polymer.

[0298] Embodiment 21. The nucleic acid polymer of any one of embodiments 14 to 20, wherein each converted nucleobase has one of the following nucleobase structures: guanine, adenine, thymine, or cytosine.

[0299] Embodiment 22. The nucleic acid polymer of any one of embodiments 14 to 21, wherein each convertible nucleobase comprises one of the following nucleobase structures: O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0300] Embodiment 23. The second leaving group of each convertible nucleobase is [ka] where X is a linker to the nucleic acid structure, the linker being one of NR2, NHR, OR, or SR, and R is a nucleobase structure. 23. The nucleic acid polymer of any one of embodiments 14 to 22, including one of:

[0301] Embodiment 24. A nucleic acid polymer encoded with data, comprising: a first plurality of converted nucleobases linked via a nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each converted nucleobase of the first plurality of converted nucleobases comprising a first nucleobase structure, each converted nucleobase of the first plurality of converted nucleobases having been converted from a first state to a second state by light energy or redox energy that releases a first leaving group from the first nucleobase structure; a second plurality of converted nucleobases linked via the nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each converted nucleobase of the second plurality of converted nucleobases comprising a second nucleobase structure, each converted nucleobase of the second plurality of converted nucleobases having been converted from a first state to a second state by light energy or redox energy causing release of a second leaving group from the second nucleobase structure; A data-encoded nucleic acid polymer comprising:

[0302] Embodiment 25. A nucleic acid polymer comprising a first plurality of convertible nucleobases linked via a nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the first plurality of convertible nucleobases comprising a first nucleobase structure and a first leaving group, the first leaving group being linked to the first nucleobase structure via a first linker; a second plurality of convertible nucleobases linked via the nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the second plurality of convertible nucleobases comprising a second nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a second linker; 25. The data-encoded nucleic acid polymer of embodiment 24, further comprising:

[0303] Embodiment 26. The nucleic acid polymer of embodiment 25, further comprising a plurality of spacer residues linked via the nucleic acid polymer backbone, the spacer residues being stochastically or randomly positioned between the nucleobases, including the converted and convertible nucleobases.

[0304] Embodiment 27. The nucleic acid polymer of any one of embodiments 24 to 26, wherein each converted nucleobase has one of the following nucleobase structures: guanine, adenine, thymine, or cytosine.

[0305] Embodiment 28. The nucleic acid polymer of any one of embodiments 25 to 27, wherein each convertible nucleobase comprises one of the following nucleobase structures: O6-guanine, N2-guanine, N7-guanine, N6-adenine, N5-adenine, O4-thymine, N3-thymine, 2-thio-thymine, 4-thio-thymine, N4-cytosine, or N3-cytosine.

[0306] Embodiment 29. The leaving group of each convertible nucleobase is [ka] where X is a linker to the nucleobase structure, the linker being one of NR2, NHR, OR, or SR, and R is a nucleobase structure. 29. The nucleic acid polymer of any one of embodiments 25 to 28, including one of:

[0307] Embodiment 30. A method of encoding data into a data-encodable nucleic acid polymer, comprising: A data-encodable nucleic acid polymer comprising a plurality of pairs of convertible nucleobases, the pairs being repeatedly spaced along the nucleic acid polymer, each convertible nucleobase being linked via a nucleic acid polymer backbone; providing a data-encodable nucleic acid polymer, each of each pair of convertible nucleobases comprising a nucleobase structure and a leaving group, the leaving group being linked to the nucleobase structure via a linker, each of each pair of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the leaving group from the nucleobase structure; utilizing a data encoding device to selectively convert at least one nucleobase of each pair of convertible nucleobases to a second state by providing light energy or redox energy to release a leaving group from the nucleobase structure of at least one nucleobase; A method comprising:

[0308] Embodiment 31. The method of embodiment 30, wherein the data encoding device comprises a plasmonic nanopore, and the method further comprises passing the data-encoding nucleic acid polymer through the plasmonic nanopore of the data encoding device, wherein the plasmonic nanopore supplies light energy or redox energy to release a leaving group from the nucleobase structure of at least one nucleobase.

[0309] Embodiment 32. The method of embodiment 31, wherein the data-encodable nucleic acid polymer further comprises a first plurality of sets of spacer residues, each spacer residue being linked via the nucleic acid polymer backbone, each set of the first plurality of sets comprising two or more spacer residues, each set of the first plurality of sets being interposed between each pair of the plurality of pairs of convertible nucleobases to provide repeating spacing between the plurality of pairs of convertible nucleobases.

[0310] Embodiment 33. The method of embodiment 31 or 32, wherein the repeating spacing between pairs of convertible nucleobases is equal to or greater than the resolution of the data encoding device.

[0311] Embodiment 34. The method of embodiment 30, wherein the data encoding device includes a plasmonic well or channel, and the method further comprises the step of transferring the data-encoding nucleic acid polymer to a plasmonic well or channel of the data encoding device, and providing light energy or redox energy by the plasmonic well or channel to release a leaving group from the nucleobase structure of at least one nucleobase.

[0312] Embodiment 35. The method of embodiment 30, wherein the data encoding device comprises a STED laser system, and the method further comprises the steps of stretching the data-encodeable nucleic acid polymer and focusing an STED laser on the stretched data-encodeable nucleic acid polymer, and providing light energy or redox energy by the STED laser to release a leaving group from the nucleobase structure of at least one nucleobase.

[0313] Embodiment 36. A method of encoding data into a data-encodable nucleic acid polymer, comprising: a first plurality of convertible nucleobases linked via a nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the first plurality of convertible nucleobases comprising a first nucleobase structure and a first leaving group, the first leaving group being linked to the first nucleobase structure via a first linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the first leaving group from the first nucleobase structure; a second plurality of convertible nucleobases linked via the nucleic acid polymer backbone at stochastic or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the second plurality of convertible nucleobases comprising a second nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a second linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the second leaving group from the second nucleobase structure; providing a data-encodeable nucleic acid polymer comprising: selectively converting a subset of the convertible nucleobases of the first plurality of convertible nucleobases and the second plurality of convertible nucleobases to a second state by utilizing the data encoding device to provide light energy or redox energy to release leaving groups from the nucleobase structures of the convertible nucleobases; A method comprising:

[0314] Embodiment 37. The method of embodiment 36, wherein the subset of convertible nucleobases of the first plurality of convertible nucleobases and the second plurality of convertible nucleobases that are selectively converted is based on an encoded data code.

[0315] Embodiment 38. The method of embodiment 37, wherein the selective conversion of nucleobases results in a nucleic acid polymer comprising a convertible nucleobase between the converted nucleobases.

[0316] Embodiment 39. The data encoding device comprises a plasmonic nanopore and the method comprises: passing a data-encodable nucleic acid polymer through a plasmonic nanopore of a data encoding device, where the plasmonic nanopore provides light energy or redox energy to release a leaving group from the nucleobase structure of the convertible nucleobase; 37. The method of embodiment 36, further comprising:

[0317] Embodiment 40. The data encoding device comprises a plasmonic well or channel, and the method comprises:

[0318] transferring the data-encodable nucleic acid polymer into a plasmonic well or channel of a data encoding device, where light energy or redox energy is provided by the plasmonic well or channel to release leaving groups from the nucleobase structures of the convertible nucleobases; 31. The method of embodiment 30, further comprising:

[0319] Embodiment 41. A data encoding device comprising a STED laser system and a method comprising: stretching the data-encodeable nucleic acid polymer and focusing STED laser energy onto the stretched data-encodeable nucleic acid polymer, where the STED laser provides light energy or redox energy to release leaving groups from the nucleobase structures of the convertible nucleobases. 31. The method of embodiment 30, further comprising:

[0320] Embodiment 42. A method for decoding data from a data-encoded nucleic acid polymer, comprising: a plurality of converted nucleobases, each converted nucleobase comprising a first nucleobase structure, the first converted nucleobase being converted from a first state to a second state by light energy or redox energy that releases a first leaving group from the first nucleobase structure; a plurality of convertible nucleobases, each convertible nucleobase comprising a nucleobase structure and a leaving group, the leaving group being linked to a second nucleobase structure via a linker, the convertible nucleobase being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the second leaving group from the second nucleobase structure; providing a plurality of overlapping copies of a data-encoded nucleic acid polymer, wherein the converted nucleobases and the convertible nucleobases are linked via a nucleic acid polymer backbone; determining a sequence of each duplicate copy of the plurality of duplicate copies; detecting a plurality of converted nucleobases and a plurality of convertible nucleobases; Decoding the data based on the detected plurality of converted nucleobases. A method comprising:

[0321] Embodiment 43. The method of embodiment 42, wherein the plurality of converted nucleobases and the plurality of convertible nucleobases are detected based on the results of sequencing overlapping copies of the data-encoded nucleic acid polymer.

[0322] Embodiment 44. The method of embodiment 43, wherein sequencing results indicating mixed nucleobase structures at a particular nucleobase indicate convertible nucleobases that are not part of the data code.

[0323] Embodiment 45. A nucleic acid polymer for encoding data, comprising: a first plurality of convertible nucleobases linked via a nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the first plurality of convertible nucleobases comprising a first nucleobase structure and a first leaving group, the first leaving group being linked to the first nucleobase structure via a first linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the first leaving group from the first nucleobase structure; a second plurality of convertible nucleobases linked via the nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the second plurality of convertible nucleobases comprising a second nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a second linker, each convertible nucleobase of the first plurality of convertible nucleobases being provided in a first state and capable of being converted from the first state to a second state by light energy or redox energy that releases the second leaving group from the second nucleobase structure; A nucleic acid polymer comprising:

[0324] Embodiment 46 The nucleic acid polymer of embodiment 45, further comprising a plurality of spacer residues linked via the nucleic acid polymer backbone, the spacer residues being positioned between the convertible nucleobases.

[0325] Embodiment 47. A nucleic acid polymer encoded with data, comprising: a first plurality of converted nucleobases linked via a nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each converted nucleobase of the first plurality of converted nucleobases comprising a first nucleobase structure, each converted nucleobase of the first plurality of converted nucleobases having been converted from a first state to a second state by light energy or redox energy that releases a first leaving group from the first nucleobase structure; a second plurality of converted nucleobases linked via the nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each converted nucleobase of the second plurality of converted nucleobases comprising a second nucleobase structure, each converted nucleobase of the second plurality of converted nucleobases having been converted from a first state to a second state by light energy or redox energy causing release of a second leaving group from the second nucleobase structure; A data-encoded nucleic acid polymer comprising:

[0326] Embodiment 48. A nucleic acid polymer comprising a first plurality of convertible nucleobases linked via a nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the first plurality of convertible nucleobases comprising a first nucleobase structure and a first leaving group, the first leaving group being linked to the first nucleobase structure via a first linker; a second plurality of convertible nucleobases linked via the nucleic acid polymer backbone at regular or irregular intervals along the nucleic acid polymer, each convertible nucleobase of the second plurality of convertible nucleobases comprising a second nucleobase structure and a second leaving group, the second leaving group being linked to the second nucleobase structure via a second linker; 48. The data-encoded nucleic acid polymer of embodiment 47, further comprising:

[0327] Embodiment 49. The nucleic acid polymer of embodiment 48, further comprising a plurality of spacer residues linked via the nucleic acid polymer backbone, the spacer residues being positioned between the nucleobase comprising the converted nucleobase and the nucleobase comprising the convertible nucleobase. Exemplary embodiments

[0328] Various examples of compositions, systems, and methods for data storage using nucleic acid polymers are described herein. Examples of writable nucleic acid polymers, methods for making such polymers, methods for writing data, and methods for reading data are provided. EXAMPLES

[0329] Example 1 Writeable DNA polymers with MeNPOC nucleobases Writeable nucleic acid molecules can be generated that contain bits, data fields, spacers, delimiters, and / or terminal identifier tags. In this example, the converted nucleobase (i.e., "1") is 5-aminopropynyl-deoxyuridine, and the unconverted nucleobase (i.e., "0") is the same molecule with a substitution at the amine group with a MeNPOC group that can be efficiently removed by light (see P. Klan, et al., Chem Rev. 2013; 113: 119-91, the disclosure of which is incorporated herein by reference). We construct writeable nucleic acids that are all composed of convertible nucleobases with a deoxyuridine base with a substitution with MeNPOC, shown as "0" in the following example: Data field: 5'-C-(A)6-0-(A)6-0-(A)6-0-(A)6-0-(A)6-0-(A)6-0-(A)6-0-(A)6-0-(A)6-(C)-3'.

[0330] The data field contains "0" bits spaced apart by six adenine nucleotides (A) to allow spatial resolution for writing by focused optical energy. The data field is shown here as 8 bits (1 "byte" in 8-bit architecture). The terminal cytosine may provide a data delimiter function, indicating a break between one 8-bit field and the next. It is understood that the spacers and delimiters are not limited to adenosine and cytidine, but can be almost any single or multiple natural or non-natural residues that are preferably detectably different from the convertible nucleic acid bases and non-reactive with the writing mechanism. It is also understood that delimiters may not be necessary to achieve efficient data encoding. In such cases, the writable nucleic acid contains repeated bits and spacers that are not contained within the delimiters. It is also understood that the spacing and number of spacers between bits can be easily altered to reflect the resolution and precision of the writing method.

[0331] A writeable nucleic acid polymer consists of a data field sequence repeated in a string. The polymer can be tagged at the 5' or 3' end with a data tag. The data tag can include a sequence of natural bases that indicates the time, date, type of data, user, or other useful identifying information. It is understood that for some applications, a data tag may not be necessary since identifying information can be written directly into the data field. Example 2 Writable nucleic acid polymers produced by rolling circle reaction

[0332] This example describes a circular DNA oligonucleotide that encodes the repeated "data field" of Example 1. The circle is selected to be complementary to the repeat unit, and in this case, the size is selected to be 57 nucleotides, which falls within the size range known to act as a good substrate for DNA polymerase-mediated rolling circle synthesis (see MG Mohsen and ET Kool, Acc Chem Res. 2016 Nov 15; 49(11): 2540-2550, the disclosure of which is incorporated herein by reference). The sequence of the circle is as follows: 5'-GTTTTTTATTTTTTATTTTTTATTTTTTATTTTTTTATTTTTTTTTTTTTTTTTTTTTG-3', with the 5' and 3' ends joining intramolecularly to form the circle.

[0333] A DNA primer is constructed with a 3' end complementary to the circle. An example of an effective primer sequence is as follows: Primer: 5'-ID sequence-AAAAAATAAAAAACCAAAAAA-3'

[0334] The ID sequence is optional. Mg supporting DNA polymerase activity 2+The DNA primers are annealed to the DNA circles in a buffer containing the DNA primers. The mixture is contacted with nucleoside triphosphates (dNTPs) that will constitute the repeats of the data field. For the data field of Example 1, the dNTPs required are 5-nitroveratryl-oxycarbonyl-aminoproynyl deoxyuridine 5'-triphosphate, dATP, and dCTP. The solution is contacted with an appropriate DNA polymerase enzyme at a temperature that supports enzyme activity to create a long repeating writable DNA polymer that contains the repeats of the data field and a DNA data identifier tag at the 5' end. Gel analysis shows that the blank tape is 10,000-50,000 nucleotides long. The blank tape is isolated from the smaller polymerases, nucleotides, and circles by size exclusion chromatography, column purification, precipitation, gel electrophoresis, or by other purification methods, and stored in the dark to avoid unintentional bit writing.

[0335] Various DNA polymerase enzymes for rolling circle synthesis have been described (see S. Ishino and Y. Ishino, Front Microbiol. 2014; 5: 465, the disclosure of which is incorporated herein by reference). Examples include phi29 and BST3.0 polymerases. High processivity polymerases allow for longer writeable DNA polymers to be produced. Polymerases that can efficiently accept modified nucleotides (such as modified deoxyuridines as described herein) as substrates can be used. Example 3 Writable nucleic acid polymers produced by synthesis and ligation

[0336] In this example, a ligase enzyme is used to assemble single- and / or double-stranded writeable DNA polymers containing O6-ortho-nitrobenzyl G (see FIG. 3D, here designated X) as a convertible nucleobase, which cannot be efficiently incorporated into DNA by most polymerase enzymes due to blocked base pairing. The repeat sequence of the designed 8-bit data field is as follows: 5'-CCT-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-CGA-3'.

[0337] A ligatable oligonucleotide containing a single 8-bit field is synthesized with the following sequence: 5'-pCCT-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-X-(A)6-(CGA)-3' (where "p" represents the terminal phosphate group). A splint for ligating this sequence is synthesized with the following sequence: 5'-TTTTTTAGGTCGTTTTTT-3'.

[0338] Contacting the splint and data field oligonucleotides with T4 DNA ligase and ATP in a buffer that supports the ligase results in the end-to-end joining of many data field oligomers, thereby giving rise to a long polymer chain. Gel analysis of the product reveals a length ladder ranging in size from 5000 to 50,000 nucleotides. If desired, a portion of the "data field" DNA product can be split and each end can be separately ligated to a different DNA identifier for separate use in data writing. Long data fields are used for writing as a mixture of lengths. Alternatively, electrophoretic gels can be used to produce blank tape DNA of uniform length by cutting out and eluting specific bands.

[0339] A double-stranded, writable DNA polymer is obtained by a similar method, in this case also using the first data field oligonucleotide, but with a different complement to form a duplex with sticky ends. The sequence of this complementary oligonucleotide is: 5'-pGTTTTTTCTTTTTTCTTTTTTCTTTTTTCTTTTTTCTTTTTTCTTTTTTCTTTTTTCTTTTTTAGGTC-3'

[0340] Complementary oligonucleotides are hybridized to the data field oligonucleotides to produce a duplex with sticky ends. Ligation with T4 DNA ligase and ATP produces a long repeated DNA double-stranded polymer. Gel analysis of the product reveals a length ladder ranging in size from 5000 to 50,000 base pairs. If desired, a portion of the data field DNA product can be split and one end can be ligated separately to a different DNA identifier for separate use in data writing. Long data fields are used for writing as mixed lengths. Alternatively, electrophoretic gels can be used to produce blank tape DNA of uniform length by cutting out and eluting specific bands. Example 4 Optical data writing

[0341] A nanopore device with a plasmonic bowtie on the exit side of the pore is used to write digital data into the writeable DNA polymer of Example 1. A nanopore with a plasmonic bowtie has been described (see X. Shi, et al., Small. 2018 May; 14 (18): e1703307, the disclosure of which is incorporated herein by reference). The writeable polymer is dissolved in an electrolyte solution and moved through the pore at a constant rate by an applied potential across the two sides of the pore. A test bit sequence "01100101" is written in repeats. This is achieved by illuminating the nanoplasmonic structure with a beam of light spaced in time to match the spacing of the bits in the data field. Subsequent analysis by nanopore sequencing reveals the sequence of "1" and "0" bits, which repeats, allowing the accuracy and errors of the bit writing to be analyzed. Statistical analysis and data correction for repeat units in the sequence confirm the intended bit sequence. Subsequent experiments with longer strings of data reveal the ability to encode more data per molecule. Comparing multiple copies of the DNA tape with the same data allows for sequence comparison and error correction. Example 5 Stretching DNA and writing data with light

[0342] In this example, data is encoded into the double-stranded, writable DNA polymer of Example 3 by combining DNA stretching or combing with localized illumination to write bits. The stretching / combing technique uses a flow to stretch individual DNA molecules with lengths of tens of thousands of nucleotides onto a slide or other solid support, and the location of the long DNA is visualized by a simple dye added to the solution (see TF Chan, et al., Nucleic Acids Res. 2006; 34:e113; and S Takahashi, M. Oshige, and S. Katsura, Molecules. 2021; 26: 1050, the disclosures of each of which are incorporated herein by reference). Progressively along the strand, light is focused at the intended "1" sites along the strand to convert the nucleobase bits from a "0" state to a "1" state. Light illumination is achieved at high resolution by using the STED technique, which uses two lasers to provide precisely localized illumination (see G. Vicidomini, P. Bianchini, and A. Diaspro, Nat Methods. 201; 15: 173-182, the disclosure of which is incorporated herein by reference).

[0343] The resulting written DNA can be stored for archiving. If the data is to be retrieved, the stored data can be read by nanopore sequencing of the DNA polymer (see Example 7).

[0344] In another embodiment, the bit nucleotides contain a fluorescent dye linked to a fluorescent quencher by a photocleavable linker. The presence of the quencher keeps the unwritten DNA non-fluorescent. "Localized illumination" of the "stretched DNA" strand results in cleavage of the linker, which causes the quencher to disappear and the localized nucleotide to fluoresce. The progression of photoexcitation light along the stretched data field DNA results in the writing of bits at data-encoding intervals. The slide is saved as the written data. When the data is to be retrieved, the strand is imaged on a slide and read by analyzing the "1" bits as fluorescent spots; the intervals indicate the presence and number of intervening "0" bits. Example 6 Writing data by oxidation-reduction

[0345] This example describes the writing of data by redox using a writeable DNA polymer containing the redox-reactive nucleotides of FIG. 3G. The experiment uses a nanopore device with electrodes in the pore. A DNA blank tape containing redox-reactive nucleobases is passed through the pore at a controlled rate. As the DNA passes, a reductive voltage potential is applied as pulses at timed intervals. This results in the reduction and disappearance of the group on the "0" bit, switching the "0" bit to an aminopropyne group, which encodes a "1". The time interval over which the reduction is applied results in a varying but predictable interval of "1" and "0" groups, thereby defining the digital data. Example 7 Reading written DNA polymers by nanopore sequencing

[0346] A typical nanopore sequencing device measures the current flow of an electrolyte while a DNA molecule passes through the pore. Since each DNA base is a different size and shape, the current is slightly altered as each different base passes through the pore. In this example, a commercially available nanopore device is used to perform experiments to read the change in current over time as a written DNA tape passes through. In this case, a single strand of written DNA polymer is used, as produced in Example 3 and written in Example 4. The "1" and "0" bits contain G and nitrobenzyl G, which are significantly different in size. Experiments with a DNA tape with all "0" bits (blank polymer) reveal a drop in current when the largest nitrobenzyl G nucleotide passes through, allowing the difference in current between these "0" bits and the spacer and delimiter to be distinguished. Separate measurements of DNA, an all "1" polymer, show the level of current observed when a "1" (G) bit passes through. These experiments provide a calibration for reading and distinguishing the current levels that indicate a "1" bit and a "0" bit. The fully written DNA polymer is then passed through. The current levels representing "1" and "0" are read and placed in the context of the current levels seen for the spacers and delimiters. If necessary, multiple reads of the same strand are used to improve the accuracy of the data read. Example 8 Dual-bit writable nucleic acid polymer

[0347] In this example, we present a writeable nucleic acid polymer design that allows for the writing of both "1" and "0" bits using an activation signal. In this design, zeros are not passively included in the data field, but rather an active switching signal is required. Photoremovable groups can be triggered with distinct wavelengths of light. Figures 13A-13C show examples of nucleotides that contain a group that is removable by irradiation at 325 nm and a different group that is removable by irradiation at 400 nm. If these two groups are placed close to each other in the data field of a blank DNA tape, a 400 nm light pulse will remove only one of the two groups in the pair. On the other hand, a 325 nm light pulse will result in the loss of both of these groups. These two outcomes are analogous to "0" and "1" for encoding data. Example 9 Building Data-Encoding DNA

[0348] A 141 nucleotide DNA strand is synthesized containing recurrently repeated pairs of convertible nucleobases (X and Y) separated by two spacer nucleobases. Each pair represents a bit of codeable data. Each pair of nucleobases is separated by 10 intervening spacer nucleobases. The total number of pairs in the strand is 11, therefore the DNA can code for 11 bits of "1" and "0" data. The sequence of this 150mer is: [ka] where X represents O6-nitrobenzylguanine and Y represents N6-coumarinylmethyl-adenine.

[0349] A complementary DNA sequence is synthesized that is complementary to the first strand and thus capable of forming a duplex. The complementary sequence can be designed to create overhanging sticky ends, and the two strands are further modified with 5' phosphate groups. The sequence of this 141mer is: [ka] It is.

[0350] Note that the bases in this complement are designed to be complementary to the converted versions of bases X and Y. Longer DNA can store more data per molecule. To generate longer nucleic acid polymers for data storage, the two DNA strands are separated by a Mg2 ion exchanger that supports hybridization and enzymatic ligation. + The data-encodeable polymers can be mixed in a buffer containing ATP and T4 DNA ligase, resulting in end-to-end joining of 150 nucleotides of DNA into long polymer chains of about 300 bp and longer in length, including about 1500 bp of DNA as analyzed by agarose gel electrophoresis. The data-encodeable DNA of the desired size can be isolated by gel electrophoresis and extraction. Thus, the data-encodeable polymers can be provided and utilized as a mixture of lengths, or as specific lengths by cutting out specific bands. Example 10 Encoding Data into Polymers

[0351] A nanopore device with a plasmonic bowtie on the exit side of the pore is used to write digital data into the data-encodeable DNA polymer of Example 9. A nanopore with a plasmonic bowtie has been described (see X. Shi, et al., Small. 2018 May; 14 (18): e1703307, the disclosure of which is incorporated herein by reference). The data-encodeable polymer is dissolved in an electrolyte solution and translocated through the pore at a constant rate by an applied potential across the two sides of the pore. The data sequence "01100101100" is encoded into the polymer (for the first 150 nucleotides). This is achieved by shining a light beam onto the nanoplasmonic structure at time intervals that match the spacing of the paired bits.

[0352] To encode bit data, light energy can be applied to the bit pair at a wavelength of 400 nm to release a coumarinylmethyl group from N6-coumarinylmethyl-adenine and convert the nucleobase to adenine. The 400 nm light energy has no effect on O6-nitrobenzylguanine, which remains unconverted. This bit pair conversion can be denoted as "0". Similarly, light energy can be applied to the bit pair at a wavelength of 365 nm to release a nitrobenzyl group from O6-nitrobenzylguanine and convert the nucleobase to guanine, and release a coumarinylmethyl group from N6-coumarinylmethyl-adenine and convert the nucleobase to adenine. This bit pair conversion can be denoted as "1". Continuing with the data encoding, the data sequence "01100101100" can be obtained, which structurally has the following nucleobase sequence: [ka] (In the sequences, X denotes O6-nitrobenzylguanine and Y denotes N6-coumarinylmethyl-adenine.) In particular, multiple copies can be encoded such that decoding can be performed by SBS, where unconverted nucleobases are read as mixed bases in the sequencing result. Example 11 Decoding data from encoded DNA

[0353] After encoding data into a 1500 bp DNA strand by using a nanopore device in combination with the use of dual wavelength light pulses, the resulting DNA is capable of providing a decoded ("read") if the data is to be retrieved. DNA can be encoded at a multiplicity of approximately 10-100 copies, with the encoded DNA containing enough copies to allow for decoding of mixed outcomes. The DNA is sequenced by using long-read single molecule sequencing by synthesis (Pacific Biosciences). The sequence output shows that convertible bases are sequenced as expected, with reads exactly as the bases were present in the original assembly with nearly 100% fidelity (98% or better). When a "0" is encoded, the coumarinyl group is removed from N6-coumarinylmethyl-adenine, thereby resulting in the formation of adenine. Thus, an enhancement of the signal of "A" over that of N6-coumarinylmethyl-adenine at this position is found. However, the sequencing signature of O6-nitrobenzylguanine at the same bit pair is read as a mixture of G and A. At the position coded as "1", both the coumarinyl and nitrobenzyl groups are removed, resulting in an enhanced A signal at the Y position of the bit as well as an enhanced adenine signal at the X position of the same bit pair. Example 12 Stochastic or irregular data encoding

[0354] In this example, the convertible nucleobases are provided at irregular intervals along the polymer. The data-encodeable polymer includes O6-nitrobenzylguanine and O4-nitrobenzylthymine along the strand. The conversion of O6-nitrobenzylguanine to guanine can be designated as a "0" and the conversion of O4-nitrobenzylthymine to thymine can be designated as a "1". As the polymer passes through the nanopore, the data is encoded by selectively converting the appropriate convertible nucleobase according to the data code. Additionally, convertible nucleobases can be skipped to ensure that the correct code is encoded. FIG. 15 illustrates a DNA polymer before and after data encoding in which the code "1010010" is encoded. In this process, some convertible nucleobases are skipped and remain unconverted. When decoding the encoded data, only the converted nucleobases are utilized to decode the data code and the unconverted bases are ignored. When using SBS, multiple overlapping encoded DNA polymers can be utilized to decipher whether a particular nucleobase is unconverted (e.g., resulting in a read of mixed nucleobase structures) or converted (e.g., resulting in a read of a single nucleobase structure). Example 13 Constructing "writeable" DNA with modified convertible nucleobases at regular intervals

[0355] The convertible base O6-coumarinyl G (G*) is synthesized as a deoxynucleoside triphosphate derivative (dG*TP). G* acts as a polymerase substrate when the DNA template is provided to contain a complementary base such as "benzi" (see, e.g., CMN Aloisi et al., J. Am. Chem. Soc 2020, 142 (15): 6962-6969). Benzi has been shown to selectively pair with O6 alkyl G modified bases.

[0356] A circular single-stranded DNA oligonucleotide is constructed with a size of 60 nucleotides, with a single "bend" nucleotide in the sequence. The other 59 nucleotides are composed of native A, C, T, and G nucleotides. A DNA primer (20 nucleotides long) (1 μM) complementary to the non-bend region of the circle is added to a solution of the circle (1 μM) in a buffer that supports the polymerase. To induce "rolling circle" DNA synthesis, Phi29 polymerase is added with 500 uM each of the five nucleotides (dATP, dGTP, dCTP, dTTP, and dG*TP) under known and suitable conditions for Phi29 polymerase activity. After 4 hours, a solution is obtained with long repeated single-stranded DNAs of various lengths, many of which are over 10 kB in length as judged by agarose gel electrophoresis with size markers. Sequencing of the single-stranded DNA in solution confirms that the repeated sequence contains one G* base per repeat, evenly spaced over 60 nucleotides.

[0357] This solution of single-stranded DNA is converted to double-stranded form using a primer complementary to the repeated sequence, along with the four native nucleoside triphosphates and phi29 polymerase, resulting in a long solution of double-stranded DNA containing a single G* modified base every 60 bp.

[0358] This polymerase approach can be used with modified DNA bases to solve the problem of incorporating photomodifiable groups into nucleobases of DNA when the photomodifiable groups are not substrates for the polymerase enzyme.

[0359] A modification of this strategy is used to construct repeated DNA containing a second modified base. The modified base T* is synthesized as a deoxynucleoside triphosphate derivative. T* is an O4-nitrophenethyl T that contains an NPE group that can be removed using light. O4-alkyl T has been shown to pair with the opposing G by polymerase. See, e.g., MK Dosanjh et al., Carcinogenesis 1993, 14 (9): 1915-1919.

[0360] A second circular DNA is constructed that contains one bend in the sequence. In this case, there is also only one C in the sequence, 10 nucleotides away from the bend; the remaining bases are G, C, and T. Using the above DNA polymerase and primers with the same five nucleotides (dATP, dGTP, dCTP, dTTP, and dG*TP) results in a long repeated DNA that contains one G* per repeat and a single G per repeat, 10 nucleotides apart. Using a DNA primer complementary to this repeat with a polymerase and nucleotides (but not dTTP, dGTP, dATP, dT*TP, dCTP) results in the synthesis of a long repeated DNA duplex that contains one G* per repeat and one T* per repeat, 10 bp away from the G*, in the opposite strand.

[0361] In this example, it is shown that nucleotides with photoremovable nucleobases (e.g., photoremovable nucleobases that are converted to natural nucleobases after photoconversion) can be used in the presence of a polymerase to synthesize writeable DNA with photoremovable nucleobases at regular intervals. This method can utilize the polymerase for the controllable creation of longer DNA strands. The DNA created using this method is significantly longer than the DNA that can be synthesized only by ligation of synthetic oligos, such as DNA with backbone modifications. Example 14 Writing "traceless" data into DNA and reading it using long-read SMRT sequencing

[0362] A 20 kb DNA is constructed to contain two modified convertible nucleobases (X and Y) that can be converted to native DNA nucleobases upon "writing" by light irradiation. The positions of all modifications are known and are spaced in repeats with a distance of about 60 base pairs (about 20 nm) between each occurrence of a given modification. That is, X is located approximately 60 base pairs (bp) from the adjacent X and Y is located approximately 60 bp from the adjacent Y. Both modifications (X and Y) are within 10 base pairs of each other, thus a given pair or pair of X / Y is exposed simultaneously to a given localized photoexcitation event. This DNA assembly is denoted "DNA blank tape". Mixed polymerases can be used to incorporate two or more modified nucleobases into the DNA blank tape.

[0363] Nucleobase X is a guanine modified with an O-nitrophenethyl (NPE) group attached directly to O-6 without a linker or side chain. Nucleobase X can be converted to native guanine (i.e., traceless) by irradiation at 360 nm. In this example, the O-6 modified guanine is the "unwritten" ("blank") form of the nucleotide, and after successful removal by irradiation, the guanine product is considered written, with its 1 or 0 interpretation depending on the state of the nearby Y modification.

[0364] Previous studies have shown that guanines modified with alkyl groups at O-6 can be read by polymerase enzymes via sequencing by synthesis. See, for example, AM Kietrys, J. Am. Chem. Soc. 2017, 139 (47); 17074-17081. Guanines modified with alkyl groups at O-6 are typically encoded with a mixture of A and G in the majority of reads of the sequence. The quantitative percentage of encoding depends on which exact modification and which polymerase is used to read, and is determined beforehand by SMRT sequencing of synthetic DNA fragments containing the modification (calibration experiment). The matching reads give the percentage of bases encoded with this modification. For example, if we re-read the same DNA fragment, we can observe that the polymerase inserts a C (interpreted as a "G") opposite the modified base in 30% of the reads, and a T (interpreted as an "A") opposite the base in 64% of the reads. The mixed signal for this single modified base is the signal (fingerprint) of the unwritten bit. If the base in that single molecule successfully undergoes photoconversion to G, then essentially 100% of the reads will be interpreted as G.

[0365] If there are multiple copies (e.g., 1000 copies) of the same DNA molecule containing this modification at one position, and the DNA is irradiated in bulk solution with 360 nm light to the extent that the NPE group is removed in 50% of the DNA, this change will remain readable by sequencing by synthesis. The corresponding reads will average 50% between the fingerprint of the modified nucleobase (i.e., O-6 nitrophenethyl substituted guanine) and the fingerprint of the native nucleobase (i.e., guanine). Thus, the user can read the light-encoded data with less than 100% complete yield.

[0366] Similarly, in this example, nucleobase Y is a thymine modified with a coumarinyl (Coum) group at O-4. Nucleobase Y can be converted to native thymine by irradiation with light at 360 nm or 400 nm in a "traceless reaction". As with the analysis of guanine above, calibration is performed using SMRT sequencing to determine the percentage of mixed encodings that are distinct from native thymine. This percentage of mixed encodings is a fingerprint that indicates unconverted Coum-thymine, such as those present in unwritten bits. When Coum-thymine undergoes photoconversion to native nucleobase thymine (T), essentially 100% of the reads are encoded as native T. For nucleobase X, the observation that the fingerprint of the modified nucleobase Y and the fingerprint of the native nucleobase T average can be interpreted as a partial conversion of multiple copies of DNA.

[0367] In this example, a "0" bit is interpreted as when T-Coum in a G-NPE / T-Coum pair is converted to T by irradiation at 400 nm. If both modifications are removed (using irradiation at 360 nm), the bit is interpreted as a "1". Again, reading multiple copies of the data can be used to interpret bits that have been converted with less than 100% maximum yield.

[0368] Localized writing of data "bits" can be achieved using localized illumination or localized excitation methods, such as translocating a STED microscope illumination beam along the DNA, or translocating the DNA through a zero-mode waveguide or a plasmonic nanopore using methods known in the art.

[0369] Note that the blank tape DNA in this example is modified with X and Y at approximately equal intervals throughout the DNA sequence. Thus, the blank tape DNA in this example contains the potential for binary data to be written everywhere. Pairs of X, Y modified groups are simply considered to lack data (i.e., not written). The same data can be written starting at any point in the DNA (assuming there is sufficient length for the complete writing process). Since the positioning of the DNA relative to the writing light can vary stochastically and the speed of translocation can also vary, data can nevertheless be written and read by interpreting strings of 0 and 1 bits with skipping over the "blank" bits. This has the advantage that the start and end sites of writing do not need to be carefully positioned, and the translocation speed does not need to be perfectly controlled. Since no pause is required for bit positioning, this method of writing is simpler and faster than methods that work by controlling the translocation and precise positioning through the nanopore of the DNA polymer.

[0370] Data encoding the letter "e" is written onto a DNA blank tape at the single molecule level using a super-resolution microscope on a DNA molecule stretched out on a slide. The 8-bit unicode binary string for the letter "e" is 01100101 using 8 pulses of 360 nm light (1) and / or 400 nm light (0) from a super-resolution microscope with a resolution of 20 nm. The writing is performed 1000 times for 1000 single molecules and the DNA is collected by washing the slide containing the DNA at the end.

[0371] This "written" DNA is subjected to SMRT sequencing. Locations that show a fingerprint of a modified nucleobase (as a G-NPE / T-Coum pair) are interpreted as blank, with no data encoded. Paired bit positions where a match in the reads shows an averaging of the fingerprints of the modified and unmodified bases are interpreted as data; selective blocking removal of the T by removal of the NPE shows a "0", and paired bit positions where a substantial conversion of both T and G shows a written bit "1". Progress along the strand to produce the bit sequence 01100101, which shows storage of data (data conversion interpreted as the letter "e").

[0372] It should be noted that data correction can be used to correct errors if necessary. For example, if the majority of single molecule DNA copies yield a 01100101 sequence, but other binary sequences are also present, then by comparing the binary data the correct conclusion can be drawn. For example, some bits may be overlooked (e.g. 0100101) or data may be missing due to the possibility of reaching the end of the DNA (e.g. 01100). However, by comparing these different sequences the correct conclusion can be drawn even in the presence of these errors. This dual bit active writing allows the user to write more quickly than would be possible if specific positioning of the DNA was required.

Claims

[Claim 1] The invention described in the specification.