Homopolymer-encoded nucleic acid memory

Homopolymer-encoded nucleic acid memory strands address the limitations of current DNA storage by using lower-fidelity sequencing and enzymatic synthesis to achieve efficient, cost-effective, and high-density data storage.

JP2025114722AInactive Publication Date: 2025-08-05MOLECULAR ASSEMBLIES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025077652
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-24
Filing Date
2025-05-07
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current DNA digital storage methods are limited by high cost, slow synthesis, and the production of toxic by-products, with high-fidelity sequencing techniques being expensive and slow, and traditional synthesis techniques producing short strands that require post-synthesis amplification and are length-limited.

Method used

Using homopolymer tracts of repeated nucleotides to encode data, enabling the use of lower-fidelity sequencing techniques like nanopore sequencing and enzymatic synthesis of long strands, which reduces costs and synthesis time while avoiding toxic by-products.

Benefits of technology

Enables high-throughput, cost-effective DNA storage with long strands that can be synthesized efficiently using template-independent enzymes, allowing for higher data density and faster readout.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114722000097
    Figure 2025114722000097
  • Figure 2025114722000098
    Figure 2025114722000098
  • Figure 2025114722000099
    Figure 2025114722000099
Patent Text Reader

Abstract

To provide a homopolymer-encoded nucleic acid memory.SOLUTION: A nucleic acid memory chain that encodes digital data by using a sequence of homopolymer tract of a repeated nucleotide provides a more inexpensive and faster alternative method for a conventional digital DNA storage technique. Using a homopolymer tract makes it possible to read out data encoded in a memory chain by a high throughput sequencing technique of lower fidelity, that is, nanopore sequencing for example. A specialized synthesis technique makes it possible to synthesize a long memory chain capable of encoding a large volume of data despite that data density obtained by a homopolymer tract is reduced compared to a conventional single nucleotide sequence.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims priority to U.S. Patent Application No. 16 / 393,510, filed April 24, 2019, the contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present invention relates to methods and devices for storing data in nucleic acid memory strands containing homopolymer tracts. [Background technology]

[0003] background DNA digital storage is a process that uses DNA base sequences to represent digital data and stores that data through DNA synthesis of polynucleotides corresponding to the base sequences encoding the data. DNA digital storage offers several advantages over traditional data storage methods and targets a multi-billion dollar market. Traditional data storage methods, including flash memory and magnetic tape recording, pose physical space requirements, dependency on scarce resources, and issues related to data integrity. DNA digital storage offers significantly greater data storage density with significantly lower energy requirements. Current methods rely on high-fidelity sequencing techniques with little tolerance for errors to accurately read the data encoded in DNA. The required sequencing methods are relatively slow and expensive to meet the fidelity requirements. An example of a current DNA digital storage technique is described in U.S. Pat. No. 9,384,320 to Church, et al. (incorporated herein by reference). To increase sequencing fidelity, current methods, such as those described by Church, encode data using sequences that avoid features that are difficult to read or write, such as sequence repeats.

[0004] The synthesis side of current DNA digital storage techniques further limits the adoption of this technology due to a lack of speed, production of toxic by-products, and high cost. Most de novo nucleic acid sequences are synthesized using solid-phase phosphoramidite techniques, which involve sequential deprotection and synthesis of sequences constructed from phosphoramidite reagents corresponding to natural (or unnatural) nucleic acid bases. Inkjet synthesis on array-based formats allows for very low-cost phosphoramidite synthesis, but the strands produced are limited to 100–200 bases in length, must sacrifice some length for sequence indexing, and require post-synthesis amplification to provide sufficient material for subsequent readout, all at subfemtomolar scales. Using conventional synthesis techniques, nucleic acids longer than 200 base pairs (bp) undergo high rates of breakdown and side reactions. Furthermore, conventional synthesis techniques produce toxic by-products, and disposal of this waste limits the availability of nucleic acid synthesizers and increases the cost of oligo production. These formidable problems associated with synthesis and readout in DNA digital storage have limited the applications for an otherwise promising technology. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] U.S. Patent No. 9,384,320 Summary of the Invention [Means for solving the problem]

[0006] Abstract The present invention provides systems and methods for storing data using sequences of homopolymer tracts that encode digital data. Using homopolymer tracts of repeated bases (e.g., 2-10 nucleotides) to represent each bit in a data sequence allows for higher-throughput, less expensive sequencing techniques to be used. Because sequence reading relies only on distinguishing transitions between homopolymer tracts and does not require reliable reading of each individual nucleotide, sequencing techniques such as nanopore sequencing, zero-mode waveguide (ZMW) single-molecule sequencing, and mass spectrometry can be used to increase speed and reduce cost.

[0007] Recording data using the homopolymer tracts described herein is most efficiently achieved using long strands (e.g., 5-10 kb) of nucleic acid. While traditional synthesis techniques are length-limited, template-independent polynucleotide synthesis using, for example, nucleotidyl transferases, allows for the synthesis of long strands at reduced cost and with less waste production. Enzymatically synthesized ssDNA memory strands require only 50% of the DNA synthesis compared to traditional phosphoramidite approaches, because ssDNA strands longer than approximately 100-200 nucleotides require complex and costly ligation or PCR techniques, and ssDNA can only be produced from dsDNA intermediates. Efcavitch, et al., 2004, incorporated herein by reference. See U.S. Patent No. 8,808,989 to et al. Data encoding can be numerical base 2, base 3, base 4 using standard nucleotides, or data density can be increased using any number of modified nucleotide analogs to generate base 8, base 10, base 12 or higher base encoding schemes.

[0008] The only limitations on the modified nucleotide analogs are that they can be incorporated using a selected synthesis technique (e.g., terminal deoxynucleotidyl transferase (TdT)) and be distinguishable from one another using a selected sequencing analysis. In some embodiments, synthesis is performed using Mn 2+ This can be achieved using polymerase theta in the presence of

[0009] Consistent homopolymer tract lengths are not essential to the systems and methods of the present invention, since only the transitions between individual tracts need to be recognized. While tract lengths may be variable, the synthesis techniques of the present invention effectively control the average homopolymer tract length by adjusting the ratio of deoxynucleotides (dNTPs) to the oligonucleotide memory strand being synthesized and controlling the exposure time of the dNTPs to the nascent memory strand. The length of the homopolymer tract can be optimized for the readout technology; the highest data storage density is achieved with single-nucleotide readout resolution, while the highest readout speed and accuracy are achieved by extending the size of the nucleotide bit to the minimum detectable length (e.g., 2-10 nucleotides) for a given sequencing technology.

[0010] Systems and methods of the invention that use nanopore sequencing may use specialized memory strand constructs, such as stoppers (e.g., hairpins or polymeric appendages) included on one or both ends of the strand. In other nanopore-based methods, memory strands may be circularized and threaded between adjacent nanopores.

[0011] Certain aspects of the present invention include a method for recording data using a nucleic acid memory strand. The steps of this method may include creating an in-silico oligonucleotide sequence representing a data set, where each nucleotide of the oligonucleotide sequence corresponds to a unit of the data set. A nucleic acid memory strand may then be synthesized that includes a plurality of homopolymer tracts, where each homopolymer tract corresponds to a nucleotide of the oligonucleotide sequence. The plurality of homopolymer tracts may include between 3 and 10 repeated nucleotides. Each unit of the data set may be represented by a base 2, base 3, base 4, or higher base, as desired for a particular application.

[0012] In certain embodiments, the nucleic acid memory strand can be at least about 200 nucleotides long to about 5,000 nucleotides long. The synthesizing step can include controlling the homopolymer length by varying the dNTP concentration. The steps of the method can include modifying a first end of the nucleic acid memory strand to prevent passage of the first end through a nanopore of a nanopore sequencing system; passing a second end of the nucleic acid memory strand through the nanopore; and modifying the second end of the nucleic acid memory strand to prevent passage of the second end through the nanopore.

[0013] Other embodiments may use memory strands composed of heteropolymer tracts of defined stoichiometry or composition to further increase the coding capacity of a set number of nucleotide analogs. Further embodiments may aim to protect the data encoded in the memory strands by using nucleotide analogs that are similar in structure but employ linkers that can be removed under different conditions, such as ultraviolet or visible light, oxidizing or reducing agents, alkaline or acidic pH, or sequence-specific nucleases, thereby obscuring the data from them without knowledge of the applicable process.

[0014] The dataset may be selected from the group consisting of a text file, an image file, and an audio file. The synthesizing step may include template-independent synthesis. In certain embodiments, a nucleotidyl transferase enzyme may be used to catalyze the template-independent synthesis. In some embodiments, polymerase theta may be used to catalyze the template-independent synthesis.

[0015] Aspects of the present invention may include a method for reading data from a nucleic acid memory strand. The steps of the method may include sequencing a nucleic acid memory strand containing a plurality of homopolymer tracts; converting the nucleic acid memory strand sequence into digitized data, where each of the plurality of homopolymer tracts represents a nucleotide corresponding to a unit of data; and converting the digitized piece of data into a readable format. The steps of the method may include displaying the readable format. The plurality of homopolymer tracts may include between about 2 and about 10 nucleotide repeats. The nucleic acid memory strand may be at least about 200 nucleotides long and between about 5,000 nucleotides long.

[0016] In various embodiments, the sequencing step may include nanopore sequencing, sequencing-by-synthesis, or mass spectrometry. The sequencing, converting, and converting steps may be repeated one or more times on the nucleic acid memory strand.

[0017] In particular embodiments, for example, the following items are provided: (Item 1) 1. A method for synthesizing a plurality of nucleic acid memory strands, comprising: providing an array of two or more substrate-linked nucleic acids with addressable delivery of activation energy to each of said substrate-linked nucleic acids; extending one or more of the substrate-linked nucleic acids with a homopolymeric tract of two or more repeating nucleotides by delivering an addressable activation energy to one or more of the substrate-linked nucleic acids in the presence of a plurality of blocked nucleotide analogs and a template-free polymerase, wherein the template-free polymerase incorporates unblocked nucleotide analogs but not blocked nucleotide analogs, and the addressable activation energy converts the blocked nucleotide analogs to unblocked nucleotide analogs; A method comprising: (Item 2) 2. The method of claim 1, wherein the blocked nucleotide analog is converted to an unblocked nucleotide analog by removal of the blocking group. (Item 3) 2. The method of claim 1, wherein the addressable activation energy comprises light, a reducing condition, a pH change, or heat. (Item 4) 3. The method of claim 2, wherein the removable blocking group is at the 3'-OH of the blocked nucleotide analog. (Item 5) 3. The method of claim 2, wherein the removable blocking group is on a purine or pyrimidine base of the nucleotide analogue. (Item 6) The blocked nucleotide analog is a deoxyribose or 2. The method of claim 1, wherein the nucleotide analog comprises a removable blocking group on the 3'-OH of the ribose or ribosomal base and a non-removable modification on the purine or pyrimidine base of the nucleotide analog. (Item 7) 2. The method of claim 1, wherein the plurality of blocked nucleotide analogs are modified nucleotides of the same nucleobase comprising a removable 3'-O-blocking group and two or more non-removable molecular modifications that allow for differentiation between modified nucleotide analogs of the same nucleobase. (Item 8) stopping the elongation; extending the homopolymer tract with an additional homopolymer tract of two or more repeating nucleotides by delivering addressable activation energy to the homopolymer tract in the presence of another plurality of blocked nucleotide analogs and the template-independent polymerase; Item 1, the method of claim 1 further comprising: (Item 9) 9. The method of claim 8, wherein the extension is stopped after a predetermined amount of time to obtain a desired length for the homopolymer tract. (Item 10) 2. The method of claim 1, wherein the rate of extension is regulated by modifications to the blocked nucleotide analog. (Item 11) 11. The method of claim 10, wherein the rate-regulating modification is removed from the homopolymer tract after elongation. (Item 12) 2. The method of item 1, wherein the rate-regulating modification is removed during elongation. (Item 13) 2. The method of claim 1, wherein the number of repeating nucleotides in the homopolymer tract is between 2 and about 10. (Item 14) 9. The method of claim 8, further comprising repeating the terminating and extending steps to synthesize a nucleic acid memory strand. (Item 15) Item 15. The method according to Item 14, wherein the nucleic acid memory strand is about 200 to about 5,000 nucleotides in length. (Item 16) 2. The method of claim 1, wherein a predetermined concentration of the blocked nucleotide analog is provided in the extending step to obtain a desired length for the homopolymer tract. (Item 17) 9. The method of claim 8, wherein the homopolymer tract and the further homopolymer tract comprise different nucleobases. (Item 18) 15. The method of claim 14, wherein the nucleic acid memory strand encodes a data set selected from the group consisting of a text file, an image file, and an audio file. (Item 19) 20. The method of claim 18, further comprising displaying a readable format of the dataset. (Item 20) Item 15. The method according to item 14, wherein the data unit is expressed in base 2. (Item 21) Item 15. The method according to item 14, wherein the data unit is expressed in base 3. (Item 22) Item 15. The method according to item 14, wherein the data unit is expressed in base 4. (Item 23) Item 15. The method according to item 14, wherein the units of data are represented by a base greater than base 4. (Item 24) Item 15. The method according to item 14, wherein the data is represented by the degree of decaging or the obtained tract length at each step of memory strand synthesis. (Item 25) 15. The method of claim 14, wherein the data encoded in the nucleic acid memory strand is read by DNA sequencing. (Item 26) 15. The method of claim 14, wherein the data encoded in the nucleic acid memory strand is read by passage of the nucleic acid memory strand through a nanopore. Other aspects of the invention will be apparent to those skilled in the art upon consideration of the following figures and detailed description. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 shows the enzymatic synthesis cycle used to form the homopolymer tract.

[0019] [Figure 2] FIG. 2 illustrates a method for synthesizing a nucleic acid memory strand having a homopolymer tract.

[0020] [Figure 3] FIG. 3 illustrates a method for reading data from a nucleic acid memory strand having a homopolymer tract.

[0021] [Figure 4] Figure 4 shows the relationship between chains / GB and homopolymer length and chain length.

[0022] [Figure 5] FIG. 5 shows the data that can be encoded in a DNA strand composed of 500 polymer tracts as a function of the number of distinguishable nucleotide analogs in the memory strand.

[0023] [Figure 6] FIG. 6 shows a nanopore-trapped nucleic acid memory strand of the present invention.

[0024] [Figure 7] FIG. 7 shows a scheme for varying the data encoded within a chain of homopolymer tracts in response to processing conditions.

[0025] [Figure 8] FIG. 8 shows a system for synthesizing nucleic acid memory strands having homopolymer tracts.

[0026] [Figure 9] FIG. 9 shows a system for parallel synthesis of nucleic acid memory strands with homopolymer tracts on an array of nanowells.

[0027] [Figure 10] FIG. 10 provides a more detailed schematic of the components that may appear in the system.

[0028] [Figure 11] FIG. 11 shows an analysis of the enzymatically mediated synthesis of homopolymer tracts of two different base compositions.

[0029] [Figure 12] FIG. 12 shows the process for converting character text into a 12-bit nucleic acid sequence.

[0030] [Figure 13] FIG. 13 shows polyacrylamide gel electrophoresis (PAGE) analysis of 12 cycles of enzymatic synthesis.

[0031] [Figure 14A] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14B] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14C] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14D] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14E] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14F] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14G] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14H] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14I] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14J] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14K]Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition. [Figure 14L] Figures 14A-L show the homopolymer distributions found experimentally for each of the 12 bits of homopolymer addition.

[0032] [Figure 15] FIG. 15 illustrates the cycle steps for template-independent DNA polymerase-mediated synthesis of homopolymer-encoded information polymers.

[0033] [Figure 16] FIG. 16 shows a side view of a typical reaction zone in a 2D array of reaction zones, showing the hydrophobic patterning that isolates the liquid reaction droplets.

[0034] [Figure 17] FIG. 17 shows a side view of a reaction zone with a controlled heating element or electrochemical element.

[0035] [Figure 18] FIG. 18 shows the reaction zone showing the light-transmitting top cover.

[0036] [Figure 19] FIG. 19 shows an exemplary device useful for addressable photoactivation of enzymatic DNA synthesis.

[0037] [Figure 20] FIG. 20 shows the kinetics of UV light decaging of 3′-O-(2-nitro)-benzyl dATP.

[0038] [Figure 21]Figure 21 shows optically controlled oligonucleotide elongation. Enzyme reaction mixtures containing 3'-orthonitrobenzyl dATP were illuminated with a low-power light source over various time intervals, such that the exposure time controlled the amount of decaging. The amount of available natural dATP formed then controlled the average length of the polynucleotide tract.

[0039] [Figure 22] FIG. 22 shows the structures of 3′-O-caging modifications that enable bichromatic light-mediated decaging.

[0040] [Figure 23] FIG. 23 shows gel electrophoresis analysis of each of 12 cycles of homopolymer synthesis using 3'-O-(2-nitro)-benzyl dATP and dCTP, dGTP, dTTP.

[0041] [Figure 24] FIG. 24 shows an electrophoretic analysis of an enzymatic synthesis reaction comparing incorporation-modulated dNTP analogs with unmodified dNTPs.

[0042] [Figure 25] Figure 25 shows electrophoretic analysis of two successive cycles of incorporation rate-controlling dNTP analogs. The modified analog was used to form an N+1 homopolymer, followed by the addition of an N+2 homopolymer. The homopolymer synthesis rate-controlling modifications were removed from the oligonucleotides prior to gel analysis.

[0043] [Figure 26] 26 shows an electropherogram of a TdT extension reaction showing homopolymeric rate-controlling analog incorporation followed by a second cycle of homopolymeric rate-controlling analog incorporation, where the modified nucleotide is followed by consecutive modified nucleotides.

[0044] [Figure 27] FIG. 27 shows the modified dNTP analogs used in the composition of the information polymer.

[0045] [Figure 28] FIG. 28 shows an electrophoretic analysis of information polymers composed of modified nucleotides.

[0046] [Figure 29] FIG. 29 shows the detection of an information polymer (P71) composed of modified nucleotides by translocation through a nanopore.

[0047] [Figure 30] FIG. 30 shows the detection of an information polymer (86B) composed of modified nucleotides by translocation through a nanopore. DETAILED DESCRIPTION OF THE INVENTION

[0048] Detailed Description The present invention provides systems and methods for writing data to and reading data from nucleic acids having homopolymer tracts corresponding to digital data units. By repeating each nucleotide in the data-encoding sequence (e.g., 3-10 times), only the transitions between homopolymer tracts need to be observed in the sequencing read, enabling lower-fidelity, higher-throughput sequencing techniques that may result in cheaper implementation of nucleic acid data storage. The advantages of synthesizing nucleic acid homopolymer tract memory strands are: 1) the ability to create very long (5-10 kb) strands, enabling the use of high-throughput, long-read DNA sequencing techniques for readout; 2) the ability to tolerate errors in the sequencing readout technique; and 3) the ability to create nucleic acid memory strands at a much lower cost than traditional chemical synthesis methods. The use of homopolymeric nucleic acid memory strands is best realized in long (e.g., 5-10 kb) strands that can be efficiently produced using template-independent TdT enzymes or polymerase theta, where the homopolymer tract length can be controlled by varying the exposure time and the ratio of dNTPs to polynucleotide memory strands.

[0049] Enzymatic synthesis of homopolymers for encoding data is easily achieved by using non-terminator natural or modified nucleotide triphosphates, enabling the simplest and most rapid method of DNA synthesis. One natural or modified nucleotide triphosphate is delivered to a reaction zone containing a nucleotidyl transferase, where the reaction occurs, and then removed by washing with a buffer solution to complete a single "write" cycle of data storage, as illustrated in Figure 1. Data strand synthesis occurs in a completely aqueous environment without toxic or harmful chemicals, thus enabling a practical device suitable for large-scale data storage centers.

[0050] FIG. 2 illustrates a method 101 for synthesizing a nucleic acid memory strand having homopolymer tracts according to certain embodiments. The method 101 includes step 103 of creating an in-silico oligonucleotide sequence representing a dataset. The dataset may include digitized data that may represent text, images, video, audio, or any other piece of information that can be digitized. The oligonucleotide sequence may include any number of natural or modified nucleotides or their analogs, and depending on the number of unique nucleotides or analogs used in the memory strand, the dataset may be encoded using a base-2, base-3, base-4, or higher base scheme. In a simple embodiment, the encoding scheme may correspond to a binary data scheme conventionally represented by a series of 0s and 1s, where one or more nucleotides or analogs may correspond to 0 and one or more other nucleotides may correspond to 1. A nucleic acid memory strand (e.g., RNA, single-stranded, or double-stranded DNA) may then be synthesized 105, comprising a series of homopolymer tracts that each correspond in order to a nucleotide in the in silico oligonucleotide sequence. In certain embodiments, further steps include step 107 of modifying one end of the memory strand, step 109 of passing the memory strand through the nanopore, and step 111 of modifying the other end of the strand to prevent the end from passing through the nanopore.

[0051] 3 shows a method 203 for reading data from a nucleic acid memory strand having homopolymer tracts. The steps of method 203 include sequencing a series of homopolymer tracts in the nucleic acid memory strand in step 203, converting the sequence into a data set in step 205, converting the data set into a readable format (e.g., an image, video, audio clip, or piece of text), and displaying the readable formatted data as needed (e.g., on a monitor or using a printer or other input / output device) in step 209.

[0052] Preferably, the systems and methods of the present invention use long strands of DNA (5-10 kb), which can be either single-stranded or double-stranded and can occur naturally or can be produced by chemical or enzymatic synthesis. In certain embodiments, nucleic acid memory strands can be enzymatically generated using TdT to create a series of homopolymer tracts that can be 2-10, 3-10, 4-10 nucleotides, or longer. Each homopolymer tract can consist of adenine (A), guanine (G), cytosine (C), or thymine (T). The sequence of alternating homopolymer tracts can be used to encode data to be stored in the memory strand.

[0053] Each nucleotide homopolymer tract can represent various amounts of data, depending on the number of bases used. The number of bits required to compose one byte (256 decimal) is defined by the following relationship: #bits / byte = 8 / (log2(n)), where n = the base used. Each tract may correspond to one bit if two-base encoding is used, or one-quarter of a byte if four-base encoding is used. In certain embodiments, DNA data strands may be composed of two- to ten-nucleotide homopolymer tracts (using a base-2 dataset representation), allowing 333 to 100 bits to be represented in memory strands between 999 and 1000 bases long. In preferred embodiments, the nucleobase encoding of data may be such that a single homopolymer tract of one nucleotide is always adjacent to a homopolymer tract of a different nucleotide. For example, the encoding may be such that an adenine homopolymer tract is never immediately preceded or followed by another adenine homopolymer tract. When two adjacent homopolymer tracts contain the same nucleotide, homopolymer tracts that are measurably longer than the average homopolymer tract representing a single nucleotide in the encoded data sequence can be synthesized. These longer tracts can be created through manipulation of the synthesis reaction described below, for example, by increasing the concentration of dNTPs in the reaction or increasing the reaction time. The exact length of the two adjacent identical homopolymer tracts only needs to be long enough to be uniquely distinguished from a single homopolymer tract using a readout device (i.e., a nanopore sequencer). In certain embodiments, a non-nucleotide homopolymer spacer can be added between the A, G, C, or T homopolymer tracts to clearly distinguish adjacent identical nucleotide homopolymer tracts from each other.The use of A, G, C, and T homopolymer tracts, rather than simply using two nucleotides (similar to 0 and 1 in binary code), allows for the creation of a 4-bit encoding space, increasing the density of data that can be stored in one continuous strand. For example, four consecutive homopolymer tracts can encode 256 numbers (i.e., one byte) when A, G, C, and T are used in a base-4 scheme. In such an embodiment, 83 or 25 bytes are represented in a 996 or 1000 nucleotide long nucleic acid memory strand when using 3 or 10 nucleotide long homopolymer tracts, respectively.

[0054] In various embodiments, base 8 or even base 12 coding schemes can be used through the incorporation of homopolymeric tracts of uniquely modified nucleotides or non-nucleotide analogs into the memory strand. These modified nucleotides or non-nucleotide analogs should generate a unique digital signal in a readout device such as a nanopore sequencer or a single-molecule ZMW sequencer. TdT, discussed below, can enhance the signal provided by a readout device such as a nanopore and thus incorporate a wide range of modified dNTP analogs that can be used to generate nucleic acid memory strands with data encoded therein. Modified nucleotides (e.g., A * , G * , C * and T * or A ** , G ** , C ** and T ** ) can be combined with TdT and modified dNTP analogs of each of the four bases (e.g., dA) to generate 8-bit or even 12-bit encoding schemes. * TP or dA **TP). Encoding of higher bases (n) allows for data compression, resulting in a reduction in the number of DNA strands required to encode a given amount of information. As illustrated in Figure 4, the relationship determining the number of DNA strands per GB of data as a function of homopolymer tract length, base (n), and synthesized strand length is defined by: # strands / GB = (8 / (log2(n)) * 10 9* Homopolymer Length * 1 / chain length.

[0055] The number of unique homopolymer tracts may only be limited by the ability of the readout technology (i.e., nanopore or ZMW single-molecule sequencing) to determine one homopolymer tract from another. Several reports exist in the literature of detecting homopolymers composed of unmodified nucleotides by detecting changes in ionic current during translocation through a nanopore (Venta et al., 2013; Feng et al., 2014). al, 2015). The modifications that alter the DNA residence time in the nanopore are identifiable and distinctive. Singer et al. (2010) and Morin et al. (2016)

[10] used non-covalently linked bis-PNA or gamma-PNA functionalized with 5 kDa or 10 kDa PEG to enhance nanopore detection.

[11] Liu et al.

[12] (2015) Adamantly 8-oxo-1, 2-dimethyl-1, 2-dimethyl-2 ...1, 2-dimethyl-1, 2-dimethyl-1, 2-dimethyl We selectively created nucleotide analogs of dATP and dCTP. Given TdT's tolerance for incorporating bulky modifications at N6 of dATP, N4 of dCTP, N2 or O6 of dGTP, and O4 or N3 of dTTP, acyl or alkyl modifications at these positions can be screened and selected to enhance the detection modalities of nanopore or ZMW single-molecule sequencing techniques. Detection can be improved through modified nucleotides that enhance differential current blockade in nanopores or enhance the residence time of modified nucleotides in the active site of DNA polymerase in ZMW single-molecule approaches. Other natural and unnatural purine and pyrimidine nucleotide analogs can be used if they generate unique digital signals on readout devices such as nanopore or single-molecule ZMW sequencers. Modifications at C5 or C7 of pyrimidines and purines, respectively, can be used if they generate unique digital signals on readout devices such as nanopore or single-molecule ZMW sequencers. Suitable modified nucleotide triphosphates are selected to be rapidly incorporated during the enzymatic extension step, providing substitution-specific dwell times with the shortest possible homopolymer length during the detection step. Examples of modifications to A, G, C, and T bases suitable for expanding the bit-encoding space include, but are not limited to, N6-benzoyl dA, N6-benzyl dA, N6-alkyl dA, N6-acyl dA, N6-substituted alkyl dA, N6-substituted acyl dA, N6-arylacyl dA, N6-substituted arylacyl dA, N2-alkyl-dG, N2-acyl dG, N2-arylacyl dG, N2-substituted alkyl dG, N2-substituted acyl dG, N Included may be 2-substituted arylacyl dG, O6 alkyl dG, O4 alkyl dT, N3 alkyl dT, N3 acyl dT, O6 substituted alkyl dG, O4 substituted alkyl dT, C5-propargylamine dT, C5-propargylamine dC, C7-propargylamine dA, C7-propargylamine dG, substituted C5-propargylamine dT, substituted C5-propargylamine dC, substituted C7-propargylamine dA, substituted C7-propargylamine dG.Preferred embodiments of substitutions include, but are not limited to, covalent bonds that are completely stable to removal under all but the most extreme chemical conditions of pH, temperature, and concentration of reactive species. Substitutions that can affect unique current blockages can include, but are not limited to, alkyl, heteroatom-substituted alkyl, aromatic hydrocarbon, alkyl-substituted aromatic hydrocarbon, heteroatom-substituted alkyl-substituted aromatic hydrocarbon, heteroatom-substituted aromatic hydrocarbon, benzyl, substituted benzyl, or combinations thereof. In some embodiments, the substitution can be polyethylene glycol composed of 2 to 450 monomer units. In some embodiments, substitutions composed of peptides or peptoids can be appropriate to increase the residence time of homopolymers in a specific and discriminable manner. The efficiency of incorporation of modified nucleotides by template-independent polymerases such as TdT can be tuned by the use of different metal ion cofactors, such as, but not limited to, Co++, Zn++, Mg++, Mn++, or a mixture of two or more different metal ions. Each modified nucleotide may require a different metal ion for optimal performance during enzymatic homopolymer synthesis.

[0056] Because long-term stability of DNA data strands is important, there is a distinct advantage to using non-purine-based homopolymers because they are subject to depurination at low pH. In some embodiments, a homopolymer bit can be composed of only a single nucleotide type (i.e., thymine) modified with two, three, four, or more different chemical groups, resulting in a homopolymer tract that each produces a unique current block. Thus, one nucleotide labeled with four unique modifiers can replace the occurrence of A, G, C, or T. Other embodiments are possible that use only one of the other three nucleotides with two, three, four, or more different chemical groups.

[0057] Another embodiment uses a single nucleotide bit instead of a homopolymer bit, where the modified nucleotide analog is sufficient to create a unique residence time for the passage of a single nucleotide through the nanopore. A single modified nucleotide bit is advantageous in that it allows for the greatest density of information per DNA strand, thus reducing the cost of DNA-based data storage.

[0058] The exact length of the homopolymer tract is not important, as long as the sequencing technology used for readout can clearly distinguish the start and end of one homopolymer tract from another. While there are clear synthesis and storage density advantages to increasing the number of unique nucleotides or bases (including modified nucleotides or non-nucleotide analogs) used and thus reducing the length required to capture a set amount of data, the lowest cost per DNA data storage synthesis can be achieved by using the four natural nucleotide dNTP monomers during enzymatic synthesis, because these reagents are widely used in the fields of molecular biology and sequencing and are produced in very large batches with the lowest manufacturing costs. The cost of producing dNTP analogs to increase the number of unique homopolymer tracts can be reduced as the use of DNA data storage increases and the scale of analog manufacturing also increases.

[0059] While any method for synthesizing homopolymer tract segments can be used with the systems and methods of the present invention, preferred embodiments use the template-independent enzyme TdT. TdT offers certain advantages in that it rapidly and inexpensively generates homopolymers with a Poisson distribution, where the average size of the homopolymers can be tightly controlled by the ratio of [dNTP] to nascent oligonucleotide memory strands. In some embodiments, Mn 2+Polymerase theta in the presence of can be used as a template-independent polymerase to synthesize homopolymer tract nucleic acid memory strands. In another embodiment, the length of the homopolymer tract segment can be controlled by delivering an excess of dNTPs to the reaction zone and then removing the reactants after a carefully controlled time interval.

[0060] TdT has demonstrated the ability to synthesize homopolymer tracts of precisely defined lengths by controlling the ratio of dNTP concentration to the concentration of the 3' end of the nucleic acid strand desired to be modified. Inkjet synthesis on an array-based format allows for very low-cost phosphoramidite synthesis, but the strands produced are limited to 100–200 bases in length, require a partial sacrifice of length for sequence indexing, and require post-synthesis amplification to provide sufficient material for subsequent readout. These are produced at subfemtomolar scales, making them primarily suitable for relatively low-efficiency short-read sequencing techniques.

[0061] Single-stranded DNA strands synthesized according to the process of the present invention may benefit from the prevention of hairpins or dsDNA either during synthesis or during readout. Hairpin formation can be prevented by modifying the exocyclic amine of one member of an A:T or G:C base pair to prevent the hydrogen bonding required for base pairing. In some embodiments, the exocyclic amine can be modified by acylation or alkylation. Any simple and stable modification of the exocyclic amine of A, G, or C that prevents base pairing can be used to prevent hairpin formation. In certain embodiments, the N6 of deoxyadenosine and the N2 of deoxyguanosine can be acetylated with an acetyl group to prevent base pairing. In some embodiments, the N6 of deoxyadenosine and the N4 of deoxycytidine can be modified to prevent base pairing and hairpin formation. In some embodiments, the O6 of deoxyguanosine or the O6 of deoxythymidine can be modified to prevent base pairing and hairpin formation. In some embodiments, the O4 or N3 of deoxythymidine can be modified to prevent base pairing and hairpin formation. In some embodiments, modifications to A, G, C, or T to generate higher-order base encoding schemes also serve the purpose of preventing base pairing and hairpin formation. In some embodiments, homopolymer bits can be composed of only a single nucleotide type (i.e., thymine) modified with two, three, four, or more different chemical groups, resulting in a unique current blockade and preventing the formation of intrastrand or interstrand duplex regions. In another embodiment, a thermostable version of TdT or another template-independent nucleotidyl transferase can be used to perform strand synthesis at elevated temperatures, thus preventing the formation of intrastrand or interstrand duplex regions.

[0062] Control of homopolymer tract length can be optimized for any of the dNTP analogs after determining and calibrating their incorporation rates to create a reproducible range of homopolymer tract lengths from 2 to 10 nucleotides in length. * , G * , C * and T* and A ** , G ** , C ** and T ** The use of homopolymer tracts allows for the creation of 8-bit or 12-bit encoding, increasing the density of data that can be stored in one continuous strand beyond simply using two nucleotides to encode a "0" and a "1." Three consecutive homopolymer tracts are A, G, C, T, A. * , G * , C * and T * can encode 256 numbers (i.e., 1 byte). In such an embodiment, if a 3-nucleotide or 10-nucleotide homopolymer tract is used, there are 111 or 33 bytes in a nucleic acid memory strand that is 999 or 990 nucleotides long. Two consecutive homopolymer tracts can be A, G, C, T, A. * , G * , C * , T * , A ** , G ** , C ** and T ** can encode 256 numbers (i.e., 1 byte). In these embodiments, there are 166 bytes or 50 bytes in a 996 or 1000 nucleotide long nucleic acid memory strand when using a 3 nucleotide long homopolymer tract or a 10 nucleotide long homopolymer tract, respectively.

[0063] In certain embodiments, data can also be encoded into memory strand heteropolymer tracts of random sequence and defined composition to achieve higher levels of data compression. Heteropolymer stretches can be generated using enzymatic reactions that use a mixture of different dNTPs, where dNTP stoichiometry is used to control the composition of the heteropolymer tract. The number and type of heteropolymer tracts are limited only by the combination of dNTP analogs and the ability of the detection modality to distinguish the composition of different tracts. For m dNTP analogs, the number of (m) for heteropolymer formation is 2 There are (-m) / 2 binary combinations. A detection modality that can distinguish between two different levels of tract composition for each binary combination (e.g., tracts where analogs A and B are present in a ratio of approximately 2:1, and tracts where they are present in a ratio of 1:2) can be obtained by detecting cardinal m from a set of m analogs. 2 This allows data to be encoded at a rate of 0.01, effectively doubling the coding capacity of the memory strand. Figure 5 illustrates the data that can be stored in a memory strand as a function of the number of available dNTP analogs using either a homopolymer- and a binary heteropolymer-based encoding scheme with two levels of tract composition.

[0064] The data-encoding strands of the present invention may not necessarily require precisely defined homopolymer lengths, since they need only be long enough (approximately 2–10 nucleotides) to allow unambiguous differentiation of the transition between homopolymer tract segments by high-throughput DNA sequencing techniques. Existing next-generation sequencing-by-synthesis (SBS) systems can easily determine the transition between two adjacent homopolymer tracts. Again, the exact length of the homopolymer tract is not critical for accurate detection of the homopolymer bit. The use of tracts of the same nucleotide offers the advantage of overcoming the most common errors in current SBS platforms: insertions and deletions. A deletion of one nucleotide in a homopolymer tract longer than 2 nt is still interpreted as a true homopolymer. Similarly, a single-nucleotide insertion in a homopolymer tract will not be misinterpreted as two adjacent homopolymers, because insertion of more than one nucleotide during SBS is an unlikely event. This sequencing error tolerance offers the advantage of reducing the sequencing depth required to ensure correct decoding of the information stored by the DNA data strand. Existing nanopore systems can easily distinguish homopolymer tracts of A, G, C, or T from one another based on their differential current blockade. In certain embodiments, single-molecule ZMW sequencing can be used to determine the linear order of homopolymer tracts on a linear strand. The use of either sequencing technique may require the DNA initiator to have properties compatible with the sequencing readout technique, such as a self-complementary hairpin at the 5' end of the synthesized single-stranded memory strand to provide a primer for single-molecule ZMW sequencing. Nanopore sequencing techniques may also require a self-complementary hairpin at the 5' end of the strand to provide a "start" data mark. In various embodiments, the readout technique can be any next-generation sequencing method, such as those provided by Illumina (San Diego, CA). In some embodiments, the readout or sequencing technique can be mass spectrometry-based.The technique-specific error rate of the readout technique is not important as long as the technique can unambiguously detect the transition between two different homopolymer tracts and / or, in the case of two identical homopolymer tracts adjacent to each other, the difference between one homopolymer tract length and one 2x length.

[0065] Certain readout techniques may be preferred over others based on the specific application of the present invention. Techniques such as nanopore sequencing may be non-destructive, leaving the nucleic acid memory strand intact and suitable for multiple readout cycles. Readout techniques that rely on sequencing-by-synthesis (SBS), such as ZMW single molecule, may generate a copy of the original template strand and require a post-readout operation (i.e., strand separation by melting) to remove the complementary strand and return the original nucleic acid memory strand to its initial state ready for subsequent cycles of readout. Other readout techniques, such as mass spectrometry, are destructive and deplete the pool of nucleic acid memory strands after repeated cycles of sampling and readout.

[0066] In various embodiments, the nucleic acid memory strand may include a "stopper." The "stopper" may be a polymeric construct that prevents the passage of single-stranded or double-stranded nucleic acids through the nanopore (Manrao, et al., 2012, Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase, Nature Biotechnology 30, 349-353, incorporated herein by reference). Proteins such as enzymes are large enough to avoid being pulled through the larger pore (approximately 6.3 nm) on the cis side of the protein nanopore. The diameter of the smaller pore is estimated to be approximately 1.2 nm wide. A stopper can be used on either the 5' or 3' end of the nucleic acid memory strand of the present invention. In some applications, it may be desirable to have a stopper at both the 3' or 5' end of the nucleic acid molecule. The stopper can consist of a hairpin (stem-loop) structure with either a protruding 5'-overhang or a 3'-overhang to which an information-encoding nucleic acid is covalently attached. When the stopper consists of a hairpin, the length of the ds stem can be long enough to resist any melting force exerted on it by the electric field used to translocate the memory strand through the nanopore. In certain embodiments, one base in the double-stranded stem region can be crosslinked to its cognate base with which it forms a base pair, such that the double-stranded stem portion of the hairpin cannot melt under the influence of the force exerted on it by the electric field that translocates the remainder of the molecule through the nanopore. To utilize a hairpin stopper for TdT-mediated nucleic acid memory synthesis according to certain embodiments, the stopper can have a 3'-overhang of sufficient length (i.e., >10 nucleotides) to allow binding of TdT for template-independent synthesis.

[0067] The stopper can consist of a non-nucleotide polymeric construct that can be added to either the 3' or 5' end of the nucleic acid molecule. The construct can be attached by direct conjugation of a polymeric species onto the 3' end of the nucleic acid by a polymerase or a transferase such as TdT (Sorensen, et al., 2013, Enzymatic Ligation of Large Biomolecules to DNA, ACS Nano, 7(9):8098-8104, incorporated herein by reference), or by incorporation of a functionalized nucleotide that allows specific modification of the nucleic acid via its functionality (Winz, et al., 2015, Nucleotidyl transferase assisted DNA labeling with different click chemistries, Nucleic Acids, incorporated herein by reference). Acids Res. 43(17):e110). 5'-end stoppers can be easily introduced at the time of chemical synthesis of the oligonucleotide adapter and used as initiators, either through direct synthesis of a hairpin or through secondary modification of a functional handle introduced as the final step of 3'-to-5' oligonucleotide synthesis. Alternatively, 5'-end stoppers can be constructed by attaching an oligonucleotide initiator via its 5' end to a magnetic or non-magnetic bead, particle, or nanoparticle, enzymatically synthesizing a memory strand containing a homopolymer tract, and then leaving the memory strand attached to the magnetic or non-magnetic bead, particle, or nanoparticle.

[0068] The stopper can be further modified to allow cleavage of the stopper from the rest of the molecule to allow the nucleic acid strand to passively diffuse out of the nanopore, or to allow the nucleic acid strand to be translocated out of the nanopore via application of a voltage, thus allowing the strand to be retrieved.

[0069] In certain embodiments, template-independent polymerases or transferases can be used to modify pre-synthesized strands of nucleic acid to enable the use of nanopore devices as "write-once, read-many" memory devices. Part of the inherent problem associated with using nanopore devices as DNA sequencers is the high error rate they generate due to the nanopore's insufficient discrimination. This may be due to the speed of translocation through the pore or the fact that the nanopore's approximate depth is 8 nm, allowing multiple bases to be present in the pore simultaneously. The homopolymer memory strands of the present invention address this issue through the use of homopolymer repeats, reducing the need for strict sequencing accuracy. In certain embodiments, the drawbacks of nanopore sequencing can be addressed by implementing a hairpin adapter at one end of the double-stranded DNA memory strand, so that during the translocation and base-calling process, each sense of the DNA memory strand can be read, such that reading each individual base and its complementary strand can offset the error rate of reading each base only once. In certain embodiments, the fidelity of nanopore sequencing can be increased by appropriate modification of each end (5' and 3') of a single- or double-stranded nucleic acid molecule with bulky appendages (e.g., proteins or solid-state materials) that do not translocate across the pore. The molecule can then be trapped within the pore and translocated forward and backward multiple times to allow multiple reads of the same molecule in the same pore, thus reducing the sequencing error rate by the square of the number of reads (if sequencing read errors are due to stochastic origins).

[0070] Transferases such as TdT can be used to append large, bulky modified nucleotide analogs to the 3' end of a DNA molecule. In certain embodiments, a WORM nanopore memory device can be generated using the following steps: (1) generating a single molecule of DNA that encodes specific information in any of the high-density encoding schemes discussed above and covalently modifying the 5' end with a bulky molecular construct that prevents complete translocation of the DNA molecule through the nanopore; (2) threading the DNA molecule through the nanopore until the 5'-modified end contacts the nanopore and cannot be further translocated; (3) using TdT and the modified nucleotide to effectively trap the molecule within the torus of the nanopore. (4) covalently attaching one (or more) bulky nucleotide analogs ("stoppers") to the 'end; (5) reversing the polarity of the current to the nanopore to remove any DNA molecules that are not 3'-modified, thus creating a pure population of "trapped" (5'- and 3'-modified) nucleic acid strands; (6) using an applied voltage to read the "trapped" DNA strands in either or both directions (potentially reading multiple times to reduce the error rate to an acceptable level). In various embodiments, step 6 can consist of a voltage-induced "read" in one direction and rapid translocation in the opposite direction to "unwind" the data-encoding nucleic acid through the nanopore, followed by another voltage-induced "read" in the original direction. This cycle of "read"-"unwind"-"read" can be repeated as many times as desired.

[0071] In some embodiments, the trapped nucleic acid strand can be read during translocation in either direction. In some embodiments, the trapped strand can be translocated to one end of the molecule (either 5' or 3') and read in the opposite direction, so the polarity of the read can provide a more accurate read.

[0072] In certain embodiments, circularized nucleic acid memory strands can be produced using the synthesis method described above followed by circularization. The circularized strands can contain bulky polymers or specific homopolymer sequences, where the ends of the synthesized strands are spliced to designate start and stop points for data reading. Start and stop homopolymer sequences can also be used in linear nucleic acid strands. The circularized strand 305 can be threaded between two adjacent nanopores (306 and 309) located on a single membrane 303, as shown in FIG. 6, so that the circularized strand 305 is physically trapped between the two nanopores (306 and 309). The circularized memory strand can encode digital information as a sequence of single nucleotides, a homopolymer tract sequence, a sequence of modified nucleotide analogs, or some combination thereof. One nanopore 309 may be used to generate an electrical signal as the information-encoding memory strand is translocated through the pore, while the other nanopore 307 may simply function as a portal to allow the DNA molecule to return to the cis side of the membrane 303 and the first nanopore 309. One advantage of this scheme is that the information-encoding strand may be recycled for repeated reading and may be read multiple times, thus reducing any possible readout errors.

[0073] Other embodiments may encode data in memory strands so that the data can only be accessed under a specific set of conditions. In such cases, the memory strands are at least partially composed of nucleotides containing modifications attached to cleavable linkers. The modifications (e.g., chemical protecting groups) and linkers may be selected so that if the polymer tract translocates through the nanopore without proper processing, the current blockade will be different from the sequence encoding the data. Figure 7 outlines a scheme using disulfide and amide-linked modifications to dG nucleotides and illustrates how the current blockade and the data encoded in the memory strands can change in response to processing conditions. * and G ** are structurally similar in size and flexibility and can produce similar current blockade on nanopore platforms, but are removed under different conditions. * and G *** Although structurally distinct, these modifications share the same removal conditions. Other embodiments may use other modifications or linkers that are cleavable using different treatments, such as specific wavelengths of light, acidic or alkaline pH, oxidative or reductive conditions, or sequence-specific nucleases. Some embodiments may use the presence or absence of memory strand modifications for encryption or as a chemical marker of previous access or modification of data. Most linker cleavage reactions are effectively irreversible; therefore, this approach is best suited to write-one-read-many systems, where a single molecule may be sufficient to encode data without redundancy.

[0074] Many possible information encoding schemes useful in the readout scheme of the present invention are possible and would be apparent to one of ordinary skill in the art based on this disclosure.

[0075] Synthesis can be achieved using acoustic delivery of droplets into the wells of a plate (e.g., a 1536-well plate of 1.5 μL each). In various embodiments, the nucleic acid memory strands can be synthesized on beads or magnetic beads or surfaces, and, depending on the application, can be left on the beads or magnetic beads or surfaces or removed from the synthesis support after full-length synthesis is complete.

[0076] In certain embodiments, a system for synthesizing long (5-10 kb) data strands may use inkjet delivery to an array of wells (e.g., multiple nanoliter-volume wells). In other embodiments, multiple pneumatically controlled actuators may be positioned above each well to simultaneously deliver reagents to each location in the array. Each actuator is served by a selector valve that selects between two or more nucleotides or modified nucleotides formulated with a template-independent polymerase used to identify bits in the DNA data strand. One or more additional selector valve ports are optionally dedicated to one or more wash reagents. The array of nanoliter-volume wells may be open at both ends, as long as the well diameter is such that the delivered liquid is trapped by capillary action within the length of the open-ended wells. After each round of nucleotide-enzyme formulation is delivered to the open-ended wells, each well is rinsed with the reaction mixture, and a rinse or enzyme quenching reagent may be flowed across and through the lower opening of the array of wells to prepare the array for the next cycle of enzymatic synthesis. In other embodiments, a vacuum source is used to rapidly remove one reagent from the capillary nanowell prior to delivery of the next reagent.

[0077] Certain embodiments may use highly parallel nanofluidic chambers with valve-controlled reagent delivery. An exemplary microfluidic nucleic acid memory strand synthesis device is shown in FIG. 8 for illustrative purposes, not to scale. A microfluidic channel 255 containing a regulator 257 couples a reservoir 253 to a reaction chamber 251, and an outlet channel 259 containing a regulator 257 removes waste from the reaction chamber 251. A microfluidic device for nucleic acid memory strand synthesis may include, for example, channel 255, reservoir 253, and / or regulator 257. Nucleic acid memory strand synthesis may occur in a microfluidic reaction chamber 251 that may contain several anchored synthesized nucleotide initiators anchored or bound to the inner surface of the reaction chamber, which may optionally include beads or other substrates capable of releasably binding polynucleotide initiators. The reaction chamber 251 may include at least one intake channel and one outlet channel 259 so that reagents can be added to and removed from the reaction chamber 254. The reaction chamber 251 must be temperature-controlled to maintain optimal and reproducible enzymatic synthesis conditions. The microfluidic device may contain a reservoir 253 for each respective dNTP or analog used in the memory strand coding scheme. Each of these reservoirs 253 may also contain an appropriate amount of TdT or any other enzyme that extends DNA or RNA strands in a template-independent manner. Additional reservoirs 253 may contain reagents for washing or other tasks.

[0078] Reservoirs 253 may be coupled to reaction chambers 254 via separate channels 255, and the flow of reagents into reaction chambers 254 via each channel 255 may be individually regulated through the use of gates, valves, pressure regulators, or other means. Flow from reaction chambers 254 via outlet channels 259 may be similarly regulated. Reservoirs 253 may hold dNTPs, modified dNTPs, or any of the above analogs suspended in a fluid at known concentrations, such that the concentration of the reagents can be precisely controlled based on the volume of reagent flowed into reaction chambers 254. Thus, the length of each homopolymer tract can be managed through control of the reagent concentrations.

[0079] In certain examples, reagents, particularly dNTP and enzyme reagents, may be recycled. Reagents may be drawn back from the reaction chambers 254 into their respective reservoirs 253 through the same channel 255 they entered by inducing backflow using gates, valves, vacuum pumps, pressure regulators, or other regulators 257. Alternatively, reagents may be returned from the reaction chambers 254 to their respective reservoirs 253 through a separate return channel. The microfluidic device may include a controller capable of operating the gates, valves, pressure, or other regulators 257 described above.

[0080] An exemplary microfluidic nucleic acid memory strand synthesis reaction may include the following steps: flowing the desired dNTP (used throughout to refer to any component molecule used to encode data in the nucleic acid memory strand of the present invention) reagent into the reaction chamber 254 at a predetermined concentration (calculated to produce the desired homopolymer length) for a predetermined amount of time before removing the NTP reagent from the reaction chamber 254 via the outlet channel 259 or return channel (not shown); flowing a wash reagent into the reaction chamber 254; removing the wash reagent from the reaction chamber 254 via the outlet channel 259; flowing the next NTP reagent into the desired memory strand sequence under conditions calculated to achieve the desired homopolymer tract ratio; and repeating until the desired nucleic acid memory strand is synthesized. After the desired nucleic acid memory strand is synthesized, it may be released from the reaction chamber anchor or substrate and collected via the outlet channel 259 or other means.

[0081] Due to the significant number of homopolymer-encoded DNA strands required to encode usable amounts of data, highly parallel methods of DNA synthesis are required. In some embodiments, as shown in Figure 9, a flow cell (9-1) containing an array of wells is formed on a suitable substrate by patterning horizontal and vertical stripes of hydrophobic material (9-2) to form multiple hydrophilic wells (?) bounded by hydrophobic regions. Typical dimensions of the hydrophilic wells can be 300 x 300 nm to 1000 x 1000 nm. This hydrophilic array, with a gap of appropriate dimensions between the floor and an optically transparent cover, forms the floor of the flow cell. A solution of cold (i.e., below the optimal enzyme-specific temperature) nucleotidyl transferase, one of native or modified nucleotide triphosphates, and any necessary cofactors is flowed into the flow cell via an inlet (9-3) such that, upon cessation of fluid flow, the enzyme-nucleotide triphosphate solution bead up into spatially defined droplets (9-4) positioned above each hydrophilic region. The IR source (9-5) projects a beam (9-7) through a shaped lens (9-6) onto a DLP (digital light projection) device (9-8), which is used to simultaneously direct the IR beam (9-9) to each of the hydrophobically constrained droplets (9-4) to which a specific nucleotide is selected for addition, resulting in rapid heating of the polymerase extension reaction mixture to a temperature for maximum enzyme activity over a defined period to synthesize a homopolymer of the desired length. After some appropriately defined reaction time, the IR source is switched off, and a cold rinse buffer is rapidly injected into the flow cell and drained through the outlet (9-10), thus quenching the reaction and completing one "write" cycle. This series of steps is repeated multiple times for each "write" cycle, such that each nascent data strand is randomly accessed according to its spatial location, and selected nucleotides are added until a full-length homopolymer data strand is completed. In some embodiments, a DLP device with 1920x1080 directional mirrors can be used to simultaneously randomly access approximately 2M synthesis locations in a synthesis flow cell.In some embodiments, the template-independent polymerase used is thermophilic, with an optimal reaction temperature well above room temperature, such that enzymatic activity is minimized in the hydrophobically confined droplets prior to rapid heating of the droplets by the IR source. In some embodiments, the bottom surface of the flow cell bearing the hydrophobically defined hydrophilic wells abuts a cooling device that maintains the droplets in the hydrophobic wells at a reduced temperature to prevent enzymatic activity until the temperature is elevated by the IR source.

[0082] In another embodiment, a flow cell is constructed from an array of wells formed on a suitable substrate by patterning horizontal and vertical stripes of hydrophobic material to form multiple hydrophilic spots bounded by hydrophobic regions. Typical dimensions of the hydrophilic spots can range from 300 x 300 nm to 1000 x 1000 nm. Each hydrophilic spot is positioned on an individually addressable CMOS heater. This hydrophilic-CMOS heater array, with a gap of appropriate dimensions between the floor and an optically transparent cover, forms the floor of the flow cell. A cold (i.e., below the optimal enzyme-specific temperature) solution of nucleotidyl transferase and one of natural or modified nucleotide triphosphates is flowed through the flow cell such that, upon cessation of fluid flow, the polymerase extension reaction solution beaded into spatially defined droplets positioned above each hydrophilic region with its associated CMOS heater. Each hydrophobically tethered droplet, to which a specific nucleotide is selected for addition, is rapidly heated to a temperature for maximum enzyme activity over a defined period of time to synthesize a homopolymer of the desired length. After some well-defined reaction time, the heater is turned off and a cold rinse buffer is rapidly injected into the flow cell, thus quenching the reaction and terminating one "write" cycle. This series of steps is repeated multiple times for each "write" cycle, such that each nascent data strand is randomly accessed according to its spatial location and selected nucleotides are added until a full-length homopolymer data strand is completed. In some embodiments, the enzyme used is thermophilic with an optimal temperature well above room temperature, so that unwanted nucleotide addition in hydrophobically constrained droplets is unlikely prior to rapid heating of the droplets by the CMOS heater.

[0083] As those skilled in the art will recognize as necessary or best suited for the systems and methods of the present invention, the systems and methods may include computing devices such as those shown in FIG. 10 , which may include one or more processors 309 (e.g., central processing unit (CPU), graphics processing unit (GPU), etc.), computer-readable storage devices 307 (e.g., main memory, static memory, etc.), or combinations thereof, in communication with each other via a bus. The computing devices may include a mobile device 101 (e.g., a mobile phone), a personal computer 901, and a server computer 511. In various embodiments, the computing devices may be configured to communicate with each other via a network 517.

[0084] A computing device can be used to control the synthesis of memory strands, the reading of sequenced memory strands, and the compilation or conversion of data between human- or machine-readable formats, digitized data, and nucleic acid sequences during other steps described herein. A computing device can be used to display data in a readable format.

[0085] Processor 309 may include any suitable processor known in the art, such as the processor sold under the trademark XEON E7 by Intel (Santa Clara, CA) or the processor sold under the trademark OPTERON 6200 by AMD (Sunnyvale, CA).

[0086] Memory 307 preferably includes at least one tangible, non-transitory medium capable of storing: one or more sets of executable instructions (e.g., software embodying any methodology or function found herein) for causing the system to perform the functions described herein; data (e.g., data encoded in a memory chain); or both. While the computer-readable storage device may be a single medium in an exemplary embodiment, the term "computer-readable storage device" should be interpreted to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store instructions or data. The term "computer-readable storage device" should be interpreted accordingly to include, without limitation, solid-state memory (e.g., subscriber identity module (SIM) cards, secure digital cards (SD cards), microSD cards, or solid-state drives (SSDs)), optical and magnetic media, hard drives, disk drives, and any other tangible storage medium.

[0087] Any suitable service, such as Amazon Web Services, memory 307 of server 511, cloud storage, another server, or other computer-readable storage may be used for storage 527. Cloud storage may refer to a data storage scheme where data is stored in a logical pool and physical storage may span multiple servers and multiple locations. Storage 527 may be owned and managed by a hosting company. Preferably, storage 527 is used to store records 399 as needed to perform and support the operations described herein.

[0088] Input / output devices 305 according to the present invention may include one or more of a video display unit (e.g., a liquid crystal display (LCD) or cathode ray tube (CRT) monitor), an alphanumeric input device (e.g., a keyboard), a cursor control device (e.g., a mouse or trackpad), a disk drive unit, a signal generating device (e.g., a speaker), a touch screen, buttons, an accelerometer, a microphone, a cellular radio frequency antenna, a network interface device which may be, for example, a network interface card (NIC), a Wi-Fi card or a cellular modem, or any combination thereof.

[0089] Those skilled in the art will recognize that any suitable development environment or programming language may be used to enable the operability described herein for the various systems and methods of the present invention. For example, the systems and methods herein may be implemented in any suitable programming environment, including Perl, Python, C++, C#, Java, JavaScript, Visual Studio, etc. It may be implemented using Basic, Ruby on Rails, Groovy and Grails, or any other suitable tools. For computing device 101, it may be preferable to use native xCode or Android Java.

[0090] Figure 11 shows polyacrylamide gel electrophoresis analysis of two different single homopolymer tracts generated via enzymatic synthesis. Lane A is a sample of the 20-mer starting oligonucleotide used in all following lanes. Lane B is a sample from a TdT reaction containing a 20-mer oligonucleotide and the irreversible terminator ddATP, showing the formation of a 21-mer. Lane C is a sample from a TdT reaction containing a 20-mer and the natural nucleotide dATP after 1 minute at 37°C. Lane D is a sample of the same reaction mixture in lane D after 5 minutes at 37°C. Lane E is a sample of the same reaction mixture in lane C after 15 minutes at 37°C. These three lanes demonstrate homopolymer length control due to consumption of input dATP at approximately 5 minutes during the TdT extension reaction, and no homopolymer growth observed between 5 and 15 minutes. Lane F is a sample from a TdT reaction containing a 20-mer oligonucleotide and the nucleotide analog N6-benzoyl-dATP after 1 minute at 37°C. Lane G is a sample of the reaction mixture in lane F after 5 minutes at 37°C. Lane H is a sample of the same reaction mixture in lane F after 15 minutes at 37°C. dA formed in lanes F-H Bz Although qualitative differences exist between the lengths of the homopolymers, the same length control has also been demonstrated with N6-modified dATP analogs.

[0091] Writing digital data into molecular storage formats based on molecular approaches offers advantages over currently used storage media such as tapes or disks. DNA-based storage has sparked interest due to the high information density that can be achieved, extremely long lifespan, and low energy consumption during rest periods. To date, most attempts to use synthetic DNA as a storage medium have involved chemical synthesis using the common phosphoramidite method.

[0092] DNA-based data storage may require a much larger number of strands than are currently synthesized for the existing research market. Any synthesis technique (chemical or enzymatic) that relies on nucleotide blocking or terminator removal imposes additional steps and complexity because reagents must be delivered to an array of synthesis features (i.e., wells or spots) in an addressable manner to direct the correct nucleotide to the correct location on the array. Array-based methods of synthesis can be performed in one of several ways: 1) bulk delivery of activated reactants followed by selective removal of blocking groups, or 2) addressable delivery of activated reactants followed by bulk blocking group removal, or 3) bulk delivery of inactive reactants followed by addressable activation. Large (10 4 ~10 6 Addressable delivery of reagents to 2-D arrays is commonly achieved using inkjet deposition to each desired location. If the data encoding scheme uses four nucleotides, four separate writing heads must be used and indexed in a complex XY mechanical manner. Furthermore, the use of inkjet delivery limits the dimensions between each feature (well or spot) to the low tens of microns. The challenge in this process, however it is achieved, is to reduce the step time for each synthesis cycle.

[0093] In certain embodiments, the systems and methods of the present invention may include delivering an inert reaction mixture to all features on a 2-D array and then selectively activating only those features requiring the addition of A, G, C, or T. Addressable methods of delivery, activation, or blocking group removal that do not rely on mechanical motion are preferred. A preferred embodiment uses bulk delivery of reactants and selective activation of specific synthetic features by removing blocking groups from nucleotide analogs, which then enables rapid DNA polymerase-mediated incorporation and formation of homopolymer bits, as shown in Figure 15, thus enabling highly parallel and rapid synthesis of homopolymer-encoded nucleic acid memory strands. Homopolymer-encoded nucleic acid information polymers offer several other advantages: 1) homopolymer-encoded bits overcome the error profile associated with next-generation sequencing; 2) the resulting polymers are non-natural nucleic acids and therefore cannot be diverted for bioterrorist activities.

[0094] Delivery and selective activation of template-independent polymerase DNA synthesis can be performed sequentially (deliver A → addressably activate and initiate homopolymer synthesis → wash; deliver C → addressably activate and initiate homopolymer synthesis → wash; deliver G → addressably activate and initiate homopolymer synthesis → wash; deliver T → addressably activate and initiate homopolymer synthesis → wash) or in parallel (deliver all four nucleotides simultaneously to all features → addressably activate A, C, G, T simultaneously or sequentially → wash). In this manner, the synthesis cycle becomes highly efficient and involves only three steps: reactant delivery, incorporation reaction activation, and then wash before initiating the next cycle. The reaction is stopped by rapid removal of the reactants with a gas or liquid or by the rapid delivery of a quenching reagent. In a preferred embodiment, the quenching reagent is a metal chelator.

[0095] A preferred design for a device that can be used to synthesize homopolymer bit-encoding memory strands consists of a flow cell composed of a 2-D array of hydrophobically patterned wells appropriately modified to support template-independent enzymatic synthesis. After delivering reactants to the 2-D array of hydrophobic wells, the liquid beaded up to define spatially distinct reaction zones, as shown in Figure 16.

[0096] In a preferred embodiment, the bottom surface of the well, which is surrounded on all four sides by hydrophobic patterning, is modified with a 5'->3'-oriented covalently bound oligonucleotide initiator, with its 5'-end attached to the bottom surface of the well.In another embodiment, the well is physically formed by etching a depression in the bottom surface of the flow cell, in which case the covalently bound oligonucleotide initiator is covalently bound to the surface (bottom and side) of the well.In some cases, the well is open at the bottom to allow liquid flow through the well.In other cases, the well is closed at the bottom.

[0097] Nucleotide analogs protected at the 3'-OH are generally inactive with commercially available or wild-type TdT enzymes (U.S. Pat. No. 10,059,929). Thus, a mixture of 3'-blocked dNTP analogs, TdT protein, and appropriate cofactors can be mixed together in the presence of an initiator oligonucleotide at 37°C with little to no homopolymer formation. Once the 3'-OH is "decased" (i.e., unblocked) by removal of the protecting or blocking group, the resulting nucleotide is available for free-running incorporation and homopolymer formation. Decasing that can be achieved by a "delivery-free" method is preferred. Analogs constructed to allow addressable 3'-OH decasing by the application of or exposure to activation energy, such as light, heat, electrochemical generation of pH changes, or the administration of or exposure to reducing agents, are preferred. Each of these decazing or unblocking reactions can be achieved in an appropriately constructed flow cell using either mechanically directable light through an optically transparent cover (e.g., Figure 17), individually addressable heaters (e.g., Figure 18), electrochemically induced pH changes, or the creation of reducing conditions.

[0098] Each of the illustrated flow cell designs is compatible with the method of addressable decaking of nucleotides contained in droplets constrained by surrounding hydrophobic patterning. Figure 19 illustrates an apparatus that can be used with a flow cell designed for the use of photodecaked nucleotide analogs. Other embodiments and configurations known to those skilled in the art are possible for this purpose.

[0099] Each decaking mechanism requires a nucleotide analogue specifically designed for the physicochemical process selected for use in the system: [ka]

[0100] In some embodiments, the nucleotide analog suitable for light-mediated decaging is 3'-O-(2-nitrobenzyl)-dNTP. In some embodiments, the nucleotide analog suitable for heat-mediated decaging is 3'-O-(tetrahydrofuranyl)-dNTP. In some embodiments, the nucleotide analog suitable for reduction-mediated decaging is 3'-O-methyl-dithiomethyl-dNTP. Many other 3'-OH protecting groups are suitable, as long as the resulting 3'-OH modified nucleotide analog is not a substrate for template-independent polymerase and is easily removed by a "delivery-free" method, for example, by light, heat, or electrochemically generated reactants. Insofar as the 3'-O modifications described herein are explicitly stated as substrates for template-independent polymerases and are referred to as reversible terminators, the composition of these dNTP analogs differs from that described in WO2016 / 034807. The subject of the present invention is the 3'-O-blocking group, which is not explicitly a substrate for the polymerase, but rather functions to cage the dNTP until it is removed to allow polymerization to proceed. Mathews AS et al (2016) demonstrate the controlled synthesis of natural non-homopolymeric oligonucleotides. describes the use of 3'-O-(2-nitrobenzyl)-dNTP analogs as reversible terminators for enzymatic synthesis. They explicitly teach the use of such nucleotides as substrates for DNA synthesis using extremely long enzymatic reaction times (approximately 1 hour), and furthermore, in contrast to the subject matter of this patent, which teaches the use of these analogs as caging groups to initiate enzymatic polymerization of multiple nucleotides, explicitly teach their use as reversible terminators. Initiation of free-running homopolymer synthesis can be achieved by methods other than caged dNTP analogs. Some embodiments may use fluid pulses of polymerase to initiate and terminate ssDNA synthesis, as described in Church 9,928,869 (2018). In Reza et al. WO2017 / 196783, template-dependent enzymatic synthesis is initiated by activation of the polymerase by an electrochemically generated pH change. Although modified nucleotide analogs containing 3'-O-reversible terminators have been described, none of the patents teach the activation of homopolymer synthesis using decaging of caged dNTP analogs. Church 9,928,869 (2018) describes free-running homopolymer synthesis using natural dNTPs, with the length of the homopolymer formed controlled by the reaction duration. Lee HR et al (2018) describe the use of mechanical delivery to deposit natural nucleotides into multiple reaction zones on a 2D array and control the length of homopolymer formation by actively destroying unreacted, unmodified dNTPs using apyrase. In some embodiments of the subject matter of this patent, homopolymer rate-controlling modified dNTP analogs with an unmodified 3'-OH are used in combination with either fluid pulse or mechanical XY delivery or a pH-activated, template-independent DNA polymerase.

[0101] Figures 20 and 21 show the kinetics of light-mediated decaging. Once converted into a substrate for the polymerase, 3'-caged dNTPs in a template-independent polymerase reaction mixture are polymerized into a nascent homopolymer strand. Complete conversion of the caged dNTP to an uncaged (substrate) dNTP is not required, as available uncaged dNTPs are readily incorporated by the polymerase. Upon decaging, the desired homopolymer length is controlled by controlling the concentration of the decaged dNTPs produced and / or the duration of the enzymatically mediated incorporation reaction.

[0102] Figure 22 shows examples of two nucleotide analogs caged with 3'-O modifications that can be de-caged with visible wavelength light and exhibit tunable photophysical properties (Peterson JA, et al J. Amer. Chem. Soc. 2018 140:7343-6), which shows that two This allows simultaneous introduction, decaching and synthesis of different homopolymer chain sequences.In some embodiments, nucleotide analogs modified with two different 3'-caging species that can be decaked by two different wavelengths of light are simultaneously delivered to a 2D array of synthesis positions.In another embodiment, four dNTP analogs modified with four different photoremovable 3'-caging species are simultaneously delivered to a 2D array of synthesis positions and decaked by four different wavelengths of light.The decaking reaction can be carried out sequentially for a set of positions on the 2D array, followed by the delivery of the two remaining dNTPs and subsequent sequential decaking, thus completing one round of increasing the length of all information polymers by one nucleotide.In an alternative embodiment, four dNTP analogs modified with four different 3'-caging species can be simultaneously delivered and decaked by simultaneously exposing each synthesis feature to one of four different wavelengths of light. A hardware configuration such as that shown in Figure 19 can be used for multi-color decaging by incorporating two or more wavelength-specific light sources and an appropriate shutter mechanism to direct one of the two or more light sources onto a 2D light directing system.

[0103] Figure 23 shows the synthesis of a 12-bit information strand consisting of successive cycles of homopolymer synthesis using dCTP, dGTP, and dTTP with UV light-decaged 3'-oNBn-dATP incorporated in cycles 1, 6, 8, 10, and 12. The incorporation reaction can be terminated by several methods, including, but not limited to, the rapid removal of enzymatic reaction components by the introduction of a bolus of either gas or liquid. In some embodiments, the liquid simply rinses away the reaction components, while in other embodiments, the rinse liquid contains an active quenching agent such as EDTA or other enzyme inhibitors.

[0104] In another embodiment, the method of caging a dNTP can be mediated by a steric mechanism involving a modification to the nucleotide base instead of the 3'-OH. In a preferred embodiment, a nucleotide with an unmodified 3'-OH is modified at N6, N4, N2, O4 of A, C, G, or T, respectively, with a removable steric blocking group (incorporation blocker) that cages the dNTP. Homopolymer synthesis can be initiated by either photo-, thermally, or reduction-mediated cleavage of the steric blocking group, thus "uncaging" the natural nucleotide suitable for free-running homopolymer synthesis.

[0105] In some embodiments, the modified nucleotide is composed of a linker containing one moiety that cages the dNTP until it is removed, thus making it incapable of incorporation, and another moiety that remains covalently bound throughout the synthesis of multiple homopolymer tracts.In some embodiments, the purine or pyrimidine base is modified at two positions, one modification that cages the nucleotide and prevents enzymatic incorporation, while the other modification remains covalently bound throughout the synthesis of the polymer; each modification can be removed by a different mechanism that allows for selective removal of each.In some embodiments, the purine or pyrimidine base is modified at two positions, and each modification can be removed by a different mechanism that allows for selective removal of each.

[0106] Depending on the sequence detection modality used to read the homopolymer bit data strand, appropriate dNTP analog modifications can be designed to generate intact or damaged nucleotides. For readout methods that rely on sequencing by synthesis (SBS), all modifications to purines or pyrimidines must be removed after homopolymer synthesis but before SBS readout. For readout methods that rely on current modulation during polymer translocation through one or more nanopores, modifications to purines and pyrimidines that are easily distinguishable from each other are desired. In some embodiments, modifications that modulate polymerase kinetics during homopolymer synthesis are made to purine and pyrimidine bases. These "rate-regulating" modifications can also act as current modulators for nanopore detection or can be removed for SBS detection.

[0107] Regardless of the activation mechanism, it is important to control the length of the free-running homopolymer synthesis. Ideally, homopolymers of 2-4 nucleotides are desirable. In the presence of natural nucleotides, TdT tends to form homopolymers based on the nucleobases corresponding to the Km of each nucleotide (A>T>G>C). According to Lee HR et al. (2018), As pointed out by [1], it is difficult to limit homopolymer growth without increasing deletion frequency. The authors reported approximately 66% deletion errors in attempts to limit homopolymer growth to 2–3 nucleotides. Although TdT is reported to operate in a dispersive manner, resulting in a Poisson distribution of homopolymer lengths during free-running synthesis, in practice, it is difficult to drive complete conversion of the initiator nucleic acid during homopolymer synthesis without long enzymatic extension reaction times. Free-running homopolymer synthesis reaction times that are too short result in bit deletion errors, as reported above. Homopolymer synthesis reaction times that are too long result in excessively long homopolymer formation and low data density. One solution is to use modified nucleotide analogs that modulate the rate of multiple base incorporation but do not require removal at every step, thereby allowing complete conversion of the starting nucleic acid, controlling the length of the homopolymer formed, and maintaining the simple two-step cycle that is the subject of this patent. Ideally, the homopolymer tracts in the information polymer are greater than two nucleotides but four or fewer nucleotides in length.

[0108] If necessary, homopolymer length-limiting modifications can be removed from all nucleotides at the end of synthesis, thus generating native DNA molecules suitable for SBS detection. Modifications chemically compatible with the decaging conditions described above are most desirable; they must be removable by chemical conditions orthogonal to those used for decaging. In some embodiments, the 3'-OH is caged with 3'-O-(2-nitrobenzyl), while the N6 of dA, the N4 of dC, the N2 of dG, and the O4 or N3 of dT are modified with non-terminating moieties that regulate multiple additions during free-running enzymatic homopolymer synthesis. Figure 24 shows a comparison of free-running homopolymer synthesis in the presence and absence of rate-controlling modifications. In some embodiments, the homopolymer synthesis rate-controlling modifications are the same for all four nucleotide analogs, while in other embodiments, the modifications are different for each of the dATP, dCTP, dGTP, and dTTP analogs. In some embodiments, the rate-limiting modifications remain in place during the entire synthesis.

[0109] Figures 25 and 26 show the results of two successive addition cycles of two differentially modified nucleotide analogs. In another embodiment, the nascent information polymer is exposed to periodic or alternating cycles with chemical conditions that result in partial or complete removal of the homopolymer rate-limiting modification. In some embodiments using Class 2 dNTP analogs, previously incorporated rate-controlling modifications are removed simultaneously with polymerase extension by including a mild reducing agent in alternating cycles of the enzymatic extension reaction mixture.

[0110] When a non-SBS method of readout (i.e., nanopore) is used, the homopolymeric rate-controlling modification can be covalently attached using a non-cleavable linkage and left in place during readout. In some embodiments, the homopolymeric rate-controlling modification can also be used to encode information during the readout process. In a preferred embodiment, the homopolymeric synthesis rate-controlling analog is also designed to regulate current during nanopore translocation and is further selected to provide two or more detectable and distinguishable current blockade levels. In another embodiment, only a single type of nucleotide (i.e., A or C or G or T) is used in information strand synthesis, and the homopolymer bit is encoded by two or more differential current blockade-producing modifications.

[0111] Figures 27-30 show an example of solid-state nanopore detection of enzymatically synthesized homopolymers of peptide-modified dNTP analogs. A dTTP analog (N-Ac-CYPEE) modified with a cleavable disulfide linkage to a five-amino acid peptide was used to generate fully modified molecules via free-running homopolymer synthesis longer than 100 nt in length. Translocation through a 2D silicon nitride 30 nm thick nanopore at 500 mV resulted in a large signal swing compared to unmodified dU homopolymers (courtesy of Goeppert LLC, Philadelphia, PA).

[0112] The following table lists six different classes of modified dNTP analogs useful for three different "delivery-free" homopolymer synthesis activation approaches and two different homopolymer-encoded nucleic acid memory strand readout technologies. [Table 1]

[0113] Examples of the four light-mediated decazing nucleotide analogs that make up Class I are shown below: [ka]

[0114] This class of analogs represents a 3'-O-caged dNTP that becomes a native, unmodified dNTP upon light-mediated decaging. Upon decaging, the rate of homopolymer formation is no different from that of natural nucleotides, and the length of the homopolymer formed must be controlled by additives to the template-independent polymerase reaction formulation. In some embodiments, tetrahydropyranyl, 4-methoxytetrahydropyranyl, tetrahydrofuranyl, acetyl, methoxyacetyl, or phenoxyacetyl modifications are useful for thermally induced decaging of 3'-O-blocked dNTP analogs. In some embodiments, 3'-O-analogues such as -O-CH2-SSR are useful for electrochemically mediated decaging under reducing conditions. Class I dNTP analogs are characterized by nucleotides that yield unmodified homopolymers suitable for readout by either classical sequencing-by-synthesis or ratchet-style nanopores. In both cases, precision in the homopolymer length is not required, but precise detection of the transition between homopolymers is.

[0115] Examples of four Class II photo-mediated decaging nucleotide analogs with orthogonally removable peptide rate-controlling modifications are shown below: [ka] [ka]

[0116] Peptide modifications for use with Class II dNTP analogs are intended to be removed at the end of synthesis of the information polymer for subsequent interrogation by SBS-dependent sequencing methods or nanopore-mediated translocation, which are sensitive enough to detect the transition from one homopolymer type (A, C, G, T) to the next. Peptides suitable for use in this application slow the rate of incorporation of modified nucleotide analogs to prevent long homopolymer formation prior to complete conversion of unmodified or modified strands of different composition. In some embodiments, the incorporation rate-regulating modifications can consist of one, two, three, or four different peptides for four nucleotide analogs. In some embodiments, the peptides are linked to the nucleotides via disulfide linkers bearing a self-immolative scar that is cleavable under mild chemical conditions. In some embodiments, the peptide may consist of, but is not limited to, Ac-EECGY, Ac-EEGCGW, Ac-EEGCGGW, Ac-EC-pNA, Ac-CWEE, Ac-CYPEE, Ac-EEGCPPW, Ac-CPYEE, Ac-CPWEE, or Ac-CWPEE. Many other peptide sequences and compositions are possible to those skilled in the art. Peptides with an overall anionic composition are likely most suitable; peptides with a cationic composition accelerate the rate of nucleotide incorporation, as previously pointed out by Finn PS et al. (2003). Peptides with covalent linkages to nucleotides other than via disulfides to cysteine amino acids are possible, as long as they maintain their ability to be removed from nucleotides of completed homopolymeric DNA strands without leaving residual damage. Disulfide cleavage-mediated self-immolative linkers, which are eliminated under mild conditions by thiolactone formation, are particularly useful.

[0117] The following are examples of four class II nucleotides with light-mediated decaging and orthogonally removable non-peptide homopolymer synthesis rate-regulating modifications: [ka]

[0118] Non-peptide modifications to N6 of adenine, N4 of cytidine, N2 of guanine, and N3 of thymine can act as incorporation rate modulators. In some embodiments, acetyl, diacetyl, isobutyryl, proprionyl, pivaloyl, benzoyl, cyclohexyl, and other organic modifiers are useful. Modifiers of this type are removed after synthesis of the information polymer, including but not limited to, by ammonolysis. Class II dNTP analogs are characterized by nucleotides that yield unmodified homopolymers suitable for readout by either classical sequencing-by-synthesis or ratchet-style nanopores. In both cases, precision in the length of the homopolymer is not required, but precise detection of the transition between homopolymers is.

[0119] Below are examples of two Class III nucleotide analogs with an unmodified 3'-OH and a removable incorporation caging modification on the purine or pyrimidine: [ka]

[0120] In some embodiments, useful caging modifications to the 2-nitrobenzyl functional group are bulky modifications, including, but not limited to, peptides, cyclic peptides, PEGs, branched PEGs, star PEGs, dendrimers, nanoparticles, etc. Class III dNTP analogs are also characterized by nucleotides that yield unmodified homopolymers suitable for readout by either classical sequencing-by-synthesis or ratchet-style nanopores. In both cases, precision in the length of the homopolymer is not required, but precise detection of the transition between homopolymers is.

[0121] An example of one Class IV nucleotide analog having an unmodified 3'-OH, a removable incorporation blocker at a position other than the 3'-OH, and an orthogonally removable homopolymer synthesis rate modulator modification that is also not located at the 3'-OH is: [ka]

[0122] Class IV nucleotide analogs are designed with one or more incorporation-blocking modifications that cage the nucleotide from enzymatic incorporation until removed by light, heat, or electrochemical means. Large, bulky modifications that render the nucleotide analog inactive to polymerase incorporation are preferred. In some embodiments, one or more modifications are derivatives of 2-nitrobenzyl, which can be removed by photolysis. Removal of the caging modification in every cycle leaves behind an incorporation rate-controlling modification that is removed by a chemical method orthogonal to the method used to de-cage the nucleotide analog upon completion of the synthesis of the information polymer. In a preferred embodiment, the incorporation-blocking modification is covalently linked to the incorporation rate-controlling modification. In some embodiments, a caging modification to the 2-nitrobenzyl functional group is useful, as it is a bulky modification and acts as a steric blocking agent. Examples of modifications that can act as steric blocking agents include, but are not limited to, peptides, proteins, peptoids, PEG, branched PEG, star PEG, dendrimers, or nanoparticles.

[0123] Examples of Class V pyrimidine nucleotides with a 3'-O-caging modification and a non-removable base modification are: [ka]

[0124] When nanopore detection by direct translocation is desired as the readout method, the nucleotide is doubly modified with a suitable removable 3'-O-caging group and a covalent, non-removable modification that acts as a blocking current modulator during direct translocation through a non-indexing nanopore. Suitable modifications can be peptides or non-peptides. The ideal embodiment of the blocking current-modulating group is the smallest possible modification that produces the most unique signal in the shortest possible homopolymer stretch. The key innovation of this class of dNTP analogs is that only one type of modified nucleotide is required, since the "sequence" of the homopolymer is encoded by the sequence of two or more current-modulating modifications, rather than by the nucleotides themselves.

[0125] The following are examples of Class VI nucleotide analogs that contain an unmodified 3'-OH, a removable uptake-blocking (caging) modification, and two or more different types of non-removable current-modulating modifications useful for detection by direct nanopore sequencing. Class VI The key innovation of dNTP analogs is that only one type of modified nucleotide is required, since the "sequence" of the homopolymer is encoded by the sequence of two or more current-modifying modifications, rather than by the nucleotides themselves. Encoding is not limited to base 2 or base 4, but is limited only by the number of current-modifying modifications that can be made to a single purine or pyrimidine nucleotide. [ka]

[0126] In some embodiments, the uptake-blocking (caging) modification is removable by either light, heat, or electrochemical means, while two or more nanopore sensing elements are non-removable and survive repeated exposure to the decaging conditions used in every cycle during homopolymer chain synthesis. The 2-nitrobenzyl caging modification, which acts as a steric blocking agent, is a bulky modification such as, but not limited to, a peptide, protein, peptoid, PEG, branched PEG, star PEG, dendrimer, or nanoparticle. [Example]

[0127] Example 1 N 6 -benzoyl-deoxyadenosine triphosphate in a vial under dry N 6A solution of 2'-benzoyl-2'-deoxyadenosine (0.055 g, 0.16 mmol) was added to a 100 ml solution of 2'-benzoyl-2'-deoxyadenosine (0.055 g, 0.16 mmol), followed by trimethyl phosphate (0.435 mL) was added. To the resulting solution, tributylamine (0.077 mL, 0.32 mmol) was added, and the reaction mixture was flushed with dry N2 for 30 minutes while maintaining the temperature at -5°C. To this vial, anhydrous phosphorus oxychloride (0.018 mL, 0.19 mmol) was added via syringe, and the reaction mixture was stirred at -5°C for 3 minutes. A second aliquot of anhydrous phosphorus oxychloride (0.009 mL, 0.10 mmol) was added via syringe, and the reaction mixture was stirred at -5°C for 8 minutes. To a second vial, tributylamine pyrophosphate (0.075 g, 0.14 mmol) was added, flushed with dry N2, and anhydrous acetonitrile (0.609 mL) was added, followed by tributylamine (0.231 mL, 0.97 mmol). The prepared tributylamine pyrophosphate mixture was cooled to -20 °C and added to the reaction mixture, allowing it to react for 10 minutes. The reaction was quenched by the dropwise addition of HO (4.35 mL). The contents of the flask were combined with 0.87 mL of HO and extracted with dichloromethane (3 × 150 mL). The aqueous phase was adjusted to pH 6.5 with concentrated NH4OH and stirred at 4 °C for 12 hours. The mixture was transferred to a 250 mL round-bottom flask with 50 mL of water and concentrated under reduced pressure. The residue was dissolved in 40 mL water and purified via ion exchange chromatography (AKTA FPLC, Fractogel DEAE 48 mL column volume, step gradient 0->70% TEAB in water, pH 7.5). Fractions containing the desired product were pooled, concentrated under reduced pressure, and purified by repeated concentration from water to dryness (5 x 50 mL) to remove residual triethylammonium bicarbonate. 6 -benzoyl-2'-deoxyadenosine triphosphate was obtained.

[0128] The controlled synthesis of homopolymer tracts composed of modified nucleotides by nucleotidyl transferase TdT was carried out in the following manner: 1 mM each of deoxyadenosine triphosphate (TriLink Biosciences) and N 6A stock solution of -benzoyl-deoxyadenosine triphosphate was prepared in H2O.

[0129] 0.5 μL (500 pmoles) of each of the different triphosphates was mixed separately with 1.5 μL (30 U) of commercially available TdT (Thermo Scientific), 0.5 μL (50 pmoles) of 5'-TAATAATAATAATTTTT-3' (IDT), 2 μL of commercially available TdT Rxn buffer (Thermo Scientific - 1 M potassium cacodylate, 0.125 M Tris, 0.05% (v / v) Triton X100, 5 mM C o The reaction mixture was combined with Cl2 (pH 7.2 at 25°C) and 4.5 μL of HO. The reaction was incubated at 37°C. 30 μL aliquots were removed after 1, 5, and 15 minutes and quenched with 20 μL of 5 mM EDTA. Each sample was dried under vacuum and reconstituted in 100 μL of HO. 10 μL of each time point was mixed with 10 μL of denaturing loading buffer (100% formamide and 0.1% Orange G) and applied to the wells of a 1 mm × 20 cm × 14 cm 20% polyacrylamide gel. After 3.5 hours of electrophoresis at 400 V, bands were visualized with Sybr Gold (Thermo Scientific) and photographed under UV illumination (UV-blocking Wratten 2A filter, 405 nm cutoff, UVP, LLC).

[0130] For the synthesis of multiple homopolymer tracts, initiator oligonucleotides can be attached to beads to allow for multiple rounds of enzymatic synthesis with integrated removal of previous reactants and washes. 5'-Biotin-TAATAATAATAATTTTT-3' (IDT) can be incubated with streptavidin-coated magnetic Sepharose microbeads (GE Healthcare Life Sciences). Oligonucleotide-loaded beads can be prepared by removing an aliquot of the bead slurry and transferring it to a filter cup. The beads can then be washed five times with 1x PBS (using 2x the volume of bead slurry) by vortexing, with each rinse being spun down. Then, 1 / 2 the bead slurry volume of 1x PBS can be added, and biotinylated oligonucleotides can be spiked in at 1 / 10 the published bead binding capacity. The mixture can be incubated at 37°C for 2 hours, vortexing every 30 minutes. After 2 hours, a small amount of supernatant can be removed and the A260 can be measured for any unbound oligonucleotide. When the A260 indicates less than 10% remaining oligonucleotide, the beads can be washed five times with MQ water. The washed beads can be brought to the desired concentration in MQ water.

[0131] Homopolymer synthesis can be carried out as described above, using 2-10x molar equivalents of the desired dNTPs relative to the bead-bound oligonucleotide. The reaction mixture containing beads, dNTPs, TdT, and buffer can be incubated at 37°C for 15 minutes. The reaction can be stopped with 10 μl EDTA and rinsed three times with water. A new cycle of homopolymer synthesis can be initiated by adding fresh TdT enzyme, dNTPs, and buffer and incubated at 37°C for 15 minutes. After quenching the reaction with EDTA and rinsing three times with water, the homopolymer synthesis cycle can be repeated as many times as desired. After the final EDTA quench and three rinses with water, the support-bound alternating homopolymer can be cleaved from the solid support using 100 μl concentrated ammonium hydroxide, and the supernatant can be dried by Gen-vac and then stored at -20°C until ready to be analyzed using a polyacrylamide gel as described above.

[0132] In another experiment, as shown in Figure 12, the two-letter message "MA" was converted to binary and then to base-2 nucleotide code. Each letter symbol was converted to a corresponding number from 1 to 26 (A → 1, ... Z → 26). Each number was then converted to a 6-bit binary number by adding two zeros to the normal binary representation of 1 to 26 (e.g., "M" = 001101; "A" = 000001). Each bit of the binary representation was converted to a base-2 nucleotide representation according to the following table: [Table 2]

[0133] Thus, "MA" is converted to 001101 000001, which is then converted to the single nucleotide string ACTGAGACACAG, which is converted to A n C n T n G n A n G n A n C n A n Cn A n G n where each nucleotide is synthesized as a homopolymer of variable length.

[0134] Synthesis of a 12-bit homopolymer encoded nucleic acid was performed using a 5'-biotinylated 39 nt long oligonucleotide initiator: 5'biotin-CAGGTCCTAUC bound to 34 um streptavidin-Sepharose beads (GE Healthcare) at approximately 20 pmol / ul of beads. GATATC This was performed using TGTGAGCTTAATGTCCTTATGT-3'.

[0135] This oligonucleotide contains two features for releasing the final product from the solid support used during synthesis: 1) a single deoxyuridine residue that allows cleavage with the USER enzyme system (New England Biolabs), and 2) an Eco RV endonuclease restriction site.

[0136] Starting with approximately 2 nmol of bead-bound initiator, homopolymers of varying lengths were enzymatically synthesized using TdT and one of four modified nucleotide triphosphates. Each reaction was performed in a total volume of 750 μl containing 40–100 μM modified dNTPs (40 μM-A; 100 μM-C; 50 μM-T; 100 μM-G), 20 U TdT (Thermo-Fisher Scientific), and 1× TdT buffer (Thermo-Fisher Scientific) with incubation times of 2.5–20 min at 37°C. After each enzymatic extension step, the reaction was quenched by adding 500 μl of 250 mM EDTA in 10 mM Tris buffer (pH 6.8). Beads were recovered by centrifugation at 10,000 × g and removal of the supernatant. Figure 13 shows PAGE analysis of each cycle of enzymatic synthesis of a 12-bit homopolymer data strand. Each lane begins with "N," indicating unreacted 39-nt initiator, and is marked with the cycle number. The black arrow indicates the size marker of the 60-nt oligonucleotide.

[0137] After removal of the full-length data strand from the solid support, NGS library preparation was performed using the ACCEL-NGS® 1S PLUS DNA LIBRARY KIT (Swift Bioscience) according to the manufacturer's instructions, followed by sequencing using an Illumina MiSeq System. Figures 14A-L show histograms of the observed base compositions of each of the 12 nucleotide additions (as labeled) generated during enzymatic synthesis.

[0138] Example 2 Procedure for synthesizing class I-purine and pyrimidine dNTP analogs Scheme for the synthesis of class I dNTP analogs Nitrobenzyl adenosine [ka] Nitrobenzyl-cytosine [ka] Nitrobenzyl-guanosine [ka] Nitrobenzyl-thymidine [ka]

[0139] Example 3 Detailed procedure for 3'-O-nitrobenzyldeoxyadenosine: [ka] 9-[β-D-5'-hydroxy-2'-deoxyribofuranosyl]-6-chloropurine (1.00 g, 3.69 mmol) and imidazole (554 mg, 8.12 mmol) were dissolved in anhydrous dimethyl formate (18 mL), followed by the addition of tert-butyldimethylsilyl chloride (611 mg, 3.93 mmol). The reaction mixture was stirred at room temperature under argon for 20 hours. The dried residue was impregnated onto silica and purified by flash column chromatography (hexane / ethyl acetate, 2:1) to give 9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]-6-chloropurine. [ka]

[0140] 9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]-6-chloropurine (1.73 g, 4.48 mmol) was dissolved in anhydrous dichloromethane (135 mL). Tetrabutylammonium bromide (722 mg, 2.24 mmol), 2-nitrobenzyl bromide (2.41 g, 11.2 mmol), and 40% aqueous sodium hydroxide (65 mL) were added to the previously prepared solution. The reaction mixture was stirred at room temperature for 1 hour and diluted with ethyl acetate (300 mL). The layers were separated. The aqueous layer was extracted with ethyl acetate (125 mL x 2). The combined organic layers were dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica gel and then purified by flash column chromatography (hexane / ethyl acetate, 2:1) to give 9-[β-D-5′-O-(tert-butyldimethylsilyl)-3′-O-(2-nitrobenzyl)-2′-deoxyribofuranosyl]-6-chloropurine. [ka]

[0141] 9-[β-D-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyribofuranosyl]-6-chloropurine (2.22 g, 3.83 mmol) was dissolved in anhydrous tetrahydrofuran (40 mL) and cooled to 0°C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (4.20 mL, 4.20 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 1 hour. After the reaction mixture was dried, the residue was dissolved in dioxane and 7N ammonia in ethanol (40 mL). The reaction mixture was stirred at 90°C for 18 hours in a sealed round-bottom tube. The reaction mixture was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 20:1) to give 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine. [ka]

[0142] 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (15 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 minutes, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 minutes. Tetrabutylammonium hydrogen diphosphate (tetrabutylammonium hydrogen diphosphate) ( A solution of N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (0.14 g, 4 eq., 0.15 mmol) was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 eq., 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0143] Example 4 Detailed procedure for 3'-O-nitrobenzyldeoxycytidine: [ka] 3',5'-Di-O-(tert-butyldimethylsilyl)-2'-deoxyuridine (1.00 g, 2.12 mmol) was dissolved in anhydrous acetonitrile (90 mL) and cooled to 0°C under argon. Phosphoryl trichloride (1.49 mL, 2.12 mmol) was added dropwise over 2 minutes. After 10 minutes, triethylamine (11.1 mL, 79.7 mmol) was added dropwise over 3 minutes. After 15 minutes, the reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was cooled to 0°C, and triazole (4.40 g, 63.7 mmol) was added as a solid in one portion. A precipitate was observed, and the suspension was stirred for 30 minutes. After stirring at room temperature for 2 hours, the reaction mixture was concentrated to dryness. The residue was dissolved in dichloromethane (30 mL) and washed with a saturated solution of sodium bicarbonate (25 mL x 2) and brine (25 mL). The organic layer was dried over anhydrous sodium sulfate and concentrated to dryness. The residue was dissolved in dichloromethane, followed by the addition of allyl alcohol (2.00 mL, 29.4 mmol) and triethylamine (2.67 mL, 18.9 mmol). The reaction mixture was stirred at 0° C. for 15 minutes. DBU (0.33 mL, 2.17 mmol) was added and stirred at room temperature for 6 hours. The reaction mixture was diluted with dichloromethane (17 mL) and washed with brine (15 mL). The organic layer was dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (hexane / ethyl acetate, 4:1) to give 4-O-allyl-3′,5′-di-O-(tert-butyldimethylsilyl)-2′-deoxyuridine. [ka]

[0144] 4-O-Allyl-3',5'-di-O-(tert-butyldimethylsilyl)-2'-deoxyuridine (3.37 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 4-O-allyl-2'-deoxyuridine. [ka]

[0145] 4-O-allyl-2'-deoxyuridine (1.18 g, 4.40 mmol) was dissolved in anhydrous pyridine (37 mL), followed by the addition of tert-butyldimethylsilyl chloride (815 mg, 5.41 mmol) under argon. The reaction mixture was stirred at room temperature for 20 hours. After concentration, the residue was dissolved in dichloromethane and impregnated onto silica. The crude product was purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 4-O-allyl-5'-O-(tert-butyldimethylsilyl)-2'-deoxyuridine. [ka]

[0146] To a mixture of 4-O-allyl-5'-O-(tert-butyldimethylsilyl)-2'-deoxyuridine (1.28 g, 3.35 mmol), tetrabutylammonium hydroxide (1.5 mL, 55-60% in water), and sodium iodide (50.0 mg, 0.335 mmol) in dichloromethane / water (20 mL, 1:1) was added 1.0 M sodium hydroxide solution (10 mL) under argon. The reaction mixture was stirred at room temperature for 10 minutes, after which 2-nitrobenzyl bromide (1.45 g, 6.70 mmol) in 10 mL dichloromethane was added over 5 minutes. After stirring at room temperature for 7 hours, the reaction mixture was diluted with dichloromethane (150 mL). The organic layer was washed with brine (20 mL) and dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (hexane / ethyl acetate, 1:1) to give 4-O-allyl-5′-O-(tert-butyldimethylsilyl)-3′-O-(2-nitrobenzyl)-2′-deoxyuridine. [ka]

[0147] 4-O-Allyl-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyuridine (1.55, 3.00 mmol) was dissolved in 7N ammonia in ethanol (55 mL) and stirred in a sealed round-bottom tube at 55° C. for 20 hours. The reaction mixture was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 20:1) to give 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine. [ka]

[0148] 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (2.22 g, 3.83 mmol) was dissolved in anhydrous tetrahydrofuran (40 mL) and cooled to 0°C. Then, 1.0 M tetrabutylammonium fluoride solution in tetrahydrofuran (4.20 mL, 4.20 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 1 hour. After the reaction, the mixture was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 8:2) to obtain 3'-O-(2-nitrobenzyl)-2'-deoxycytidine. [ka]

[0149] 3'-O-(2-nitrobenzyl)-2'-deoxycytidine (14 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first 6 μL aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 min, a second 5 μL aliquot was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0150] Example 5 Detailed procedure for 3'-O-nitrobenzyldeoxyguanosine: [ka] 2-Amino-6-chloro-9-[β-D-2'-deoxyribofuranosyl]purine (1.00 g, 3.50 mmol) and imidazole (715 mg, 10.50 mmol) were dissolved in anhydrous dimethyl formate (18 mL), followed by the addition of tert-butyldimethylsilyl chloride (686 mg, 4.60 mmol). The reaction mixture was stirred at room temperature under argon for 12 hours. The dried residue was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 2-amino-6-chloro-9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]purine. [ka]

[0151] 2-Amino-6-chloro-9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]purine (1.19 g, 3.00 mmol) was dissolved in anhydrous tetrahydrofuran (8.0 mL), and then N,N-dimethylformamide dimethyl acetal (3.10 mL, 18.0 mmol) was added at room temperature. The reaction mixture was stirred at 40°C for 3 hours. The reaction mixture was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 6-chloro-N 2 -[(dimethylaminomethylene)amino]-9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]purine was obtained. [ka]

[0152] 6-Chloro-N 2-[(Dimethylaminomethylene)amino]-9-[β-D-5'-O-(tert-butyldimethylsilyl)-2'-deoxyribofuranosyl]purine (1.09 g, 2.40 mmol) was dissolved in anhydrous acetonitrile (3.5 mL), followed by the addition of sodium hydride powder (60%) in mineral oil (122 mg, 4.80 mmol) at 0 °C. After stirring at room temperature for 1 hour, a solution of 2-nitrobenzyl bromide (1.04 g, 4.80 mmol) in anhydrous acetonitrile (1.5 mL) was added. After stirring at room temperature for 2 hours, the reaction mixture was filtered. The filtrate was dried, and the resulting residue was dissolved in ethyl acetate (100 mL). The organic layer was washed with saturated sodium bicarbonate solution (50 mL), brine (50 mL), and dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (hexane / ethyl acetate 4:6) to give 6-chloro-N 2 -[(dimethylaminomethylene)amino]-9-[β-D-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyribofuranosyl]purine was obtained. [ka]

[0153] 6-Chloro-N 2-[(Dimethylaminomethylene)amino]-9-[β-D-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyribofuranosyl]purine (1.02 g, 1.73 mmol) was dissolved in dimethyl formate anhydride (15 mL), followed by the addition of cesium acetate (996 mg, 5.19 mmol), 1,4-diazabicyclo[2.2.2]octane (194 mg, 1.73 mmol), and triethylamine (0.72 mL, 5.19 mmol) under argon. The reaction mixture was stirred at room temperature for 18 hours. Acetic anhydride (5 mL) was added and stirred for 0.5 hours. The reaction mixture was quenched with water (100 mL) and extracted with ethyl acetate (100 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 20:1) to give 5'-O-(tert-butyldimethylsilyl)-N 2 -[(dimethylamino)methylene]-3'-O-(2-nitrobenzyl)-2'-deoxyguanosine was obtained. [ka]

[0154] 5'-O-(tert-butyldimethylsilyl)-N 2 -[(Dimethylamino)methylene]-3'-O-(2-nitrobenzyl)-2'-deoxyguanosine (732 mg, 1.28 mmol) was dissolved in anhydrous tetrahydrofuran (8 mL) under argon and cooled to 0°C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (2.56 mL, 2.56 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was poured into cold water (50 mL) and extracted with ethyl acetate (50 mL x 2). The organic layers were combined and dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 10:1) to give N 2 -[(dimethylamino)methylene]-3'-O-(2-nitrobenzyl)-2'-deoxyguanosine was obtained. [ka]

[0155] N 2 -[(Dimethylamino)methylene]-3'-O-(2-nitrobenzyl)-2'-deoxyguanosine (17 mg, 1 equiv., 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first 6 μL aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equiv., 0.11 mmol) was added. After 5 min, a second 5 μL aliquot was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv., 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0156] Example 6 Detailed procedure for 3'-O-nitrobenzyldeoxythymidine: [ka] Thymidine (2.50 g, 10.3 mmol) was suspended in dimethylformamide (60 mL) at room temperature under argon. To this suspension, imidazole (4.22 g, 61.9 mmol) and tert-butylchlorodimethylsilane (4.66 g, 30.1 mmol) were added. After stirring for 2 hours, the reaction mixture was quenched with methanol (8 mL) and diluted with ethyl acetate (200 mL). The organic layer was washed with water (100 mL × 2), saturated sodium bicarbonate solution (100 mL), and brine (100 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 8:2) to give 3',5'-di-O-(tert-butyldimethylsilyl)thymidine. [ka]

[0157] 3',5'-Di-O-(tert-butyldimethylsilyl)thymidine (4.34 g, 9.22 mmol) and dimethyl-4-aminopyridine (1.12 g, 9.22 mmol) were dissolved in anhydrous dichloromethane (140 mL). Triethylamine (5.14 mL, 36.9 mmol) was added, and the reaction mixture was cooled to 0°C. Benzyl chloride (3.21 mL, 27.7 mmol) was added dropwise and allowed to warm up to room temperature. After stirring for 14 hours, a saturated solution of sodium bicarbonate (80 mL) was added, and the layers were separated. The aqueous layer was extracted with dichloromethane (200 mL x 2). The combined organic layers were washed with water (300 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 8:2) to give 3-N-benzoyl-3′5′-di-O-(tert-butyldimethylsilyl)thymidine. [ka]

[0158] 3-N-benzoyl-3',5'-di-O-(tert-butyldimethylsilyl)thymidine (3.03 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 3-N-benzoylthymidine. [ka]

[0159] 3-N-benzoylthymidine (XX g, 4.40 mmol) was dissolved in anhydrous pyridine (37 mL), and then tert-butyldimethylsilyl chloride (815 mg, 5.41 mmol) was added under argon. The reaction mixture was stirred at room temperature for 20 hours. After concentration, the residue was dissolved in dichloromethane and impregnated onto silica. The crude product was purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 3-N-benzoyl-5'-O-(tert-butyldimethylsilyl)thymidine. [ka]

[0160] To 3-N-benzoyl-5'-O-(tert-butyldimethylsilyl)thymidine (1.18 g, 2.56 mmol) was added an aqueous solution of tetrabutylammonium hydroxide (10 mL, 60%), followed by the addition of sodium iodide (76.7 mg, 0.51 mmol), dichloromethane (10 mL), water (10 mL), and 1 M aqueous sodium hydroxide solution (10 mL). This mixture was added dropwise to a solution of 2-nitrobenzyl bromide (718 mg, 3.32 mmol) in dichloromethane (10 mL). The reaction mixture was stirred at room temperature for 6 hours, and then water (10 mL) was added. The aqueous layer was extracted with dichloromethane (50 mL x 3). The organic layers were combined, washed with brine (100 mL), and dried over anhydrous sodium sulfate. After filtration and concentration, the residue was dissolved in ethyl acetate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 6:4) to give 3-N-benzoyl-5′-O-(tert-butyldimethylsilyl)-3′-O-(2-nitrobenzyl)thymidine. [ka]

[0161] 3-N-benzoyl-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)thymidine (1.24 g, 2.08 mmol) was dissolved in ethanol (15 mL), followed by the addition of 30% ammonium hydroxide solution (1.5 mL, 12.3 mmol). The reaction mixture was stirred at room temperature for 1 hour, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 6:4) to give 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)thymidine. [ka]

[0162] 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)thymidine (681 mg, 1.39 mmol) was dissolved in anhydrous tetrahydrofuran (12 mL) under argon and cooled to 0°C. Then, 1.0 M tetrabutylammonium fluoride solution in tetrahydrofuran (2.78 mL, 2.78 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was poured into cold water (50 mL) and extracted with ethyl acetate (50 mL × 3). The organic layers were combined and dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 10:1) to give 3'-O-(2-nitrobenzyl)thymidine. [ka]

[0163] 3'-O-(2-nitrobenzyl)thymidine (15 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 min, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0164] Example 6 Class II—Procedures for synthesizing purine and pyrimidine dNTP analogs. Scheme for the synthesis of class II non-peptide dNTP analogs Non-peptide precursor - adenosine [ka] Non-peptide precursor - cytosine [ka] Scheme for the synthesis of class II peptide-dNTP analogs Disulfide Peptide Precursor - Adenosine [ka] Disulfide Peptide Precursor - Cytosine [ka]

[0165] Example 7 Detailed procedure for synthesis of class II non-peptide dATP constructs: [ka] 9-[β-D-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (0.24 mmol) was dissolved in anhydrous tetrahydrofuran (4 mL) at room temperature under argon. Sodium hydride (48 mg, 1.21 mmol) was added in one portion. After stirring for 1 hour, the reaction mixture was cooled to 0°C and acyl chloride (3 equivalents, 0.723 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 18 hours. The reaction mixture was poured into a cold solution of saturated sodium bicarbonate and dichloromethane. The layers were separated. The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 2:8) to give the desired product. [ka]

[0166] The N-substituted nucleoside (5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine. [ka]

[0167] N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (1 equiv., 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 equiv., 0.11 mmol) was added. After 5 min, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv., 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0168] Example 8 Detailed procedure for synthesis of class II non-peptide dCTP constructs: [ka] 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (524.3 mg, 1.10 mmol) and carboxylic acid (1.32 mmol) were dissolved in dimethyl formate anhydride under argon. N,N-Diisopropylethylamine (0.48 mL, 2.74 mmol) and 2-(3H-[1,2,3]triazolo[4,5-b]pyridin-3-yl)-1,1,3,3-tetramethylisouronium hexafluorophosphate (V) (417 mg, 1.10 mmol) were added. The reaction mixture was stirred at room temperature for 18 hours and diluted with ethyl acetate (30 mL). The organic layer was washed with saturated sodium bicarbonate solution (30 mL), dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 4:6) to give N-substituted 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine. [ka]

[0169] N-substituted 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (3.37 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxycytidine. [ka]

[0170] N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxycytidine (1 equivalent, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 equivalents, 0.11 mmol) was added. After 5 minutes, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equivalents, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0171] Example 9 Detailed procedure for the synthesis of peptide-dATP conjugates: [ka] 9-[β-D-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (524.3 mg, 1.10 mmol) and 4-(pyridin-2-yldisulfaneyl)butanoic acid (302.7 mg) N,N-Diisopropylethylamine (0.48 mL, 2.74 mmol) and 2-(3H-[1,2,3]triazolo[4,5-b]pyridin-3-yl)-1,1,3,3-tetramethylisouronium hexafluorophosphate (V) (417 mg, 1.10 mmol) were added under argon. The reaction mixture was stirred at room temperature for 18 hours and diluted with ethyl acetate (30 mL). The organic layer was washed with saturated sodium bicarbonate solution (30 mL), dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 4:6) to give N-(4-(pyridin-2-yldisulfanyl)butanyryl)-5'-O-(tert-butyldimethylsilyl)-3'-O. -(2-nitrobenzyl)-2'-deoxyadenosine was obtained. [ka]

[0172] N-(4-(pyridin-2-yldisulfanyl)butanyl)-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (3.75 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give N-(4-(pyridin-2-yldisulfanyl)butanilyl)-3′-O-(2-nitrobenzyl)-2′-deoxyadenosine. [ka]

[0173] N-(4-(pyridin-2-yldisulfanyl)butanyl)-3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (23 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 min, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 10 minutes after the addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0174] Example 10 Detailed procedure for the synthesis of peptide-dCTP conjugates: [ka] 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (524.3 mg, 1.10 mmol) and 4-(pyridin-2-yldisulfanyl)butanoic acid (1.32 mmol) were dissolved in anhydrous dimethyl formate under argon. N,N-Diisopropylethylamine (0.48 mL, 2.74 mmol) and 2-(3H-[1,2,3]triazolo[4,5-b]pyridin-3-yl)-1,1,3,3-tetramethylisouronium hexafluorophosphate (V) (417 mg, 1.10 mmol) were added. The reaction mixture was stirred at room temperature for 18 hours and diluted with ethyl acetate (30 mL). The organic layer was washed with saturated sodium bicarbonate solution (30 mL), dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 4:6) to give N-(4-(pyridin-2-yldisulfanyl)butanilyl)-5′-O-(tert-butyldimethylsilyl)-3′-O-(2-nitrobenzyl)-2′-deoxycytidine. [ka]

[0175] N-(4-(pyridin-2-yldisulfanyl)butanyl)-5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (3.62 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, followed by triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated onto silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give N-(4-(pyridin-2-yldisulfanyl)butanilyl)-3′-O-(2-nitrobenzyl)-2′-deoxycytidine. [ka]

[0176] N-(4-(pyridin-2-yldisulfanyl)butanyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (22 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 min, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 10 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0177] Example 11 Procedures for synthesizing class III—purine and pyrimidine dNTP analogs. Scheme for the synthesis of class III dNTP analogs [ka]

[0178] Example 12 Detailed Procedure for Class III-Purine dNTP Analogs: [ka] To a solution of phosgene (6.46 g, 2 eq, 65.3 mmol) in dry toluene (100 mL) at 23° C. was added (2-nitrophenyl)methanol (5.00 g, 1 eq, 32.6 mmol) in 20 mL dry THF. The reaction was stirred at 23° C. for 24 h. The reaction was then concentrated to dryness under vacuum with a trap of aqueous NaOH. The residual amber oil was used directly in the reaction without further purification. [ka]

[0179] To a solution of 9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-9H-purin-6-amine (12.2 g, 1.1 equiv., 25.5 mmol) in dry DMF (100 mL) at 0° C. was added N,N-diisopropylethylamine (3.60 g, 4.9 mL, 1.2 equiv., 27.8 mmol). The reaction was stirred for 30 minutes, then 2-nitrobenzyl carbonochloridate (5.00 g, 1 equiv., 23.2 mmol) was added slowly dropwise over 30 minutes, maintaining the temperature below 5° C. The reaction was then warmed to room temperature and stirred overnight. The reaction was poured into a cooled solution of 5% Na2CO3 and EtOAc. The EtOAc layer was dried over sodium sulfate and then concentrated to dryness. The crude product was chromatographed on silica gel using hexane / EtOAc mixtures to give the purified product, which was used in the next reaction. [ka]

[0180] 2-Nitrobenzyl (9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-9H-purin-6-yl)carbamate (5.00 g, 1 equiv., 7.59 mmol) was dissolved in THF at room temperature and then cooled to 0°C under a blanket of dry argon. To the mixture was then added tetrabutylammonium fluoride (4.96 g, 2.5 equiv., 19.0 mmol). The mixture was stirred at 0°C for 2 hours and then warmed to 23°C for 1 hour. The solution was poured into a cold solution of 10% NaHCO3 and extracted with DCM. The DCM layer was concentrated and the crude product was purified on silica gel eluting with 5-50% DCM / MeOH to give the product suitable for triphosphorylation. [ka]

[0181] 2-Nitrobenzyl (9-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-9H-purin-6-yl)carbamate (28 mg, 1 equiv., 65 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (30 mg, 18 μL, 3 equiv., 0.20 mmol) was added. After 5 min, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.23 g, 4 equiv., 0.26 mmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 s, with rxn t=35 min. Immediately, pre-weighed N1,N1,N8,N8-tetramethyl-naphthalene-1,8-diamine (56 mg, 4 equivalents, 0.26 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0182] Example 13 Detailed Procedure for Class III-Pyrimidine dNTP Analogs: [ka] To a solution of 4-amino-1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)-methyl)tetrahydrofuran-2-yl)pyrimidin-2(1H)-one (9.30 g, 1.1 equiv., 20.4 mmol) in dry DMF (100 mL) at 0° C. was added N,N-diisopropylethylamine (2.88 g, 3.9 mL, 1.2 equiv., 22.3 mmol). The reaction was stirred for 30 minutes, then 2-nitrobenzyl carbonochloridate (4.00 g, 1 equiv., 18.6 mmol) was added slowly dropwise over 30 minutes, maintaining the temperature below 5° C. The reaction was then warmed to room temperature and stirred overnight. The reaction was poured into a cooled solution of 5% Na2CO3 and EtOAc. The EtOAc layer was dried over sodium sulfate and then concentrated to dryness. The crude product was chromatographed on silica gel using hexane / EtOAc mixtures to give the purified product, which was used in the next reaction. [ka]

[0183] 2-Nitrobenzyl (1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)-tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamate (5.00 g, 1 equiv., 7.88 mmol) was dissolved in 25 mL THF at room temperature and then cooled to 0°C under a blanket of dry argon. To the mixture was then added tetrabutylammonium fluoride (5.15 g, 2.5 equiv., 19.7 mmol). The mixture was stirred at 0°C for 2 hours and then warmed to 23°C for 1 hour. The solution was poured into a cold solution of 10% NaHCO3 and extracted with DCM. The DCM layer was concentrated and the crude product was purified on silica gel eluting with 5-50% DCM / MeOH to give the product suitable for triphosphorylation. [ka]

[0184] 2-Nitrobenzyl (1-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamate (35.0 mg, 1 eq, 86.1 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (39.6 mg, 3 eq, 258 μmol) was added. After 5 min, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 eq, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise over 30 s to the reaction mixture at rxn t=35 min. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 equivalents, 345 μmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0185] Example 14 Procedure for synthesizing class IV-dCTP analogs. Scheme for the synthesis of class IV dCTP analogs [ka] Detailed procedure for Class IV dCTP analogs: [ka]

[0186] Methyl 3,4,5-trihydroxybenzoate (10 g, 1 eq, 54 mmol) was dissolved in 50 mL of acetone. Sodium iodide (0.81 g, 0.1 eq, 5.4 mmol) and potassium carbonate (38 g, 5 eq, 270 mmol) were added as solids at ambient temperature. 1-(Chloromethyl)-2-nitrobenzene (34 g, 3.6 eq, 0.20 mol) was added dropwise over 10 minutes as a solution in 40 mL of acetone. The mixture was stirred for 1 hour and then heated to 50°C for 6 hours. The mixture was cooled to ambient temperature, and the bulk of the solvent was removed on a rotary evaporator. The residue was suspended in 200 mL of EtOAc, which was washed successively with 200 mL portions of water and saturated aqueous NaCl. The EtOAc layer was dried over sodium sulfate and evaporated. The crude product was chromatographed on silica using hexane / EtOAc mixtures to give a purified product that could be used in the next reaction. [ka]

[0187] Methyl 3,4,5-tris((2-nitrobenzyl)oxy)benzoate (25 g, 1 equivalent, 42 mmol) was dissolved in 300 mL of THF. Aqueous 2 M NaOH (105 mL, 210 mmol) was added and the mixture was stirred at ambient temperature for 18 hours. The bulk of the THF was removed on a rotary evaporator and the residue was diluted with 6 M NaOH until a pH of 1 or less was reached. The resulting solid was filtered, washed thoroughly with water, and dried on the filter funnel for 5 hours, then dried under high vacuum for 18 hours. The product was carried forward to the next reaction without further purification. [ka]

[0188] 4-Amino-1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)pyrimidin-2(1H)-one (12 g, 1 equivalent, 26 mmol) and 3,4,5-tris((2-nitrobenzyl)oxy)benzoic acid (15 g, 1 equivalent, 26 mmol) were dissolved in 50 mL of dry DMF at ambient temperature under an argon atmosphere. N-Ethyl-N-isopropylpropan-2-amine (5.1 g, 6.8 mL, 1.5 equiv., 39 mmol) was added, followed by the dropwise addition of a solution of 1-((dimethylamino)(dimethyliminio)methyl)-1H-[1,2,3]triazolo[4,5-b]pyridine 3-oxide hexafluorophosphate (V) (12 g, 1.2 equiv., 31 mmol) in 10 mL of dry DMF over 5 minutes at ambient temperature. The mixture was stirred at ambient temperature for 18 hours. The mixture was dissolved in 300 mL of EtOAc, which was washed successively with 200 mL portions of water (2×) and saturated aqueous NaCl. The EtOAc was dried over sodium sulfate and evaporated first on a rotary evaporator and then under high vacuum for 18 hours. The residue was chromatographed on silica using a mixture of dichloromethane and methanol to give the desired product as a colorless foam. [ka]

[0189] N-(1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)-3,4,5-tris((2-nitrobenzyl)oxy)benzamide (5 g, 1 equivalent, 5 mmol) was dissolved in 25.0 mL of dry THF at ambient temperature under argon. Triethylamine (4 g, 6 mL, 8 equivalents, 4e+1 mmol) was added rapidly, followed by triethylammonium fluoride dihydrofluoride (5 g, 5 mL, 6 equivalents, 3e+1 mmol), also at ambient temperature. The mixture was stirred at ambient temperature for 24 hours. Silica gel (20 g) was added and the mixture was evaporated on a rotary evaporator to a fine powder, then loaded onto a 100 g silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam. [ka]

[0190] N-(1-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)-3,4,5-tris((2-nitrobenzyl)oxy)benzamide (30 mg, 1 eq, 38 μmol) was coevaporated with pyridine (1 mL × 3) and dried overnight under high vacuum. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of 6 μL of phosphoryl trichloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 min, a second aliquot of 5 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq., 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 eq., 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in the ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was performed by reverse-phase HPLC.

[0191] Example 15 Procedures for synthesizing Class V peptides and non-peptide analogs. Scheme for the synthesis of peptide-thymidine dNTP conjugates [ka] Scheme for the synthesis of non-peptide-thymidine dNTP conjugates [ka]

[0192] Example 16 Detailed procedure for peptide-dTTP analogs: [ka] 4-(Bromomethyl)benzenethiol (5.00 g, 1 equiv., 24.6 mmol) was dissolved in methanol (50 mL) and cooled to 0 °C. To the mixture was then added 1,2-di(pyridin-2-yl)disulfane (5.42 g, 1 equiv., 24.6 mmol) and stirred at 0 °C for 18 h. The reaction was then directly concentrated and purified on silica gel eluting with hexane / ethyl acetate (0-100% EtOAc) to give the product. Obtained as a white solid which was used directly in the next reaction. [ka]

[0193] 1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (5.00 g, 1 eq, 10.6 mmol) was dissolved in 50 mL DMF and then cooled to 0 C. The reaction was stirred at 0 C for 30 minutes and then sodium hydride (306 mg, 1.2 eq, 12.7 mmol) was added. The reaction was then stirred at 0 C for an additional 30 minutes and then warmed to 23 C. To the mixture was then added 2-((4-(bromomethyl)phenyl)-disulfanyl)-pyridine (3.32 g, 1 eq, 10.6 mmol) and stirring was continued at 23 C for an additional 2 hours. The reaction was then poured into a cold solution of 10% NaHCO3 and DCM. The DCM layer was separated, dried over sodium sulfate, and concentrated to dryness. The mixture was then purified on silica gel eluting with 0-20% DCM / methanol to give the desired product. [ka]

[0194] 1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)-methyl)tetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (2.00 g, 1 eq, 2.85 mmol) was dissolved in THF and cooled to 0° C. To the mixture was then added tetrabutylammonium fluoride (2.23 mg, 3 eq, 8.55 mmol) at 0° C. The reaction was continued to stir at 0° C. for 2 h and then warmed to 23° C. for an additional 1 h. The reaction was then cooled to 0° C. again and added to a pre-cooled solution of 10% NaHCO and DCM at 0° C. The DCM layer was then separated, dried over sodium sulfate, concentrated and purified on silica gel eluting with 5-50% DCM / methanol to give the pure product. [ka]

[0195] 1-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (5.00 g, 1 eq, 10.6 mmol) was dissolved in 20 mL of THF at 23° C. To the mixture was then added triethylamine (1.07 g, 1.5 mL, 1 eq, 10.6 mmol) and cooled to 0° C. To the mixture was then added TBS-Cl (1.59 g, 1 eq, 10.6 mmol) and stirring was continued at 0° C. for an additional 2 h. The mixture was then added to a pre-cooled mixture of 10% aqueous NaCl and DCM. The DCM layer was dried over sodium sulfate, concentrated, and dried to give an amber oil. The crude product was then purified on silica gel eluting with 5-50% DCM / methanol to give the pure product. [ka]

[0196] 1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-hydroxytetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (2.00 g, 1 eq, 3.40 mmol) was dissolved in 20 mL DMF and then cooled to 0° C. To the mixture was then added sodium hydride (98.0 mg, 1.2 eq, 4.08 mmol) and stirring was continued at 0° C. for an additional 30 minutes. To the reaction was then added 1-(bromomethyl)-2-nitrobenzene (735 mg, 1 eq, 3.40 mmol) and stirring was continued at 0° C. for an additional 1 hour. The reaction was then added to a pre-cooled mixture of 10% NaCl and EtOAc. The EtOAc layer was separated, dried over sodium sulfate and concentrated to dryness. The crude product was purified on silica gel eluting with 0-50 hexanes / EtOAc to give the desired product. [ka]

[0197] 1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-((2-nitrobenzyl)oxy)tetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (1.00 g, 1 equiv., 1.38 mmol) was dissolved in THF and cooled to 0° C. To the reaction was then added TBAF (362 mg, 1 equiv., 1.38 mmol) at 0° C. and stirred at 0° C. for 1 h, then warmed to room temperature over 2 h. The mixture was then poured into a pre-cooled solution of 10% NaHCO3 and DCM. The DCM layer was separated and dried over sodium sulfate. The crude product was purified on silica gel eluting with 5-25% DCM / methanol to give the desired product. [ka]

[0198] 4-(pyridin-2-yl)benzyl (1-((2R,4S,5R)-5-(hydroxymethyl)-4-((2-nitrobenzyl)oxy)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamate (35.0 mg, 1 eq, 61.0 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (39.6 mg, 3 eq, 258 μmol) was added. After 5 min, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 eq, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 seconds at rxn t=35 minutes. Immediately, pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 equivalents, 345 μmol) was added in one portion as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0199] Example 17 Detailed procedure for non-peptide-dTTP analogs: [ka] 1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)-oxy)methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (5.00 g, 1 eq, 10.6 mmol) was dissolved in 50 mL DMF and then cooled to 0° C. The reaction was stirred at 0° C. for 30 minutes and then sodium hydride (306 mg, 1.2 eq, 12.7 mmol) was added. The reaction was stirred at 0° C. for an additional 30 minutes and then warmed to 23° C. To the mixture was then added (bromomethyl)benzene (1.82 g, 1 eq, 10.6 mmol) and stirring was continued at 23° C. for an additional 2 hours. The reaction was then poured into a cold solution of 10% NaHCO and DCM. The DCM layer was separated, dried over sodium sulfate and concentrated to dryness. The mixture was then purified on silica gel eluting with 0-20% DCM / methanol to give the desired product. [ka]

[0200] 3-Benzyl-1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)-methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (4.00 g, 1 eq, 7.13 mmol) was dissolved in THF and cooled to 0° C. To the mixture was then added tetrabutylammonium fluoride (3.72 g, 2 eq, 14.26 mmol) at 0° C. The reaction was continued to stir at 0° C. for 2 h and then warmed to 23° C. for an additional 1 h. The reaction was then cooled again to 0° C. and added to a pre-cooled solution of 10% NaHCO and DCM at 0° C. The DCM layer was then separated, dried over sodium sulfate, concentrated and purified on silica gel eluting with 5-50% DCM / methanol to give the pure product. [ka]

[0201] 3-Benzyl-1-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (1.00 g, 1 eq, 3.01 mmol) was dissolved in 20 mL of THF at 23° C. To the mixture was then added triethylamine (304 mg, 0.42 mL, 1 eq, 3.01 mmol) and cooled to 0° C. To the mixture was then added TBS-Cl (453 mg, 1 eq, 3.01 mmol) and stirring was continued at 0° C. for an additional 2 h. The mixture was then added to a pre-cooled mixture of 10% aqueous NaCl and DCM. The DCM layer was dried over sodium sulfate, concentrated, and dried to give an amber oil. The crude product was then purified on silica gel eluting with 5-50% DCM / methanol to give the pure product. [ka]

[0202] 3-Benzyl-1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-hydroxytetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (6.00 g, 1 eq, 13.4 mmol) was dissolved in 20 mL DMF and then cooled to 0° C. To the mixture was then added sodium hydride (387 mg, 1.2 eq, 16.1 mmol) and stirring was continued at 0° C. for an additional 30 minutes. To the reaction was then added 1-(bromomethyl)-2-nitrobenzene (2.90 g, 1 eq, 13.4 mmol) and stirring was continued at 0° C. for an additional 1 hour. The reaction was then added to a pre-cooled mixture of 10% NaCl and EtOAc. The EtOAc layer was separated, dried over sodium sulfate, and concentrated to dryness. The crude product was purified on silica gel eluting with 0-50 hexanes / EtOAc to give the desired product. [ka]

[0203] 3-Benzyl-1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-((2-nitrobenzyl)oxy)-tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (1.50 g, 1 equiv., 2.58 mmol) was dissolved in THF and cooled to 0° C. To the reaction was then added TBAF (674 mg, 1 equiv., 2.58 mmol) at 0° C. and stirred at 0° C. for 1 h, then warmed to room temperature over 2 h. The mixture was then poured into a pre-cooled solution of 10% NaHCO3 and DCM. The DCM layer was separated and dried over sodium sulfate. The crude product was purified on silica gel eluting with 5-25% DCM / methanol to give the desired product. [ka]

[0204] Benzyl (1-((2R,4S,5R)-5-(hydroxymethyl)-4-((2-nitrobenzyl)oxy)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamate (35.0 mg, 1 eq, 70.5 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (39.6 mg, 3 eq, 258 μmol) was added. After 5 min, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 eq, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 s, with rxn t=35 min. Immediately, preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 equivalents, 345 μmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0205] Example 18 Procedure for synthesizing Class VI peptide and non-peptide dGTP analogs. Scheme for synthesis of class VI-dGTP constructs [ka] Detailed procedure for non-peptide dGTP analogs: [ka]

[0206] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1,9-dihydro-6H-purin-6-one (0.50 g, 1 eq, 1.0 mmol) was dissolved in 5.0 mL of dry dimethylacetamide under argon. Oxirane (0.13 g, 3 eq, 3.0 mmol) was added at ambient temperature, followed by sodium hydroxide (40 mg, 1 eq, 1.0 mmol) as a solid. The mixture was stirred at ambient temperature for 4 hours. The mixture was diluted with 50 mL of EtOAc, which was washed successively with 100 mL of water and 100 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was chromatographed on 40 g of silica using a dichloromethane / methanol mixture as eluent to give a white foam. [ka]

[0207] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1-(2-hydroxyethyl)-1,9-dihydro-6H-purin-6-one (200 mg, 1 eq, 370 μmol) was suspended in 5 mL of dry pyridine at ambient temperature under argon. Chloroformate (1 eq) was added as a solid. The mixture was heated to 95° C. for 8 hours and cooled to ambient temperature. The solvent was removed in vacuo and the residue was diluted with 50 mL of EtOAc, which was washed successively with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was chromatographed on 40 g of silica using dichloromethane / methanol mixtures as eluent to give a white foam. [ka]

[0208] The alcohol starting material was dissolved in dry THF under argon at ambient temperature. Two equivalents of triethylamine were added. A solution of acyl chloride in THF was added dropwise at ambient temperature, and the mixture was stirred for 18 hours. The solvent was removed in vacuo, and the residue was diluted with 50 mL of EtOAc, which was washed successively with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over Na2SO4 and evaporated to leave a light brown solid. This was chromatographed on a silica column using a dichloromethane / methanol mixture as the eluent to give the corresponding ester. [ka]

[0209] The bis-silyl ether (1 equivalent) was dissolved in dry THF under argon at ambient temperature. Triethylamine (8 equivalents) was added rapidly, followed by triethylammonium fluoride dihydrofluoride (6 equivalents), also at ambient temperature. The mixture was stirred at ambient temperature for 24 hours. Silica gel was added, and the mixture was evaporated to a fine powder on a rotary evaporator. It was then loaded onto a silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam. [ka]

[0210] The nucleoside was coevaporated with pyridine (1 mL x 3) and dried under high vacuum overnight. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (1.5 equivalents) was added. After 5 minutes, a second aliquot of 1.5 equivalents was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (4 equivalents) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed amount of N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (4 equivalents) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was performed by reverse-phase HPLC.

[0211] Detailed procedure for peptide-dGTP analogs: [ka]

[0212] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1,9-dihydro-6H-purin-6-one (1.00 g, 1 eq, 2.02 mmol) was dissolved in 30 mL of dry N,N-dimethylacetamide under argon. 4-Bromobutanoic acid (337 mg, 1 eq, 2.02 mmol) was added at ambient temperature, followed by sodium hydroxide (161 mg, 2 eq, 4.03 mmol) as a solid. The mixture was heated to 80° C. and stirred for 12 hours. The mixture was cooled to ambient temperature and diluted with 100 mL of EtOAc, which was washed successively with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over Na2SO4 and evaporated to leave a light brown solid, which was chromatographed on a silica column using a dichloromethane / methanol mixture as eluent to give 4-(2-amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-6-oxo-6,9-dihydro-1H-purin-1-yl)butanoic acid as a white solid. [ka]

[0213] 4-(2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-6-oxo-6,9-dihydro-1H-purin-1-yl)butanoic acid (1 equivalent) was suspended in 5 mL of dry pyridine at ambient temperature under argon. Chloroformate (1 equivalent) was added as a solid. The mixture was heated to 95° C. for 8 hours and cooled to ambient temperature. The solvent was removed in vacuo and the residue was diluted with 50 mL of EtOAc, which was washed successively with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was chromatographed on 40 g of silica using dichloromethane / methanol mixtures as eluent to give a white foam. [ka]

[0214] The carboxylic acid was dissolved or suspended in dry THF at ambient temperature. To this solution, 1.3 equivalents of triethylamine were added, followed by 1.1 equivalents of diphenylphosphoryl azide. The mixture was heated to reflux for 20 hours and cooled to ambient temperature. Silica gel was added to the mixture, and the solvent was evaporated to give a fine powder. This was loaded onto a silica gel column and eluted with a mixture of EtOAc and dichloromethane to give the desired isocyanate as a colorless oil. [ka]

[0215] The bis-silyl ether (1 equivalent) was dissolved in dry THF under argon at ambient temperature. Triethylamine (8 equivalents) was added rapidly, followed by triethylammonium fluoride dihydrofluoride (6 equivalents), also at ambient temperature. The mixture was stirred at ambient temperature for 24 hours. Silica gel was added, and the mixture was evaporated to a fine powder on a rotary evaporator. It was then loaded onto a silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam. [ka]

[0216] The nucleoside was coevaporated with pyridine (1 mL x 3) and dried under high vacuum overnight. It was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (1.5 equivalents) was added. After 5 minutes, a second aliquot of 1.5 equivalents was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (4 equivalents) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed amount of N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (4 equivalents) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was performed immediately after the EtOAc extraction. Final purification was performed by reverse-phase HPLC.

[0217] Decaging of 3'-O-(2-nitrobenzyl)-dATP and homopolymer synthesis are shown in Figures 20 and 21. 25 μM 3'-O-(2-nitrobenzyl)-dATP (TriLink Technologies, San Diego, CA) was mixed with 1 μM oligonucleotide initiator (5'-biotin-TTTTTTGGCCTTTTUTAATAATAATAATAATTTTT, IDT) along with 1× TdT reaction buffer (Thermo-Fisher), 2 U / μL terminal deoxynucleotidyl transferase (Thermo-Fisher), and 0.002 U / μL inorganic pyrophosphatase (Thermo-Fisher). The reaction volume was subjected to 20-22 mW / cm² of 365 nm light over various time intervals and then placed at 37°C for 30 minutes. After quenching by the addition of 0.1 M EDTA, each time point was mixed with an equal volume of 2× Novex TBE-urea gel loading buffer (Thermo-Fisher) and analyzed by polyacrylamide gel electrophoresis (15%) and Sybr The sections were stained with Gold (Thermo-Fisher) and photographed using an ultraviolet transilluminator.

[0218] Incorporation by Reference References and citations to other documents, such as patents, patent applications, patent publications, journals, books, articles, web content, etc. have been made throughout this disclosure. All such documents are incorporated herein by reference in their entirety for all purposes.

[0219] equivalent Various modifications of the invention and many further embodiments thereof, in addition to those shown and described herein, will become apparent to those skilled in the art from the entire contents of this document, including the references to the scientific and patent literature cited herein. The subject matter of this specification contains important information, exemplification and guidance that can be adapted to the practice of this invention in its various embodiments and equivalents thereof.

Claims

[Claim 1] The invention described in the specification.

Citation Information

Patent Citations

  • Methods of storing information using nucleic acids

    US9384320B2