Homopolymer-Encoded Nucleic Acid Memory

By encoding data with homopolymer tracts and employing enzymatic synthesis, DNA digital memory achieves faster, more accurate, and cost-effective data storage with reduced toxic by-products, overcoming limitations of conventional methods.

JP7710993B2Active Publication Date: 2025-07-22MOLECULAR ASSEMBLIES INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021563351
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-24
Filing Date
2020-04-24
Publication Date
2025-07-22
Estimated Expiration
2040-04-24

AI Technical Summary

Technical Problem

Current DNA digital memory technologies face limitations in speed, cost, and production of toxic by-products, with conventional synthesis techniques producing strands shorter than 200 base pairs and requiring post-synthesis amplification, and high-fidelity sequencing methods being slow and expensive.

Method used

The use of homopolymer tracts of repeated bases (2-10 nucleotides) to encode data, allowing for high-throughput sequencing techniques like nanopore sequencing and mass spectrometry, and enzymatic synthesis methods such as template-independent polynucleotide synthesis to create long strands (5-10 kb) with reduced waste and cost.

Benefits of technology

Enables faster and less expensive data storage with improved sequencing accuracy and tolerance for errors, using long nucleic acid strands synthesized enzymatically, which can be read multiple times without degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710993000097
    Figure 0007710993000097
  • Figure 0007710993000098
    Figure 0007710993000098
  • Figure 0007710993000099
    Figure 0007710993000099
Patent Text Reader

Abstract

Nucleic acid memory strands that encode digital data using sequences of homopolymer tracts of repeated nucleotides offer a cheaper and faster alternative to conventional digital DNA storage techniques. The use of homopolymer tracts allows the data encoded in the memory strands to be read using lower fidelity high-throughput sequencing techniques, such as nanopore sequencing. Specialized synthesis techniques enable the synthesis of long memory strands capable of encoding large volumes of data, despite the reduced data density afforded by homopolymer tracts compared to conventional single nucleotide sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority to U.S. Patent Application No. 16 / 393,510, filed on April 24, 2019, the content of which is incorporated herein by reference.

[0002] Field of the Invention The present invention relates to methods and apparatuses for storing data in nucleic acid memory strands containing homopolymer tracts.

Background Art

[0003] Background DNA digital memory is a process of representing digital data using the base sequence of DNA and storing that data via DNA synthesis of polynucleotides corresponding to the base sequences encoding the data. DNA digital memory offers several advantages over conventional data storage methods and targets a multi-hundred-billion-dollar market. Conventional data storage methods, including recording on flash memory and magnetic tape, pose problems related to physical space requirements, dependence on scarce resources, and data integrity. DNA digital memory provides a significantly lower energy requirement and a fairly large data storage density. Current methods rely on high-fidelity sequencing techniques that have little tolerance for errors in order to accurately read the data encoded in DNA. The required sequencing methods are relatively slow and expensive to meet the fidelity requirements. An example of current DNA digital memory techniques is described in U.S. Patent No. 9,384,320 to Church, et al. (incorporated herein by reference). To increase the fidelity of sequencing, current methods, such as those described by Church, encode data using sequences that avoid features that are difficult to read or write, such as sequence repeats.

[0004] On the synthesis side of current DNA digital memory techniques, the lack of speed, the production of toxic by-products, and the high cost further limit the adoption of this technology. Most de novo nucleic acid sequences are synthesized using solid-phase phosphoramidite techniques that involve sequential deprotection and synthesis of sequences constructed from phosphoramidite reagents corresponding to natural (or unnatural) nucleobases. Inkjet synthesis on an array-based format enables very low-cost phosphoramidite synthesis, but the strands produced are limited to 100 - 200 bases in length, a portion of which must be sacrificed to index the sequence, and require post-synthesis amplification to provide sufficient material for subsequent readout, and are produced on a scale below femtomolar concentrations. Using conventional synthesis techniques, nucleic acids longer than 200 base pairs (bp) in length undergo breakage and side reactions at a high rate. Additionally, conventional synthesis techniques produce toxic by-products, and the disposal of this waste limits the availability of nucleic acid synthesizers and increases the cost of oligo production. These troublesome problems associated with synthesis and readout in DNA digital memory have limited the application for otherwise promising technologies.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0006] Abstract The present invention provides a system and method for storing data using an array of homopolymer tracts that encode digital data. By using homopolymer tracts of repeated bases (e.g., 2 - 10 nucleotides) to represent each bit in a data array, it becomes possible to use more high - throughput and less expensive sequencing techniques. Array reading depends only on distinguishing transitions between homopolymer tracts and does not require reliable reading of each individual nucleotide, so sequencing techniques such as nanopore sequencing, zero - mode waveguide (ZMW) single - molecule sequencing, and mass spectrometry can be used to increase speed and reduce cost.

[0007] Recording of data using the homopolymer tracts described herein is most efficiently achieved using long strands (e.g., 5 - 10 kb) of nucleic acid. Conventional synthetic techniques have length limitations, but for example, template - independent polynucleotide synthesis using nucleotidyl transferase can synthesize long strands with reduced cost and lower waste generation. Enzymatically synthesized ssDNA memory strands require only 50% of the DNA synthesis compared to the conventional phosphoramidite approach, because ssDNA strands longer than about 100 - 200 nucleotides require complex and costly ligation or PCR techniques and can only produce ssDNA from dsDNA intermediates. See U.S. Patent No. 8,808,989 to Efcavitch, et al., which is incorporated herein by reference. Data encoding can be in base 2, base 3, base 4 using standard nucleotides, or the data density can be increased using any number of modified nucleotide analogs to generate an encoding scheme with a base of 8, 10, 12, or a larger base.

[0008] The only limitation on modified nucleotide analogs is that they can be incorporated using the selected synthetic technique (e.g., terminal deoxynucleotidyl transferase (TdT)) and can be distinguished from each other using the selected sequencing analysis. In some embodiments, synthesis can be achieved using polymerase theta in the presence of Mn 2+ The presence of can be achieved using polymerase theta.

[0009] A consistent homopolymer tract length is not essential for the systems and methods of the present invention, because it is only the transitions between individual tracts that need to be recognized. The tract length can potentially vary, but the synthetic techniques of the present invention can effectively control the average homopolymer tract length by adjusting the ratio of deoxynucleotides (dNTPs) to the oligonucleotide memory strand being synthesized and controlling the exposure time of dNTPs to the nascent memory strand. The length of the homopolymer tract can be optimized for the readout technique; the highest data storage density is achieved with single-nucleotide readout resolution, but the highest readout speed and accuracy are achieved by expanding the size of the nucleotide bits to the minimum detectable length (e.g., 2-10 nucleotides) for a given sequencing technique.

[0010] The systems and methods of the present invention using nanopore sequencing can use specialized memory strand constructs such as stoppers (e.g., hairpins or polymeric appendages) included on one or both ends of the strand. In other nanopore-based methods, the memory strand can be circularized and passed through between adjacent nanopores.

[0011] Certain embodiments of the present invention include a method of recording data using nucleic acid memory strands. The steps of this method can include creating an in-silico oligonucleotide sequence representing a data set, where each nucleotide of the oligonucleotide sequence corresponds to a unit of the data set. A nucleic acid memory strand can then be synthesized that includes a plurality of homopolymer tracts, where each homopolymer tract corresponds to a nucleotide of the oligonucleotide sequence. The plurality of homopolymer tracts can include from 3 to 10 repeating nucleotides. Each unit of the data set can be represented, optionally for a particular application, in base 2, base 3, base 4, or a higher base.

[0012] In certain embodiments, the nucleic acid memory strand can be at least about 200 nucleotides in length to about 5,000 nucleotides in length. The synthesizing step can include controlling the homopolymer length by varying the dNTP concentration. The steps of this method can include modifying a first end of the nucleic acid memory strand to prevent passage of the first end through a nanopore of a nanopore sequencing system; passing a second end of the nucleic acid memory strand through the nanopore; and modifying the second end of the nucleic acid memory strand to prevent passage of the second end through the nanopore.

[0013] Other embodiments can use memory strands composed of heteropolymer tracts of defined stoichiometry or composition to further increase the coding capacity of a set number of nucleotide analogs. Further embodiments can use nucleotide analogs with linkers that can be removed under structurally similar but different conditions, such as ultraviolet or visible light, oxidizing or reducing agents, alkaline or acidic pH, or sequence-specific nucleases, thereby aiming to protect the data encoded in the memory strand by obscuring the data from them without knowledge of the applicable process.

[0014] The dataset can be selected from the group consisting of text files, image files, and audio files. The synthesizing step may include template-independent synthesis. In certain embodiments, a nucleotidyl transferase enzyme may be used to catalyze the template-independent synthesis. In some embodiments, polymerase theta may be used to catalyze the template-independent synthesis.

[0015] Aspects of the invention may include a method of reading data from a nucleic acid memory strand. The steps of this method may include sequencing a nucleic acid memory strand that includes a plurality of homopolymer tracts, and converting the nucleic acid memory strand sequence into digitized data, wherein each of the plurality of homopolymer tracts indicates a nucleotide corresponding to a unit of data, and converting the digitized piece of data into a readable format. The steps of this method may include displaying the readable format. The plurality of homopolymer tracts may include repeats between about 2 nucleotides and about 10 nucleotides. The nucleic acid memory strand may be between at least about 200 nucleotides in length and about 5,000 nucleotides in length.

[0016] In various embodiments, the sequencing step may include nanopore sequencing, sequencing by synthesis, or mass spectrometry. The sequencing step, the converting step, and the transforming step may be repeated one or more times on the nucleic acid memory strand.

[0017] In certain embodiments, for example, the following items are provided. (Item 1) A method of synthesizing a plurality of nucleic acid memory strands, comprising: providing addressable delivery of activation energy to each of the substrate-linked nucleic acids in an array of two or more substrate-linked nucleic acids; extending one or more of the substrate-linked nucleic acids with a homopolymer tract of two or more repeating nucleotides by delivering addressable activation energy to one or more of the substrate-linked nucleic acids in the presence of a plurality of blocked nucleotide analogs and a template-independent polymerase, wherein the template-independent polymerase incorporates unblocked nucleotide analogs but does not incorporate blocked nucleotide analogs, and wherein the addressable activation energy converts the blocked nucleotide analogs to unblocked nucleotide analogs; A method comprising. (Item 2) The method according to item 1, wherein the blocked nucleotide analog is converted to an unblocked nucleotide analog by removal of a blocking group. (Item 3) The method according to item 1, wherein the addressable activation energy comprises light, reducing conditions, pH change or heat. (Item 4) The method according to item 2, wherein the removable blocking group is at the 3'-OH of the blocked nucleotide analog. (Item 5) The method according to item 2, wherein the removable blocking group is on the purine or pyrimidine base of the nucleotide analog. (Item 6) The method according to item 1, wherein the blocked nucleotide analogs comprise a removable blocking group on the 3'-OH of the deoxyribose or ribose of a nucleotide triphosphate and a non-removable modification on the purine or pyrimidine base of the nucleotide analog. (Item 7) The method according to item 1, wherein the plurality of blocked nucleotide analogs are modified nucleotides of the same nucleobase, comprising a removable 3'-O-blocking group and two or more non-removable molecular modifications that allow discrimination between modified nucleotide analogs of the same nucleobase. (Item 8) The step of terminating the extension; In the presence of a plurality of other blocked nucleotide analogs and said template-independent polymerase, extending said homopolymer tract with a further homopolymer tract of two or more repeating nucleotides by delivering an addressable activation energy to said homopolymer tract The method according to item 1, further comprising. (Item 9) The method according to item 8, wherein said extension is stopped after a predetermined length of time to obtain a desired length for said homopolymer tract. (Item 10) The method according to item 1, wherein the rate of extension is regulated by modification to said blocked nucleotide analog. (Item 11) The method according to item 10, wherein the rate-regulating modification is removed from said homopolymer tract after extension. (Item 12) The method according to item 1, wherein the rate-regulating modification is removed during extension. (Item 13) The method according to item 1, wherein said repeating nucleotides of said homopolymer tract are between 2 and about 10. (Item 14) The method according to item 8, further comprising repeating said stopping step and said extending step to synthesize a nucleic acid memory strand. (Item 15) The method according to item 14, wherein said nucleic acid memory strand is about 200 nucleotides to about 5,000 nucleotides in length. (Item 16) The method according to item 1, wherein a predetermined concentration of said blocked nucleotide analog is provided in said extending step to obtain a desired length for said homopolymer tract. (Item 17) The method according to item 8, wherein said homopolymer tract and said further homopolymer tract contain different nucleobases. (Item 18) The method according to item 14, wherein said nucleic acid memory strand encodes a data set selected from the group consisting of a text file, an image file, and an audio file. (Item 19) The method according to item 18, further comprising the step of displaying a readable format of said data set. (Item 20) The method according to item 14, wherein the unit of data is represented in base 2. (Item 21) The method according to item 14, wherein the unit of data is represented in base 3. (Item 22) The method according to item 14, wherein the unit of data is represented in base 4. (Item 23) The method according to item 14, wherein the unit of data is represented in a base larger than base 4. (Item 24) The method according to item 14, wherein the data is indicated by the degree of decay or the obtained tract length in individual steps of memory strand synthesis. (Item 25) The method according to item 14, wherein the data encoded in the nucleic acid memory strand is read by DNA sequencing. (Item 26) The method according to item 14, wherein the data encoded in the nucleic acid memory strand is read by passing the nucleic acid memory strand through a nanopore. Other aspects of the present invention will be apparent to those skilled in the art in view of the following drawings and detailed description.

Brief Description of the Drawings

[0018]

Figure 1

[0019]

Figure 2

[0020]

Figure 3

[0021]

Figure 4

[0022]

Figure 5

[0023]

Figure 6

[0024]

Figure 7

[0025]

Figure 8

[0026]

Figure 9

[0027]

Figure 10

[0028]

Figure 11

[0029]

Figure 12

[0030]

Figure 13

[0031]

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 14G

Figure 14H

Figure 14I

Figure 14J

Figure 14K

Figure 14L

[0032]

Figure 15

[0033]

Figure 16

[0034]

Figure 17

[0035]

Figure 18

[0036]

Figure 19

[0037]

Figure 20

[0038]

Figure 21

[0039]

Figure 22

[0040]

Figure 23

[0041]

Figure 24

[0042]

Figure 25

[0043]

Figure 26

[0044]

Figure 27

[0045]

Figure 28

[0046]

Figure 29

[0047]

Figure 30

Mode for Carrying Out the Invention

[0048] Detailed Description The present invention presents a system and method for writing data to and reading data from nucleic acids having homopolymer tracts corresponding to units of digital data. By repeating each nucleotide in a data encoding sequence (e.g., 3 to 10 times), in sequence determination reads, only transitions between homopolymer tracts need to be observed, which enables lower fidelity and higher throughput sequencing techniques that can result in less expensive implementation in nucleic acid data storage. The advantages of synthesizing nucleic acid homopolymer tract memory strands are as follows: 1) the ability to create very long (5 to 10 kb) strands that enable the use of high throughput long read DNA sequencing technology for reading, 2) the ability to tolerate errors in sequencing read technologies, and 3) the ability to create nucleic acid memory strands at a much lower cost than those of conventional chemical synthesis methods. The use of homopolymer nucleic acid memory strands is best realized in long (e.g., 5 to 10 kb) strands that can be efficiently produced using a template-independent TdT enzyme or polymerase theta, where the homopolymer tract length can be controlled by varying the exposure time and the ratio of dNTPs to the polynucleotide memory strand.

[0049] The synthesis of homopolymers for encoding data by an enzymatically mediated approach is readily achieved by using natural or modified nucleotide triphosphates that are not terminators, enabling the simplest and fastest method of DNA synthesis. One natural or modified nucleotide triphosphate is delivered to a reaction zone having a nucleotidyl transferase, the reaction occurs, and then it is removed by washing with buffer to complete one "write" cycle of data storage illustrated in FIG. 1. Data strand synthesis occurs in a completely aqueous environment without the accompaniment of toxic or hazardous chemicals, thus enabling a practical device suitable for large-scale data storage centers.

[0050] Figure 2 shows a method 101 for synthesizing a nucleic acid memory strand having a homopolymer tract according to a particular embodiment. Method 101 includes a step 103 of creating an in-silico oligonucleotide sequence representing a data set. The data set can include digitized data that can represent text, images, videos, audio, or any other piece of information that can be digitized. The oligonucleotide sequence can include any number of natural or modified nucleotides or their analogs, and depending on the number of unique nucleotides or analogs used in the memory strand, can encode the data set using a base-2, base-3, base-4, or higher base scheme. In a simple embodiment, the encoding scheme can correspond to a binary data scheme conventionally represented by a series of 0s and 1s, where one or more nucleotides or analogs can correspond to 0 and one or more other nucleotides can correspond to 1. A nucleic acid memory strand (e.g., RNA, single-stranded or double-stranded DNA) including a series of homopolymer tracts each corresponding in order to the nucleotides in the in silico oligonucleotide sequence can then be synthesized 105. In a particular embodiment, further steps include a step 107 of modifying one end of the memory strand, a step 109 of passing the memory strand through a nanopore, and a step 111 of modifying the other end of the strand to prevent the end from passing through the nanopore.

[0051] Figure 3 shows a method 203 for reading data from a nucleic acid memory strand having a homopolymer tract. The steps of method 203 include a step 203 of sequencing a series of homopolymer tracts in the nucleic acid memory strand, a step 205 of converting the sequence into a data set, a step of converting the data set into a readable format (e.g., an image, a video, an audio clip, or a piece of text), and a step 209 of displaying the data in the readable format as needed (e.g., on a monitor, or using a printer or other input / output device).

[0052] Preferably, the system and method of the present invention use long strands of DNA (5-10 kb) that can be either single-stranded or double-stranded and can occur naturally or be generated by chemical or enzymatic synthesis. In certain embodiments, the nucleic acid memory strand can be enzymatically generated using TdT to create a series of homopolymer tracts that may be 2-10, 3-10, 4-10 nucleotides or longer. Each homopolymer tract can consist of adenine (A), guanine (G), cytosine (C), or thymine (T). An alternating sequence of homopolymer tracts can be used to encode the data stored in the memory strand.

[0053] Each nucleotide homopolymer tract can represent various amounts of data depending on the number of bases used. The number of bits required to make up 1 byte (256 in decimal) is defined by the following relationship: # bits / byte = 8 / (log2(n)), where n = the base used. Each tract can correspond to 1 bit if a 2-base encoding is used, or 1 / 4 of a byte if a 4-base encoding is used. In certain embodiments, the DNA data strand can be composed of homopolymer tracts of 2 to 10 nucleotides (using a base-2 data set representation), which allows 333 bits to 100 bits to be represented in the memory strand between 999 bases long and 1000 bases long. In a preferred embodiment, the nucleic acid base encoding of the data can be an encoding such that a single homopolymer tract of one nucleotide is always adjacent to homopolymer tracts of different nucleotides. For example, the encoding can be such that a homopolymer tract of adenine is not preceded or followed immediately by another adenine homopolymer tract. When two adjacent homopolymer tracts contain the same nucleotide, a homopolymer tract measurably longer than the average homopolymer tract representing a single nucleotide in the encoded data sequence can be synthesized. These longer tracts can be created via manipulation of the synthesis reaction described below, for example, by increasing the concentration of dNTPs during the reaction or by increasing the reaction time. The exact length of two adjacent identical homopolymer tracts need only be long enough to be uniquely identified from a single homopolymer tract using a readout device (i.e., a nanopore sequencer). In certain embodiments, non-nucleotide homopolymer spacers can be added between A, G, C, or T homopolymer tracts to clearly distinguish adjacent identical nucleotide homopolymer tracts from each other.The use of A, G, C, and T homopolymer tracts allows for the creation of a 4-bit encoding space that increases the density of data that can be stored in a single continuous strand, rather than simply using two nucleotides (similar to 0 and 1 in binary code). For example, four consecutive homopolymer tracts can encode 256 numbers (i.e., 1 byte) when A, G, C, and T are used in a base-4 scheme. In such embodiments, when using 3-nucleotide long homopolymer tracts or 10-nucleotide long homopolymer tracts respectively, 83 bytes or 25 bytes are represented in a 996 or 1000 nucleotide long nucleic acid memory strand.

[0054] In various embodiments, base-8 or even base-12 coding schemes can be used through the incorporation of uniquely modified nucleotides or non-nucleotide analog homopolymer tracts into the memory strand. These modified nucleotides or non-nucleotide analogs should generate unique digital signals with a readout device such as a nanopore sequencer or a single molecule ZMW sequencer. TdT, discussed below, can enhance the signal provided by a readout device such as a nanopore and thus can be used to incorporate a wide range of modified dNTP analogs that can generate nucleic acid memory strands encoded with data therein. Modified nucleotides (e.g., A * , G * , C * and T * or A ** , G ** , C ** and T ** ) homopolymers can be used with TdT and modified dNTP analogs of each of the four bases (e.g., dATP or dA * TP or dA **It can be synthesized using TP). Higher base (n) encodings enable data compression and result in a reduction in the number of DNA strands required to encode a given amount of information. The relationship determining the number of DNA strands per 1 GB of data, the base (n), and the synthesized strand length as a function of the homopolymer tract length, as illustrated in Figure 4, is defined by: # strands / GB = (8 / (log2(n)) * 10 9* Homopolymer length * 1 / strand length.

[0055] The number of unique homopolymer tracts can be limited only by the ability of the readout technology (i.e., nanopore or ZMW single-molecule sequencing) to distinguish one homopolymer tract from another. There are several reports in the literature of the detection of homopolymers composed of unmodified nucleotides by detecting changes in the ionic current during translocation through a nanopore (Venta et al, 2013; Feng et al, 2015). Modifications that alter the dwell time of DNA in a nanopore generate distinguishable and characteristic ionic current signals. Singer et al 2010 and Morin et al 2016 use non-covalently linked bisPNA or γPNA functionalized with 5 kDa or 10 kDa PEG to enhance detection by a nanopore. Liu et al 2015 selectively created adamantly 8-oxoG analogs to modify the dwell time and generate unique signals. Considering the tolerance of TdT for incorporating bulky modifications at N6 of dATP, N4 of dCTP, N2 or O6 of dGTP, and O4 or N3 of dTTP, acyl or alkyl modifications at these positions can be screened and selected to enhance the detection modality of nanopore or ZMW single-molecule sequencing technologies. Detection can be improved through modified nucleotides that enhance differential current blockade in a nanopore or enhance the dwell time of modified nucleotides in the active site of DNA polymerase in a ZMW single-molecule approach. Other natural and non-natural purine and pyrimidine nucleotide analogs can be used if they generate unique digital signals on readout devices such as nanopore sequencers or single-molecule ZMW sequencers. Modifications at C5 or C7 of pyrimidines and purines, respectively, can be used if they generate unique digital signals on readout devices such as nanopore sequencers or single-molecule ZMW sequencers.Suitable modified nucleotide triphosphates are selected to be incorporated rapidly during the enzymatic extension step and provide substitution-specific residence times with the shortest possible homopolymers during the detection step. Examples of modifications to the A, G, C, and T bases suitable for expanding the bit encoding space include, but are not limited to, N6-benzoyl dA, N6-benzyl dA, N6-alkyl dA, N6-acyl dA, N6-substituted alkyl dA, N6-substituted acyl dA, N6-aryl acyl dA, N6-substituted aryl acyl dA, N2-alkyl-dG, N2-acyl dG, N2-aryl acyl dG, N2-substituted alkyl dG, N2-substituted acyl dG, N2-substituted aryl acyl dG, O6 alkyl dG, O4 alkyl dT, N3 alkyl dT, N3 acyl dT, O6-substituted alkyl dG, O4-substituted alkyl dT, C5-propynylamine dT, C5-propynylamine dC, C7-propynylamine dA, C7-propynylamine dG, substituted C5-propynylamine dT, substituted C5-propynylamine dC, substituted C7-propynylamine dA, substituted C7-propynylamine dG. Preferred embodiments of the substitution include, but are not limited to, covalent bonds that are completely stable to removal except under the most extreme chemical conditions of pH, temperature, and concentration of the reactive species. Substitutions that can affect the unique current blocking include, but are not limited to, alkyl, heteroatom-substituted alkyl, aromatic hydrocarbon, alkyl-substituted aromatic hydrocarbon, heteroatom-substituted alkyl-substituted aromatic hydrocarbon, heteroatom-substituted aromatic hydrocarbon, benzyl, substituted benzyl, or combinations thereof. In some embodiments, the substitution can be polyethylene glycol composed of 2 to 450 monomer units. In some embodiments, substitutions composed of peptides or peptoids can be suitable for increasing the residence time of homopolymers in a specific and distinguishable manner. The efficiency of incorporation of modified nucleotides by template-independent polymerases such as TdT can be regulated by the use of different metal ion cofactors such as, but not limited to, Co++, Zn++, Mg++, Mn++, or a mixture of two or more different metal ions.Each modified nucleotide may require different metal ions for optimal performance during enzymatic homopolymer synthesis.

[0056] Since the long-term stability of the DNA data strand is important, there are distinct advantages to using non-purine-based homopolymers, because they are subject to depurination at low pH. In some embodiments, the homopolymer bits can consist of only a single nucleotide type (i.e., thymine) modified with two, three, four, or more different chemical groups, creating homopolymer tracts each causing a unique current blockage. Thus, one nucleotide labeled with four unique modifiers can replace the presence of A, G, C, T. Other embodiments are possible using only one of the other three nucleotides having two, three, four, or more different chemical groups.

[0057] Another embodiment uses single nucleotide bits instead of homopolymer bits when the modified nucleotide analog creates a unique dwell time for the passage of a single nucleotide through the nanopore. The single modified nucleotide bit is advantageous in that it allows for the maximum density of information per DNA strand, thus reducing the cost of DNA-based data storage.

[0058] As long as the array determination technique used for reading can clearly distinguish the start and stop of one homopolymer tract from another homopolymer tract, the exact length of the homopolymer tract is not important. Increasing the number of unique nucleotides or bases used (including modified nucleotides or non-nucleotide analogs), and thus reducing the length required to capture a set amount of data, has distinct advantages in synthesis and storage density, but the lowest cost per DNA data storage synthesis can be achieved by using the four natural nucleotide dNTP monomers during enzymatic synthesis, because these reagents are widely used in the fields of molecular biology and sequencing and are produced in very large batches at the lowest manufacturing costs. The cost of producing dNTP analogs to increase the number of unique homopolymer tracts can be reduced as the use of DNA data storage increases and the manufacturing scale of the analogs also increases.

[0059] Any method of synthesizing homopolymer tract segments can be used with the systems and methods of the present invention, but a preferred embodiment uses the template-independent enzyme TdT. TdT provides certain benefits as long as it rapidly and inexpensively generates homopolymers with a Poisson distribution, where the average size of the homopolymer can be precisely controlled by the ratio of [dNTP] to the nascent oligonucleotide memory strand. In some embodiments, polymerase theta in the presence of Mn 2+ can be used as a template-independent polymerase for synthesizing homopolymer tract nucleic acid memory strands. In another embodiment, the length of the homopolymer tract segment can be controlled by delivering an excess amount of dNTP to the reaction zone and then removing the reactants after carefully controlled time intervals.

[0060] TdT has demonstrated the ability to synthesize homopolymer tracts of a properly defined length by controlling the ratio of the dNTP concentration to the concentration of the 3' end of the nucleic acid strand to be modified. Inkjet synthesis on an array-based format enables very low-cost phosphoramidite synthesis, but the strands produced are limited to 100 - 200 bases in length, sacrificing part of the length to index the sequence and requiring post-synthesis amplification to provide sufficient material for subsequent reads, are produced on the femtomolar concentration scale and are mainly suitable for relatively low-efficiency short-read sequencing readout technologies.

[0061] The strands of single-stranded DNA synthesized according to the process of the present invention can benefit from the prevention of hairpins or dsDNA, either during synthesis or during reading. Hairpin formation can be prevented by modifying the exocyclic amine of one member of an A:T or G:C base pair to prevent the hydrogen bonding necessary for base pairing. In some embodiments, the exocyclic amine can be modified by acylation or alkylation. Any simple and stable modification of the exocyclic amine of A, G, or C that prevents base pairing can be used to prevent hairpin formation. In certain embodiments, the N6 of deoxyadenosine and the N2 of deoxyguanosine can be acetylated with an acetyl group that prevents base pairing. In some embodiments, the N6 of deoxyadenosine and the N4 of deoxycytidine can be modified to prevent base pairing and hairpin formation. In some embodiments, the O6 of deoxyguanosine or the O6 of thymidine can be modified to prevent base pairing and hairpin formation. In some embodiments, the O4 of thymidine or the N3 of thymidine can be modified to prevent base pairing and hairpin formation. In some embodiments, modifications to A, G, C, or T to generate higher-order base encoding schemes are also suitable for the purpose of preventing base pairing and hairpin formation. In some embodiments, homopolymer bits can be composed of only a single nucleotide type (i.e., thymine) modified with two, three, four, or more different chemical groups, resulting in unique current blocking and preventing the formation of intra- or intermolecular double-stranded regions. In another embodiment, a thermostable version of TdT or another template-independent nucleotidyl transferase can be used to perform strand synthesis at elevated temperatures, thus preventing the formation of intra- or intermolecular double-stranded regions.

[0062] Control of homopolymer tract length can be optimized for any of the above analogs after determination and calibration of the incorporation rate of dNTP analogs to create a reproducible range of homopolymer tract lengths from 2 to 10 nucleotides in length. A * , G * , C * and T* and A ** 、G ** 、C ** and T ** The use of homopolymer tracts enables the creation of 8-bit or 12-bit encodings, increasing the density of data that can be stored in a single continuous strand compared to simply using two nucleotides to encode "0" and "1". Three consecutive homopolymer tracts, A, G, C, T, A * 、G * 、C * and T * when used, can encode the number 256 (i.e., 1 byte). In such embodiments, when a 3-nucleotide long homopolymer tract or a 10-nucleotide long homopolymer tract is used respectively, there are 111 bytes or 33 bytes in a 999 or 990 nucleotide long nucleic acid memory strand. Two consecutive homopolymer tracts, A, G, C, T, A * 、G * 、C * 、T * 、A ** 、G ** 、C ** and T ** when used, can encode the number 256 (i.e., 1 byte). In these embodiments, when a 3-nucleotide long homopolymer tract or a 10-nucleotide long homopolymer tract is used respectively, there are 166 bytes or 50 bytes in a 996 or 1000 nucleotide long nucleic acid memory strand.

[0063] In certain embodiments, data may also be encoded into random arrays and memory strand heteropolymer tracts of defined composition to achieve a higher level of data compression. Heteropolymer stretches can be generated using enzymatic reactions that use mixtures of different dNTPs, where dNTP stoichiometry is used to control the composition of the heteropolymer tract. The number and type of heteropolymer tracts are limited only by the combination of dNTP analogs and the ability of the detection modality to distinguish between the compositions of different tracts. For m dNTP analogs, there are (m 2 -m) / 2 binary combinations for heteropolymer formation. A detection modality that can distinguish between two different levels of tract composition for each binary combination (e.g., a tract in which analogs A and B are present in a ratio of approximately 2:1, respectively, and a tract in which they are present in a ratio of 1:2) enables data to be encoded at a rate of base m 2 from a set of m analogs, effectively doubling the coding capacity of the memory strand. Figure 5 illustrates the data that can be stored in a memory strand as a function of the number of available dNTP analogs, using either a homopolymer or a binary heteropolymer-based encoding scheme with two levels of tract composition.

[0064] The data encoding strands of the present invention may not necessarily require precisely defined homopolymer lengths, provided that they are long enough (about 2-10 nucleotides) to enable unambiguous discrimination of transitions between homopolymer tract segments by high-throughput DNA sequencing techniques. Existing next-generation synthesis-based sequencing (SBS) systems can readily determine the transition between two adjacent homopolymer tracts. Again, the exact length of the homopolymer tract is not critical for accurate detection of homopolymer bits. The use of tracts of the same nucleotide provides an advantage in overcoming the most common errors in current SBS platforms: insertions and deletions. A deletion of one nucleotide in a homopolymer tract longer than 2nt is still interpreted as a true homopolymer. Similarly, a single nucleotide insertion in a homopolymer tract is not misinterpreted as two adjacent homopolymers, as insertions of more than one nucleotide during SBS are low-probability events. This sequencing error tolerance provides the advantage of reducing the sequencing depth required to ensure correct decoding of the information stored by the DNA data strand. Existing nanopore systems can readily distinguish homopolymer tracts of A, G, C, or T from one another based on their differential current blockades. In certain embodiments, single molecule ZMW sequencing can be used to determine the linear order of homopolymer tracts on a strand. The use of any sequencing technique may require that the DNA initiator have properties compatible with the sequencing readout technique, such as a self-complimentary hairpin at the 5' end of the synthesized single-stranded memory strand, to provide a primer for single molecule ZMW sequencing. Nanopore sequencing techniques may also require a self-complimentary hairpin at the 5' end of the strand to provide a "start" data mark. In various embodiments, the readout technique can be any next-generation sequencing method, such as those provided by Illumina (San Diego, CA). In some embodiments, the readout or sequencing technique can be mass spectrometry-based.The technology-specific error rate of a readout technique is not important as long as the technique can unambiguously detect transitions between two different homopolymer tracts and / or can unambiguously detect the difference between one homopolymer tract length and one 2× length when two identical homopolymer tracts are adjacent to each other.

[0065] Certain readout techniques may be more preferred than others based on the specific application of the present invention. Techniques such as nanopore sequencing can be non-destructive, can leave the nucleic acid memory strand intact, and can be suitable for multiple readout cycles. Readout techniques that rely on synthesis-based sequencing (SBS), such as ZMW single molecule, generate a copy of the original template strand and then require post-read operations (i.e., strand separation by melting) to return the original nucleic acid memory strand to its initial state where it is ready for subsequent cycles of readout and removal of the complementary strand. Other readout techniques, such as mass spectrometry, are destructive and deplete the pool of nucleic acid memory strands after repeated cycles of sampling and readout.

[0066] In various embodiments, the nucleic acid memory strand can include a "stopper". The "stopper" can be a polymeric construct that prevents the passage of single-stranded or double-stranded nucleic acids through the nanopore (Manrao, et al., 2012, Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase, Nature Biotechnology 30, 349-353, which is incorporated herein by reference). Proteins such as phi29 DNA polymerase are large enough not to be pulled through the larger pore (about 6.3 nm) on the cis side of the protein nanopore. The diameter of the smaller side pore is estimated to be about 1.2 nm wide. A stopper can be used either at the 5' end or the 3' end of the nucleic acid memory strand of the present invention. In some applications, it may be desirable to have stoppers at both the 3' and 5' ends of the nucleic acid molecule. The stopper can consist of a hairpin (stem-loop) structure having either a protruding 5'-overhang or 3'-overhang to which the information-encoding nucleic acid is covalently attached. When the stopper consists of a hairpin, the length of the ds stem can be made long enough to resist any melting force exerted on it by the electric field used to translocate the memory strand through the nanopore. In certain embodiments, one base of the double-stranded stem region can be cross-linked to its cognate base that forms the base pair such that it is impossible for the double-stranded stem portion of the hairpin to melt under the influence of the force exerted on it by the electric field that translocates the remainder of the molecule through the nanopore. For utilization of the hairpin stopper for TdT-mediated nucleic acid memory synthesis according to certain embodiments, the stopper can have a 3'-overhang that is long enough (i.e., >10 nucleotides) to allow binding of TdT for template-independent synthesis.

[0067] The stopper can consist of a non-nucleotide polymer construct that can be appended to either the 3'- or 5'-end of a nucleic acid molecule. The construct can be synthesized by direct conjugation of a polymer species onto the 3'-end of a nucleic acid by a polymerase or a transferase such as TdT (Sorensen, et al., 2013, Enzymatic Ligation of Large Biomolecules to DNA, ACS Nano, 7(9):8098-8104, incorporated herein by reference), or by incorporation of a functionalized nucleotide that enables specific modification of the nucleic acid via its functionality (Winz, et al., 2015, Nucleotidyl transferase assisted DNA labeling with different click chemistries, Nucleic Acids Res. 43(17):e110, incorporated herein by reference). The 5'-end stopper can be readily introduced at the time of chemical synthesis of an oligonucleotide adapter via direct synthesis of a hairpin or via secondary modification of a functional handle introduced as the last step of 3'-to-5' oligonucleotide synthesis and can be used as an initiator. Alternatively, the 5'-end stopper can be constructed by binding an oligonucleotide initiator to magnetic or non-magnetic beads or particles or nanoparticles via the 5'-end, enzymatically synthesizing a memory strand containing a homopolymer tract, and then leaving the memory strand bound to the magnetic or non-magnetic beads or particles or nanoparticles.

[0068] The stopper can be further modified to enable cleavage of the stopper from the rest of the molecule to allow the nucleic acid strand to passively diffuse out of the nanopore or to be translocated out of the nanopore via application of a voltage, thus allowing the strand to be retrieved.

[0069] In certain embodiments, a template-independent polymerase or transferase can be used to modify pre-synthesized strands of nucleic acids in order to enable the use of nanopore devices as "write-once-read-many" type memory devices. Some of the unique problems associated with the use of nanopore devices as DNA sequencers are the high error rates they introduce due to poor discrimination by the nanopore. This can be due to the speed of translocation through the pore, or the fact that the approximate depth of the nanopore is 8 nm, allowing multiple bases to be present in the pore simultaneously. The homopolymer memory strands of the present invention address this problem through the use of homopolymer repeats, reducing the need for exact sequencing accuracy. In certain embodiments, the drawbacks of nanopore sequencing can be addressed by implementing hairpin adapters at one end of the double-stranded DNA memory strand, such that during the translocation and base calling processes, each sense of the DNA memory strand can be read such that reading the individual bases and their complementary strands can cancel out the error rate when each base is read only once. In certain embodiments, the fidelity of nanopore sequencing can be increased by appropriate modification of each end (5' and 3') of a single-stranded or double-stranded nucleic acid molecule having a bulky appendage (e.g., a protein or solid state) that does not translocate across the pore. The molecule can then be trapped within the pore and translocated forward and backward multiple times to allow multiple reads of the same molecule in the same pore, and thus the sequencing error rate can be reduced by the square of the number of reads (if the sequencing read errors are due to stochastic origins).

[0070] Transferases such as TdT can be used to append large, bulky modified nucleotide analogs to the 3' end of a DNA molecule. In certain embodiments, a WORM nanopore memory device can be generated using the following steps: (1) generating a single molecule of DNA encoding specific information in any high-density encoding scheme as discussed above and covalently modifying the 5' end with a bulky molecular construct that prevents complete translocation of the DNA molecule through the nanopore; (2) passing the DNA molecule through the nanopore until the 5'-modified end contacts the nanopore and can no longer translocate; (3) using TdT and modified nucleotides to covalently add one (or more) bulky nucleotide analogs ("stoppers") to the 3' end of the DNA molecule to effectively trap the molecule within the torus of the nanopore; (4) reversing the polarity of the current to the nanopore to remove any unmodified 3'-DNA molecules and thus create a pure population of "trapped" (5'- and 3'-modified) nucleic acid strands; (5) removing any untrapped nucleic acids from the vicinity of the nanopore via washing or other means; (6) using the applied voltage to read the "trapped" DNA strand in either direction or both directions (potentially reading multiple times to reduce the error rate to an acceptable level). In various embodiments, step 6 can consist of a voltage-induced "read" in one direction and a rapid translocation in the opposite direction to "rewind" the data-encoding nucleic acid through the nanopore followed by another voltage-induced "read" in the original direction. This cycle of "read" - "rewind" - "read" can be repeated any number of times as desired.

[0071] In some embodiments, the trapped nucleic acid strand can be read during translocation in either direction. In some embodiments, the trapped strand can be translocated to one end (either 5' or 3') of the molecule and read in the opposite direction, such that the reading polarity can provide a higher accuracy of reading.

[0072] In certain embodiments, the circularized nucleic acid memory strand can be generated using the above synthesis method and subsequent circularization. The circularized strand can include a bulky polymer or a specific homopolymer sequence, where the ends of the synthesized strand are joined to specify a starting point and a stopping point for data reading. The start and stop homopolymer sequences can also be used in linear nucleic acid strands. The circularized strand 305 can pass through between two adjacent nanopores (306 and 309) such that the circularized strand 305 is physically trapped between two nanopores (306 and 309) located on a single membrane 303 as shown in FIG. 6. The circularized memory strand can encode digital information as any one of a sequence of single nucleotides, a homopolymer tract sequence, a sequence of modified nucleotide analogs, or a combination of some of them. One nanopore 309 can be used to generate an electrical signal when the information-encoded memory strand is translocated through the pore, while the other nanopore 307 can simply function as a portal to allow the DNA molecule to return to the cis side of the membrane 303 and the first nanopore 309. One advantage of this scheme is that the information-encoding strand can be recycled for repeated reading, can be read multiple times, and thus can reduce any possible read errors.

[0073] Other embodiments may encode data in a memory strand such that the data can only be accessed under a specific set of conditions. In such cases, the memory strand is at least partially composed of nucleotides containing modifications, linked by a cleavable linker. The modifications (e.g., chemical protecting groups) and linkers can be selected such that current blockade is different from the sequence encoding the data when the polymer tract translocates through the nanopore without the correct processing. Figure 7 outlines a scheme using disulfide and amide-linked modifications to dG nucleotides and illustrates how current blockade and the data encoded in the memory strand can vary in response to processing conditions. G * and G ** are structurally similar in size and flexibility and can produce similar current blockade on the nanopore platform but are removed under different conditions. G * and G *** are structurally different, but these modifications share the same removal conditions. Other embodiments may use other cleavable modifications or linkers using different treatments such as light of a specific wavelength, acidic or alkaline pH, oxidative or reductive conditions, or sequence-specific nucleases. Some embodiments may use the presence or absence of memory strand modifications for encryption or as a chemical marker of previous access or alteration to the data. Most linker cleavage reactions are effectively irreversible, and thus this approach may be best suited for a write-one read many system where a single molecule may be sufficient to encode data without redundancy.

[0074] Many possible information encoding schemes useful in the readout scheme of the present invention are possible and may be apparent to those skilled in the art based on the present disclosure.

[0075] Synthesis can be achieved using acoustic delivery of droplets into the wells of a plate (e.g., a 1536 well plate with 1.5 μL each). In various embodiments, the nucleic acid memory strands can be synthesized on beads or magnetic beads or on a surface and, depending on the application, can remain on the beads or magnetic beads or on the surface after full-length synthesis is complete, or can be removed from the synthesis support.

[0076] In certain embodiments, a system for synthesis of long (5 - 10 kb) data strands can use inkjet delivery into an array of wells (e.g., wells of multiple nanoliter volume). In other embodiments, a plurality of air pressure controlled actuators can be positioned above each well to simultaneously deliver reagents to each position in the array. Each actuator is served by a selector valve that selects between each of two or more nucleotides or modified nucleotides formulated with a template-independent polymerase used to specify bits of the DNA data strand. One or more additional selector valve ports are dedicated to one or more wash reagents as needed. The array of nanoliter volume wells can be open at both ends as long as the diameter of the wells is such that the delivered liquid is trapped by capillary action within the length of the open-ended well. After each round of nucleotide-enzyme formulation is delivered to the open-ended wells, each well is rinsed with a reaction mixture and then a rinse reagent or enzyme reaction stop reagent can be flowed across and through the lower opening of the well array to prepare the array for the next cycle of enzymatic synthesis. In other embodiments, a vacuum source is used to rapidly remove one reagent from the capillary nanowell prior to delivery of the next reagent.

[0077] Certain embodiments may use a highly parallel nanofluidic chamber with valve-controlled reagent delivery. An exemplary microfluidic nucleic acid memory strand synthesis device is shown in FIG. 8 for illustrative purposes and is not to scale. A microfluidic channel 255 including a regulator 257 couples a reservoir 253 to a reaction chamber 251, and an outlet channel 259 including a regulator 257 removes waste from the reaction chamber 251. A microfluidic device for nucleic acid memory strand synthesis may include, for example, channel 255, reservoir 253, and / or regulator 257. Nucleic acid memory strand synthesis may occur in a microfluidic reaction chamber 251 that may include several anchored synthetic nucleotide initiators that can releasably bind to a polynucleotide initiator that is anchored or bound to the inner surface of the reaction chamber and that may include beads or other substrates that can releasably bind to a polynucleotide initiator as needed. The reaction chamber 251 may include at least one inlet channel and one outlet channel 259 such that reagents can be added to and removed from the reaction chamber 254. The reaction chamber 251 must be temperature controlled to maintain optimal and reproducible enzymatic synthesis conditions. The microfluidic device may include a reservoir 253 for each respective dNTP or analog used in the memory strand coding scheme. Each of these reservoirs 253 may also include an appropriate amount of TdT or any other enzyme that elongates DNA or RNA strands in a template-independent manner. Additional reservoirs 253 may include reagents for washing or other tasks.

[0078] Reservoir 253 can be coupled to reaction chamber 254 via separate channels 255, and the flow of reagent into reaction chamber 254 via each channel 255 can be individually regulated via the use of gates, valves, pressure regulators or other means. The flow from reaction chamber 254 via outlet channel 259 can be similarly regulated. Reservoir 253 can hold dNTPs, modified dNTPs or any of their analogs as described above, suspended in a fluid at a known concentration, such that the concentration of the reagent can be precisely controlled based on the volume of reagent flowing into reaction chamber 254. Thus, the length of each homopolymer tract can be managed via control of the reagent concentration.

[0079] In certain examples, reagents, particularly dNTPs and enzyme reagents, can be recycled. The reagents can be drawn back from reaction chamber 254 into their respective reservoirs 253 via the same channels 255 through which they entered, by inducing backflow using gates, valves, vacuum pumps, pressure regulators or other regulators 257. Alternatively, the reagents can be returned from reaction chamber 254 to their respective reservoirs 253 via separate return channels. The microfluidic device can include a controller capable of operating the gates, valves, pressure or other regulators 257 described above.

[0080] An exemplary microfluidic nucleic acid memory strand synthesis reaction involves flowing a desired dNTP (used throughout to refer to any component molecule used to encode data in the nucleic acid memory strand of the present invention) reagent at a predetermined concentration for a predetermined amount of time (calculated to result in a desired homopolymer length) into reaction chamber 254 prior to removing NTP reagent from reaction chamber 254 via outlet channel 259 or a return channel (not shown); flowing a wash reagent into reaction chamber 254; removing the wash reagent from reaction chamber 254 via outlet channel 259; flowing the next NTP reagent to a desired memory strand sequence under conditions calculated to achieve a desired homopolymer tract ratio; and repeating until a desired nucleic acid memory strand is synthesized. After the desired nucleic acid memory strand is synthesized, it may be released from the reaction chamber anchor or substrate and collected via outlet channel 259 or other means.

[0081] Due to the fairly large number of homopolymer-encoded DNA strands required to encode a useful amount of data, a highly parallel method of DNA synthesis is required. In some embodiments, as shown in FIG. 9, a flow cell (9-1) containing an array of wells forms a plurality of hydrophilic wells (?) bounded by hydrophobic regions by patterning horizontal and vertical stripes (9-2) of hydrophobic material on a suitable substrate. Typical dimensions of the hydrophilic wells can be from 300×300 nm to 1000×1000 nm. This hydrophilic array forms the floor of the flow cell with a gap of appropriate dimensions between the floor and the optically transparent cover. A solution of cold (i.e., lower than the optimal enzyme-specific temperature) nucleotidyl transferase, one of the natural or modified nucleotide triphosphates, and any necessary cofactors is flowed into the flow cell via an inlet (9-3) such that upon cessation of fluid flow, the enzyme-nucleotide triphosphate solution forms into a bead-like spatially defined droplet (9-4) above each hydrophilic region. An IR source (9-5) projects a beam (9-7) onto a DLP (Digital Light Projection) device (9-8) via a shaping lens (9-6), which is used to simultaneously direct an IR beam (9-9) to each of the hydrophobic-bound droplets (9-4) selected to have a particular nucleotide added, resulting in a rapid heating of the polymerase extension reaction formulation to the temperature for maximum enzyme activity over a defined period for synthesizing a homopolymer of the desired length. After some appropriately defined reaction time, the IR source is switched off and cold rinse buffer is rapidly injected into the flow cell and drained via an outlet (9-10), thus quenching the reaction and completing one “write” cycle. This series of steps is repeated multiple times for each “write” cycle such that each nascent data strand is randomly accessed according to its spatial position and the selected nucleotide is added until a full-length homopolymer data strand is completed. In some embodiments, a DLP device with a 1920×1080 steerable mirror can be used to simultaneously and randomly access approximately 2M synthesis positions in the synthesis flow cell.In some embodiments, the template-independent polymerase used is thermophilic at an optimal reaction temperature well above room temperature, such that enzyme activity is minimized in the hydrophobically confined droplets prior to rapid heating of the droplets by the IR source. In some embodiments, the bottom surface of the flow cell carrying hydrophobically defined hydrophilic wells abuts a cooling device that maintains the droplets in the hydrophobic wells at a reduced temperature to prevent enzyme activity until the temperature is raised by the IR source.

[0082] In another embodiment, a flow cell composed of an array of wells is formed on a suitable substrate by patterning horizontal and vertical stripes of a hydrophobic material that form a plurality of hydrophilic spots bounded by hydrophobic regions. Typical dimensions of the hydrophilic spots can be from 300×300 nm to 1000×1000 nm. Each hydrophilic spot is positioned on an individually addressable CMOS heater. This hydrophilic-CMOS heater array forms the floor of the flow cell with a gap of appropriate dimensions between the floor and an optically transparent cover. A solution of cold (i.e., lower than the optimal enzyme-specific temperature) nucleotide transferase and one of the natural or modified nucleotide triphosphates is flowed into the flow cell such that, upon cessation of fluid flow, the polymerase extension reaction solution balls up into spatially defined droplets positioned above each hydrophilic region having an associated CMOS heater. Each of the hydrophobicity-bound droplets, selected to have a particular nucleotide added, is rapidly heated to the temperature for maximum enzyme activity over a defined period for synthesizing a homopolymer of a desired length. After some appropriately defined reaction time, the heater is switched off and a cold rinse buffer is rapidly injected into the flow cell, thus quenching the reaction and completing one “write” cycle. This series of steps is repeated multiple times for each “write” cycle such that each nascent data strand is randomly accessed according to its spatial position and the selected nucleotide is added until a full-length homopolymer data strand is completed. In some embodiments, the enzyme used is thermophilic at an optimal temperature well above room temperature, such that prior to the rapid heating of the droplets by the CMOS heater, the likelihood of unwanted nucleotide addition in the hydrophobicity-bound droplets is low.

[0083] Those skilled in the art will recognize the systems and methods of the present invention as necessary or best suited. The systems and methods of the present invention may include one or more processors 309 (e.g., central processing unit (CPU), graphics processing unit (GPU), etc.) that communicate with each other via a bus, a computer-readable storage device 307 (e.g., main memory, static memory, etc.), or a combination thereof, and may include a computing device as shown in FIG. 10. Computing devices may include mobile devices 101 (e.g., mobile phones), personal computers 901, and server computers 511. In various embodiments, the computing devices may be configured to communicate with each other via a network 517.

[0084] The computing device may be used to control the synthesis of memory strands, the reading of sequenced memory strands, and the compilation or conversion of nucleic acid sequences between data in human or machine-readable formats, digitized data, and other steps described herein. The computing device may be used to display data in a readable format.

[0085] Processor 309 may include any suitable processor known in the art, such as a processor sold under the trademark XEON E7 by Intel (Santa Clara, CA), or a processor sold under the trademark OPTERON 6200 by AMD (Sunnyvale, CA).

[0086] Memory 307 preferably includes at least one tangible non-transitory medium capable of storing the following: one or more sets of executable instructions (e.g., software embodying any methodology or function found herein) for causing the system to implement the functions described herein; data (e.g., data encoded in a memory chain); or both. The computer-readable storage device may be a single medium in an exemplary embodiment, but the term "computer-readable storage device" should be interpreted to include a single medium or multiple media (e.g., centralized or distributed databases as well as / or associated caches and servers) that store instructions or data. The term "computer-readable storage device" should be construed to include, without limitation, solid-state memory (e.g., subscriber identity module (SIM) cards, secure digital cards (SD cards), micro SD cards or solid-state drives (SSDs)), optical and magnetic media, hard drives, disk drives, and any other tangible storage medium.

[0087] Any suitable service such as Amazon Web Services, the memory 307 of server 511, cloud storage, another server, or other computer-readable storage may be used for storage 527. Cloud storage may refer to a data storage scheme where data is stored in a logical pool and physical storage may span multiple servers and multiple locations. Storage 527 may be owned and managed by a hosting company. Preferably, storage 527 is used to store record 399 to implement and support the operations described herein as needed.

[0088] The input / output device 305 according to the present invention may include one or more of a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT) monitor), an alphanumeric input device (e.g., a keyboard), a cursor control device (e.g., a mouse or a trackpad), a disk drive unit, a signal generation device (e.g., a speaker), a touch screen, buttons, an accelerometer, a microphone, a cellular radio frequency antenna, a network interface device such as a network interface card (NIC), a Wi-Fi card or a cellular modem, or any combination thereof.

[0089] One of ordinary skill in the art will recognize that any suitable development environment or programming language may be used to enable the operations described herein for the various systems and methods of the present invention. For example, the systems and methods herein may be implemented using Perl, Python, C++, C#, Java®, JavaScript®, Visual Basic, Ruby on Rails, Groovy and Grails, or any other suitable tools. For the computing device 101, it may be preferable to use native xCode or Android Java®.

[0090] Figure 11 shows polyacrylamide gel electrophoresis analysis of two different single homopolymer tracts made via enzymatic synthesis. Lane A is a sample of the 20-mer starting oligonucleotide used in all the following lanes. Lane B is a sample from a TdT reaction containing the 20-mer oligonucleotide and the irreversible terminator ddATP, showing the formation of 21-mers. Lane C is a sample from a TdT reaction containing the 20-mer and the natural nucleotide dATP after 1 minute at 37°C. Lane D is a sample of the same reaction mixture in Lane D after 5 minutes at 37°C. Lane E is a sample of the same reaction mixture in Lane C after 15 minutes at 37°C. These three lanes show homopolymer length control due to consumption of the input dATP in about 5 minutes during the TdT extension reaction, and that no homopolymer growth is observed between 5 and 15 minutes. Lane F is a sample from a TdT reaction containing the 20-mer oligonucleotide and the nucleotide analog N6-benzoyl-dATP after 1 minute at 37°C. Lane G is a sample of the reaction mixture in Lane F after 5 minutes at 37°C. Lane H is a sample of the same reaction mixture in Lane F after 15 minutes at 37°C. There is a qualitative difference in the length of the dA homopolymers formed in Lanes F–H, but the same length control has also been demonstrated when using N6-modified dATP analogs. Bz Although there is a qualitative difference in the length of the homopolymers, the same length control has also been demonstrated when using N6-modified dATP analogs.

[0091] Writing digital data into a molecular memory format based on a molecular approach offers advantages over currently used memory media such as tapes or disks. DNA-based memory has attracted interest because of its achievable high information density, extremely long lifespan, and low energy consumption during dormant periods. To date, most attempts to use synthetic DNA as a memory medium have involved chemical synthesis using the common phosphoramidite method.

[0092] DNA-based data storage can require a very large number of strands compared to what is currently synthesized for existing research markets. Any synthetic technique (chemical or enzymatic) that relies on nucleotide blocking or the removal of terminators imposes additional steps and complexity because the reagents need to be delivered in an addressable manner to an array of synthetic features (i.e., wells or spots) in order to direct the correct nucleotides to the correct positions on the array. Array-based methods of synthesis can be carried out in one of several ways: 1) bulk delivery of activated reactants followed by selective removal of blocking groups, or 2) addressable delivery of activated reactants followed by bulk blocking group removal, or 3) bulk delivery of inactive reactants followed by addressable activation. The addressable delivery of reagents to large (10 4 ~10 6 ) 2-D arrays is generally achieved using inkjet deposition to each desired location. If the data encoding scheme uses four nucleotides, four separate write heads are used and must be indexed in a complex X-Y mechanical manner. Additionally, the use of inkjet delivery limits the dimensions between each feature (well or spot) to the low tens of microns. The challenge in this process is to reduce the step time for each synthesis cycle, however it is achieved.

[0093] In certain embodiments, the systems and methods of the invention can include delivering an inert reaction mixture to all features on a 2-D array and then selectively activating only those features that require the addition of A or G or C or T. Addressable methods of delivery, activation, or blocking group removal that do not rely on mechanical motion are preferred. Preferred embodiments use bulk delivery of reactants and selective activation of specific synthesis features by removal of blocking groups from nucleotide analogs, which then allows rapid DNA polymerase-mediated incorporation and formation of homopolymer bits as shown in FIG. 15, and thus enables highly parallel and rapid synthesis of homopolymer-encoded nucleic acid memory strands. There are several other advantages to homopolymer-encoded nucleic acid information polymers: 1) homopolymer-encoded bits overcome the error profiles associated with next-generation sequencing, 2) the resulting polymers are non-natural nucleic acids and thus cannot be diverted for bioterrorist activities.

[0094] Delivery and selective activation of template-independent polymerase DNA synthesis can be performed sequentially (deliver A -> addressably activate and initiate homopolymer synthesis -> wash; deliver C -> addressably activate and initiate homopolymer synthesis -> wash; deliver G -> addressably activate and initiate homopolymer synthesis -> wash; deliver T -> addressably activate and initiate homopolymer synthesis -> wash) or in parallel (deliver all four nucleotides to all features simultaneously -> activate A, C, G, T addressably, either simultaneously or sequentially -> wash). In this manner, the synthesis cycle becomes very efficient and includes only the following three steps: reactant delivery, activation of the incorporation reaction, and then washing prior to the start of the next cycle. The reaction is stopped by rapid removal of reactants by gas or liquid, or by rapid delivery of a quenching reagent. In preferred embodiments, the quenching reagent is a metal chelator.

[0095] A preferred design for an apparatus that can be used to synthesize homopolymer bit-encoding memory strands consists of a flow cell composed of a 2-D array of hydrophobic, patterned wells that are appropriately modified to support template-independent enzymatic synthesis. After delivering reactants to the 2-D array of hydrophobic wells, the liquid forms into beads that define spatially distinct reaction zones as shown in FIG. 16.

[0096] A preferred embodiment is one in which the bottom surface of the well, surrounded on four sides by hydrophobic patterning, is modified with a 5’->3’ oriented covalently bound oligonucleotide initiator with the 5’ end attached to the bottom surface of the well. In another embodiment, the well is physically formed by etching a depression in the bottom surface of the flow cell, in which case the covalently bound oligonucleotide initiator is covalently attached to the surface of the well (bottom and sides). In some cases, the well is open at the bottom to allow for liquid flow through the well. In other cases, the well is closed at the bottom.

[0097] Nucleotide analogs protected at the 3'-OH are generally inactive with commercially available or wild-type TdT enzymes (U.S. Patent No. 10,059,929). Thus, a mixture of 3'-blocked dNTP analogs, TdT protein, and appropriate cofactors can be mixed together at 37 °C in the presence of an initiator oligonucleotide with little to no homopolymer formation. When the 3'-OH is "decaged" (i.e., made unblocked) by removal of the protecting or blocking group, the resulting nucleotide is available for free running incorporation and homopolymer formation. Decaging that can be achieved by a "deliveryless" method is preferred. Analogs constructed to enable addressable 3'-OH decaging by activation energy, e.g., light, heat, electrochemical generation of a pH change, or administration or exposure to a reducing agent, are preferred. Each of these decaging or unblocking reactions can be achieved in a properly constructed flow cell using mechanically steerable light (e.g., FIG. 17) through a light-transmissive cover, individually addressable heaters (e.g., FIG. 18), electrochemically induced pH changes, or generation of reducing conditions.

[0098] Each of the illustrated flow cell designs is compatible with a method of addressable decaging of nucleotides contained in droplets confined by a hydrophobic patterning around them. FIG. 19 illustrates an apparatus that can be used with a flow cell designed for the use of photo-decaged nucleotide analogs. Other embodiments and configurations known to those of skill in the art are possible for this purpose.

[0099] Each decaging mechanism requires a nucleotide analog specifically designed for the physicochemical process selected for use in the system:

Chemical formula

[0100] In some embodiments, a nucleotide analog suitable for photo-mediated de-caging is 3’-O-(2-nitro-benzyl)-dNTP. In some embodiments, a nucleotide analog suitable for heat-mediated de-caging is 3’-O-(tetrahydrofuranyl)-dNTP. In some embodiments, a nucleotide analog suitable for reductive-mediated de-caging is 3’-O-methyl-dithiomethyl-dNTP. Many other 3’-OH protecting groups are suitable as long as the resulting 3’-OH modified nucleotide analog is not a substrate for template-independent polymerase and is readily removed by a “delivery-free” method, such as a reactant generated photochemically, thermally or electrochemically. The 3’-O modifications described herein are explicitly stated to be substrates for template-independent polymerase, and as long as they are called reversible terminators, the composition of these dNTP analogs is different from that described in WO2016 / 034807. The subject of the present invention is a 3’-O-blocking group that is not explicitly a substrate for polymerase and functions to cage dNTP until it is removed to allow polymerization to proceed. Mathews AS et al (2016) describe the use of 3’-O-(2-nitrobenzyl)-dNTP analogs as reversible terminators for the controlled enzymatic synthesis of natural non-homopolymeric oligonucleotides. They explicitly teach the use of such nucleotides as substrates for DNA synthesis by using extremely long (about 1 hour) enzyme reaction times, and further teach the use of these analogs as caging groups for initiating the enzymatic polymerization of multiple nucleotides. In contrast to the subject of this patent, which explicitly teaches their use as reversible terminators, their use as reversible terminators is explicitly taught. The initiation of free-running homopolymer synthesis can be achieved by methods other than caged dNTP analogs. Some embodiments may use fluid pulses of polymerase to initiate and stop ssDNA synthesis, as described in Church 9,928,869 (2018). In Reza et al WO2017 / 196783, template-dependent enzymatic synthesis is initiated by activation of polymerase by an electrochemically generated pH change.Modified nucleotide analogs containing 3'-O-reversible terminators have been described, but no patents teach the activation of homopolymer synthesis using the uncaging of caged dNTP analogs. Church 9,928,869 (2018) describes free-running homopolymer synthesis using natural dNTPs and controlling the length of the formed homopolymer by the reaction duration. Lee HR et al (2018) describe the use of mechanical delivery to deposit natural nucleotides in multiple reaction zones on a 2D array and actively destroy unreacted unmodified dNTPs using apyrase to control the length of homopolymer formation. In some embodiments of the subject matter of this patent, homopolymer rate-modified dNTP analogs having an unmodified 3'-OH are used in combination with either a fluid pulse or mechanical XY delivery or pH-activated template-independent DNA polymerase.

[0101] Figures 20 and 21 show the kinetics of photo-mediated uncaging. 3'-caged dNTPs in a template-independent polymerase reaction mixture are polymerized into nascent homopolymer chains when converted to substrates for the polymerase. Since available uncaged dNTPs are readily incorporated by the polymerase, complete conversion of the caged dNTPs to uncaged (substrate) dNTPs is not required. When uncaged, the length of the desired homopolymer is controlled by controlling the concentration of the generated uncaged dNTPs and / or the time of the enzymatically mediated incorporation reaction.

[0102] Figure 22 shows examples of two 3’-O-modified caged nucleotide analogs that can be decaged with visible wavelength light and exhibit tunable photophysical properties (Peterson J.A., et al J. Amer. Chem. Soc. 2018 140:7343-6), which enable the simultaneous introduction, decaging, and synthesis of two different homopolymer strand sequences. In some embodiments, nucleotide analogs modified with two different 3’-caging species that can be decaged by two different wavelengths of light are simultaneously delivered to a 2D array of synthesis sites. In another embodiment, four dNTP analogs modified with four different photocleavable 3’-caging species are simultaneously delivered to a 2D array of synthesis sites and decaged with four different wavelengths of light. The decaging reactions can be performed sequentially for one set of positions on the 2D array, followed by delivery of the two remaining dNTPs and subsequent sequential decaging, thus completing one round that increases the length of all informational polymers by one nucleotide. In an alternative embodiment, four dNTP analogs modified with four different 3’-caging species can be delivered simultaneously and decaged by simultaneous exposure of each synthesis feature to one of four different wavelengths of light. Hardware configurations such as those shown in Figure 19 can be used for multicolor decaging by incorporation of two or more wavelength-specific light sources, and a suitable shutter mechanism for directing one of two or more of the light sources onto the 2D light directing system.

[0103] Figure 23 shows the synthesis of a 12-bit information strand consisting of consecutive cycles of homopolymer synthesis using dCTP, dGTP, and dTTP incorporated with 3'-oNBn-dATP that was UV-light decayed in cycles #1, 6, 8, 10, and 12. The incorporation reaction can be terminated by several methods including, but not limited to, the rapid removal of enzymatic reaction components by the introduction of a bolus of either gas or liquid. In some embodiments, the liquid simply rinses away the reaction components, while in other embodiments, the rinse liquid contains an active quencher such as EDTA or other enzyme inhibitors.

[0104] In another embodiment, the method of caging dNTPs can be mediated by a steric mechanism involving modification to the nucleotide base instead of the 3'-OH. In a preferred embodiment, nucleotides with an unmodified 3'-OH are modified at N6, N4, N2, or O4 of A, C, G, or T, respectively, with a removable steric blocking group (incorporation blocking factor) that cages the dNTP. Homopolymer synthesis can be initiated by cleavage of the steric blocking group by either photo-mediated, thermo-mediated, or reduction-mediated means and thus by "uncaging" of the natural nucleotide appropriate for free-running homopolymer synthesis.

[0105] In some embodiments, the modified nucleotide is composed of a linker containing one moiety that keeps the dNTP caged and thus non-incorporable until removed, and another moiety that remains covalently attached throughout multiple homopolymer tract syntheses. In some embodiments, the purine or pyrimidine base is modified at two positions, one modification caging the nucleotide and preventing enzymatic incorporation, while the other modification remains covalently attached throughout polymer synthesis; each modification is removable by a different mechanism allowing for each selective removal. In some embodiments, the purine or pyrimidine base is modified at two positions, and each modification is removable by a different mechanism allowing for each selective removal.

[0106] Depending on the array detection modality used to read homopolymer bit data strands, appropriate dNTP analog modifications can be designed to yield nucleotides without or with lesions. For readout methods that rely on sequencing by synthesis (SBS), all modifications to purines or pyrimidines must be removed after homopolymer synthesis but before SBS readout. For readout methods that rely on current modulation during polymer translocation through one or more nanopores, modifications to purines and pyrimidines that are readily distinguishable from each other are desired. In some embodiments, modifications that modulate polymerase kinetics during homopolymer synthesis are made to purine and pyrimidine bases. These "speed-regulating" modifications can also act as current modulators for nanopore detection or can be removed for SBS detection.

[0107] Regardless of the activation mechanism, it is important to control the length of free continuous homopolymer synthesis. Ideally, homopolymers of 2 - 4 nucleotides are desirable. In the presence of natural nucleotides, TdT has a tendency to form homopolymers based on the nucleobases corresponding to the Km (A > T > G > C) of individual nucleotides. As pointed out by Lee HR et al (2018), it is difficult to limit homopolymer growth without increasing the deletion frequency. The above authors reported deletion errors of about 66% in an attempt to limit homopolymer growth to 2 - 3 nucleotides. TdT has been reported to operate in a distributive manner, resulting in a Poisson distribution of homopolymer lengths in free continuous synthesis. However, in practice, without a long enzymatic elongation reaction time, it is difficult to drive the complete conversion of the initiator nucleic acid during homopolymer synthesis. A free continuous homopolymer synthesis reaction time that is too short results in bit deletion errors, as reported above. A homopolymer synthesis reaction time that is too long results in overly long homopolymer formation and low - efficiency data density. One solution is to use modified nucleotide analogs that regulate the rate of multiple base incorporations but do not require removal at all steps, thus enabling complete conversion of the starting nucleic acid, control of the length of the homopolymer formed, and maintaining the simple two - step cycle that is the subject of this patent. Ideally, the homopolymer tract in the informational polymer is greater than 2 nucleotides but 4 or fewer nucleotides in length.

[0108] Optionally, the modification that limits the homopolymer length can be removed from all nucleotides at the end of synthesis, thus generating a native DNA molecule suitable for SBS detection. Modifications that are chemically compatible with the above de-caging conditions are most desirable; they must be removable by chemical conditions orthogonal to those used for de-caging. In some embodiments, the 3'-OH is caged by 3'-O-(2-nitrobenzyl), while N6 of dA, N4 of dC, N2 of dG, and O4 or N3 of dT are modified with non-terminating moieties that regulate multiple additions during free continuous enzymatic homopolymer synthesis. Figure 24 shows a comparison of free continuous homopolymer synthesis in the presence and absence of rate-regulating modifications. In some embodiments, the homopolymer synthesis rate-regulating modification is the same for all four nucleotide analogs, while in other embodiments, the modification is different for each of the dATP, dCTP, dGTP, and dTTP analogs. In some embodiments, the rate-limiting modification remains in place throughout the synthesis.

[0109] Figures 25 and 26 show the results of two consecutive addition cycles of two differently modified nucleotide analogs. In another embodiment, the nascent information polymer is exposed to periodic or alternating cycles having chemical conditions that result in partial or complete removal of the homopolymer rate-limiting modification. In some embodiments of the use of class 2 dNTP analogs, the previously incorporated rate-regulating modification is removed simultaneously with polymerase extension by including a mild reducing agent in alternating cycles of the enzymatic extension reaction mixture.

[0110] When non-SBS methods of readout (i.e., nanopores) are used, homopolymer rate-modulating modifications can be covalently linked using non-cleavable linkages and left in place during readout. In some embodiments, homopolymer rate-modulating modifications can also be used to encode information during the readout process. In preferred embodiments, homopolymer synthesis rate-modulating analogs are also designed to modulate the current during nanopore translocation and are further selected to provide two or more distinguishable and resolvable current blockade levels. In another embodiment, only a single type of nucleotide (i.e., A or C or G or T) is used in information strand synthesis and homopolymer bits are encoded by two or more differential current-blockade generating modifications.

[0111] Figures 27 - 30 show examples of detection of enzymatically synthesized homopolymers of peptide-modified dNTP analogs by solid-state nanopores. Using a dTTP analog (N-Ac-CYPEE) modified by a cleavable disulfide linkage to a 5-amino acid peptide, fully modified molecules were generated via free-running homopolymer synthesis longer than 100 nt in length. Translocation through a 2D silicon nitride 30 nm thick nanopore at 500 mV resulted in a large signal swing compared to an unmodified dU homopolymer (courtesy of Goeppert LLC, Philadelphia, PA).

[0112] The following table describes six different classes of modified dNTP analogs useful for three different "delivery-free" homopolymer synthesis activation approaches and two different homopolymer-encoded nucleic acid memory strand readout techniques.

Table 1

[0113] Examples of four photo-mediated degrading nucleotide analogs that make up Class I are shown below:

Chem.

[0114] Analogs of this class represent 3'-O-caged dNTPs that become native unmodified dNTPs upon photo-mediated uncaging. Upon uncaging, the rate of homopolymer formation is no different from that of native nucleotides, and the length of the homopolymer formed needs to be adjusted by additives to the template-independent polymerase reaction mixture. In some embodiments, tetrahydropyranyl, 4-methoxy-tetrahydropyranyl, tetrahydrofuranyl, acetyl, methoxyacetyl or phenoxyacetyl modifications are useful for thermally induced uncaging of 3'-O-blocked dNTP analogs. In some embodiments, 3'-O-analogs such as -O-CH2-S-S-R are useful for electrochemically mediated uncaging under reducing conditions. Class I dNTP analogs are characterized by nucleotides that give rise to unmodified homopolymers suitable for readout by either classical synthetic sequencing or a ratchet-style nanopore. In both cases, accuracy in the length of the homopolymer is not required, but accurate detection of transitions between homopolymers is required.

[0115] Examples of four class II photo-mediated uncaging nucleotide analogs with orthogonally removable peptide rate-modulating modifications are shown below:

Chemical formula

Chemical formula

[0116] Peptide modifications for use with class II dNTP analogs are intended to be removed at the end of informational polymer synthesis for SBS-dependent sequencing methods that are sensitive enough to detect transitions from one homopolymer type (A, C, G, T) to the next, or for subsequent interrogation by translocation through nanopores. Peptides suitable for use in the present application slow down the rate of incorporation of modified nucleotide analogs to prevent the formation of long homopolymers prior to complete conversion of unmodified or modified strands of different compositions. In some embodiments, the incorporation rate modulating modification can consist of 1, 2, 3, or 4 different peptides for the 4 nucleotide analogs. In some embodiments, the peptide is linked to the nucleotide via a disulfide linker having a self-sacrificial moiety cleavable under mild chemical conditions. In some embodiments, the peptide can consist of, but is not limited to, Ac-EECGY, Ac-EEGCGW, Ac-EEGCGGW, Ac-EC-pNA, Ac-CWEE, Ac-CYPEE, Ac-EEGCPPW, Ac-CPYEE, Ac-CPWEE, Ac-CWPEE. Many other peptide sequences and compositions are possible to those skilled in the art. Peptides with an overall anionic composition are thought to be most suitable; peptides with a cationic composition accelerate the rate of nucleotide incorporation, as previously pointed out by Finn PS et al 2003. Peptides with covalent linkages to nucleotides other than via disulfides to cysteine amino acids are possible, so long as the ability to remove them from the nucleotides of a completed homopolymer DNA strand without leaving residual moieties is maintained. Disulfide-cleavage mediated self-sacrificial linkers that are eliminated under mild conditions by formation of thiolactones are particularly useful.

[0117] The following are examples of four nucleotides of class II with photo-mediated decay and orthogonally removable non-peptide homopolymer synthesis rate modulating modifications:

Chemical formula

[0118] Non-peptide modifications to N6 of adenine, N4 of cytidine, N2 of guanine, and N3 of thymine can act as incorporation rate modulators. In some embodiments, acetyl, diacetyl, isobutyryl, propionyl, pivaloyl, benzoyl, cyclohexyl, and other organic modifiers are useful. This type of modifier is removed after information polymer synthesis, by ammonolysis, among other things. Class II dNTP analogs are characterized by nucleotides that result in unmodified homopolymers suitable for either classical synthesis-based sequencing or readout by a ratchet-style nanopore. In both cases, accuracy in homopolymer length is not required, but accurate detection of transitions between homopolymers is required.

[0119] The following are examples of two Class III nucleotide analogs having an unmodified 3'-OH and a removable incorporation caging modification on a purine or pyrimidine:

Chemical formula

[0120] In some embodiments, useful caging modifications to the 2-nitrobenzyl functionality are bulky modifications such as, but not limited to, peptides, cyclic peptides, PEG, branched PEG, star PEG, dendrimers, nanoparticles, etc. Class III dNTP analogs are also characterized by nucleotides that result in unmodified homopolymers suitable for either classical synthesis-based sequencing or readout by a ratchet-style nanopore. In both cases, accuracy in homopolymer length is not required, but accurate detection of transitions between homopolymers is required.

[0121] An example of one Class IV nucleotide analog having an unmodified 3'-OH, a removable incorporation blocking factor at a position other than 3'-OH, and an orthogonally removable homopolymer synthesis rate modulating modification that also does not reside at the 3'-OH is as follows: [Chem.]

[0122] Class IV nucleotide analogs are designed to have one or more incorporation-blocking modifications that cage the nucleotide from enzymatic incorporation until removed by light, heat, or electrochemical means. Large, bulky modifications that render the nucleotide analog inert to polymerase incorporation are preferred. In some embodiments, one or more of the modifications are derivatives of 2-nitrobenzyl that can be removed by photolysis. Removal of the caging modification in all cycles leaves an incorporation rate-modulating modification that is removed by a chemical method orthogonal to the method used to de-cage the nucleotide analog at the completion of informational polymer synthesis. In preferred embodiments, the incorporation-blocking modification is covalently linked to the incorporation rate-modulating modification. In some embodiments, caging modifications to the 2-nitrobenzyl functional group, which are bulky and act as a steric blocking factor, are useful. Examples of modifications that can act as a steric blocking factor include, but are not limited to, peptides, proteins, peptoids, PEG, branched PEG, star PEG, dendrimers, or nanoparticles.

[0123] Examples of class V pyrimidine nucleotides having 3'-O-caging modifications and non-removable base modifications are as follows: [Chem.]

[0124] When nanopore detection by direct translocation is desired as a readout method, the nucleotides are doubly modified with an appropriate removable 3'-O-caging group and a covalently attached non-removable modification that acts as a blocking current modulator during direct translocation through an unindexed nanopore. The appropriate modification can be a peptide or non-peptide. The ideal embodiment of the blocking current modulating group is the smallest possible modification that produces the most unique signal with the shortest possible homopolymer stretch. An important innovation of this class of dNTP analogs is that only one type of modified nucleotide is required, because the "sequence" of the homopolymer is encoded not by the nucleotide itself, but by the sequence of two or more current modulating modifications.

[0125] The following are examples of Class VI nucleotide analogs that contain an unmodified 3'-OH, a removable incorporation blocking (caging) modification, and two or more different types of non-removable current modulating modifications useful for detection by direct nanopore sequencing. An important innovation of Class VI dNTP analogs is that only one type of modified nucleotide is required, because the "sequence" of the homopolymer is encoded not by the nucleotide itself, but by the sequence of two or more current modulating modifications. The encoding is not limited to base 2 or base 4 and is only limited by the number of current modulating modifications that can be made to a single purine or pyrimidine nucleotide.

Chemical Structure

[0126] In some embodiments, the caging modification is removable by any of light, heat, or electrochemical means, while two or more nanopore detection elements are non-removable and withstand repeated exposure to the decaging conditions used in all cycles during homopolymer chain synthesis. The caging modification to 2-nitrobenzyl, which acts as a steric blocking factor, is a bulky modification such as, but not limited to, peptides, proteins, peptoids, PEG, branched PEG, star PEG, dendrimers, or nanoparticles.

Example

[0127] (Example 1) N 6 -benzoyl-deoxyadenosine triphosphate was placed in a vial under dry N2 with N 6-Benzoyl-2'-deoxyadenosine (0.055 g, 0.16 mmol) was added to prepare, and trimethyl phosphate (0.435 mL) was added. To the resulting solution, tributylamine (0.077 mL, 0.32 mmol) was added, and dry N2 was passed through the reaction mixture for 30 minutes while maintaining at -5 °C. To this vial, phosphorus oxychloride anhydride (0.018 mL, 0.19 mmol) was added via syringe, and the reaction mixture was stirred at -5 °C for 3 minutes. A second aliquot of phosphorus oxychloride anhydride (0.009 mL, 0.10 mmol) was added via syringe, and the reaction mixture was stirred at -5 °C for 8 minutes. To a second vial, tributylammonium pyrophosphate (0.075 g, 0.14 mmol) was added, dry N2 was passed through, acetonitrile anhydride (0.609 mL) was added, and then tributylamine (0.231 mL, 0.97 mmol) was added. The prepared tributylammonium pyrophosphate mixture was cooled to -20 °C, added to the reaction mixture, and reacted for 10 minutes. The reaction was quenched by the dropwise addition of H2O (4.35 mL). The contents of the flask were combined with 0.87 mL of H2O and extracted with dichloromethane (3 × 150 mL). The aqueous phase was adjusted to pH 6.5 using concentrated NH4OH and stirred at 4 °C for 12 hours. The mixture was transferred to a 250 mL round-bottom flask with 50 mL of water and concentrated under reduced pressure. The residue was dissolved in 40 mL of water and purified via ion-exchange chromatography (AKTA FPLC, Fractogel DEAE 48 mL column volume, step gradient 0->70% TEAB in water, pH 7.5). The fractions containing the desired product were pooled, concentrated by A260 concentration, and residual triethylammonium bicarbonate was removed via repeated concentration from water (5 × 50 mL) until dry to obtain N 6 -benzoyl-2'-deoxyadenosine triphosphate was obtained.

[0128] Controlled synthesis of homopolymer tracts composed of modified nucleotides by the nucleotide transferase TdT was carried out in the following manner. Deoxyadenosine triphosphate (TriLink Biosciences) and N at 1 mM each 6A stock solution of -benzoyl-deoxyadenosine triphosphate was prepared in H2O.

[0129] 0.5 μL (500 pmol) of each of the different triphosphates was separately combined with 1.5 μL (30 U) of commercially available TdT (Thermo Scientific), 0.5 μL (50 pmol) of 5’-TAATAATAATAATTTTT-3’ (IDT), 2 μL of commercially available TdT Rxn buffer (Thermo Scientific - 1M potassium cacodylate, 0.125 M Tris, 0.05% (v / v) Triton X100, 5 mM C o Cl2 (pH 7.2 at 25 °C)) and 4.5 μL of H2O. The reaction was incubated at 37 °C. 30 μL aliquots were removed at 1 minute, 5 minutes and 15 minutes and quenched with 20 μL 5 mM EDTA each. Each sample was dried under vacuum and reconstituted in 100 μL H2O. 10 μL of each time point was mixed with 10 μL of denaturing loading buffer (100% formamide and 0.1% Orange G) and applied to the wells of a 20% polyacrylamide gel 1 mm × 20 cm × 14 cm. After electrophoresis at 400 V for 3.5 hours, the bands were visualized with Sybr Gold (Thermo Scientific) and photographed under UV illumination (UV-blocking Wratten 2A filter 405 nm cut-off, UVP, LLC).

[0130] For the synthesis of multiple homopolymer tracts, an initiator oligonucleotide can be conjugated to beads to enable multiple rounds of enzymatic synthesis incorporating removal of previous reactants and wash solutions. 5'-Biotin-TAATAATAATAATTTTT-3' (IDT) can be incubated with streptavidin-coated magnetic sepharose microbeads (GE Healthcare Life Sciences). Beads with added oligonucleotide can be prepared by taking an aliquot of the bead slurry and transferring it to a filter cup. The beads can then be washed 5 times with 1×PBS (using 2× volume of bead slurry) by vortexing, and each rinse solution can be spun down. Then, 1 / 2 bead slurry volume of 1×PBS can be added, and the biotinylated oligonucleotide can be spiked at 1 / 10 of the published bead binding capacity. The mixture can be incubated at 37°C for 2 hours while vortexing every 30 minutes. After 2 hours, a small amount of supernatant can be removed, and the A260 can be measured for any unbound oligonucleotide. Where the remaining oligonucleotide shows an A260 of less than 10%, the beads can be washed 5 times with MQ water. The washed beads can be brought to the desired concentration in MQ water.

[0131] Homopolymer synthesis can be carried out as described above using 2 to 10× molar equivalents of the desired dNTP with respect to the bead-bound oligonucleotide. A reaction mixture containing beads, dNTP, TdT and buffer can be incubated at 37 °C for 15 minutes. The reaction can be stopped with 10 μl EDTA and rinsed three times with water. A new cycle of homopolymer synthesis can be initiated by adding fresh TdT enzyme, dNTP, buffer and incubated at 37 °C for 15 minutes. After stopping the reaction with EDTA and rinsing three times with water, the cycle of homopolymer synthesis can be repeated as many times as desired. After the final EDTA quench and three rinses with water, the support-bound alternating homopolymers can be cleaved from the solid support using 100 μl concentrated ammonium, the supernatant can be dried by Gen-vac and then stored at -20 °C until ready to be analyzed using polyacrylamide gel as described above.

[0132] In another experiment, as shown in Figure 12, the two-letter message "MA" was converted to binary and then to a base-2 nucleotide code. Each letter symbol was converted to the corresponding number from 1 to 26 (A→1, … Z→26). Each number was then converted to a 6-bit binary number by adding two zeros to the normal binary representation from 1 to 26 (e.g., "M" = 001101; "A" = 000001). Each bit of the binary representation was converted to a base-2 nucleotide representation according to the following table:

Table 2

[0133] Thus, "MA" was converted to 001101 000001 and then to the single nucleotide string ACTGAGACACAG, which was n C n T n G n A n G n A n C n A n Cn A n G n were synthesized as, where each nucleotide was synthesized as a homopolymer of variable length.

[0134] Synthesis of 12-bit homopolymer-encoded nucleic acids was performed with beads at approximately 20 pmol / ul of 5'-biotinylated 39 nt-long oligonucleotide initiator: 5'-biotin-CAGGTCCTAUC GATATC TGTGAGCTTAATGTCCTTATGT-3' bound to 34 um streptavidin-sepharose beads (GE Healthcare).

[0135] This oligonucleotide contains two features for releasing the final product from the solid support used during synthesis: 1) a single deoxyuridine residue that allows cleavage by the USER enzyme system (New England Biolabs), and 2) an EcoR V endonuclease restriction site.

[0136] Starting with approximately 2 nmol of bead-bound initiator, each homopolymer of variable length was enzymatically synthesized using TdT and one of four modified nucleotide triphosphates. Each reaction was carried out in a total volume of 750 ul containing 40 - 100 uM of modified dNTP (40 uM - A; 100 uM - C; 50 uM - T; 100 uM - G), 20 U TdT (Thermo-Fisher Scientific), 1×TdT buffer (Thermo-Fisher Scientific) with an incubation at 37 °C for 2.5 - 20 minutes. After each enzymatic extension step, the reaction was quenched by adding 500 ul of 250 mM EDTA in 10 mM Tris buffer (pH 6.8). The beads were recovered by centrifugation at 10000×g and removal of the supernatant. Figure 13 shows a PAGE analysis of each cycle of the enzymatic synthesis of the 12-bit homopolymer data strand. Each lane starts with "N" indicating the unreacted 39 nt initiator and is marked with the cycle number. The black arrow indicates the size marker for the 60 nt oligonucleotide.

[0137] After removal of the full-length data strands from the solid support, NGS library preparation was performed using the ACCEL-NGS® 1S PLUS DNA LIBRARY KIT (Swift Bioscience) according to the manufacturer's instructions, followed by sequencing using the Illumina MiSeq System. Figures 14A - L show histograms of the observed base composition for each of the 12 nucleotide additions (as labeled) generated during enzymatic synthesis.

[0138] (Example 2) Procedure for synthesizing class I - purine and pyrimidine dNTP analogs Scheme for the synthesis of class I dNTP analogs Nitrobenzyl - adenosine [Chemical formula] Nitrobenzyl - cytosine [Chemical formula] Nitrobenzyl - guanosine [Chemical formula] Nitrobenzyl - thymidine [Chemical formula]

[0139] (Example 3) Detailed procedure for 3’ - O - nitrobenzyldeoxyadenosine: [Chemical formula] 9-[β-D-5’-Hydroxy-2’-deoxyribofuranosyl]-6-chloropurine (1.00 g, 3.69 mmol) and imidazole (554 mg, 8.12 mmol) were dissolved in dimethyl formate anhydride (18 mL), and then tert-butyldimethylsilyl chloride (611 mg, 3.93 mmol) was added. The reaction mixture was stirred at room temperature for 20 hours under argon. The dried residue was impregnated on silica and purified by flash column chromatography (hexane / ethyl acetate, 2:1) to obtain 9-[β-D-5’-O-(tert-butyldimethylsilyl)-2’-deoxyribofuranosyl]-6-chloropurine.

Chemical formula

[0140] 9-[β-D-5’-O-(tert-butyldimethylsilyl)-2’-deoxyribofuranosyl]-6-chloropurine (1.73 g, 4.48 mmol) was dissolved in dichloromethane anhydride (135 mL). Tetrabutylammonium bromide (722 mg, 2.24 mmol), 2-nitrobenzyl bromide (2.41 g, 11.2 mmol) and 40% aqueous sodium hydroxide (65 mL) were added to the previously prepared solution. The reaction mixture was stirred at room temperature for 1 hour and diluted with ethyl acetate (300 mL). The layers were separated. The aqueous layer was extracted with ethyl acetate (125 mL × 2). The combined organic layers were dried over anhydrous sodium sulfate. The organic layer was impregnated on silica gel and then purified by flash column chromatography (hexane / ethyl acetate, 2:1) to obtain 9-[β-D-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyribofuranosyl]-6-chloropurine.

Chemical formula

[0141] 9-[β-D-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyribofuranosyl]-6-chloropurine (2.22 g, 3.83 mmol) was dissolved in anhydrous tetrahydrofuran (40 mL) and cooled to 0 °C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (4.20 mL, 4.20 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 1 hour. After drying the reaction mixture, the residue was dissolved in dioxane and 7N ammonia in ethanol (40 mL). The reaction mixture was stirred in a sealed round bottom at 90 °C for 18 hours. The reaction mixture was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 20:1) to give 3’-O-(2-nitrobenzyl)-2’-deoxyadenosine.

Chemical formula

[0142] 3'-O-(2-Nitrobenzyl)-2'-deoxyadenosine (15 mg, 1 equiv, 38 μmol) was co-evaporated with pyridine (1 mL×3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equiv, 0.11 mmol) was added. After 5 minutes, a 5 μL aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equiv, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0143] (Example 4) Detailed procedure for 3'-O-nitrobenzyldeoxycytidine: [Chemical formula] 3’,5’-Di-O-(tert-butyldimethylsilyl)-2’-deoxyuridine (1.00 g, 2.12 mmol) was dissolved in anhydrous acetonitrile (90 mL) and cooled to 0 °C under argon. Phosphoryl trichloride (1.49 mL, 2.12 mmol) was added dropwise over 2 minutes. After 10 minutes, triethylamine (11.1 mL, 79.7 mmol) was added dropwise over 3 minutes. After 15 minutes, the reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was cooled to 0 °C and triazole (4.40 g, 63.7 mmol) was added as a solid all at once. A precipitate was observed and the suspension was stirred for 30 minutes. After stirring at room temperature for 2 hours, the reaction mixture was concentrated to dryness. The residue was dissolved in dichloromethane (30 mL) and washed with saturated sodium bicarbonate solution (25 mL × 2) and brine (25 mL). The organic layer was dried over anhydrous sodium sulfate and concentrated to dryness. The residue was dissolved in dichloromethane, and then allyl alcohol (2.00 mL, 29.4 mmol) and triethylamine (2.67 mL, 18.9 mmol) were added. The reaction mixture was stirred at 0 °C for 15 minutes. DBU (0.33 mL, 2.17 mmol) was added and the mixture was stirred at room temperature for 6 hours. The reaction mixture was diluted with dichloromethane (17 mL) and washed with brine (15 mL). The organic layer was dried over anhydrous sodium sulfate. The organic layer was impregnated on silica and purified by flash column chromatography (hexane / ethyl acetate, 4:1) to give 4-O-allyl-3’,5’-di-O-(tert-butyldimethylsilyl)-2’-deoxyuridine. [Chemical formula]

[0144] 4-O-Allyl-3’,5’-di-O-(tert-butyldimethylsilyl)-2’-deoxyuridine (3.37 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 h and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL×1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain 4-O-allyl-2’-deoxyuridine.

Chem.

[0145] 4-O-Allyl-2’-deoxyuridine (1.18 g, 4.40 mmol) was dissolved in anhydrous pyridine (37 mL), and then tert-butyldimethylsilyl chloride (815 mg, 5.41 mmol) was added under argon. The reaction mixture was stirred at room temperature for 20 h. After concentration, the residue was dissolved in dichloromethane and impregnated on silica. The crude product was purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain 4-O-allyl-5’-O-(tert-butyldimethylsilyl)-2’-deoxyuridine.

Chem.

[0146] A mixture of 4-O-allyl-5’-O-(tert-butyldimethylsilyl)-2’-deoxyuridine (1.28 g, 3.35 mmol), tetrabutylammonium hydroxide (1.5 mL, 55 - 60% in water) and sodium iodide (50.0 mg, 0.335 mmol) in dichloromethane / water (20 mL, 1:1) was added with 1.0 M sodium hydroxide solution (10 mL) under argon. The reaction mixture was stirred at room temperature for 10 minutes, and then 2-nitrobenzyl bromide (1.45 g, 6.70 mmol) in 10 mL of dichloromethane was added over 5 minutes. After stirring at room temperature for 7 hours, the reaction mixture was diluted with dichloromethane (150 mL). The organic layer was washed with brine (20 mL) and dried over anhydrous sodium sulfate. The organic layer was impregnated on silica and purified by flash column chromatography (hexane / ethyl acetate, 1:1) to obtain 4-O-allyl-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyuridine.

Chem.

[0147] 4-O-allyl-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyuridine (1.55, 3.00 mmol) was dissolved in 7N ammonia in ethanol (55 mL) and stirred at 55 °C for 20 hours in a sealed round bottom. The reaction mixture was impregnated on silica and purified by flash column chromatography (dichloromethane / methanol, 20:1) to obtain 5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxycytidine.

Chem.

[0148] 5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxycytidine (2.22 g, 3.83 mmol) was dissolved in anhydrous tetrahydrofuran (40 mL) and cooled to 0 °C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (4.20 mL, 4.20 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 1 hour. After the reaction, the mixture was impregnated on silica and purified by flash column chromatography (dichloromethane / methanol, 8:2) to obtain 3’-O-(2-nitrobenzyl)-2’-deoxycytidine. [Chemical formula]

[0149] 3’-O-(2-nitrobenzyl)-2’-deoxycytidine (14 mg, 1 equivalent, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl chloride (18 mg, 11 μL, 3 equivalents, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equivalents, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added all at once as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0150] (Example 5) Detailed procedure for 3’-O-nitrobenzyl deoxyguanosine:

Chem.

Chem.

[0151] 2-Amino-6-chloro-9-[β-D-5’-O-(tert-butyldimethylsilyl)-2’-deoxyribofuranosyl]purine (1.19 g, 3.00 mmol) was dissolved in anhydrous tetrahydrofuran (8.0 mL), and then N,N-dimethylformamide dimethyl acetal (3.10 mL, 18.0 mmol) was added at room temperature. The reaction mixture was stirred at 40 °C for 3 h. The reaction mixture was impregnated on silica and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give 6-chloro-N 2 -[(dimethylaminomethylene)amino]-9-[β-D-5’-O-(tert-butyldimethylsilyl)-2’-deoxyribofuranosyl]purine.

Chem.

[0152] 6-Chloro-N 2-[(Dimethylaminomethylene)amino]-9-[β-D-5’-O-(tert-butyldimethylsilyl)-2’-deoxyribofuranosyl]purine (1.09 g, 2.40 mmol) was dissolved in anhydrous acetonitrile (3.5 mL), and then sodium hydride powder (60%) (122 mg, 4.80 mmol) in mineral oil was added at 0 °C. After stirring at room temperature for 1 hour, a solution of 2-nitrobenzyl bromide (1.04 g, 4.80 mmol) in anhydrous acetonitrile (1.5 mL) was added. After stirring at room temperature for 2 hours, the reaction mixture was filtered. The filtrate was dried, and the resulting residue was dissolved in ethyl acetate (100 mL). The organic layer was washed with saturated sodium bicarbonate solution (50 mL) and brine (50 mL), and dried over anhydrous sodium sulfate. The organic layer was impregnated on silica and purified by flash column chromatography (hexane / ethyl acetate 4:6) to obtain 6-chloro-N 2 -[(Dimethylaminomethylene)amino]-9-[β-D-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyribofuranosyl]purine. [Chemical formula]

[0153] 6-chloro-N 2-[(Dimethylaminomethylene)amino]-9-[β-D-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyribofuranosyl]purine (1.02 g, 1.73 mmol) was dissolved in dimethyl formate anhydride (15 mL), and then, under argon, cesium acetate (996 mg, 5.19 mmol), 1,4-diazabicyclo[2.2.2]octane (194 mg, 1.73 mmol) and triethylamine (0.72 mL, 5.19 mmol) were added. The reaction mixture was stirred at room temperature for 18 h. Acetic anhydride (5 mL) was added and stirred for 0.5 h. The reaction mixture was quenched with water (100 mL) and extracted with ethyl acetate (100 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 20:1) to give 5’-O-(tert-butyldimethylsilyl)-N 2 -[(dimethylamino)methylene]-3’-O-(2-nitrobenzyl)-2’-deoxyguanosine. [Chemical formula]

[0154] 5’-O-(tert-butyldimethylsilyl)-N 2 -[(dimethylamino)methylene]-3’-O-(2-nitrobenzyl)-2’-deoxyguanosine (732 mg, 1.28 mmol) was dissolved in anhydrous tetrahydrofuran (8 mL) under argon and cooled to 0 °C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (2.56 mL, 2.56 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 2 h. The reaction mixture was poured into cold water (50 mL) and extracted with ethyl acetate (50 mL × 2). The combined organic layers were dried over anhydrous sodium sulfate. The organic layer was impregnated on silica and purified by flash column chromatography (dichloromethane / methanol, 10:1) to give N 2 -[(dimethylamino)methylene]-3’-O-(2-nitrobenzyl)-2’-deoxyguanosine. [Chem.]

[0155] N 2 -[(Dimethylamino)methylene]-3'-O-(2-nitrobenzyl)-2'-deoxyguanosine (17 mg, 1 equiv, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equiv, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equiv, 0.15 mmol) was added all at once as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was carried out immediately after the EtOAc extraction. Final purification was by reverse phase HPLC.

[0156] (Example 6) Detailed procedure for 3'-O-nitrobenzyldeoxythymidine: [Chem.] Thymidine (2.50 g, 10.3 mmol) was suspended in dimethylformamide (60 mL) at room temperature under argon. To this suspension were added imidazole (4.22 g, 61.9 mmol) and tert-butyldimethylsilyl chloride (4.66 g, 30.1 mmol). After stirring for 2 hours, the reaction mixture was quenched with methanol (8 mL) and diluted with ethyl acetate (200 mL). The organic layer was washed with water (100 mL × 2), saturated sodium bicarbonate solution (100 mL), and brine (100 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (hexane / ethyl acetate, 8:2) to obtain 3′,5′-di-O-(tert-butyldimethylsilyl)thymidine. [Chemical formula]

[0157] 3′,5′-di-O-(tert-butyldimethylsilyl)thymidine (4.34 g, 9.22 mmol) and dimethyl-4-aminopyridine (1.12 g, 9.22 mmol) were dissolved in anhydrous dichloromethane (140 mL). Triethylamine (5.14 mL, 36.9 mmol) was added and the reaction mixture was cooled to 0 °C. Benzyl chloride (3.21 mL, 27.7 mmol) was added dropwise and the mixture was warmed up to room temperature. After stirring for 14 hours, saturated sodium bicarbonate solution (80 mL) was added and the layers were separated. The aqueous layer was extracted with dichloromethane (200 mL × 2). The combined organic layers were washed with water (300 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (hexane / ethyl acetate, 8:2) to obtain 3-N-benzoyl-3′,5′-di-O-(tert-butyldimethylsilyl)thymidine. [Chemical formula]

[0158] 3-N-benzoyl-3',5'-di-O-(tert-butyldimethylsilyl)thymidine (3.03 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain 3-N-benzoylthymidine.

Chemical formula

[0159] 3-N-benzoylthymidine (XX g, 4.40 mmol) was dissolved in anhydrous pyridine (37 mL), and then tert-butyldimethylsilyl chloride (815 mg, 5.41 mmol) was added under argon. The reaction mixture was stirred at room temperature for 20 hours. After concentration, the residue was dissolved in dichloromethane and impregnated on silica. The crude product was purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain 3-N-benzoyl-5'-O-(tert-butyldimethylsilyl)thymidine.

Chemical formula

[0160] To 3-N-benzoyl-5’-O-(tert-butyldimethylsilyl)thymidine (1.18 g, 2.56 mmol) was added an aqueous solution of tetrabutylammonium hydroxide (10 mL, 60%), followed by the addition of sodium iodide (76.7 mg, 0.51 mmol), dichloromethane (10 mL), water (10 mL), and an aqueous solution of 1 M sodium hydroxide (10 mL). This mixture was added dropwise to a solution of 2-nitrobenzyl bromide (718 mg, 3.32 mmol) in dichloromethane (10 mL). The reaction mixture was stirred at room temperature for 6 hours, after which water (10 mL) was added. The aqueous layer was extracted with dichloromethane (50 mL × 3). The combined organic layers were washed with brine (100 mL) and dried over anhydrous sodium sulfate. After filtration and concentration, the residue was dissolved in ethyl acetate, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 6:4) to obtain 3-N-benzoyl-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)thymidine.

Chem.

[0161] 3-N-benzoyl-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)thymidine (1.24 g, 2.08 mmol) was dissolved in ethanol (15 mL), followed by the addition of 30% ammonium hydroxide solution (1.5 mL, 12.3 mmol). The reaction mixture was stirred at room temperature for 1 hour, impregnated onto silica, and purified by flash column chromatography (hexane / ethyl acetate, 6:4) to obtain 5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)thymidine.

Chem.

[0162] 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)thymidine (681 mg, 1.39 mmol) was dissolved in anhydrous tetrahydrofuran (12 mL) under argon and cooled to 0 °C. Then, a 1.0 M solution of tetrabutylammonium fluoride in tetrahydrofuran (2.78 mL, 2.78 mmol) was added dropwise. The reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was poured into cold water (50 mL) and extracted with ethyl acetate (50 mL × 3). The combined organic layers were dried over anhydrous sodium sulfate. The organic layer was impregnated onto silica and purified by flash column chromatography (dichloromethane / methanol, 10:1) to obtain 3'-O-(2-nitrobenzyl)thymidine. [Chemical formula]

[0163] 3'-O-(2-nitrobenzyl)thymidine (15 mg, 1 equivalent, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl chloride (18 mg, 11 μL, 3 equivalents, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equivalents, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid all at once. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0164] (Example 6) Procedure for synthesizing class II-purine and pyrimidine dNTP analogs. Scheme for the synthesis of class II non-peptide dNTP analogs Non-peptide precursor - adenosine

Chem.

Chem.

Chem.

Chem.

[0165] (Example 7) Detailed procedure for the synthesis of class II non-peptide dATP constructs:

Chem.

[0166] N-substituted nucleoside (5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL × 1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine. [Chemistry]

[0167] N-Replaced 3'-O-(2-nitrobenzyl)-2'-deoxyadenosine (1 equivalent, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equivalents, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equivalents, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added as a solid all at once. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse phase HPLC.

[0168] (Example 8) Detailed procedure for the synthesis of class II non-peptidic dCTP constructs:

Chemical formula

Chemical formula

[0169] N-substituted 5'-O-(tert-butyldimethylsilyl)-3'-O-(2-nitrobenzyl)-2'-deoxycytidine (3.37 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 h and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL×1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain N-substituted 3'-O-(2-nitrobenzyl)-2'-deoxycytidine.

Chemical formula

[0170] N-Replaced 3'-O-(2-nitrobenzyl)-2'-deoxycytidine (1 equivalent, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl chloride (18 mg, 11 μL, 3 equivalents, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equivalents, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equivalents, 0.15 mmol) was added all at once as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0171] (Example 9) Detailed procedure for the synthesis of the peptide-dATP conjugate: [Chemical formula] 9-[β-D-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyadenosine (524.3 mg, 1.10 mmol) and 4-(pyridin-2-yl disulfanyl)butanoic acid (302.7 mg, 1.32 mmol) were dissolved in dimethyl formate anhydride under argon. N,N-Diisopropylethylamine (0.48 mL, 2.74 mmol) and 2-(3H-[1,2,3]triazolo[4,5-b]pyridin-3-yl)-1,1,3,3-tetramethyluronium hexafluorophosphate(V) (417 mg, 1.10 mmol) were added. The reaction mixture was stirred at room temperature for 18 h and diluted with ethyl acetate (30 mL). The organic layer was washed with saturated sodium bicarbonate solution (30 mL), dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (hexane / ethyl acetate, 4:6) to obtain N-(4-(pyridin-2-yl disulfanyl)butanoyl)-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyadenosine.

Chemical Structure

[0172] N-(4-(Pyridin-2-yldisulfanyl)butanoyl)-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxyadenosine (3.75 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 h and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL×1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to give N-(4-(pyridin-2-yldisulfanyl)butanoyl)-3’-O-(2-nitrobenzyl)-2’-deoxyadenosine.

Chemical formula

[0173] N-(4-(Pyridin-2-yl disulfanyl) butanoyl)-3’-O-(2-nitrobenzyl)-2’-deoxyadenosine (23 mg, 1 equiv, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equiv, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equiv, 0.15 mmol) was added as a solid in one portion. The mixture was stirred for 10 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse phase HPLC.

[0174] (Example 10) Detailed procedure for the synthesis of the peptide-dCTP conjugate:

Chemical Structure

[0175] N-(4-(Pyridin-2-yldisulfanyl)butanoyl)-5’-O-(tert-butyldimethylsilyl)-3’-O-(2-nitrobenzyl)-2’-deoxycytidine (3.62 g, 5.27 mmol) was dissolved in dry tetrahydrofuran (50 mL). Triethylamine (1.98 mL, 14.2 mmol) was added, and then triethylammonium fluoride dihydrofluoride (2.32 mL, 14.2 mmol) was added under argon. The reaction mixture was stirred at room temperature for 29 hours and then concentrated. The residue was dissolved in dichloromethane (100 mL) and washed with 1.5 M ammonium carbonate (75 mL×1) and brine (75 mL). The organic layer was dried over anhydrous sodium sulfate, impregnated on silica, and purified by flash column chromatography (dichloromethane / methanol, 9:1) to obtain N-(4-(pyridin-2-yldisulfanyl)butanoyl)-3’-O-(2-nitrobenzyl)-2’-deoxycytidine. [Chemical formula]

[0176] N-(4-(Pyridin-2-yl disulfanyl) butanoyl)-3’-O-(2-nitrobenzyl)-2’-deoxycytidine (22 mg, 1 equiv, 38 μmol) was co-evaporated with pyridine (1 mL×3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl trichloride (18 mg, 11 μL, 3 equiv, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 equiv, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 equiv, 0.15 mmol) was added all at once as a solid. The mixture was stirred for 10 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, and this separation was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0177] (Example 11) Procedure for synthesizing class III-purine and pyrimidine dNTP analogs. Scheme for the synthesis of class III dNTP analogs

Chem.

[0178] (Example 12) Detailed procedure for class III-purine dNTP analogs:

Chem.

Chem.

[0179] To a solution of 9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-9H-purin-6-amine (12.2 g, 1.1 equiv, 25.5 mmol) in dry DMF (100 mL) was added N,N-diisopropylethylamine (3.60 g, 4.9 mL, 1.2 equiv, 27.8 mmol) at 0 °C. The reaction mixture was stirred for 30 min, then 2-nitrobenzyl carbonochloridate (5.00 g, 1 equiv, 23.2 mmol) was slowly added dropwise over 30 min while maintaining the temperature below 5 °C. The reaction mixture was then warmed to room temperature and stirred overnight. The reaction mixture was poured into a cooled solution of 5% Na2CO3 and EtOAc. The EtOAc layer was dried over sodium sulfate and then concentrated to dryness. The crude product was purified by chromatography on silica gel using a hexane / EtOAc mixture to give the purified product, which was used in the next reaction.

Chem.

[0180] (9 - ((2R,4S,5R)-4 - ((tert-butyldimethylsilyl)oxy)-5 - (((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-9H-purin-6-yl)carbamic acid 2-nitrobenzyl (5.00 g, 1 equivalent, 7.59 mmol) was dissolved in THF at room temperature and then cooled to OC under a blanket of dry argon. Then, tetrabutylammonium fluoride (4.96 g, 2.5 equivalents, 19.0 mmol) was added to the mixture. The mixture was stirred at 0°C for 2 hours and then warmed to 23°C for 1 hour. The solution was poured into a cold solution of 10% NaHCO3 and extracted with DCM. The DCM layer was concentrated and the crude product was purified on silica gel eluting with 5 - 50% DCM / MeOH to afford the product suitable for triphosphorylation. [Chemical formula]

[0181] (9 - ((2R,4S,5R)-4 - hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-9H-purin-6-yl)carbamic acid 2-nitrobenzyl (28 mg, 1 equivalent, 65 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled under argon in an ice bath. The first aliquot of phosphoryl trichloride (30 mg, 18 μL, 3 equivalents, 0.20 mmol) was added. After 5 minutes, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.23 g, 4 equivalents, 0.26 mmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 seconds at rxn t = 35 minutes. Immediately, the pre-weighed N1,N1,N8,N8-tetramethyl-naphthalene-1,8-diamine (56 mg, 4 equivalents, 0.26 mmol) was added as a solid in one portion. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0182] (Example 13) Detailed procedure for class III - pyrimidine dNTP analogs:

Chem.

Chem.

[0183] (1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)-tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamic acid 2-nitrobenzyl (5.00 g, 1 equivalent, 7.88 mmol) was dissolved in 25 mL of THF at room temperature and then cooled to OC under a blanket of dry argon. Then, tetrabutylammonium fluoride (5.15 g, 2.5 equivalents, 19.7 mmol) was added to the mixture. The mixture was stirred at 0C for 2 hours and then warmed to 23C for 1 hour. The solution was poured into a cold solution of 10% NaHCO3 and extracted with DCM. The DCM layer was concentrated and the crude product was purified on silica gel eluting with 5 - 50% DCM / MeOH to give the product suitable for triphosphorylation.

Chemical formula

[0184] ((1 - ((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamic acid 2-nitrobenzyl (35.0 mg, 1 equivalent, 86.1 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled under argon in an ice bath. A first aliquot of phosphoryl trichloride (39.6 mg, 3 equivalents, 258 μmol) was added. After 5 minutes, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 equivalents, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 seconds at rxn t = 35 minutes. Immediately, a preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 equivalents, 345 μmol) was added all at once as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.)

[0185] (Example 14) Procedure for synthesizing class IV-dCTP analogs. Scheme for the synthesis of class IV dCTP analogs [Chemical Structure] Detailed procedure for class IV dCTP analogs: [Chemical Structure]

[0186] Methyl 3,4,5-trihydroxybenzoate (10 g, 1 equivalent, 54 mmol) was dissolved in 50 mL of acetone. Sodium iodide (0.81 g, 0.1 equivalent, 5.4 mmol) and potassium carbonate (38 g, 5 equivalents, 270 mmol) were added as solids at ambient temperature. 1-(Chloromethyl)-2-nitrobenzene (34 g, 3.6 equivalents, 0.20 mol) was added dropwise over 10 minutes as a solution in 40 mL of acetone. The mixture was stirred for 1 hour and then heated to 50 °C for 6 hours. The mixture was cooled to ambient temperature and the bulk solvent was removed on a rotary evaporator. The residue was suspended in 200 mL of EtOAc and this was washed sequentially with 200 mL portions of water and saturated aqueous NaCl. The EtOAc layer was dried over sodium sulfate and evaporated. The crude product was subjected to chromatography on silica using a hexane / EtOAc mixture to give a purified product that could be used in the next reaction.

Chemical formula

[0187] Methyl 3,4,5-tris((2-nitrobenzyl)oxy)benzoate (25 g, 1 equivalent, 42 mmol) was dissolved in 300 mL of THF. Aqueous 2M NaOH (105 mL, 210 mmol) was added and the mixture was stirred at ambient temperature for 18 hours. The bulk THF was removed on a rotary evaporator and the residue was slowly acidified with 6M HCl until a pH of 1 or less was reached. The resulting solid was filtered, washed thoroughly with water, dried on a filter funnel for 5 hours and then dried under high vacuum for 18 hours. The product was carried forward to the next reaction without further purification.

Chemical formula

[0188] 4-Amino-1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)pyrimidin-2(1H)-one (12 g, 1 equiv, 26 mmol) and 3,4,5-tris((2-nitrobenzyl)oxy)benzoic acid (15 g, 1 equiv, 26 mmol) were dissolved in 50 mL of dry DMF at ambient temperature under an argon atmosphere. N-Ethyl-N-isopropylpropan-2-amine (5.1 g, 6.8 mL, 1.5 equiv, 39 mmol) was added, and then a solution of 1-((dimethylamino)(dimethyliminio)methyl)-1H-[1,2,3]triazolo[4,5-b]pyridine 3-oxide hexafluorophosphate (V) (12 g, 1.2 equiv, 31 mmol) in 10 mL of dry DMF was added dropwise over 5 minutes at ambient temperature. The mixture was stirred at ambient temperature for 18 hours. The mixture was dissolved in 300 mL of EtOAc, and this was washed sequentially with 200 mL portions of water (2×) and saturated aqueous NaCl. The EtOAc was dried over sodium sulfate, evaporated first on a rotary evaporator and then under high vacuum for 18 hours. The residue was subjected to chromatography on silica using a mixture of dichloromethane and methanol to give the desired product as a colorless foam.

Chemical formula

[0189] N-(1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)-3,4,5-tris((2-nitrobenzyl)oxy)benzamide (5 g, 1 equivalent, 5 mmol) was dissolved in 25.0 mL of dry THF at ambient temperature under argon. Triethylamine (4 g, 6 mL, 8 equivalents, 4e+1 mmol) was added rapidly, followed by the rapid addition of triethylammonium fluoride dihydrofluoride (5 g, 5 mL, 6 equivalents, 3e+1 mmol) also at ambient temperature. The mixture was stirred at ambient temperature for 24 h. Silica gel (20 g) was added and the mixture was evaporated on a rotary evaporator to a fine powder and then loaded onto a 100 g silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam. [Chemical formula]

[0190] N-(1-((2R,4S,5R)-4-Hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)-3,4,5-tris((2-nitrobenzyl)oxy)benzamide (30 mg, 1 eq, 38 μmol) was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A 6 μL first aliquot of phosphoryl chloride (18 mg, 11 μL, 3 eq, 0.11 mmol) was added. After 5 minutes, a 5 μL second aliquot was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (0.14 g, 4 eq, 0.15 mmol) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 seconds. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (33 mg, 4 eq, 0.15 mmol) was added as a solid all at once. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was carried out immediately after the EtOAc extraction. Final purification was by reverse phase HPLC.

[0191] (Example 15) Procedure for synthesizing class V peptides and non-peptide analogs. Scheme for the synthesis of peptide-thymidine dNTP conjugates [Chemical formula] Scheme for the synthesis of non-peptide-thymidine dNTP conjugates [Chemical formula]

[0192] (Example 16) Detailed Procedure for Peptide-dTTP Analog:

Chem.

Chem.

[0193] 1-((2R,4S,5R)-4-((tert-Butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (5.00 g, 1 equiv, 10.6 mmol) was dissolved in 50 mL of DMF and then cooled to 0 °C. The reaction was stirred at 0 °C for 30 min, then sodium hydride (306 mg, 1.2 equiv, 12.7 mmol) was added. Then, the reaction was stirred at 0 °C for an additional 30 min and then warmed to 23 °C. Then, 2-((4-(Bromomethyl)phenyl)disulfanyl)-pyridine (3.32 g, 1 equiv, 10.6 mmol) was added to the mixture, and stirring was continued at 23 °C for an additional 2 h. Then, the reaction was poured into a cold solution of 10% NaHCO3 and DCM. The DCM layer was separated, dried over sodium sulfate, and concentrated until dry. Then, the mixture was purified on silica gel eluting with 0 - 20% DCM / methanol to give the desired product.

Chem.

[0194] 1-((2R,4S,5R)-4-((tert-Butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)-methyl)tetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-ylsulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (2.00 g, 1 equivalent, 2.85 mmol) was dissolved in THF and cooled to 0 °C. Then, tetrabutylammonium fluoride (2.23 mg, 3 equivalents, 8.55 mmol) was added to the mixture at 0 °C. The reaction was continued to stir at 0 °C for 2 hours and then warmed to 23 °C for an additional 1 hour. Then, the reaction was cooled to 0 °C again and added to a pre-cooled solution of 10% NaHCO3 and DCM at 0 °C. Then, the DCM layer was separated, dried over sodium sulfate, concentrated, and purified on silica gel eluting with 5 - 50% DCM / methanol to obtain the pure product.

Chemical formula

[0195] 1-((2R,4S,5R)-4-Hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-5-methyl-3-(4-(pyridin-2-ylsulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (5.00 g, 1 equivalent, 10.6 mmol) was dissolved in 20 mL of THF at 23 °C. Then, triethylamine (1.07 g, 1.5 mL, 1 equivalent, 10.6 mmol) was added to the mixture and cooled to 0 °C. Then, TBS-Cl (1.59 g, 1 equivalent, 10.6 mmol) was added to the mixture and stirring was continued at 0 °C for an additional 2 hours. Then, the mixture was added to a pre-cooled mixture of 10% aqueous NaCl and DCM. The DCM layer was dried over sodium sulfate, concentrated, and dried to obtain an amber oil. Then, the crude product was purified on silica gel eluting with 5 - 50% DCM / methanol to obtain the pure product.

Chemical formula

[0196] 1 - ((2R,4S,5R)-5 - (((tert-butyldimethylsilyl)oxy)methyl)-4 - hydroxyoxolane-2-yl)-5 - methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (2.00 g, 1 equivalent, 3.40 mmol) was dissolved in 20 mL of DMF and then cooled to 0 °C. Then, sodium hydride (98.0 mg, 1.2 equivalents, 4.08 mmol) was added to the mixture and stirring was continued at 0 °C for an additional 30 minutes. Then, 1-(bromomethyl)-2-nitrobenzene (735 mg, 1 equivalent, 3.40 mmol) was added to the reaction and stirring was continued at 0 °C for an additional 1 hour. Then, the reaction was added to a pre-cooled mixture of 10% NaCl and EtOAc. The EtOAc layer was separated, dried over sodium sulfate, and concentrated until dry. The crude product was purified on silica gel eluting with 0 - 50 hexane / EtOAc to afford the desired product.

Chem.

[0197] 1 - ((2R,4S,5R)-5 - (((tert-butyldimethylsilyl)oxy)methyl)-4 - ((2-nitrobenzyl)oxy)oxolane-2-yl)-5 - methyl-3-(4-(pyridin-2-yldisulfanyl)benzyl)pyrimidine-2,4(1H,3H)-dione (1.00 g, 1 equivalent, 1.38 mmol) was dissolved in THF and cooled to 0 °C. Then, TBAF (362 mg, 1 equivalent, 1.38 mmol) was added to the reaction at 0 °C and stirred at 0 °C for 1 hour and then warmed to room temperature over 2 hours. Then, the mixture was poured into a pre-cooled solution of 10% NaHCO3 and DCM. The DCM layer was separated and dried over sodium sulfate. The crude product was purified on silica gel eluting with 5 - 25% DCM / methanol to afford the desired product.

Chem.

[0198] (1-((2R,4S,5R)-5-(Hydroxymethyl)-4-((2-nitrobenzyl)oxy)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamic acid 4-(pyridin-2-yl)benzyl (35.0 mg, 1 eq, 61.0 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (39.6 mg, 3 eq, 258 μmol) was added. After 5 minutes, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 eq, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 seconds at rxn t = 35 minutes. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 eq, 345 μmol) was added all at once as a solid. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0199] (Example 17) Detailed procedure for non-peptide-dTTP analog: [Chemical formula] 1-((2R,4S,5R)-4-((tert-Butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (5.00 g, 1 equivalent, 10.6 mmol) was dissolved in 50 mL of DMF and then cooled to 0 °C. The reaction mixture was stirred at 0 °C for 30 minutes and then sodium hydride (306 mg, 1.2 equivalents, 12.7 mmol) was added. The reaction mixture was stirred at 0 °C for an additional 30 minutes and then warmed to 23 °C. Then, (bromomethyl)benzene (1.82 g, 1 equivalent, 10.6 mmol) was added to the mixture and stirring was continued at 23 °C for an additional 2 hours. The reaction mixture was then poured into a cold solution of 10% NaHCO3 and DCM. The DCM layer was separated, dried over sodium sulfate and concentrated until dry. The mixture was then purified on silica gel eluting with 0 - 20% DCM / methanol to afford the desired product.

Chemical Structure

[0200] 3-Benzyl-1-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (4.00 g, 1 equivalent, 7.13 mmol) was dissolved in THF and cooled to 0 °C. Then, tetrabutylammonium fluoride (3.72 g, 2 equivalents, 14.26 mmol) was added to the mixture at 0 °C. The reaction mixture was stirred at 0 °C for 2 hours and then warmed to 23 °C for an additional 1 hour. The reaction mixture was then cooled back to 0 °C and added at 0 °C to a pre-cooled solution of 10% NaHCO3 and DCM. The DCM layer was separated, dried over sodium sulfate, concentrated and purified on silica gel eluting with 5 - 50% DCM / methanol to afford the pure product.

Chemical Structure

[0201] 3-Benzyl-1-((2R,4S,5R)-4-hydroxy-5-(hydroxymethyl)tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (1.00 g, 1 equivalent, 3.01 mmol) was dissolved in 20 mL of THF at 23 °C. Then, triethylamine (304 mg, 0.42 mL, 1 equivalent, 3.01 mmol) was added to the mixture, and the mixture was cooled to 0 °C. Then, TBS-Cl (453 mg, 1 equivalent, 3.01 mmol) was added to the mixture, and stirring was continued at 0 °C for an additional 2 hours. Then, the mixture was added to a pre-cooled mixture of 10% aqueous NaCl and DCM. The DCM layer was dried over sodium sulfate, concentrated, and dried to obtain an amber oil. The crude product was then purified on silica gel eluting with 5 - 50% DCM / methanol to obtain the pure product.

Chemical formula

[0202] 3-Benzyl-1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-hydroxytetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (6.00 g, 1 equivalent, 13.4 mmol) was dissolved in 20 mL of DMF and then cooled to 0 °C. Then, sodium hydride (387 mg, 1.2 equivalents, 16.1 mmol) was added to the mixture, and stirring was continued at 0 °C for an additional 30 minutes. Then, 1-(bromomethyl)-2-nitrobenzene (2.90 g, 1 equivalent, 13.4 mmol) was added to the reaction, and stirring was continued at 0 °C for an additional 1 hour. Then, the reaction was added to a pre-cooled mixture of 10% NaCl and EtOAc. The EtOAc layer was separated, dried over sodium sulfate, and concentrated until dry. The crude product was purified on silica gel eluting with 0 - 50 hexane / EtOAc to obtain the desired product.

Chemical formula

[0203] 3-Benzyl-1-((2R,4S,5R)-5-(((tert-butyldimethylsilyl)oxy)methyl)-4-((2-nitrobenzyl)oxy)-tetrahydrofuran-2-yl)-5-methylpyrimidine-2,4(1H,3H)-dione (1.50 g, 1 equivalent, 2.58 mmol) was dissolved in THF and cooled to 0 °C. Then, TBAF (674 mg, 1 equivalent, 2.58 mmol) was added to the reaction at 0 °C, and the mixture was stirred at 0 °C for 1 hour and then warmed to room temperature over 2 hours. The mixture was then poured into a pre-cooled solution of 10% NaHCO3 and DCM. The DCM layer was separated and dried over sodium sulfate. The crude product was purified on silica gel eluting with 5 - 25% DCM / methanol to afford the desired product. [Chemical formula]

[0204] (1-((2R,4S,5R)-5-(Hydroxymethyl)-4-((2-nitrobenzyl)oxy)tetrahydrofuran-2-yl)-2-oxo-1,2-dihydropyrimidin-4-yl)carbamic acid benzyl (35.0 mg, 1 equivalent, 70.5 μmol) was dissolved in trimethyl phosphate (1.5 mL) and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot of phosphoryl trichloride (39.6 mg, 3 equivalents, 258 μmol) was added. After 5 minutes, a second aliquot of 10 μL was added. The mixture was stirred for an additional 30 minutes. A solution of tetrabutylammonium hydrogen diphosphate (311 mg, 4 equivalents, 345 μmol) in dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the reaction mixture over 30 seconds at rxn t = 35 minutes. Immediately, a pre-weighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (73.8 mg, 4 equivalents, 345 μmol) was added as a solid all at once. The mixture was stirred for 30 minutes after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred for 30 minutes and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation.

[0205] (Example 18) Procedure for synthesizing class VI peptides and non-peptide dGTP analogs. Scheme for the synthesis of class VI-dGTP constructs [Chemical formula] Detailed procedure for non-peptide dGTP analogs: [Chemical formula]

[0206] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1,9-dihydro-6H-purin-6-one (0.50 g, 1 equivalent, 1.0 mmol) was dissolved in 5.0 mL of dry dimethylacetamide under argon. Oxirane (0.13 g, 3 equivalents, 3.0 mmol) was added at ambient temperature, and then sodium hydroxide (40 mg, 1 equivalent, 1.0 mmol) was added as a solid. The mixture was stirred at ambient temperature for 4 hours. The mixture was diluted with 50 mL of EtOAc and this was washed sequentially with 100 mL of water and 100 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was chromatographed on 40 g of silica using a dichloromethane / methanol mixture as eluent to give a white foam. [Chemical formula]

[0207] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1-(2-hydroxyethyl)-1,9-dihydro-6H-purin-6-one (200 mg, 1 equivalent, 370 μmol) was suspended in 5 mL of dry pyridine at ambient temperature under argon. Chloroformate (1 equivalent) was added as a solid. The mixture was heated to 95 °C for 8 hours and cooled to ambient temperature. The solvent was removed in vacuo, and the residue was diluted with 50 mL of EtOAc, which was washed sequentially with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was subjected to chromatography on 40 g of silica using a dichloromethane / methanol mixture as eluent to give a white foam.

Chemical formula

[0208] The alcohol starting material was dissolved in dry THF at ambient temperature under argon. 2 equivalents of triethylamine were added. A solution of acyl chloride in THF was added dropwise at ambient temperature, and the mixture was stirred for 18 hours. The solvent was removed in vacuo, and the residue was diluted with 50 mL of EtOAc, which was washed sequentially with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over Na2SO4 and evaporated to leave a light brown solid. This was subjected to chromatography on a silica column using a dichloromethane / methanol mixture as eluent to give the corresponding ester.

Chemical formula

[0209] Bis-silyl ether (1 equivalent) was dissolved in dry THF at ambient temperature under argon. Triethylamine (8 equivalents) was added rapidly, followed by rapid addition of triethylammonium fluoride dihydrofluoride (6 equivalents) also at ambient temperature. The mixture was stirred at ambient temperature for 24 h. Silica gel was added and the mixture was evaporated on a rotary evaporator to a fine powder, then loaded onto a silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam.

Chem.

[0210] The nucleoside was co-evaporated with pyridine (1 mL × 3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL of trimethyl phosphate and 0.60 mL of dry pyridine and cooled in an ice bath under argon. A first aliquot (1.5 equivalents) of phosphoryl trichloride was added. After 5 min, a second aliquot of 1.5 equivalents was added. The mixture was stirred for a further 30 min. A solution of tetrabutylammonium hydrogen diphosphate (4 equivalents) in 1.5 mL of dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, a pre-weighed amount of N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (4 equivalents) was added as a solid in one portion. The mixture was stirred for 30 min after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 min and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was carried out immediately after the EtOAc extraction. Final purification was by reverse phase HPLC.

[0211] Detailed procedure for the peptide-dGTP analog:

Chem.

[0212] 2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-1,9-dihydro-6H-purin-6-one (1.00 g, 1 equivalent, 2.02 mmol) was dissolved in 30 mL of dry N,N-dimethylacetamide under argon. 4-Bromobutanoic acid (337 mg, 1 equivalent, 2.02 mmol) was added at ambient temperature, and then sodium hydroxide (161 mg, 2 equivalents, 4.03 mmol) was added as a solid. The mixture was heated to 80 °C and stirred for 12 hours. The mixture was cooled to ambient temperature, diluted with 100 mL of EtOAc, and this was washed successively with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over Na2SO4 and evaporated to leave a light brown solid. This was subjected to chromatography on a silica column using a dichloromethane / methanol mixture as eluent to give 4-(2-amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-6-oxo-6,9-dihydro-1H-purin-1-yl)butanoic acid as a white solid. [Chemical formula]

[0213] 4-(2-Amino-9-((2R,4S,5R)-4-((tert-butyldimethylsilyl)oxy)-5-(((tert-butyldimethylsilyl)oxy)methyl)tetrahydrofuran-2-yl)-6-oxo-6,9-dihydro-1H-purin-1-yl)butanoic acid (1 equivalent) was suspended in 5 mL of dry pyridine at ambient temperature under argon. Chloroformate (1 equivalent) was added as a solid. The mixture was heated to 95 °C for 8 h and cooled to ambient temperature. The solvent was removed in vacuo and the residue was diluted with 50 mL of EtOAc, which was sequentially washed with 50 mL of water and 50 mL of brine. The EtOAc layer was dried over sodium sulfate and evaporated to leave a yellow oil. This was subjected to chromatography on 40 g of silica using a dichloromethane / methanol mixture as eluent to give a white foam.

Chemical formula

[0214] The carboxylic acid was dissolved or suspended in dry THF at ambient temperature. 1.3 equivalents of triethylamine were added to this solution, followed by 1.1 equivalents of diphenylphosphoryl azide. The mixture was heated to reflux for 20 h and cooled to ambient temperature. Silica gel was added to the mixture and the solvent was evaporated to give a fine powder. This was loaded onto a column of silica gel and eluted with a mixture of EtOAc and dichloromethane to give the desired isocyanate as a colorless oil.

Chemical formula

[0215] Bis-silyl ether (1 equivalent) was dissolved in dry THF at ambient temperature under argon. Triethylamine (8 equivalents) was added rapidly, followed by rapid addition of triethylammonium fluoride dihydrofluoride (6 equivalents) also at ambient temperature. The mixture was stirred at ambient temperature for 24 h. Silica gel was added and the mixture was evaporated on a rotary evaporator to a fine powder, then loaded onto a silica column and eluted with a mixture of dichloromethane and methanol to give the nucleoside as a slightly yellow foam.

Chem.

[0216] The nucleoside was co-evaporated with pyridine (1 mL×3) and dried under high vacuum overnight. This was then dissolved in 1.5 mL trimethyl phosphate and 0.60 mL dry pyridine and cooled in an ice bath under argon. The first aliquot (1.5 equivalents) of phosphoryl trichloride was added. After 5 min, a second aliquot of 1.5 equivalents was added. The mixture was stirred for an additional 30 min. A solution of tetrabutylammonium hydrogen diphosphate (4 equivalents) in 1.5 mL dry DMF was prepared under Ar and cooled in an ice bath. This was added dropwise to the rxn mixture over 30 s. Immediately, the preweighed N1,N1,N8,N8-tetramethylnaphthalene-1,8-diamine (4 equivalents) was added as a solid in one portion. The mixture was stirred for 30 min after this addition and quenched with 8 mL of cold 0.1 M TEAB buffer. The mixture was stirred in an ice bath for 10 min and then transferred to a separatory funnel. The solution was extracted once with 10 mL of EtOAc. The aqueous layer was transferred to a small tube for FPLC separation, which was carried out immediately after the EtOAc extraction. Final purification was by reverse-phase HPLC.

[0217] The decay of 3’-O-(2-nitro-benzyl)-dATP and homopolymer synthesis are shown in FIGS. 20 and 21. 25 uM of 3’-O-(2-nitro-benzyl)-dATP (TriLink Technologies, San Diego, CA) was mixed with 1×TdT reaction buffer (Thermo-Fisher), 2 U / uL of terminal deoxynucleotidyl transferase (Thermo-Fisher) and 0.002 U / uL inorganic pyrophosphatase (Thermo-Fisher), together with 1 uM of oligonucleotide initiator (5’-biotin-TTTTTTGGCCTTTTUTAATAATAATAATAATTTTT, IDT). The reaction volume was subjected to 365 nm light at 20 - 22 mW / cm2 for various intervals and then incubated at 37 oC for 30 minutes. After quenching by the addition of 0.1 M EDTA, each time point was mixed with an equal volume of 2×Novex TBE-urea gel loading buffer (Thermo-Fisher), analyzed by polyacrylamide gel electrophoresis (15%), stained with Sybr Gold (Thermo-Fisher), and photographed with an ultraviolet transilluminator.

[0218] Incorporation by reference References and citations to other documents such as patents, patent applications, patent publications, journals, books, papers, web content, etc. have been made throughout this disclosure. All such documents are hereby incorporated by reference in their entirety for all purposes.

[0219] Equivalents In addition to what is shown and described herein, various modifications of the invention and many additional embodiments thereof will become apparent to those skilled in the art from the entire contents of this document, including references to scientific and patent literature cited herein. The subject matter of this specification includes important information, exemplification and guidance that can be adapted for the practice of the invention in its various embodiments and their equivalents.

Claims

**Claim 1** A method for synthesizing a plurality of nucleic acid memory strands, comprising: providing an array of two or more nucleic acids linked to a substrate; providing blocked nucleotide analogs and a template-independent polymerase to the array; delivering an addressable activation energy to a position on the array to remove a blocking group at the 3'-OH of the blocked nucleotide analog, thereby converting the blocked nucleotide analog to an unblocked nucleotide analog, wherein the activation energy includes light, reducing conditions, pH change, or heat; extending one or more of the nucleic acids of the array with a homopolymer tract of at least two repeating unblocked nucleotide analogs; and wherein the blocking group is a 3'-O-blocking group that cages the nucleotide analog until removed to allow polymerization to proceed. **Claim 2** The method of claim 1, wherein the blocked nucleotide analog includes a removable blocking group on the 3'-OH of the deoxyribose or ribose of a nucleotide triphosphate and a non-removable modification on the purine or pyrimidine base of the nucleotide analog. **Claim 3** The method of claim 1, wherein the plurality of blocked nucleotide analogs are modified nucleotides of the same nucleic acid base, including a removable 3'-O-blocking group and two or more non-removable molecular modifications that allow discrimination between modified nucleotide analogs of the same nucleic acid base. **Claim 4** The method of claim 1, further comprising stopping the extension; and delivering an addressable activation energy to the homopolymer tract in the presence of another plurality of blocked nucleotide analogs and the template-independent polymerase to extend the homopolymer tract with a further homopolymer tract of two or more repeating nucleotides. **Claim 5** The method of claim 4, wherein the extension is stopped after a predetermined length of time to obtain a desired length for the homopolymer tract. **Claim 6** The method of claim 1, wherein the repeating nucleotides of the homopolymer tract are between 2 and 10. **Claim 7** ​ The method according to claim 4, further comprising the step of repeating the steps of stopping and extending to synthesize a nucleic acid memory strand.

8. The method according to claim 7, wherein the nucleic acid memory strand is 200 nucleotides to 5,000 nucleotides in length.

9. The method according to claim 1, wherein a predetermined concentration of the blocked nucleotide analog is provided to obtain a desired length for the homopolymer tract.

10. The method according to claim 4, wherein the homopolymer tract and the additional homopolymer tract contain different nucleobases.

11. The method according to claim 7, wherein the nucleic acid memory strand encodes a data set selected from the group consisting of a text file, an image file, and an audio file.

12. The method according to claim 11, further comprising the step of displaying a human-readable format of the data set on a monitor or using a printer.

13. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is represented in base 2.

14. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is represented in base 3.

15. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is represented in base 4.

16. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is represented in a base greater than base 4.

17. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is indicated by the degree of decay or the tract length obtained in individual steps of memory strand synthesis.

18. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is read by DNA sequencing.

19. The method according to claim 7, wherein the data encoded in the nucleic acid memory strand is read by passing the nucleic acid memory strand through a nanopore.

Citation Information

Patent Citations

  • Information storage methods using nucleic acids

    JP2015533077A

  • Methods of analyzing nucleic acids associated with single cells using nucleic acid barcodes

    JP2017506877A

  • Methods and apparatus for nucleic acid synthesis

    JP2017525391A

  • Systems for nucleic acid-based data storage

    JP2020507168A

  • Homopolymer-encoded nucleic acid memory

    JP2020522257A