Methods of recording digital data using DNA
By employing DNA polymerases to alternate nucleotide solutions for forming standard or non-standard base pairs with template strands, the method addresses accuracy and scalability issues in DNA oligonucleotide synthesis, achieving efficient, sustainable, and high-density data storage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE UNIV OF NORTH CAROLINA AT GREENSBORO
- Filing Date
- 2025-11-19
- Publication Date
- 2026-05-28
AI Technical Summary
Traditional chemical synthesis methods for DNA oligonucleotides suffer from insufficient synthesis accuracy, leading to lower data storage reliability and scalability issues, and generate significant toxic waste, which are unsustainable for high-density long-term storage solutions.
The use of DNA polymerases to generate new DNA molecules by alternating solutions of nucleotide triphosphates that form standard or non-standard base pairs with template strands, allowing for discrete and scalable synthesis of nucleic acid data storage compounds, including methods like rolling circle amplification to extend strands to tens of thousands of nucleotides.
This approach increases synthesis accuracy, reduces environmental impact, and enables scalable, high-density data storage with faster writing speeds and lower costs, compatible with parallel synthesis on microscale chips.
Smart Images

Figure US2025056168_28052026_PF_FP_ABST
Abstract
Description
METHODS OF RECORDING DIGITAL DATA USING DNARELATED APPLICATION DATA
[0001] The present application claims priority pursuant to 35 U.S.C. § 119(e) to United States Provisional Patent Application Number 63 / 722,754 filed November 20, 2024, which is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with Government support under Grant No. GM133483 awarded by the National Institutes of Health. This invention was also made with Government support under Grant No. DMR-2027738 awarded by the National Science Foundation. The Government has certain rights in the invention.FIELD
[0003] The present disclosure relates to systems and methods of recording digital data through generation of new DNA molecules via environmentally friendly DNA polymerase and nucleotides.BACKGROUND
[0004] Digital information production and storage is increasing exponentially with the demand for cloud storage and external data storage devices, as well as higher adoption and use of smartphones, laptops, and personal computers with online storage capabilities. For instance, consumers generate large amounts of data and media files when using smartphone cameras and viewing increasingly accessible audio and video content on the Internet. However, the digital storage industry is approaching physical limits regarding silicon extraction and production for microchips, leading to increasingly unmet needs for alternative high-density long-term storage solutions.
[0005] DNA molecules have emerged as an alternative archival storage medium that offers several potential advantages, including higher density and retention, and lower energy consumption compared to state-of-the-art memory materials. However, the traditional chemicalsynthesis methods used to produce data-containing DNA oligonucleotides suffer from insufficient synthesis accuracies at a nucleotide level, leading to lower accuracy of data storage using long DNA oligonucleotides. Thus, improvements are desired in increasing accuracy in synthesis of long DNA oligonucleotides for data storage applications.SUMMARY
[0006] Generally, the present disclosure is directed to methods and systems for storing information using nucleic acid molecules. In one instance, the present disclosure is directed to a nucleic acid data storage compound. The nucleic acid storage compound includes a nucleic acid new strand having at least one first distinct nucleic acid segment and at least one second distinct nucleic acid segment. One or more of the at least one first distinct nucleic acid segment and the at least one second distinct nucleic acid segment is mutated when generated using chemically- altered nucleotide triphosphates that form non-standard based pairs with a template strand.
[0007] In some embodiments, each of the at least one first distinct nucleic acid segment and at least second distinct segment corresponds to a binary value of ‘ 1’ if mutated and a binary value of ‘0’ if not mutated. In some embodiments, the nucleic acid new strand is surface-anchored. In some instances, the nucleic acid new strand includes at least 10,000 nucleotides, or at least 100,000 nucleotides.
[0008] The present disclosure is further directed to a nucleic acid data storage system. The system includes a nucleic acid template having at least two segments, where the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides. The system further includes a first solution of three nucleotide triphosphates configured to base pair with the first group of three nucleotides of the alpha segment. The first solution optionally includes chemically-altered nucleotide triphosphates configured to form non-standard based pairs. Additionally, the system includes a second solution of three nucleotide triphosphates that at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment. The second solution optionally includes chemically-altered nucleotide triphosphates configured to form non-standard based pairs. The system includes a DNA polymerase configured to generate a nucleic acid new strand with at least one first distinctsegment corresponding to the alpha segment and at least one second distinct segment corresponding to the beta segment. The DNA polymerase is configured to introduce mutations to one or more of the at least one first distinct segment and the at least one second distinct segment when chemically-altered nucleotide triphosphates are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates, respectively.
[0009] In some embodiments, the nucleic acid template is a circularized nucleic acid template. In some embodiments, the DNA polymerase is from phage phi29. In some embodiments, the DNA polymerase is configured to participate in rolling circle amplification. In some embodiments, the DNA polymerase is configured to halt generation of the nucleic new acid strand when at least one nucleotide triphosphate corresponding to a nucleotide of the nucleic acid template is not present.
[0010] The present disclosure is further directed to methods of storing information using a nucleic acid strand. The method includes generating a nucleic acid template having at least two segments, where the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides. The nucleic acid template is contacted with a DNA polymerase and, altematingly, a first solution and second solution of nucleotide triphosphates. The first solution of three nucleotide triphosphates is configured to base pair with the first group of three nucleotides of the alpha segment, while the second solution of three nucleotide triphosphates at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment. A nucleic acid new strand that base pairs with the nucleic acid template is generated using the DNA polymerase. The nucleic acid new strand includes at least one first distinct segment corresponding to the alpha segment and at least one second distinct segment corresponding to the beta segment, where mutations are selectively introduced to one or more of the at least one first distinct segment and the at least one second distinct segment when chemically-altered nucleotide triphosphates configured to form non-standard based pairs are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates.
[0011] In some embodiments, the method further includes sequencing the nucleic acid new strand, where each of the at least one first distinct segment and at least second distinct segment is evaluated for mutations. In some embodiments, the method further includes assigning each of the at least one first distinct segment and at least second distinct segment a binary value of ‘1’ if mutated and a binary value of ‘0’ if not mutated. In some embodiments, the nucleic acid new strand is surface-anchored. In some instances, the nucleic acid new strand includes at least 10,000 nucleotides, or at least 100,000 nucleotides. In some embodiments, the method includes washing the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates between cycles of alternating contacting. In some embodiments, the nucleic acid template is a circularized nucleic acid template. In some embodiments, the DNA polymerase is from phage phi29. In some embodiments, the method includes generating the nucleic acid new strand occurs through rolling circle amplification.BRIEF DESCRIPTION OF THE FIGURES
[0012] FIG. 1A is a schematic representation of recording digital information with DNA using discretized template strands. Each segment depicted is an alpha or beta segment and uses only three of the four nucleotides. When DNA polymerase (dark large circle) encounters a nucleotide of the template strand for which there is no corresponding free nucleotide triphosphate in the solution, it is halted.
[0013] FIG. IB is a schematic representation of recording digital information with DNA using discretized template strands. Following the halting in FIG. 1A, when the original three nucleotides in the first solution are washed and then replaced with second solution of three of the four nucleotides corresponding for the beta segment, the polymerase extends the new strand until it again reaches the next segment (here, alpha) without a corresponding free nucleotide triphosphate.
[0014] FIG. 1C is a schematic representation of recording digital information with DNA using discretized template strands. When extension occurs under mutagenic conditions (with dashed nucleotide triphosphate) that can be incorrectly incorporated during extension, the segment is scored as a ‘ 1’ after sequencing, while unmutated segments generated under standard conditions are scored a ‘O’.
[0015] FIG. 2A is a schematic representation of circularized extension with discrete segments. When a circularized template is used, segments can still be extended in a discretized manner based on the identity of nucleotide triphosphates in the solution.
[0016] FIG. 2B is a schematic representation of circularized extension with discrete segments. When DNA polymerase (dark large circle) returns to a segment that it has previously used as a template, it can continue extension based on the identity of nucleotide triphosphate in the solution.
[0017] FIG. 2C is a schematic representation of circularized extension with discrete segments. When DNA polymerase (dark large circle) returns to a segment that it has previously used as a template, it can extend the strand under mutagenic conditions with nucleotides that can be incorporated incorrectly during the extension, allowing the mutated segment (light grey) to be scored as a ‘ 1 ’ after nanopore sequencing and allowing unmutated segments (black) to be scored as ‘0’ after nanopore sequencing.DETAILED DESCRIPTION
[0018] Embodiments described herein can be understood more readily by reference to the following detailed description and examples. Elements, apparatus and methods described herein, however, are not limited to the specific embodiments presented in the detailed description and examples. It should be recognized that these embodiments are merely illustrative of the principles of the present disclosure. Numerous modifications and adaptations will be readily apparent to those of skill in the art without departing from the spirit and scope of the disclosure.
[0019] In addition, all ranges disclosed herein are to be understood to encompass any and all subranges subsumed therein. For example, a stated range of “1.0 to 10.0” should be considered to include any and all subranges beginning with a minimum value of 1.0 or more and ending with a maximum value of 10.0 or less, e.g., 1.0 to 5.3, or 4.7 to 10.0, or 3.6 to 7.9.
[0020] All ranges disclosed herein are also to be considered to include the end points of the range, unless expressly stated otherwise. For example, a range of “between 5 and 10,” “from 5 to 10,” or “5-10” should generally be considered to include the end points 5 and 10.
[0021] Further, when the phrase “up to” is used in connection with an amount or quantity, it is to be understood that the amount is at least a detectable amount or quantity. For example, amaterial present in an amount “up to” a specified amount can be present from a detectable amount and up to and including the specified amount.
[0022] The terms “about” and “approximately” shall generally mean an acceptable degree of error or variation for the quantity measured given the nature or precision of the measurements. Typical, exemplary degrees of error or variation are within 20 percent (%), preferably within 10%, more preferably within 5%, and still more preferably within 1% of a given value or range of values. Numerical quantities given in this description are approximate unless stated otherwise, meaning that the term “about” or “approximately” can be inferred when not expressly stated.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0024] The terms “first”, “second”, and the like are used herein to describe various features or elements, but these features or elements should not be limited by these terms. These terms are only used to distinguish one feature or element from another feature or element. Thus, a first feature or element discussed below could be termed a second feature or element, and similarly, a second feature or element discussed below could be termed a first feature or element without departing from the teachings of the present disclosure.
[0025] The term “consisting essentially of’ means that, in addition to the recited elements, what is claimed may also contain other elements (steps, structures, ingredients, components, etc.) that do not adversely affect the operability of what is claimed for its intended purpose as stated in this disclosure. Importantly, this term excludes such other elements that adversely affect the operability of what is claimed for its intended purpose as stated in this disclosure, even if such other elements might enhance the operability of what is claimed for some other purpose.
[0026] DNA molecules may be used in data storage applications owing to their potential advantages, including higher density and retention, and lower energy consumption compared to the conventional memory materials. Chemical synthesis of DNA molecules has enabled the development of DNA-based memory technologies that use the nucleotide sequence of a DNA oligonucleotide polymer, with the assignment of the four nucleotides (A, G, C, and T) to different digital values (00, 01, 10, 11, respectively, for example). When a DNA molecule is sequenced, the corresponding digital data (01001010111010101011 , for example) can berecovered from the chemical structure of the DNA double helix and converted from binary data to the intended data.
[0027] The traditional chemical synthesis of these data-containing DNA oligonucleotides entails the individual addition of nucleotides to the extending polymer. For instance, a nucleotide that is chemically protected on one side and reactive on the other side is chemically coupled to a surface-anchored oligonucleotide, extending that oligonucleotide. Next, the solution of the remaining, unreacted nucleotides is washed from the anchored oligonucleotide and the previously added nucleotide on the anchored oligonucleotide is de-protected so that it is reactive. The cycle is then repeated with addition of another solution of nucleotides having one protected side and another reactive side, allowing the anchored oligonucleotide to be extended one nucleotide at a time. Notably, these synthesis methods can allow simultaneous writing of oligonucleotide-based data in discrete regions of a substrate, such as within microscale wells of specialized microchips.
[0028] However, this typical method of DNA oligonucleotide chemical synthesis is inefficient, with an approximately 0.5% probability of failure with each nucleotide addition cycle. That is, there is an approximately 60% chance of accurate and correct synthesis for every 100 nucleotides, an approximately 37% chance of accurate and correct synthesis for every 200 nucleotides, and less than 1% chance of accurate and correct synthesis for every 1,000 nucleotides. Therefore, the lack of accuracy in synthesizing longer DNA oligonucleotides using traditional methods imposes a maximum size limitation for DNA-based memory storage.
[0029] It is predicted that advances in DNA synthesis technologies will increase the speed and reduce the costs of these methods. Further, while automation of writing and reading processes can improve the portability and scalability of DNA memory, traditional de novo synthesis methods are unlikely to meet the scalability needs of high-volume memory manufacturing. Moreover, these traditional methods generate large and unsustainable amounts of toxic waste. For instance, the storage of five minutes of a 1080p video stream using commercially acquired DNA would cost more than $7 million dollars, consume over 100 kilowatt-hours (KWh) of energy, require over four days to store, and generate more than 15 liters of toxic waste. By extension, the synthetic DNA storage of approximately 6 x 1023Bytes of information predicted to be generated by 2030 would produce nearly 85 petaliters of toxic waste using traditionalmethods, an amount surpassing the volume of water that the Mississippi River empties into the Gulf of Mexico over 40 years.
[0030] Alternative methods of DNA-based digital storage have been developed to address some of these issues. For instance, sustainability has been increased using methods of introducing mutations or changes into nucleotide sequences of previously existing DNA molecules rather than by synthesizing DNA storage molecules anew. See Sadremomtaz, A., Glass, R.F., Guerrero, J.E. et al. Digital data storage on DNA tape using CRISPR base editors. Nat Commun 14, 6472 (2023). In those methods, a DNA molecule is engineered with a number of 20-nucleotide sequences that are recognizable by Cas9, a CRISPR enzyme. When Cas9 recognizes the sequence and binds the DNA molecule, a mutagen that is also present chemically changes C nucleotides to T nucleotides only within those sequences where Cas9 binds. Thus, the DNA molecules can be mutated at different 20-nucleotide positions along the length of the preexisting DNA by introducing Cas9 targeting different sequences. Then, the mutated DNA molecule is sequenced using nanopore sequencing techniques that can process molecules of up to hundreds of thousands of nucleotides in length. Where known Cas9 target locations are mutated, a score of ‘ 1’ is assigned, and where the known Cas9 target locations are unmutated, the value is a ‘0’ in binary code. This system of assigning binary values allows digital information to be recorded into DNA molecules using CRISPR enzymes targeted to different sequences.
[0031] One problem with the CRISPR approach to digital storage is that each site of the preexisting DNA molecule must have a unique CRISPR target, limiting the method’s scalability. Therefore, methods for digital storage are desired that are both scalable and sustainable to accommodate growing digital storage needs.
[0032] The present disclosure is directed to approaches for writing digital data by generating DNA molecules of indefinite lengths with a “green” DNA polymerase technique. The presented methods and systems include the use of DNA polymerases that add solution-based, free nucleotide triphosphates individually to the 3 ’-end of a DNA molecule (the “new strand”) that is base-paired with another DNA strand (the “template strand”) based on complementary pairing of nucleotides (e.g., A added opposite a T on the template strand; T added opposite A on the template strand; C added opposite G on the template strand; G added opposite C on the template strand). However, when there is no free nucleotide triphosphate to form a complementary pair with the template strand, the DNA polymerase halts new strand extension, while remainingbound to the 3 ’-end of the new strand. As shown in FIG. 1A, the methods include template strand designs with regions having only three of the four nucleotides in the template sequence (“alpha segments”), and other regions with a different set of three of the four nucleotides in the template sequence (“beta segments”). Due to these different segments each lacking a different nucleotide, the DNA polymerase may be utilized to discretely synthesize the new strand.
[0033] For example, template strand, which may be anchored to a surface, may include at least one alpha segment with a first group of three of the four nucleotides and at least one beta segment with a second group of three of the four nucleotides, where the first group and second group differ in the identity of one nucleotide. The DNA polymerase then extends a new strand, which is base-paired to the surface-anchored template strand. In the presence of a first solution of three free nucleotide triphosphates complementary to those of the alpha segment of the template strand, the DNA polymerase will extend the new strand until it reaches a beta segment that requires a nucleotide triphosphate that is not present in the first solution. At this point, the DNA polymerase halts extension while remaining bound to the new strand (FIG. 1 A). At this point, the first solution of free nucleotide triphosphates is washed or otherwise removed from the surface- tethered template strand, and a second solution of free nucleotide triphosphates is added, the second solution containing the three nucleotide triphosphates complementary to those of the beta segment of the template strand. Thus, the second solution contains one nucleotide that differs in identity from the three nucleotides of the first solution. When the second solution is present, the DNA polymerase resumes extension of the new strand until it reaches another alpha segment that requires a nucleotide triphosphate that is not present in the second solution. At this point, the DNA polymerase again halts while remaining bound to the new strand (FIG. IB), and the second solution is washed or otherwise removed from the surface-tethered template strand. The first solution of free nucleotide triphosphates is again added, and the extension can continue as described previously. Therefore, with alternating addition and removal cycles of first and second free nucleotide solutions, the DNA polymerase extends the new strand in discrete segments.
[0034] Using the present methods of discrete strand extension, the extension conditions can be varied for particular segment extension to create a readable code for digital storage. For instance, extension can occur under standard conditions (FIG. 1 A and FIG. IB) or altered conditions, such as mutagenic conditions (FIG. 1C). Standard conditions include standard nucleotide triphosphate solutions, where the nucleotides in solution form complementary base pairs. Mutagenicconditions include chemical variations of at least one of the nucleotide triphosphates in at least one of the first solution or the second solution. The chemical variations allow base pairing with the template strand nucleotides in a non-standard way, which can be read by a DNA sequencer as a “mutation.” As shown in FIG. 1C, when the DNA polymerase extends the new strand along the desired segment (shown as an alpha segment, though beta segments can be used) under mutagenic conditions, the extended region base-pairs in a non-standard way. Upon sequencing, DNA written under the mutagenic conditions is read by the sequencer as a mutation and scored as a ‘ 1’, while DNA written under standard conditions include no mutations for detection by the sequencer and are scored as a ‘0’ in binary code. In other instances, DNA written under the mutagenic conditions may be scored as a ‘O’, while DNA written under standard conditions may be scored as a ‘ 1 ’ in binary code.
[0035] The presented methods and systems further include embodiments using a “circularized” template strand. This approach uses rolling circle amplification (RCA), which can generate longer new strands than would typically be possible with a linear template strand. The circularized template strand approach is shown in FIG. 2A, where segments can still be extended in a discretized manner based on the identity of nucleotides in segments of the circularized template strand and the identity of nucleotide triphosphates in solution. A circularized template strand is a single-stranded oligonucleotide with its 5’-end connected to its 3’-end, typically through chemical connection. Certain DNA polymerases, such as the DNA polymerase from phage phi29, can extend a new strand repeatedly in the presence of the circularized template strand. In such cases, when the DNA polymerase arrives back at its original, starting nucleotide, and with portions of the new strand in front of it (FIG. 2B), the portions are displaced from the circularized template strand so that the DNA polymerase can continue extending the new strand using the RCA process. The DNA polymerase from phi29 can extend the new strand to a length of tens or hundreds of thousands of nucleotides during RCA, improving the scalability compared to strand extension techniques using a linear template strand.
[0036] Similar to the discretized strand synthesis in FIG. 1A-FIG. 1C, a new strand can be synthesized in discrete segments using rolling circle amplification (RCA) in FIG. 2A-FIG. 2C. This approach uses a “circularized” template, which is a single-stranded oligonucleotide where its 5’-end is chemically connected to its 3’ end, and thus is less limited by template strand size relative to linear template approaches. RCA is compatible with certain DNA polymerases, likethe DNA polymerase from phage phi29, which will extend a new strand repeatedly based on the circularized template. Phi29 can extend new strands that are tens or hundreds of thousands of nucleotides long during RCA.
[0037] In FIG. 2A, the circularized template strand includes an alpha segment with a first group of three of the four nucleotides and a beta segment with a second group of three of the four nucleotides, where the first group and second group differ in the identity of one nucleotide. The DNA polymerase that is compatible with RCA then extends a new strand, which is base-paired to the circularized template strand. In the presence of a first solution of three free nucleotide triphosphates complementary to those of the alpha segment of the circularized template strand, the DNA polymerase will extend the new strand until it reaches the beta segment that requires a nucleotide triphosphate that is not present in the first solution. At this point, the DNA polymerase halts extension while remaining bound to the new strand (FIG. 2A). The first solution of free nucleotide triphosphates is washed or otherwise removed from the circularized template strand, and a second solution of free nucleotide triphosphates is added, the second solution containing the three nucleotide triphosphates complementary to those of the beta segment of the circularized template strand. Thus, the second solution contains one nucleotide that differs in identity from the three nucleotides of the first solution. When the second solution is present, the DNA polymerase resumes extension of the new strand until it again reaches the alpha segment and requires a nucleotide triphosphate that is not present in the second solution. At this point, the DNA polymerase again halts while remaining bound to the new strand (FIG. 2B), and the second solution is washed or otherwise removed from the surface-tethered template strand. The first solution of free nucleotide triphosphates is again added, and the extension can continue as described previously, with the DNA polymerase displacing extended regions of the new strand from the circularized template strand to accommodate the continued extension. Throughout this process, the new strand may be anchored to a surface. Therefore, with alternating addition and removal cycles of first and second free nucleotide solutions, the DNA polymerase extends the new strand in discrete segments.
[0038] Similar to extension using a linearized template strand, extension using a circularized template strand can occur under standard conditions (FIG. 2A and FIG. 2B) or altered conditions, such as mutagenic conditions (FIG. 2C). As shown in FIG. 2C, when the DNA polymerase extends the new strand along the desired segment (shown as a beta segment, thoughalpha segments can be used) under mutagenic conditions, the extended region base-pairs in a non-standard way. The single-stranded DNA generated during RCA can be converted to doublestranded DNA readily by virtue of its repeating patterns (of alpha and beta segments) and then stored long-term for sequencing or sequenced directly using nanopore sequencing. Upon sequencing, DNA written under the mutagenic conditions is read by the sequencer as a mutation and scored as a ‘ 1’, while DNA written under standard conditions include no mutations for detection by the sequencer and are scored as a ‘0’ in binary code.
[0039] The disclosed methods and systems present the ability to write digital data by generating new DNA in a sustainable, green, and scalable manner. The speed of these methods and systems can be orders of magnitude faster in writing new information, can have higher data density at a per nucleotide level, and can generate DNA molecules with longer strings of data than traditional chemical DNA synthesis methods permit. Additionally, the present methods and systems are scalable and compatible with parallel synthesis using distinct substrate synthesis regions, such as wells of a microscale chip.
[0040] “Binding region” as used herein refers to the region within a nuclease target region that is recognized and bound by the nuclease, such as Cas9.
[0041] “Complement” or “complementary” as used herein means a nucleic acid can mean Watson-Crick (e.g., A-T / U and C-G) or Hoogsteen base pairing between nucleotides or nucleotide analogs of nucleic acid molecules.
[0042] ‘ ‘Identical” or “identity” as used herein in the context of two or more nucleic acids or polypeptide sequences means that the sequences have a specified percentage of residues that are the same over a specified region. The percentage may be calculated by optimally aligning the two sequences, comparing the two sequences over the specified region, determining the number of positions at which the identical residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the specified region, and multiplying the result by 100 to yield the percentage of sequence identity. In cases where the two sequences are of different lengths or the alignment produces one or more staggered ends and the specified region of comparison includes only a single sequence, the residues of single sequence are included in the denominator but not the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) may be consideredequivalent. Identity may be performed manually or by using a computer sequence algorithm such as BLAST or BLAST 2.0.
[0043] “Mismatched” as used herein refers to mismatched bases that include a G / T or A / C pairing.
[0044] “Nucleic acid” or “oligonucleotide” or “polynucleotide” as used herein means at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary strand. Thus, a nucleic acid also encompasses the complementary strand of a depicted single strand. Many variants of a nucleic acid may be used for the same purpose as a given nucleic acid. Thus, a nucleic acid also encompasses substantially identical nucleic acids and complements thereof. A single strand provides a probe that may hybridize to a target sequence under stringent hybridization conditions. Thus, a nucleic acid also encompasses a probe that hybridizes under stringent hybridization conditions.
[0045] Nucleic acids may be single stranded or double stranded or may contain portions of both double stranded and single stranded sequence. The nucleic acid may be DNA, both genomic and cDNA, RNA, or a hybrid, where the nucleic acid may contain combinations of deoxyribo- and ribo-nucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine and isoguanine. Nucleic acids may be obtained by chemical synthesis methods or by recombinant methods.
[0046] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. For example, any nomenclatures used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those that are well known and commonly used in the art. The meaning and scope of the terms should be clear; in the event however of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities, and plural terms shall include the singular.EMBODIMENTS
[0047] Some additional non-limiting example Embodiments are as follows.
[0048] Embodiment 1. A nucleic acid data storage compound comprising a nucleic acid new strand comprising at least one first distinct nucleic acid segment and at least one second distinct nucleic acid segment, wherein one or more of the at least one first distinct nucleic acid segment and the at least one second distinct nucleic acid segment is mutated when generated using chemically-altered nucleotide triphosphates that form non-standard based pairs with a template strand.
[0049] Embodiment 2. The nucleic acid data storage compound according to Embodiment 1, wherein each of the at least one first distinct nucleic acid segment and at least second distinct segment corresponds to a binary value of ‘ 1 ’ if mutated and a binary value of ‘O’ if not mutated.
[0050] Embodiment 3. The nucleic acid data storage compound of Embodiments 1 or 2, wherein the nucleic acid new strand is surface-anchored.
[0051] Embodiment 4. The nucleic acid data storage compound of Embodiments 1, 2, or 3, wherein the nucleic acid new strand includes at least 10,000 nucleotides.
[0052] Embodiment 5. The nucleic acid data storage compound of Embodiments 1, 2, 3, or 4, wherein the nucleic acid new strand includes at least 100,000 nucleotides.
[0053] Embodiment 6. A nucleic acid data storage system comprising: a nucleic acid template having at least two segments, the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides; a first solution of three nucleotide triphosphates configured to base pair with the first group of three nucleotides of the alpha segment, the first solution optionally including chemically-altered nucleotide triphosphates configured to form non-standard based pairs; a second solution of three nucleotide triphosphates that at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment, the second solution optionally including chemically-altered nucleotide triphosphates configured to form non-standard based pairs; and a DNA polymerase configured to generate a nucleic acid new strand comprising at least one first distinct segment corresponding to the alpha segment and at least one second distinct segment corresponding to the beta segment, the DNA polymerase configured to introduce mutations to one or more of the at least one firstdistinct segment and the at least one second distinct segment when chemically-altered nucleotide triphosphates are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates, respectively.
[0054] Embodiment 7. The nucleic acid data storage system according to Embodiment 6, wherein the nucleic acid template is a circularized nucleic acid template.
[0055] Embodiment 8. The nucleic acid data storage system of Embodiments 6 or 7, wherein the DNA polymerase is from phage phi29.
[0056] Embodiment 9. The nucleic acid data storage system of Embodiments 6, 7, or 8, wherein the DNA polymerase is configured to participate in rolling circle amplification.
[0057] Embodiment 10. The nucleic acid data storage system of Embodiments 6, 7, 8, or 9, wherein the DNA polymerase is configured to halt generation of the nucleic new acid strand when at least one nucleotide triphosphate corresponding to a nucleotide of the nucleic acid template is not present.
[0058] Embodiment 11. A method for storing information using a nucleic acid strand comprising: generating a nucleic acid template having at least two segments, the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides; contacting the nucleic acid template with a DNA polymerase and, alternatingly: a first solution of three nucleotide triphosphates configured to base pair with the first group of three nucleotides of the alpha segment; and a second solution of three nucleotide triphosphates that at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment; generating a nucleic acid new strand that base pairs with the nucleic acid template using the DNA polymerase, the nucleic acid new strand comprising at least one first distinct segment corresponding to the alpha segment and at least one second distinct segment corresponding to the beta segment, wherein mutations are selectively introduced to one or more of the at least one first distinct segment and the at least one second distinct segment when chemically-altered nucleotide triphosphates configured to form non-standard based pairs are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates.
[0059] Embodiment 12. The method according to Embodiment 11, further comprising sequencing the nucleic acid new strand, wherein each of the at least one first distinct segment and at least second distinct segment is evaluated for mutations.
[0060] Embodiment 13. The method of Embodiments 11 or 12, further comprising assigning each of the at least one first distinct segment and at least second distinct segment a binary value of ‘ 1’ if mutated and a binary value of ‘0’ if not mutated.
[0061] Embodiment 14. The method of Embodiments 11, 12, or 13, wherein the nucleic acid new strand is surface-anchored.
[0062] Embodiment 15. The method of Embodiments 11, 12, 13, or 14, wherein the nucleic acid new strand includes at least 10,000 nucleotides.
[0063] Embodiment 16. The method of any of Embodiments 11-15, wherein the nucleic acid new strand includes at least 100,000 nucleotides.
[0064] Embodiment 17. The method of any of Embodiments 11-16, including washing the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates between cycles of alternating contacting.
[0065] Embodiment 18. The method of any of Embodiments 11-17, wherein the nucleic acid template is a circularized nucleic acid template.
[0066] Embodiment 19. The method of any of Embodiments 11-18, wherein the DNA polymerase is from phage phi29.
[0067] Embodiment 20. The method of any of Embodiments 11-19, wherein generating the nucleic acid new strand occurs through rolling circle amplification.
Claims
CLAIMS1. A nucleic acid data storage compound comprising: a nucleic acid new strand comprising at least one first distinct nucleic acid segment and at least one second distinct nucleic acid segment, wherein one or more of the at least one first distinct nucleic acid segment and the at least one second distinct nucleic acid segment is mutated when generated using chemically-altered nucleotide triphosphates that form nonstandard based pairs with a template strand.
2. The nucleic acid data storage compound of claim 1, wherein each of the at least one first distinct nucleic acid segment and at least second distinct segment corresponds to a binary value of ‘ 1’ if mutated and a binary value of ‘0’ if not mutated.
3. The nucleic acid data storage compound of claim 1, wherein the nucleic acid new strand is surface-anchored.
4. The nucleic acid data storage compound of claim 1, wherein the nucleic acid new strand includes at least 10,000 nucleotides.
5. The nucleic acid data storage compound of claim 4, wherein the nucleic acid new strand includes at least 100,000 nucleotides.
6. A nucleic acid data storage system comprising: a nucleic acid template having at least two segments, the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides; a first solution of three nucleotide triphosphates configured to base pair with the first group of three nucleotides of the alpha segment, the first solution optionally including chemically-altered nucleotide triphosphates configured to form non-standard based pairs; a second solution of three nucleotide triphosphates that at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment, the second solution optionally including chemically-altered nucleotide triphosphates configured to form non-standard based pairs; and a DNA polymerase configured to generate a nucleic acid new strand comprising at least one first distinct segment corresponding to the alpha segment and at least one seconddistinct segment corresponding to the beta segment, the DNA polymerase configured to introduce mutations to one or more of the at least one first distinct segment and the at least one second distinct segment when chemically-altered nucleotide triphosphates are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates, respectively.
7. The nucleic acid data storage system of claim 6, wherein the nucleic acid template is a circularized nucleic acid template.
8. The nucleic acid data storage system of claim 6, wherein the DNA polymerase is from phage phi29.
9. The nucleic acid data storage system of claim 6, wherein the DNA polymerase is configured to participate in rolling circle amplification.
10. The nucleic acid data storage system of claim 6, wherein the DNA polymerase is configured to halt generation of the nucleic new acid strand when at least one nucleotide triphosphate corresponding to a nucleotide of the nucleic acid template is not present.
11. A method for storing information using a nucleic acid strand comprising: generating a nucleic acid template having at least two segments, the at least two segments including an alpha segment comprising a first group of three nucleotides and a beta segment comprising a second group of three nucleotides that at least partially differs from the first group of three nucleotides; contacting the nucleic acid template with a DNA polymerase and, altematingly: a first solution of three nucleotide triphosphates configured to base pair with the first group of three nucleotides of the alpha segment; and a second solution of three nucleotide triphosphates that at least partially differs from the first solution of three nucleotide triphosphates and is configured to base pair with the second group of three nucleotides of the beta segment; generating a nucleic acid new strand that base pairs with the nucleic acid template using the DNA polymerase, the nucleic acid new strand comprising at least one first distinct segment corresponding to the alpha segment and at least one second distinct segment corresponding to the beta segment, wherein mutations are selectively introduced to one or more of the at least one first distinct segment and the at least one second distinct segment when chemically-alterednucleotide triphosphates configured to form non-standard based pairs are included in one or more of the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates.
12. The method of claim 11, further comprising sequencing the nucleic acid new strand, wherein each of the at least one first distinct segment and at least second distinct segment is evaluated for mutations.
13. The method of claim 12, further comprising assigning each of the at least one first distinct segment and at least second distinct segment a binary value of ‘ 1 ’ if mutated and a binary value of ‘0’ if not mutated.
14. The method of claim 11, wherein the nucleic acid new strand is surface-anchored.
15. The method of claim 11, wherein the nucleic acid new strand includes at least 10,000 nucleotides.
16. The method of claim 15, wherein the nucleic acid new strand includes at least 100,000 nucleotides.
17. The method of claim 11, including washing the first solution of three nucleotide triphosphates and the second solution of three nucleotide triphosphates between cycles of alternating contacting.
18. The method of claim 11, wherein the nucleic acid template is a circularized nucleic acid template.
19. The method of claim 11, wherein the DNA polymerase is from phage phi29.
20. The method of claim 11, wherein generating the nucleic acid new strand occurs through rolling circle amplification.