Nucleic Acid Security and Authentication
By encoding digital information using the presence or absence of unique nucleic acid sequences, the method addresses the inefficiencies and costs of base-by-base synthesis, providing a cost-effective and error-reduced method for nucleic acid data storage and authentication.
Patent Information
- Application Number
- JP2022521744
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-11
- Filing Date
- 2020-10-13
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2040-10-13
AI Technical Summary
Current methods for encoding digital information in nucleic acid sequences are error-prone and costly due to the need for base-by-base synthesis, making them inefficient for commercial applications.
Encoding digital information by designating bit positions with unique nucleic acid sequences and using their presence or absence to represent bit values, allowing for the creation of nucleic acid libraries that can be used for authentication and security without the need for de novo synthesis for each base.
This approach reduces encoding costs and improves efficiency by reusing nucleic acid sequences, enabling robust authentication and secure data storage with reduced errors.
Smart Images

Figure 0007726874000006 
Figure 0007726874000007 
Figure 0007726874000008
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 914,086, entitled "DNA STORAGE FOR SECURITY AND AUTHENTICATION," filed October 11, 2019, the entire contents of which are incorporated herein by reference. [Background technology]
[0002] background Nucleic acid digital data storage is a stable method for encoding and storing information over long periods of time, with data being stored at higher densities than magnetic tape or hard drive storage systems. Additionally, digital data stored in nucleic acid molecules stored in cold and dry conditions can be retrieved after as many as 60,000 years or more.
[0003] Current methods rely on encoding digital information (e.g., binary code) into nucleic acid sequences base by base, directly converting the relationship between bases in the sequence into digital information (e.g., binary code). Sequencing digital data stored in base-by-base sequences, which can be read into a bitstream or byte of digitally encoded information, can be error-prone, and encoding costs can be high because the cost of de novo nucleic acid synthesis for each base can be expensive. Opportunities for new methods of implementing nucleic acid digital data storage can provide approaches for data encoding and retrieval that are less expensive and easier to commercially implement. Summary of the Invention [Means for solving the problem]
[0004] Provided herein are methods and systems for encoding digital information into nucleic acid (e.g., deoxyribonucleic acid, DNA) molecules without base-by-base synthesis by encoding bit-value information based on the presence or absence of a unique nucleic acid sequence within a pool, comprising designating each bit position in a bitstream with a unique nucleic acid sequence, and designating the bit value at that position based on the presence or absence of a corresponding unique nucleic acid sequence within a pool. These encoded nucleic acid molecules are particularly useful for encoding confidential information or tagging artifacts with information using very small chemical quantities. By associating an artifact with a quantity of nucleic acid molecules that encode information, the artifact can be uniquely tagged in a manner that is not readily apparent to an external user, so that the artifact can be used for robust authentication or tracing the origin of the artifact.
[0005] In certain aspects, provided herein are methods for tagging fluids for tracking or authentication. The methods include obtaining a library of nucleic acid molecules representing digital information and combining the fluid with a tag comprising the library of nucleic acid molecules to obtain a tagged fluid for tracking or authentication. Because the library can be designed to uniquely identify the fluid, its origin, its production date, or any other characteristic of the fluid, tagging the fluid can be advantageous for authenticating the authenticity of fluids such as valuable fuels or pharmaceuticals.
[0006] In some implementations, the method further includes sampling the tagged fluid to obtain a sample containing at least a portion of the library of nucleic acid molecules. The sampling step can involve swabbing or removing a volume from the tag or tagged fluid. In some implementations, the method further involves sequencing the nucleic acid molecules of the sample to obtain a sequencing readout. The sequencing readout can be compared to a reference sequence to determine the presence of matching sequences. Thus, information encoded by the library can be determined, and the fluid can be authenticated or identified.
[0007] The fluid may be any one of oil, ink, compressed gas, or a drug. In some implementations, the method further includes measuring the concentration of the tag in the tagged fluid to determine the amount of dilution, which is useful in determining whether the fluid has been tampered with.
[0008] In some implementations, the tag includes a molecular barcode specific to the tag. The information can include a message or a monetary value. The information can include at least 1 kilobit of information. In some implementations, the method further includes accessing the tagged fluid, thereby attenuating the tag in the tagged fluid. In some implementations, the tag is part of a two-factor authentication system.
[0009] In some implementations, the library is generated randomly. In some implementations, the library is generated by selecting a subset of nucleic acid molecules from a pool of nucleic acid molecules. In some implementations, the information includes a plurality of symbols, each symbol represented by a distinct sequence of nucleic acid molecules in the library. In some implementations, the information is represented by a library of nucleic acid molecules using an encoding scheme, where the information is mapped to a plurality of symbols having one of two possible symbol values, where if a symbol has a first of the two possible symbol values, a symbol of the plurality of symbols is represented by the presence of a distinct nucleic acid molecule in the library, and if a symbol has a second of the two possible symbol values, the symbol is represented by the absence of the distinct nucleic acid molecule.
[0010] In another aspect, provided herein is a method for preparing a library of nucleic acid molecules for use in security and authentication, the method comprising obtaining a library of nucleic acid molecules representing security tokens and applying chemical manipulations to the library representing the security tokens to obtain a hashed library of nucleic acid molecules representing the hashed tokens. This method has the advantage over prior methods that air-gap the values of the security tokens by hashing the tokens before reading the library of nucleic acid molecules so that the sequence of the pre-hashed library is not revealed.
[0011] In some implementations, the chemical operation results in one or more Boolean functions on the security token. For example, the one or more Boolean functions apply a hash function to the security token to obtain a hashed token represented by a hashed library. In some implementations, the hashed library is a subset of the library.
[0012] In some implementations, the method further includes sequencing at least a portion of the nucleic acid molecules in the hashed library to obtain a sequencing readout. The sequencing readout is compared to a database or lookup table to determine the presence or absence of a matching sequence. The method may further include granting or denying access to the secured asset or location based on the determined presence or absence, respectively, of the matching sequence. For example, the sequencing may include any one of high-throughput sequencing, shotgun sequencing, or nanopore sequencing.
[0013] In some implementations, the method further includes applying an additional chemical operation to the hashed library to produce an output molecule if the hashed token matches the reference sequence, and determining the presence or absence of the output molecule by an assay. For example, the assay is one of polymerase chain reaction (PCR), real-time PCR, reverse transcription PCR (RT-PCR), fluorometry, and gel electrophoresis. The output molecule is a distinguishable nucleic acid molecule of the hashed library. In some implementations, the method further involves granting or denying access to the secured asset or location based on the presence of the output molecule. This implementation of chemically verifying the hashed token to produce the output molecule is advantageous in that it eliminates the need for sequencing the hashed library; rather, the hashed library is subjected to yet another chemical operation, which may be cheaper or faster than sequencing, to determine the authenticity of the security token.
[0014] In some implementations, the library includes a unique molecular barcode. In some implementations, the security token includes a randomly generated key. In some implementations, the security token is part of a two-factor authentication system. In some implementations, the library is collocated with the artifact, and the security token is unique to the artifact. For example, the artifact is a fluid. The fluid can be any one of oil, ink, compressed gas, or a drug. In some implementations, the method further involves measuring the concentration of the library in the fluid to determine the amount of dilution. As another example, the artifact is a living organism or a document. In some implementations, the library is contained in any one of a well, a droplet, a spot, a sealed container, a gel, a suspension, or a solid matrix. In some implementations, the library is lyophilized.
[0015] In some implementations, the library is generated by selecting a subset of nucleic acid molecules from a pool of nucleic acid molecules. In some implementations, the security token includes a plurality of symbols, each symbol represented by a distinct sequence of a nucleic acid molecule in the library. In some implementations, the library is randomly generated. In some implementations, the security token is represented by a library of nucleic acid molecules using an encoding scheme, and the security token is mapped to a plurality of symbols having one of two possible symbol values, where if a symbol has a first of the two possible symbol values, a symbol of the plurality of symbols is represented by the presence of a distinct nucleic acid molecule in the library, and if a symbol has a second of the two possible symbol values, a symbol is represented by the absence of the distinct nucleic acid molecule. In some implementations, the security token includes at least 1 kilobit of information. In some implementations, the security token is unique to a user.
[0016]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein merely illustrative implementations of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different implementations, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0017] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event that the publications and patents or patent applications incorporated by reference conflict with the present disclosure contained herein, the present specification is intended to supersede and / or take precedence over any such conflicting material.
[0018] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative implementations, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 illustrates a schematic overview of the process for encoding, writing, accessing, querying, reading, and decoding digital information stored in nucleic acid sequences according to an exemplary implementation.
[0020] [Figure 2A-2B] Figures 2A and 2B schematically illustrate examples of how objects or identifiers (e.g., nucleic acid molecules) can be used to encode digital data, referred to as "data at an address." Figure 2A illustrates combining a rank object (or address object) with a byte value object (or data object) to create an identifier, according to an exemplary implementation. Figure 2B illustrates implementing a data at address scheme in which the rank object and byte value object are themselves combinatorial concatenations of other objects, according to an exemplary implementation.
[0021] [Figure 3A-3B] 3A and 3B schematically illustrate examples of methods for encoding digital information using objects or identifiers (e.g., nucleic acid sequences). Fig. 3A illustrates encoding digital information using a rank object as an identifier, according to an exemplary implementation. Fig. 3B illustrates an implementation of an encoding method in which an address object is itself a combinatorial concatenation of other objects, according to an exemplary implementation.
[0022] [Figure 4]FIG. 4 shows a contour plot in logarithmic space of the relationship between the combinatorial space of possible identifiers (C, x-axis) and the average number of identifiers (k, y-axis) that can be constructed to store information of a given size (contours) according to an exemplary implementation.
[0023] [Figure 5] FIG. 5 illustrates a schematic overview of a method for writing information into a nucleic acid sequence (eg, deoxyribonucleic acid) according to an exemplary implementation.
[0024] [Figures 6A-6B] Figures 6A and 6B illustrate an example of a method, termed a "product scheme," for constructing an identifier (e.g., a nucleic acid molecule) by combinatorially assembling distinguishable components (e.g., nucleic acid sequences). Figure 6A illustrates the architecture of an identifier constructed using the product scheme, according to an exemplary implementation. Figure 6B illustrates an example of the combinatorial space of identifiers that can be constructed using the product scheme, according to an exemplary implementation.
[0025] [Figure 7A] 7A-7C are schematic diagrams illustrating an example of a method for accessing a portion of information stored in a nucleic acid sequence by accessing several specific identifiers from a larger number of identifiers. Figure 7A shows an example of a method for accessing identifiers containing specified components using polymerase chain reaction, affinity-tagged probes, and degradation targeting probes, according to an exemplary implementation. Figure 7B shows an example of a method for accessing identifiers containing multiple specified components using polymerase chain reaction to perform an "OR" or "AND" operation, according to an exemplary implementation. Figure 7C shows an example of a method for accessing identifiers containing multiple specified components using affinity tags to perform an "OR" or "AND" operation, according to an exemplary implementation. [Figure 7B]7A-7C are schematic diagrams illustrating an example of a method for accessing a portion of information stored in a nucleic acid sequence by accessing several specific identifiers from a larger number of identifiers. Figure 7A shows an example of a method for accessing identifiers containing specified components using polymerase chain reaction, affinity-tagged probes, and degradation targeting probes, according to an exemplary implementation. Figure 7B shows an example of a method for accessing identifiers containing multiple specified components using polymerase chain reaction to perform an "OR" or "AND" operation, according to an exemplary implementation. Figure 7C shows an example of a method for accessing identifiers containing multiple specified components using affinity tags to perform an "OR" or "AND" operation, according to an exemplary implementation. [Figure 7C] 7A-7C are schematic diagrams illustrating an example of a method for accessing a portion of information stored in a nucleic acid sequence by accessing several specific identifiers from a larger number of identifiers. Figure 7A shows an example of a method for accessing identifiers containing specified components using polymerase chain reaction, affinity-tagged probes, and degradation targeting probes, according to an exemplary implementation. Figure 7B shows an example of a method for accessing identifiers containing multiple specified components using polymerase chain reaction to perform an "OR" or "AND" operation, according to an exemplary implementation. Figure 7C shows an example of a method for accessing identifiers containing multiple specified components using affinity tags to perform an "OR" or "AND" operation, according to an exemplary implementation.
[0026] [Figure 8] FIG. 8 illustrates a computer system programmed or otherwise configured to perform the methods provided herein, according to an exemplary implementation.
[0027] [Figure 9]FIG. 9 shows an example of two source bitstreams and a universal identifier library prepared for computation using defined operations in an identifier pool according to an exemplary implementation.
[0028] [Figure 10] FIG. 10 shows the inputs to and results of three example logical operations performed on a pool of identifiers, illustrating how an identifier library can be used as a platform for in vitro computation according to an exemplary implementation.
[0029] [Figure 11] FIG. 11 illustrates an example method for generating entropy that can be used to create a random bit string, according to an exemplary implementation.
[0030] [Figures 12A-12C] 12A-12C show an example of a method for generating and storing entropy (random bit strings) according to an exemplary implementation.
[0031] [Figures 13A-13B] 13A-13B show an example of a method for organizing and accessing a random bit string using an input, according to an exemplary implementation.
[0032] [Figure 14] FIG. 14 illustrates an example of a method for securing and authenticating access to an artifact using a physical DNA key, according to an exemplary implementation.
[0033] [Figure 15] FIG. 15 shows a flowchart describing a method for preparing a nucleic acid library for authentication, according to an exemplary implementation.
[0034] [Figure 16]FIG. 16 shows a flowchart describing a method for tagging a fluid with a nucleic acid tag for tracking or authentication, according to an exemplary implementation.
[0035] [Figures 17A-17B] Figures 17A and 17B show examples of encoding, writing, and reading data encoded in a nucleic acid molecule. Figure 17A shows an example of encoding, writing, and reading 5,856 bits of data according to an exemplary implementation. Figure 17B shows an example of encoding, writing, and reading 62,824 bits of data according to an exemplary implementation. DETAILED DESCRIPTION OF THE INVENTION
[0036] Detailed Description While various implementations of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such implementations are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art that do not depart from the invention. It should be understood that various alternatives to the implementation of the invention described herein may be utilized.
[0037] The term "symbol," as used herein, generally refers to a representation of a unit of digital information. Digital information can be divided or converted into strings of symbols. In one example, a symbol can be a bit, and a bit can have a value of "0" or "1."
[0038] The terms "distinguishable" or "unique," as used herein, generally refer to an object that can be distinguished from other objects in a group. For example, a distinguishable or unique nucleic acid sequence can be a nucleic acid sequence that does not have the same sequence as any other nucleic acid sequence. A distinguishable or unique nucleic acid molecule can not have the same sequence as any other nucleic acid molecule. A distinguishable or unique nucleic acid sequence or molecule can also share regions of similarity with another nucleic acid sequence or molecule.
[0039] The term "component," as used herein, generally refers to a nucleic acid sequence. A component may be a distinct sequence. A component may be linked or assembled with one or more other components to produce another nucleic acid sequence or molecule.
[0040] The term "stratum," as used herein, generally refers to a group or pool of components. Each stratum may contain a set of distinguishable components such that the components in one stratum differ from the components in another stratum. Components from one or more stratums may be assembled to generate one or more identifiers.
[0041] The term "identifier," as used herein, generally refers to a nucleic acid molecule or sequence that represents the position and value of a bit string within a larger bit string. More generally, an identifier can refer to any object that represents or corresponds to a symbol in a symbol string. In some implementations, an identifier may include one or more concatenated components.
[0042] The term "combinatorial space," as used herein, generally refers to the set of all possible distinguishable identifiers that can be generated from a starting set of objects, such as components, and an allowable set of rules for how to modify those objects to form identifiers. The size of the combinatorial space of identifiers created by assembling or linking components can depend on the number of layers of components, the number of components in each layer, and the particular assembly method used to generate the identifiers.
[0043] The term "identifier rank," as used herein, generally refers to a relationship that defines the ordering of identifiers within a set.
[0044] The term "identifier library," as used herein, generally refers to a collection of identifiers that correspond to symbols in a string representing digital information. In some implementations, the absence of a given identifier in an identifier library can indicate a symbol value at a particular position. One or more identifier libraries can be combined in a pool, group, or set of identifiers. Each identifier library may also include a unique barcode that identifies the identifier library.
[0045] The term "nucleic acid," as used herein, generally refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or variants thereof. Nucleic acids can contain one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Nucleotides can include A, C, G, T, or U, or variants thereof. Nucleotides can include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits can be A, C, G, T, or U, or any other subunit that can be specific to one of more complementary A, C, G, T, or U, or that can be complementary to a purine (i.e., A or G, or variants thereof) or a pyrimidine (i.e., C, T, or U, or variants thereof). In some cases, nucleic acids can be single-stranded or double-stranded, and in some cases, nucleic acid molecules are circular.
[0046] The term "nucleic acid molecule" or "nucleic acid sequence," as used herein, generally refers to a polymeric form of nucleotides, or polynucleotides, which may be of various lengths and are either deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs thereof. The term "nucleic acid sequence" can refer to the alphabetic representation of a polynucleotide, or the term can be applied to the physical polynucleotide itself. This alphabetic representation can be entered into a database in a computer having a central processing unit and used to encode digital information, mapping the nucleic acid sequence or nucleic acid molecule to symbols or bits. A nucleic acid sequence or oligonucleotide may also contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0047] "Oligonucleotide," as used herein, generally refers to a single-stranded nucleic acid sequence, typically composed of a specific sequence of the four nucleotide bases adenine (A), cytosine (C), guanine (G), and thymine (T), or, if the polynucleotide is RNA, adenine (A), cytosine (C), guanine (G), and uracil (U).
[0048] Examples of modified nucleotides include diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylcuosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, Examples of suitable uracils include, but are not limited to, queuosin, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueuosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid(v), wybutoxocin, pseudouracil, queuosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine. Nucleic acid molecules may have modified base moieties (e.g., one or more atoms normally available to form hydrogen bonds with a complementary nucleotide and / or one or more atoms normally unavailable to form hydrogen bonds with a complementary nucleotide), modified sugar moieties, or modified phosphate backbones. Nucleic acid molecules may also contain amine-modifying groups such as aminoallyl-dUTP (aa-dUTP) and aminohexhylacrylamide-dCTP (aha-dCTP) to allow covalent attachment of amine-reactive moieties such as N-hydroxysuccinimide ester (NHS).
[0049] The term "primer," as used herein, generally refers to a nucleic acid strand that serves as the starting point for nucleic acid synthesis, such as in polymerase chain reaction (PCR). In one example, during replication of a DNA sample, an enzyme that catalyzes replication initiates replication at the 3' end of a primer bound to the DNA sample and copies the opposite strand. For more information regarding PCR, including details on primer design, see Chemical Methods Section D.
[0050] The term "polymerase" or "polymerase enzyme," as used herein, generally refers to any enzyme that can catalyze a polymerase reaction. Examples of polymerases include, but are not limited to, nucleic acid polymerases. Polymerases can be naturally occurring or synthetic. An example of a polymerase is Φ29 polymerase or its derivatives. In some cases, transcriptases or ligases (i.e., enzymes that catalyze the formation of bonds) are used in conjunction with or as a substitute for polymerases to construct new nucleic acid sequences. Examples of polymerases include DNA polymerases, RNA polymerases, thermostable polymerases, wild-type polymerases, modified polymerases, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, and Platinum Examples of polymerases that can be used with PCR include Taq polymerase, Tbr polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3' to 5' exonuclease activity, and variants, modified products, and derivatives thereof. For additional polymerases that can be used with PCR, and for details on how polymerase properties can affect PCR, see Chemical Methods Section D.
[0051] The term "species," as used herein, generally refers to one or more DNA molecules of the same sequence. When "species" is used in the plural sense, it can be assumed that all species in the plurality of species have distinguishable sequences, which can sometimes be clarified by writing "distinguishable species" instead of "species."
[0052] The terms "about" and "approximately" should be understood to mean within plus or minus 20% of the value that follows said term.
[0053] Digital information, such as computer data, in the form of binary code may include a sequence or string of symbols. Binary code may, for example, encode or represent text or computer processor instructions using a binary system having two binary symbols called bits, usually 0 and 1. Digital information may be represented in the form of a non-binary code, which may include a sequence of non-binary symbols. Each encoded symbol may be reassigned to a unique bit sequence (or "byte"), and the unique bit sequence or byte may be arranged into a byte sequence or byte stream. The bit value for a given bit may be one of two symbols (e.g., 0 or 1). A byte may contain a string of N bits, for a total of 2 N For example, a byte containing 8 bits can have a total of 2 8 Or it can give rise to 256 possible unique byte values, each of which can correspond to one of 256 possible distinct symbols, characters, or instructions that can be encoded in the byte. Raw data (e.g., text files and computer instructions) can be represented as a string of bytes or a byte stream. Zip files, or compressed data files containing raw data, can also be stored in a byte stream; these files can be stored in compressed form as a byte stream and then restored to raw data before being read by a computer.
[0054] The disclosed methods and systems can be used to encode computer data or information with multiple identifiers, each of which can represent one or more bits of primary information. In some examples, the disclosed methods and systems encode data or information using identifiers, each of which represents two bits of primary information.
[0055] Previous methods for encoding digital information into nucleic acids have relied on the base-by-base synthesis of nucleic acids, which can be expensive and time-consuming. Alternative methods can improve the efficiency and commercial viability of digital information storage by reducing the reliance on base-by-base nucleic acid synthesis to encode digital information, eliminating the need for de novo synthesis of distinct nucleic acid sequences for every new information storage requirement.
[0056] Rather than relying on base-by-base or de novo nucleic acid synthesis (e.g., phosphoramidite synthesis), the new methods can encode digital information (e.g., binary code) into multiple identifiers or nucleic acid sequences that contain combinatorial sequences of components. Thus, the new strategies can generate a first set of distinct nucleic acid sequences (or components) for a first request for information storage, and then reuse the same nucleic acid sequences (or components) for subsequent information storage requests. These approaches can significantly reduce the cost of DNA-based information storage by reducing the role of de novo synthesis of nucleic acid sequences in the process of encoding and writing information to DNA. Furthermore, unlike implementations of base-by-base synthesis (e.g., phosphoramidite chemistry-based or template-free polymerase-based nucleic acid extension), which may use cyclic delivery of each base to each extended nucleic acid, the new methods of writing information to DNA using identifier construction from components are highly parallelizable processes that do not necessarily use cyclic nucleic acid extension. Therefore, the new methods can increase the speed at which digital information is written to DNA compared to traditional methods.
[0057] A method for encoding information into a nucleic acid sequence is provided herein. The method for encoding information into a nucleic acid sequence may include (a) converting the information into a symbol string, (b) mapping the symbol string to a plurality of identifiers, and (c) constructing an identifier library including at least a subset of the plurality of identifiers. Each identifier of the plurality of identifiers may include one or more components. Each component of the one or more components may include a nucleic acid sequence. Each symbol at each position in the symbol string may correspond to a distinct identifier. Each identifier may correspond to each symbol at each position in the symbol string. Furthermore, one symbol at each position in the symbol string may correspond to the absence of an identifier. For example, each occurrence of a "0" in a string of binary symbols (e.g., bits) of "0" and "1" may correspond to the absence of an identifier.
[0058] In another aspect, the present disclosure provides a method for nucleic acid-based computer data storage. The method for nucleic acid-based computer data storage may include (a) receiving computer data, (b) synthesizing nucleic acid molecules including nucleic acid sequences that encode the computer data, and (c) storing the nucleic acid molecules having the nucleic acid sequences. The computer data may be encoded into at least a subset of the synthesized nucleic acid molecules, but may not be encoded into the sequence of each of the nucleic acid molecules.
[0059] In another aspect, the present disclosure provides a method for writing and storing information in a nucleic acid sequence. The method may include the steps of: (a) receiving or encoding a virtual identifier library representing the information; (b) physically constructing the identifier library; and (c) storing one or more physical copies of the identifier library in one or more separate locations. Each identifier in the identifier library may include one or more components. Each of the one or more components may include a nucleic acid sequence.
[0060] In another aspect, the present disclosure provides a method for nucleic acid-based computer data storage. The method for nucleic acid-based computer data storage may include: (a) receiving computer data; (b) synthesizing a nucleic acid molecule comprising at least one nucleic acid sequence encoding the computer data; and (c) storing the nucleic acid molecule comprising at least one nucleic acid sequence. The step of synthesizing the nucleic acid molecule may be a step in the absence of base-by-base nucleic acid synthesis.
[0061] In another aspect, the present disclosure provides a method for writing and storing information in a nucleic acid sequence. The method for writing and storing information in a nucleic acid sequence may include the steps of (a) receiving or encoding a virtual identifier library representing the information, (b) physically constructing the identifier library, and (c) storing one or more physical copies of the identifier library in one or more separate locations. Each identifier in the identifier library may include one or more components. Each of the one or more components may include a nucleic acid sequence.
[0062] FIG. 1 shows an overview of a process for encoding information into a nucleic acid sequence, writing information into a nucleic acid sequence, reading information written into a nucleic acid sequence, and decoding the read information. Digital information, or data, can be converted into one or more symbol sequences. In one example, the symbols are bits, and each bit can have a value of either "0" or "1". Each symbol can be mapped or encoded to an object (e.g., an identifier) that represents that symbol. Each symbol can be represented by a distinguishable identifier. The distinguishable identifier can be a nucleic acid molecule composed of components. The components can be nucleic acid sequences. Digital information can be written into a nucleic acid sequence by generating an identifier library corresponding to that information. The identifier library can be physically generated by physically constructing identifiers corresponding to each symbol of the digital information. All or any part of the digital information can be accessed simultaneously. In one example, a subset of identifiers is accessed from the identifier library. The subset of identifiers can be read by sequencing or identifying the identifiers. The identified identifiers can be associated with their corresponding symbols to decode the digital data.
[0063] A method of encoding and reading information using the approach of FIG. 1 can include, for example, receiving a bitstream, and mapping each 1-bit (a bit having a bit value of "1") in the bitstream to a distinguishable nucleic acid identifier using an identifier rank or nucleic acid index. Constructing a nucleic acid sample pool or identifier library that includes a copy of the identifier corresponding to the 1-bit value (and not including identifiers for 0-bit values). Reading the sample can include using molecular biology methods (e.g., sequencing, hybridization, PCR, etc.) to determine which identifiers in the identifier library are represented, and assigning the bit value of "1" to the bits corresponding to these identifiers and the bit value of "0" elsewhere (referring to the identifier rank again to identify the bits in the original bitstream corresponding to each identifier), thus decoding the information into the original bitstream in which the information was encoded.
[0064] Encoding a string of N distinct bits may use the same number of unique nucleic acid sequences as possible identifiers. This information encoding technique may use de novo synthesis of an identifier (e.g., a nucleic acid molecule) for each new item of information (string of N bits) to store. In another example, the cost of synthesizing identifiers (equal to or less than N in number) de novo for each new item of information to store can be reduced by a one-time de novo synthesis and subsequent maintenance of all possible identifiers, such that encoding a new item of information may involve mechanically selecting and mixing together pre-synthesized (or pre-made) identifiers to form an identifier library. In another example, either (1) the cost of de novo synthesis of up to N identifiers for each new item of information to store, or (2) the cost of maintaining and selecting from N possible identifiers for each new item of information to store, or any combination thereof, can be reduced by synthesizing nucleic acid sequences, maintaining that number (less than N, and in some cases much less than N), and then modifying these sequences by enzymatic reactions to generate up to N identifiers for each new item of information to store.
[0065] Identifiers can be rationally designed and selected to facilitate read, write, access, copy, and delete operations. Identifiers can be designed and selected to minimize write errors, mutations, degradation, and read errors. See Chemical Methods Section H for rational design of DNA sequences, including synthetic nucleic acid libraries (e.g., identifier libraries).
[0066] 2A and 2B schematically illustrate an example of a method, called "data-at-address," for encoding digital data in objects or identifiers (e.g., nucleic acid molecules). FIG. 2A illustrates encoding a bitstream into an identifier library, where each identifier is constructed by concatenating or assembling a single component that specifies an identifier rank and a single component that specifies a byte value. In general, the data-at-address method uses identifiers that modularly encode information by including two objects: a "byte value object" (or "data object"), which is one object that identifies a byte value, and a "rank object" (or "address object"), which is one object that identifies the identifier rank (or relative position of the byte in the original bitstream). FIG. 2B illustrates an example of a data-at-address method in which each rank object is combinatorially constructed from a set of components, and each byte value object can be combinatorially constructed from a set of components. This combinatorial construction of rank objects and byte value objects allows more information to be written to an identifier than if the object were created from only a single component (e.g., FIG. 2A).
[0067] 3A and 3B schematically illustrate another example of a method for encoding digital information in an object or identifier (e.g., a nucleic acid sequence). FIG. 3A illustrates encoding a bit stream into an identifier library, where the identifier is constructed from a single component that specifies the identifier rank. The presence of the identifier at a particular rank (or address) specifies a bit value of "1," while the absence of the identifier at a particular rank (or address) specifies a bit value of "0." This type of encoding uses identifiers that simply encode rank (the relative position of the bit in the original bit stream), and the presence or absence of these identifiers in the identifier library can be used to encode bit values of "1" or "0," respectively. Reading and decoding the information can involve identifying identifiers present in the identifier library, assigning bit values of "1" to their corresponding rank, and assigning bit values of "0" elsewhere. FIG. 3B illustrates an example of an encoding method in which each identifier can be combinatorially constructed from a set of components, with each possible combinatorial construction thus specifying a rank. Such combinatorial construction allows more information to be written into the identifier than if the identifier were created from only a single component (e.g., FIG. 3A). For example, a component set may include five distinct components. The five distinct components can be assembled to generate ten distinct identifiers, each containing two of the five components. The ten distinct identifiers may each have a rank (or address) corresponding to the position of a bit in the bitstream. An identifier library may include a subset of these ten possible identifiers that correspond to positions of bit value "1" and exclude a subset of these ten possible identifiers that correspond to positions of bit value "0" in a bitstream of length 10.
[0068] 4 shows a contour plot, in logarithmic space, of the relationship between the combinatorial space of possible identifiers (C, x-axis) and the average number of identifiers (k, y-axis) physically constructed to store information of a given original size in bits (D, contour) using the encoding method shown in FIGS. 3A and 3B. This plot assumes that primary information of size D is re-encoded into a string of C bits (C can be greater than D) in which several, i.e., k, bits have a bit value of "1." This plot further assumes that encoding of information into nucleic acids is done with the re-encoded bit string, and that identifiers are constructed for positions where the bit value is "1" and no identifiers are constructed for positions where the bit value is "0." According to these assumptions, the combinatorial space of possible identifiers has size C to identify every position in the re-encoded bit string, and the number of identifiers used to encode a bit string of size D is such that D=log2(Cchoosek), where Cchoosek can be a mathematical formula for the number of ways to choose k unordered outcomes from C possibilities. Thus, as the combinatorial space of possible identifiers increases beyond the size (in bits) of a given item of information, the number of physically constructed identifiers that can be used to store the given information decreases.
[0069] Figure 5 shows an overview of how information can be written to a nucleic acid sequence. Before writing, the information can be converted into a symbol string and encoded into multiple identifiers. Writing information can include initiating a reaction to generate a potential identifier. A reaction can be initiated by placing an input into a compartment. The input can include a nucleic acid, a component, a template, an enzyme, or a chemical reagent. A compartment can be a well, a tube, a location on a surface, a chamber in a microfluidic device, or a droplet in an emulsion. Multiple reactions can be initiated in multiple compartments. Reactions can proceed to generate identifiers by programmed temperature incubation or cycling. Reactions can be selectively or universally removed (e.g., deleted). Reactions can also be selectively or universally interrupted, consolidated, and purified to collect the identifiers into a single pool. Identifiers from multiple identifier libraries can be collected into the same pool. Individual identifiers can include a barcode or tag to identify which identifier library they belong to. Alternatively, or in addition, the barcode can include metadata about the encoded information. Supplemental nucleic acids or identifiers can also be included in the identifier pool along with the identifier library. The supplemental nucleic acids or identifiers may contain metadata for the encoded information or may serve to obfuscate or hide the encoded information.
[0070] The identifier rank (e.g., nucleic acid index) may include a method or key for determining the ordering of identifiers. The method may include a lookup table with all identifiers and their corresponding ranks. The method may also include a lookup table with the ranks of all components that make up the identifier and a function for determining the ordering of any identifier that includes a combination of these components. Such a method may be called lexicographic ordering and may be similar to the way words in a dictionary are ordered alphabetically. In a data encoding method in an address, the identifier rank (encoded by the identifier rank object) may be used to determine the position of a byte (encoded by the identifier byte value object) in the bit stream. In an alternative method, the identifier rank of an existing identifier (encoded by the entire identifier itself) may be used to determine the position of a "1" bit value in the bit stream.
[0071] The key can assign distinct bytes to unique subsets of identifiers (e.g., nucleic acid molecules) in a sample. For example, in a simple form, the key can assign each bit in the byte to a unique nucleic acid sequence that specifies the bit's position, and the presence or absence of that nucleic acid sequence in the sample can then assign a bit value of 1 or 0, respectively. Reading the encoded information from the nucleic acid sample can involve any number of molecular biology techniques, including sequencing, hybridization, or PCR. In some implementations, reading the encoded dataset can involve reconstructing a portion of the dataset or the entire encoded dataset from each nucleic acid sample. When the sequence can be read, the nucleic acid index can be used along with the presence or absence of unique nucleic acid sequences to decode the nucleic acid sample into a bitstream (e.g., individual bit strings, byte(s), bytes, or byte strings).
[0072] Identifiers can be constructed by combinatorially assembling component nucleic acid sequences. For example, information can be encoded using a set of nucleic acid molecules (e.g., identifiers) from a defined group of molecules (e.g., combinatorial space). Each possible identifier for a defined group of molecules can be an assembly of nucleic acid sequences (e.g., components) from a predefined set of components that can be separated into layers. Each individual identifier can be constructed by concatenating one component from every layer in a fixed order. For example, if there are M layers, each with n components, then the maximum C=n M You can construct up to 2 unique identifiers. C different information items or C bits can be encoded and stored. For example, storing a megabit of information requires 1×10 6 distinct identifiers, or size C = 1 x 10 6 The combinatorial space available is n=1×10. The identifier in this example can be assembled from various components organized in different ways. 3 Alternatively, an assembly can be made from M=2 pre-formed layers each containing n=1×10 2 An assembly can be created from M=3 layers, each containing components. As this example illustrates, it may be possible to encode the same amount of information using a greater number of layers, thereby reducing the total number of components. From a writing cost perspective, it may be advantageous to use a smaller number of total components.
[0073] The nucleic acid sequences (e.g., components) in each layer can include a unique (or distinguishable) sequence, or barcode, in the middle, a common hybridization region at one end, and another common hybridization region at the other end. The barcode can contain a sufficient number of nucleotides to uniquely identify every sequence in the layer. For example, there are typically four possible nucleotides at each base position in the barcode. Thus, a three-base barcode can have four possible nucleotides. 3=64 nucleic acid sequences can be uniquely identified. Barcodes can be designed to be randomly generated. Alternatively, barcodes can be designed to avoid sequences that may complicate the identifier construction chemistry or sequencing. In addition, barcodes can be designed so that each has a minimum Hamming distance from other barcodes, thereby reducing the likelihood that base-degrading mutations or reading errors may interfere with proper identification of the barcode.
[0074] The hybridization region at one end of a nucleic acid sequence (e.g., a component) can be different from layer to layer, but the hybridization region can be the same for each member within a layer. Adjacent layers have complementary hybridization regions on their components that allow them to interact with each other. For example, any component from layer X can be capable of binding to any component from layer Y because they can have complementary hybridization regions. The hybridization region at the opposite end can serve the same purpose as the hybridization region at the first end. For example, any component from layer Y can bind to any component of layer X at one end and to any component of layer Z at the opposite end.
[0075] Figures 6A and 6B show an example of a method, called a "product scheme," for constructing an identifier (e.g., a nucleic acid molecule) by combinatorially assembling distinguishable components (e.g., nucleic acid sequences) from each layer in a fixed order. Figure 6A shows the architecture of an identifier constructed using a product scheme. An identifier can be constructed by combining a single component from each layer in a fixed order. For M layers, each with N components, N MThere are 27 possible identifiers. Figure 6B shows an example of a combinatorial space of identifiers that can be constructed using a product scheme. In one example, a combinatorial space can be generated from three layers, each containing three distinguishable components. These components can be combined such that one component from each layer can be combined in a fixed order. The total combinatorial space for this assembly method can include 27 possible identifiers.
[0076] The identifier may be one of the many applications disclosed in U.S. Patent No. 10,650,312, filed December 21, 2017, entitled "NUCLEIC ACID-BASED DATA STORAGE," which describes encoding digital information in DNA; U.S. Patent Application No. 16 / 461,774, filed May 16, 2019, and published as U.S. Patent Application Publication No. 2019 / 0362814, entitled "SYSTEMS FOR NUCLEIC ACID-BASED DATA STORAGE," which describes encoding schemes for DNA-based data storage; and U.S. Patent Application No. 16 / 461,774, filed May 16, 2019, and published as U.S. Patent Application Publication No. 2019 / 0351673, entitled "PRINTER-FINISHER SYSTEM FOR DATA STORAGE IN A SYSTEM-BASED DATA STORAGE" which describes encoding schemes for DNA-based data storage. U.S. Patent Application No. 16 / 414,752 entitled "COMPOSITIONS AND METHODS FOR NUCLEIC ACID-BASED DATA STORAGE," filed May 16, 2019, and published as U.S. Patent Application Publication No. 2020 / 0193301, describing advanced assembly methods for DNA-based data storage; U.S. Patent Application No. 16 / 532,077 entitled "SYSTEMS AND METHODS FOR STORING AND READING NUCLEIC ACID-BASED DATA WITH ERROR PROTECTION," filed August 5, 2019, and published as U.S. Patent Application Publication No. 2020 / 0185057, describing data structures and error protection and correction for DNA encoding; ... U.S. Patent Application No. 16 / 872,129, "STRUCTURES AND OPERATIONS FOR SEARCHING, COMPUTING, AND INDEXING IN DNA-BASED DATA STORAGE" (describing data structures and operations for accessing, ranking, and searching);and U.S. Patent Application No. 17 / 012,909, filed September 4, 2020, entitled "CHEMICAL METHODS FOR NUCLEIC ACID-BASED DATA STORAGE," which describes chemical methods for encoded DNA assembly, each of which is hereby incorporated by reference in its entirety.
[0077] In some examples, all or part of the combinatorial space of possible identifiers can be constructed before encoding or writing the digital information, and the writing process can thus involve mechanically selecting and pooling identifiers (that encode the information) from a set that already exists. In other examples, the identifiers can be constructed at a point that may be after one or more steps of the data encoding or writing process have occurred (i.e., while the information is being written).
[0078] Barcodes can facilitate information indexing when the amount of digital information to be encoded exceeds the amount that can fit into a single pool. For example, by layering the approach disclosed in FIG. 3 by including tags with unique nucleic acid sequences encoded using nucleic acid indexes, longer bit strings and / or information containing multiple bytes can be encoded. An information cassette or identifier library can include nitrogenous bases or nucleic acid sequences containing unique nucleic acid sequences that provide position and bit value information, in addition to barcodes or tags that indicate the component(s) of the bit stream to which a given sequence corresponds. An information cassette can include one or more unique nucleic acid sequences and barcodes or tags. The barcodes or tags on the information cassette can provide a reference to the information cassette and any sequences contained in the information cassette. For example, the tags or barcodes on the information cassette can indicate which portion of the bit stream or which bit component of the bit stream the unique sequence encodes information (e.g., bit value and bit position information).
[0079] Barcodes can be used to encode information in bits into pools that exceed the size of the combinatorial space of possible identifiers. For example, a 10-bit sequence can be divided into two sets of bytes, each containing five bits. Each byte can be mapped to a set of five possible distinguishable identifiers. Initially, the identifiers generated for each byte may be the same, but they can be kept in separate pools; otherwise, a reader of the information may not be able to tell which byte a particular nucleic acid sequence belongs to. However, by barcoding or tagging each identifier with a label corresponding to the byte to which the encoded information applies (e.g., barcode 1 can be attached to a sequence in the nucleic acid pool to provide the first five bits, and barcode 2 can be attached to a sequence in the nucleic acid pool to provide the second five bits), the identifiers corresponding to these two bytes can then be combined into a pool (e.g., a "hyperpool" or one or more identifier libraries). Each identifier library of one or more combined identifier libraries can contain a distinguishable barcode that identifies a given identifier as belonging to the given identifier library.
[0080] A nucleic acid sample pool, hyperpool, identifier library, group of identifier libraries, or well containing a nucleic acid sample pool or hyperpool may contain a unique nucleic acid molecule corresponding to a bit of information (e.g., an identifier) and multiple complementary nucleic acid sequences. The complementary nucleic acid sequence may not correspond to encoded data (e.g., not correspond to a bit value). The complementary nucleic acid sample can mask or conceal the information stored in the sample pool. The complementary nucleic acid sequence may be derived from a biological source or synthetically generated. Complementary nucleic acid sequences derived from a biological source may include randomly fragmented nucleic acid sequences or rationally fragmented sequences. Biologically derived complementary nucleic acids can hide or obscure data-containing nucleic acids in a sample pool by providing natural genetic information together with synthetically encoded information, particularly when the synthetically encoded information (e.g., the combinatorial space of the identifier) is created to resemble natural genetic information (e.g., a fragmented genome). In one example, the identifier is derived from a biological source and the complementary nucleic acid is derived from a biological source. A sample pool may contain multiple sets of identifiers and complementary nucleic acid sequences. Each set of identifiers and complementary nucleic acid sequences may be derived from different organisms. In one example, the identifiers are derived from one or more organisms and the complementary nucleic acid sequences are derived from a single, different organism. The complementary nucleic acid sequences may be derived from one or more organisms, or the identifiers may be derived from a single organism that is different from the organism from which the complementary nucleic acids are derived. Both the identifiers and complementary nucleic acid sequences may be derived from multiple different organisms. A key may be used to distinguish between the identifiers and complementary nucleic acid sequences.
[0081] The supplemental nucleic acid sequence can store metadata about the written information. The metadata may include additional information for determining and / or authorizing the primary source and / or intended recipient of the primary information. The metadata may include additional information about the format of the primary information, the device and method used to encode and write the primary information, and the date and time the primary information was written to the identifier. The metadata may include additional information about the format of the primary information, the device and method used to encode and write the primary information, and the date and time the primary information was written to the nucleic acid sequence. The metadata may include additional information about modifications made to the primary information after the information was written to the nucleic acid sequence. The metadata may include annotations to the primary information or one or more references to external information. Alternatively, or in addition, the metadata may be stored in one or more barcodes or tags associated with the identifier.
[0082] The identifiers in the identifier pool can have the same, similar, or different lengths from each other. The complementary nucleic acid sequences can have a length that is less than the length of the identifiers, a length that is substantially equal to the length of the identifiers, or a length that is longer than the length of the identifiers. The complementary nucleic acid sequences can have an average length that is within 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, or more of the average length of the identifiers. In one example, the complementary nucleic acid sequences are the same or substantially the same length as the identifiers. The concentration of the complementary nucleic acid sequences can be less than, substantially equal to, or higher than the concentration of the identifiers in the identifier library. The concentration of the complementary nucleic acid can be about 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1×10 of the concentration of the identifiers. 4 %, 1×10 5 %, 1×10 6 %, 1×10 7 %, 1×10 8% or less, or about 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1×10 concentration of the identifier. 4 %, 1×10 5 %, 1×10 6 %, 1×10 7 %, 1×10 8 The concentration of the supplemental nucleic acid may be equal to or less than about 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1×10 of the concentration of the identifier. 4 %, 1×10 5 %, 1×10 6 %, 1×10 7 %, 1×10 8 % or more, or about 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1×10 4 %, 1×10 5 %, 1×10 6 %, 1×10 7 %, 1×10 8 % or more. Higher concentrations may be beneficial for obfuscation or data hiding. In one example, the concentration of the capture nucleic acid sequences is substantially higher than the concentration of the identifiers in the identifier pool (e.g., 1 x 10 8 %expensive).
[0083] PCR-based methods can be used to access and copy data from identifiers or nucleic acid sample pools. Common primer binding sites adjacent to identifiers within a pool or hyperpool can be used to easily copy information-containing nucleic acids. Alternatively, other nucleic acid amplification techniques, such as isothermal amplification, can be used to easily copy data from a sample pool or hyperpool (e.g., an identifier library). See Chemical Methods Section D for nucleic acid amplification. In examples where a sample contains a hyperpool, a primer that binds in the forward direction to a specific barcode on one edge of the identifier can be used with another primer that binds in the reverse direction to a common sequence on the opposite edge of the identifier to access and obtain a specific subset of information (e.g., all nucleic acids associated with a particular barcode). Various readout methods can be used to extract information from the encoded nucleic acids, for example, microarrays (or any type of fluorescent hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can further be used to read out the encoded sequences and, by extension, the digitally encoded data.
[0084] Access to information stored in nucleic acid molecules (e.g., identifiers) can be achieved by selectively removing a portion of non-targeted identifiers from an identifier library or pool of identifiers, or, for example, by selectively removing all identifiers in an identifier library from a pool of multiple identifier libraries. As used herein, "access" and "query" can be used interchangeably. Access to data can also be achieved by selectively capturing targeted identifiers from an identifier library or pool of identifiers. Targeted identifiers may correspond to data of interest within a longer information item. The pool of identifiers may also include supplemental nucleic acid molecules. Supplemental nucleic acid molecules may contain metadata about the encoded information and may be used to conceal or mask identifiers corresponding to the information. Supplemental nucleic acid molecules may or may not be extracted during access to targeted identifiers. Figures 7A-7C schematically outline an example method for accessing a portion of information stored in a nucleic acid sequence by accessing certain specific identifiers from a larger number of identifiers. Figure 7A shows an example method for accessing identifiers containing specified components using polymerase chain reaction, affinity-tagged probes, and degradation-targeted probes. For PCR-based access, the pool of identifiers (e.g., an identifier library) can include identifiers with a common sequence at each end, a variable sequence at each end, or either a common sequence or a variable sequence at each end. The common sequence or variable sequence can be a primer binding site. One or more primers can bind to the common or variable regions at the edges of the identifiers. Primed identifiers can be amplified by PCR. Amplified identifiers can significantly outnumber unamplified identifiers. Amplified identifiers can be identified during reading. Identifiers from an identifier library can contain a sequence at one or both of their ends that is distinguishable from that library, thus allowing selective access to a single library from a pool or group of more than one identifier library.
[0085] In the case of affinity tag-based access, a process sometimes referred to as nucleic acid capture, the components constituting the identifiers in the pool may share complementarity with one or more probes. The one or more probes may bind or hybridize to the identifiers to be accessed. The probes may include affinity tags. The affinity tags may bind to beads to form complexes comprising the beads, at least one probe, and at least one identifier. The beads may be magnetic, and together with a magnet, the beads may collect and isolate the identifiers to be accessed. Prior to reading, the identifiers may be removed from the beads under denaturing conditions. Alternatively, or in addition, the beads may collect non-targeted identifiers and sequester them from the remainder of the pool, which may be washed, transferred to a separate container, and read. The affinity tags may be bound to a column. The identifiers to be accessed may be bound to the capture column. The column-bound identifiers may then be eluted or denatured from the column prior to reading. Alternatively, non-targeted identifiers may be selectively targeted to a column, while targeted identifiers may flow through the column. Accessing the targeted identifiers may involve applying one or more probes to the pool of identifiers simultaneously, or may involve applying one or more probes to the pool of identifiers sequentially.
[0086] In the case of degradation-based access, the components constituting the identifiers in the pool may share complementarity with one or more degradation-targeting probes. The probes can bind or hybridize to distinguishable components of the identifiers. The probes can be targeted by degradative enzymes such as endonucleases. In one example, one or more identifier libraries can be combined. A set of probes can hybridize with one of the identifier libraries. The set of probes can include RNA, which can guide a Cas9 enzyme. The Cas9 enzyme can be introduced into one or more identifier libraries. Identifiers hybridized with the probes can be degraded by the Cas9 enzyme. Identifiers to be accessed may not be degraded by the degradative enzyme. In another example, the identifiers can be single-stranded, and the identifier library can be combined with a single-strand-specific endonuclease, such as S1 nuclease, that selectively degrades identifiers that are not to be accessed. Identifiers to be accessed can be hybridized with a complementary set of identifiers to protect them from degradation by the single-strand-specific endonuclease. The identifier to be accessed can be separated from the degradation products by size selection, such as size-selective chromatography (e.g., agarose gel electrophoresis). Alternatively, or in addition, the undegraded identifier can be selectively amplified (e.g., using PCR), so that the degradation products are not amplified. The undegraded identifier can be amplified using primers that hybridize to each end of the undegraded identifier, and therefore do not hybridize to each end of the degraded or cleaved identifier.
[0087] 7B shows an example of a method for using polymerase chain reaction to perform an "OR" or "AND" operation to access identifiers containing multiple components. In one example, if two forward primers bind to distinct sets of identifiers on the left end, "OR" amplification of the intersection of these sets of identifiers can be achieved by using the two forward primers together in a multiplex PCR reaction with a reverse primer that binds to all of the identifiers on the right end. In another example, if one forward primer binds to a set of identifiers on the left end and one reverse primer binds to a set of identifiers on the right end, "AND" amplification of the intersection of these two sets of identifiers can be achieved by using the forward and reverse primers together as a primer pair in a PCR reaction.
[0088] 7C shows an example of a method for using affinity tags to perform an "OR" or "AND" operation to access identifiers containing multiple components. In one example, if affinity probe "P1" captures all identifiers with component "C1" and another affinity probe "P2" captures all identifiers with component "C2," then P1 and P2 can be used simultaneously to capture the set of all identifiers with C1 or C2 (corresponding to an "OR" operation). In another example, using the same components and probes, P1 and P2 can be used sequentially to capture the set of all identifiers with C1 and C2 (corresponding to an "AND" operation).
[0089] In another aspect, the present disclosure provides a method for reading information encoded in a nucleic acid sequence. The method for reading information encoded in a nucleic acid sequence may include the steps of: (a) providing an identifier library; (b) identifying identifiers present in the identifier library; (c) generating a symbol string from the identifiers present in the identifier library; and (d) compiling information from the symbol string. The identifier library may include a subset of a plurality of identifiers from the combinatorial space. Each individual identifier of the subset of identifiers may correspond to an individual symbol in the symbol string. The identifier may include one or more components. The components may include a nucleic acid sequence.
[0090] Information can be written to one or more identifier libraries as described elsewhere herein. Identifiers can be constructed using any of the methods described elsewhere herein. Stored data can be copied and accessed using any of the methods described elsewhere herein.
[0091] The identifier may include information about the location of the encoded symbol, the value of the encoded symbol, or both the location and value of the encoded symbol. The identifier may include information about the location of the encoded symbol, and the presence or absence of the identifier in the identifier library can indicate the value of the symbol. The presence of the identifier in the identifier library can indicate a first symbol value (e.g., a first bit value) in the binary string, and the absence of the identifier in the identifier library can indicate a second symbol value (e.g., a second bit value) in the binary string. Basing the bit value on the presence or absence of the identifier in the identifier library in a binary system can reduce the number of identifiers to be assembled and therefore reduce write time. In one example, the presence of the identifier can indicate a bit value of "1" at the mapped location, and the absence of the identifier can indicate a bit value of "0" at the mapped location.
[0092] Generating a symbol (e.g., a bit value) for a piece of information can include identifying the presence or absence of an identifier to which the symbol (e.g., a bit) can be mapped or encoded. Determining the presence or absence of the identifier can include sequencing the identifier or using a hybridization array to detect the presence of the identifier. In one example, decoding and reading the encoded sequence can be performed using a sequencing platform. An example of a sequencing platform is described in U.S. Patent Application No. 16 / 532,077, entitled "SYSTEMS AND METHODS FOR STORING AND READING NUCLEIC ACID-BASED DATA WITH ERROR PROTECTION," filed August 5, 2019, and published as U.S. Patent Application Publication No. 2020 / 0185057, which is incorporated herein by reference in its entirety.
[0093] In one example, decoding of nucleic acid-encoded data can be accomplished by base-by-base sequencing of a nucleic acid strand, such as Illumina® Sequencing, or by utilizing sequencing techniques that indicate the presence or absence of a particular nucleic acid sequence, such as capillary electrophoretic fragmentation analysis. Sequencing may also utilize the use of reversible terminators. Sequencing may also utilize the use of natural or unnatural (e.g., engineered) nucleotides or nucleotide analogs. Alternatively, or in addition, decoding of nucleic acid sequences can be performed using a variety of analytical techniques, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal. A variety of sequencing techniques can be used, including but not limited to polymerase chain reaction (PCR), digital PCR, Sanger sequencing, high-throughput sequencing, sequencing-by-synthesis, single molecule sequencing, sequencing-by-ligation, RNA-Seq (Illumina), next-generation sequencing, digital gene expression (Helicos), clonal single microarray (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, or massively parallel sequencing.
[0094] Various readout methods can be used to extract information from the encoded nucleic acids. In one example, microarrays (or any type of fluorescent hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can also be used to read out the encoded sequence, and by extension, the digitally encoded data.
[0095] The identifier library may further include complementary nucleic acid sequences that provide metadata about the information, that conceal or mask the information, or that both provide metadata and mask the information. The complementary nucleic acids can be identified simultaneously with the identification of the identifier. Alternatively, the complementary nucleic acids can be identified before or after the identifier is identified. In one example, the complementary nucleic acid sequences are not identified during reading of the encoded information. The complementary nucleic acid sequences may be indistinguishable from the identifier. An identifier index or key can be used to differentiate complementary nucleic acid molecules from the identifier.
[0096] Data encoding and decoding efficiency can be improved by re-encoding input bit strings to allow for the use of fewer nucleic acid molecules. For example, if an encoding method receives an input string with a high occurrence of "111" subsequences that can map to three nucleic acid molecules (e.g., identifiers), it can be re-encoded to a "000" subsequence that can map to an empty set of nucleic acid molecules. Alternative input subsequences of "000" can also be re-encoded to "111." This re-encoding method can reduce the total amount of nucleic acid molecules used to encode the data because the number of "1"s in the dataset can be reduced. In this example, the total size of the dataset can be increased to accommodate a codebook that specifies new mapping instructions. An alternative method for improving encoding and decoding efficiency can be to re-encode input strings to shorten their variable length. For example, "111" can be re-encoded to "00," which can reduce the size of the dataset and the number of "1"s in the dataset.
[0097] By specifically designing identifiers to facilitate detection, the speed and efficiency of decoding nucleic acid-encoded data can be controlled (e.g., increased). For example, a nucleic acid sequence (e.g., an identifier) designed for easy detection can include a nucleic acid sequence containing a majority of nucleotides that are easier to call and detect based on their optical, electrochemical, chemical, or physical properties. Engineered nucleic acid sequences can be single-stranded or double-stranded. Engineered nucleic acid sequences can also contain synthetic or non-natural nucleotides that improve the detectable properties of the nucleic acid sequence. Engineered nucleic acid sequences can contain all natural nucleotides, all synthetic or non-natural nucleotides, or a combination of natural, synthetic, and non-natural nucleotides. Synthetic nucleotides can include nucleotide analogs, such as peptide nucleic acids, locked nucleic acids, glycol nucleic acids, and threose nucleic acids. Non-natural nucleotides can include dNaM, an artificial nucleoside containing a 3-methoxy-2-naphthyl group, and d5SICS, an artificial nucleoside containing a 6-methylisoquinol-1-thion-2-yl group. The engineered nucleic acid sequences may be designed for a single enhanced property, such as an enhanced optical property, or the engineered nucleic acid sequences may be designed with multiple enhanced properties, such as enhanced optical and electrochemical properties or enhanced optical and chemical properties.
[0098] Engineered nucleic acid sequences may contain reactive natural, synthetic, and non-natural nucleotides that do not enhance the optical, electrochemical, chemical, or physical properties of the nucleic acid sequence. The reactive components of the nucleic acid sequence may allow for the addition of chemical moieties that confer enhanced properties to the nucleic acid sequence. Each nucleic acid sequence may contain a single chemical moiety or multiple chemical moieties. Examples of chemical moieties include, but are not limited to, fluorescent moieties, chemiluminescent moieties, acidic or basic moieties, hydrophobic or hydrophilic moieties, and moieties that alter the oxidation state or reactivity of the nucleic acid sequence.
[0099] Sequencing platforms can be specifically designed for decoding and reading information encoded in nucleic acid sequences. Sequencing platforms can be dedicated to sequencing single-stranded or double-stranded nucleic acid molecules. Sequencing platforms can decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of the entire nucleic acid sequence (e.g., component) incorporated into a nucleic acid molecule (e.g., an identifier). Sequencing platforms can include the use of promiscuous reagents, extended read lengths, and the addition of detectable chemical moieties to detect specific nucleic acid sequences. The use of more promiscuous reagents during sequencing can increase read efficiency by enabling faster base calling, thereby reducing sequencing time. The use of extended read lengths can allow longer sequences of encoded nucleic acids to be decoded per read. The addition of detectable chemical moiety tags can allow the presence or absence of a nucleic acid sequence to be detected by the presence or absence of the chemical moiety. For example, each nucleic acid sequence encoding a bit of information can be tagged with a chemical moiety that generates a unique optical, electrochemical, or chemical signal. The presence or absence of that unique optical, electrochemical, or chemical signal can indicate a "0" or "1" bit value. A nucleic acid sequence may contain a single chemical moiety, or it may contain multiple chemical moieties. A chemical moiety can be added to a nucleic acid sequence prior to using the nucleic acid sequence to encode data. Alternatively, or in addition, a chemical moiety can be added to a nucleic acid sequence after encoding the data but before decoding the data. A chemical moiety tag can be added directly to a nucleic acid sequence, or the nucleic acid sequence can contain a synthetic or non-natural nucleotide anchor to which a chemical moiety tag can be added.
[0100] A unique code can be applied to minimize or detect encoding and decoding errors. Encoding and decoding errors can occur due to false negatives (nucleic acid molecules or identifiers not included in random sampling). One example of an error detection code can be a checksum sequence that counts the number of identifiers in a consecutive set of possible identifiers included in an identifier library. During reading of the identifier library, the checksum can indicate the expected number of identifiers obtained from that consecutive set, and identifiers can continue to be sampled for reading until the expected number is met. In some implementations, a checksum sequence can be included for every R consecutive sets of identifiers, where R can be equal to or greater than 1, 2, 5, 10, 50, 100, 200, 500, or 1000 in size, or less than 1000, 500, 200, 100, 50, 10, 5, or 2. The smaller the value of R, the better the error detection. In some implementations, the checksum can be a complementary nucleic acid sequence. For example, a set containing seven nucleic acid sequences (e.g., components) can be divided into two groups: nucleic acid sequences for constructing identifiers in a product scheme (components X1-X3 in layer X and Y1-Y3 in layer Y) and nucleic acid sequences for complementary checksums (X4-X7 and Y4-Y7). Checksum sequences X4-X7 can indicate whether 0, 1, 2, or 3 sequences from layer X assemble with each member of layer Y. Alternatively, checksum sequences Y4-Y7 can indicate whether 0, 1, 2, or 3 sequences from layer Y assemble with each member of layer X. In this example, the original identifier library with identifiers {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3} can be supplemented to include checksums to result in the following pool: {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3, X1Y6, X2Y7, X3Y4, X6Y1, X5Y2, X6Y3}. Checksum sequences can also be used for error correction. For example, the absence of X1Y1 in the above dataset and the presence of X1Y6 and X6Y1 can allow for the inference that the X1Y1 nucleic acid molecule is missing from the dataset.The checksum sequence can indicate whether the identifier is missing from the sampled or accessed portion of the identifier library. If the checksum sequence is missing, an accessing method such as PCR or affinity-tagged probe hybridization can amplify and / or isolate it. In some implementations, the checksum may not be a complementary nucleic acid sequence. In that case, the checksum can be directly encoded into the information so that it is represented by the identifier.
[0101] The noise of data encoding and decoding can be reduced by constructing identifiers as palindromes, for example, by using palindromic pairs of components instead of single components in a product scheme. Pairs of components from different layers can then be assembled together in a palindromic manner (e.g., YXY instead of XY for components X and Y). This palindromic method can be extended to a larger number of layers (e.g., ZYXYZ instead of XYZ), and this palindromic method can enable the detection of erroneous cross-reactions between identifiers.
[0102] Adding excess (e.g., large excess) complementary nucleic acid sequences to the identifier can hinder recovery of the encoded identifier by sequencing. Prior to decoding the information, the identifier can be enriched by complementary nucleic acid sequences. For example, the identifier can be enriched by a nucleic acid amplification reaction using primers specific to the identifier end. Thus, only entities possessing the identifier-specific primer or the sequence of the identifier-specific primer will be able to enrich the encoded identifier for recovery by sequencing. Alternatively, or in addition, sequencing using specific primers (e.g., sequencing-by-synthesis) can decode the information without enriching the sample pool. In both decoding methods, it can be difficult to enrich or decode the information without the decoding key or knowing something about the composition of the identifier. Alternative access methods, such as the use of affinity tag-based probes, can also be utilized.
[0103] Systems for encoding digital information into nucleic acids (e.g., DNA) can include systems, methods, and devices for converting files and data (e.g., raw data, compressed zip files, integer data, and other forms of data) into bytes and encoding the bytes into nucleic acids, typically segments or sequences of DNA, or combinations thereof.
[0104] A non-limiting implementation of a method for using the system for encoding digital data can include receiving digital information in the form of a byte stream, parsing the byte stream into individual bytes, mapping bit positions within the bytes using a nucleic acid index (or identifier rank), and encoding sequences corresponding to either a bit value of 1 or a bit value of 0 into an identifier. Obtaining the digital data includes sequencing a nucleic acid sample or pool containing sequences of nucleic acids (e.g., identifiers) mapped to one or more bits, referencing the identifier rank to determine whether the identifier is present in the nucleic acid pool, and decoding the position and bit value information for each sequence into bytes containing the sequence of digital information.
[0105] A system for encoding, writing, copying, accessing, reading, and decoding information encoded or written to nucleic acid molecules may be a single integrated unit or multiple units configured to perform one or more of the above-mentioned operations. A system for encoding and writing information to nucleic acid molecules (e.g., identifiers) may include a device and one or more computer processors. The one or more computer processors may be programmed to parse the information into a symbol string (e.g., a string of bits). The computer processor may generate a rank of the identifiers. The computer processor may categorize the symbols into two or more categories. One category may include symbols represented by the presence of a corresponding identifier in the identifier library, and the other category may include symbols represented by the absence of a corresponding identifier in the identifier library. The computer processor may direct the device to assemble an identifier corresponding to the symbol represented by the presence of the identifier in the identifier library. A suitable system is described in U.S. Patent Application Serial No. 16 / 414,752, entitled "PRINTER-FINISHER SYSTEM FOR DATA STORAGE IN DNA," filed May 16, 2019, and published as U.S. Patent Application Publication No. 2019 / 0351673.
[0106] The device may include multiple regions, sections, or partitions. Reagents and components for assembling the identifiers can be stored in one or more regions, sections, or partitions of the device. Layers can be stored in separate regions of a section of the device. A layer can contain one or more unique components. Components within one layer can be unique and do not overlap with components in another layer. A region or section can include a container, and a partition can include a well. Each layer can be stored in a separate container or partition. Each reagent or nucleic acid sequence can be stored in a separate container or partition. Alternatively, or in addition, reagents can be combined to form a master mix for identifier construction. The device can transfer reagents, components, and templates from one section of the device to another section for assembly. The device can provide conditions for completing the assembly reaction. For example, the device can provide heating, agitation, and detection of reaction progress. The constructed identifiers can be directed to one or more subsequent reactions to add barcodes, common sequences, variable sequences, or tags to one or more ends of the identifiers. The identifiers can then be directed to regions or partitions to generate an identifier library. One or more identifier libraries can be stored in each region, section, or individual partition of the device. The device can transfer fluids (e.g., reagents, components, templates) using pressure, vacuum, or suction.
[0107] The identifier library may be stored on the device, transferred to a separate database, or transferred to a composition or container suitable for tagging / tracking artifacts. The database may contain one or more identifier libraries. The database may provide conditions for long-term storage of the identifier library (e.g., conditions to reduce identifier degradation). The identifier library may be stored in powder, liquid, or solid form. Aqueous solutions of identifiers may be lyophilized for more stable storage. The database may provide UV photoprotection, reduced temperature (e.g., refrigeration or freezing), and protection from degradative chemicals and enzymes. The identifier library may be lyophilized or frozen before being transferred to the database or functionalized with the artifact. The identifier library may contain ethylenediaminetetraacetic acid (EDTA) to inactivate nucleases and / or buffers to maintain the stability of the nucleic acid molecules.
[0108] The database may be coupled to, include, or be separate from a device that writes, copies, accesses, or reads information to the identifiers. A portion of the identifier library can be removed from the database before copying, accessing, or reading. The device that copies information from the database can be the same device or a different device from the device that writes the information. The device that copies the information can extract an aliquot of the identifier library from the device and combine the aliquot with reagents and components to amplify some or all of the identifier library. The device can control the temperature, pressure, and agitation of the amplification reaction. The device can include partitions, and one or more amplification reactions can occur in the partition containing the identifier library. The device can simultaneously copy more than one pool of identifiers.
[0109] The accessed data can be read in the same device, or the accessed data can be transferred to another device. The reading device can include a detection unit for detecting and identifying the identifier. The detection unit can be part of a sequencer, hybridization array, or other unit for identifying the presence or absence of an identifier. The sequencing platform can be specifically designed for decoding and reading information encoded in a nucleic acid sequence. The sequencing platform can be dedicated to sequencing single-stranded or double-stranded nucleic acid molecules. The sequencing platform can decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of an entire nucleic acid sequence (e.g., component) incorporated within the nucleic acid molecule (e.g., identifier). Alternatively, the sequencing platform can be a system such as Illumina® Sequencing or capillary electrophoretic fragmentation analysis. Alternatively, or in addition, decoding of the nucleic acid sequence can be performed using various analysis techniques implemented by the device, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal.
[0110] Storing information in nucleic acid molecules can have various applications, including, but not limited to, long-term information storage, confidential information storage, one-time access code storage, and medical information storage. In one example, a person's medical information (e.g., medical history and medical records) can be stored in a nucleic acid molecule and kept by the person. The information can be stored outside the body (e.g., in a wearable device) or inside the body (e.g., in a subcutaneous capsule). When a patient is brought to a clinic or hospital, a sample can be obtained from the device or capsule, and the information can be decoded using a nucleic acid sequencer. Storing personal medical records in nucleic acid molecules can provide an alternative to computer- and cloud-based storage systems. Storing personal medical records in nucleic acid molecules can reduce the incidence or prevalence of medical record hacking. The nucleic acid molecules used in capsule-based medical record storage can be derived from human genome sequences. Using human genome sequences can reduce the immunogenicity of the nucleic acid sequences in the event that the capsule is damaged and leaks.
[0111] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 8 shows a computer system 801 programmed or otherwise configured to encode digital information into a nucleic acid sequence and / or read (e.g., decode) information derived from a nucleic acid sequence. The computer system 801 is capable of adjusting various aspects of the encoding and decoding procedures of the present disclosure, such as, for example, bit value and bit position information for a given bit or byte from an encoded bitstream or bytestream.
[0112] Computer system 801 includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") 805, which may be a single-core processor or a multi-core processor, or multiple processors for parallel processing. Computer system 801 also includes memory or memory locations 810 (e.g., random access memory, read-only memory, flash memory), electronic storage 815 (e.g., a hard disk), a communication interface 820 (e.g., a network adapter) for communicating with one or more other systems, and peripherals 825, such as cache, other memory, data storage, and / or electronic display adapters. Memory 810, storage 815, interface 820, and peripherals 825 communicate with CPU 805 through a communication bus (solid lines), such as a motherboard. Storage 815 may be a data storage unit (or data repository) for storing data. Computer system 801 may be operably coupled to a computer network ("network") 830 utilizing communication interface 820. Network 830 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Network 830 may, in some cases, be a telecommunications and / or data network. Network 830 may include one or more computer servers, thereby enabling distributed computing such as cloud computing. Network 830 may, in some cases, utilize computer system 801 to implement a peer-to-peer network, thereby enabling devices coupled to computer system 801 to act as clients or servers.
[0113] The CPU 805 is capable of executing sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 810. The instructions may be directed to the CPU 805, which may then cause the CPU 805 to be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 805 may include fetch, decode, execute, and writeback.
[0114] The CPU 805 may be part of a circuit, such as an integrated circuit. One or more other components of the system 801 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0115] The storage device 815 may store files, such as drivers, libraries, and saved programs. The storage device 815 may store user data, such as user preferences and user programs. The computer system 801 may, in some cases, include one or more additional data storage units that are external to the computer system 801, such as located on a remote server that communicates with the computer system 801 over an intranet or the Internet.
[0116] Computer system 801 is capable of communicating with one or more remote computer systems over network 830. For example, computer system 801 is capable of communicating with a user's remote computer system or other devices and / or mechanisms (e.g., a sequencer or other system for chemically determining the order of nitrogenous bases in a nucleic acid sequence) that the user can use in the process of analyzing data encoded or decoded in a nucleic acid sequence. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android®-enabled devices, Blackberry®), or personal digital assistants. A user can access computer system 801 via network 830.
[0117] The methods described herein can be implemented by machine-executable (e.g., computer processor) code stored in an electronic storage location of the computer system 801, such as, for example, memory 810 or electronic storage 815. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 805. In some cases, the code can be retrieved from storage 815 and stored in memory 810 for immediate access by the processor 805. In some situations, the electronic storage 815 can be omitted, and the machine-executable instructions can be stored in memory 810. The computer system 801 can be operatively coupled to any one of a sequencing machine, a barcode scanner, a retinal scanner, a fingerprint scanner, a keypad entry device, a swabbing device, and an automated liquid handling unit configured to perform any of the chemical methods and operations described herein. The computer system 801 can be configured to lock and unlock physical access to secured locations or deposits.
[0118] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or it may be compiled at run time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or as-compiled fashion.
[0119] Aspects of the systems and methods presented herein, such as computer system 801, can be embodied in programming. Various aspects of the technology can be considered "products" or "articles of manufacture," generally in the form of machine- (or processor-) executable code and / or associated data carried or embodied on some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of a computer's tangible memory, processor, etc., or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage at any time for software programming. All or portions of the software can be communicated from time to time over the Internet or various other telecommunications networks. Such communication, for example, allows the software to be loaded from one computer or processor to another, e.g., from a management server or host computer to an application server's computer platform. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices through wired and optical landline networks and through various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media bearing software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0120] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as storage devices in any computer(s), such as those that can be used to implement databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire, including the electrical wires that comprise the busbars within a computer system, and fiber optics. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tape, any other physical storage media with patterns of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links which transport such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0121] The computer system 801 may include or communicate with an electronic display 835 including a user interface (UI) 840 to provide sequence output data, including, for example, chromatographs, sequences, and bits, bytes, or bitstreams encoded or read by a machine or computer system encoding or decoding nucleic acids, raw data, files, and compressed or uncompressed zip files, encoded or read by a machine or computer system that is encoding or decoding the nucleic acid, raw data, files, and compressed or uncompressed zip files. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces. The methods and systems of the present disclosure may be implemented via one or more algorithms. The algorithms may be implemented via software when executed by the central processing unit 805. Prior to encoding the digital information, the algorithms may be used, for example, with the DNA index and the raw data or the data compressed or uncompressed into zip files, to determine a customized method for coding the digital information into the raw data or the data compressed or uncompressed into zip files.
[0122] The chemical methods involved in the systems and methods described herein are described in U.S. patent application Ser. No. 16 / 414,758, entitled "COMPOSITIONS AND METHODS FOR NUCLEIC ACID-BASED DATA STORAGE," filed May 16, 2019, and published as U.S. Patent Application Publication No. 2020 / 0193301; and U.S. patent application Ser. No. 17 / 012,909, entitled "CHEMICAL METHODS FOR NUCLEIC ACID-BASED DATA STORAGE," filed September 4, 2020, each of which is hereby incorporated by reference in its entirety.
[0123] Ligation can be used to attach sequencing adapters to a library of nucleic acids. For example, ligation can be performed using a common cohesive end or staple at the end of each member of the nucleic acid library. Sequencing adapters can be ligated asymmetrically when the cohesive end or staple at one end of the nucleic acid is distinct from that at the other end. For example, a forward sequencing adapter can be ligated to one end of a member of the nucleic acid library, and a reverse sequencing adapter can be ligated to the other end of the member of the nucleic acid library. Alternatively, blunt-ended ligation can be used to attach adapters to a library of blunt-ended double-stranded nucleic acids. Forked adapters can be used to asymmetrically attach adapters to a nucleic acid library with either blunt or cohesive ends that are equivalent at each end (e.g., A-tails, etc.).
[0124] Nucleic acid amplification can be carried out using polymerase chain reaction, or PCR. In PCR, a starting pool of nucleic acids (referred to as a template pool or template) can be combined with polymerase, primers (short nucleic acid probes), nucleotide triphosphates (e.g., dATP, dTTP, dCTP, dGTP, and analogs or variants thereof), and additional cofactors and additives such as betaine, DMSO, and magnesium ions. The template can be a single-stranded or double-stranded nucleic acid. The primer can be a short nucleic acid sequence synthetically constructed to be complementary to and hybridize with the target sequence in the template pool. "PCR" can generally refer specifically to this type of reaction, but can also be used more generally to refer to any nucleic acid amplification reaction.
[0125] High-throughput single-molecule PCR can be useful for amplifying pools of distinguishable nucleic acids that may interfere with each other. For example, if multiple distinguishable nucleic acids share a common sequence region, recombination between the nucleic acids along this common region may occur during the PCR reaction, resulting in a new, recombined nucleic acid. In single-molecule PCR, distinguishable nucleic acid sequences are compartmentalized with each other and therefore cannot interact, preventing this potential amplification error. Single-molecule PCR can be particularly useful for preparing nucleic acids for sequencing. Single-molecule PCR can also be useful for absolute quantification of several targets in a template pool. For example, digital PCR (or dPCR) uses the frequency of distinguishable single-molecule PCR amplification signals to estimate the number of starting nucleic acid molecules in a sample.
[0126] In some implementations of PCR, primers to primer binding sites common to all nucleic acids can be used to non-discriminately amplify a group of nucleic acids. For example, primers to primer binding sites flank all nucleic acids in a pool. These common sites can be used for general amplification to create or assemble a synthetic nucleic acid library. However, in some implementations, PCR can be used to selectively amplify a subset of targeted nucleic acids from a pool, for example, by using primers with primer binding sites that are present only in the targeted subset of nucleic acids. Synthetic nucleic acid libraries can be created or assembled so that all nucleic acids belonging to a potential sub-library of interest share a common primer binding site at their edge (common within the sub-library but distinct from other sub-libraries) to selectively amplify the sub-library from a more general library.
[0127] Affinity-tagged nucleic acids can be used as sequence-specific probes for nucleic acid capture. The probes can be designed to be complementary to a target sequence in a pool of nucleic acids. The probes are then incubated with the pool of nucleic acids and allowed to hybridize with their targets.
[0128] Synthetic nucleic acid libraries can be created or assembled with common probe binding sites for general nucleic acid capture. These common sites can be used to selectively capture fully assembled or potentially fully assembled nucleic acids from an assembly reaction, thereby filtering out partially assembled or misassembled (or unintended or undesired) by-products. For example, assembly can include assembling nucleic acids with probe binding sites at each edge sequence such that only fully assembled nucleic acid products contain the required two probe binding sites necessary to pass through a series of two capture reactions using each probe. To increase stringency, each component of the assembly can include a common probe binding site. In some implementations, nucleic acid capture can be used to selectively capture a subset of targeted nucleic acids from a pool, for example, by using probes with binding sites present only in the subset of targeted nucleic acids. Synthetic nucleic acid libraries can be created or assembled such that all of the nucleic acids belonging to a potential sub-library of interest share a common probe binding site (common within the sub-library but distinct from other sub-libraries) in order to selectively capture the sub-library from a more general library.
[0129] In some implementations, the library of nucleic acids can be freeze-dried, for example, for storage. Freeze-drying is a dehydration process. Both nucleic acids and enzymes can be freeze-dried. The freeze-dried material may have a longer shelf life. Additives such as chemical stabilizers can be used to maintain functional products (e.g., active enzymes) throughout the freeze-drying process. Disaccharides such as sucrose and trehalose can be used as chemical stabilizers.
[0130] Nucleic acids can be designed to facilitate sequencing. For example, nucleic acids can be designed to avoid typical sequencing complications such as secondary structures, stretches of homopolymers, repetitive sequences, and sequences with too high or too low GC content. Certain sequencers or sequencing methods can be error-prone. Nucleic acid sequences (or components) constituting a synthetic library (e.g., an identifier library) can be designed at a specific Hamming distance from each other. In this way, even if a high rate of base resolution errors occurs in sequencing, error-containing sequence stretches can still be mapped back to their most likely nucleic acid (or component). Nucleic acid sequences can be designed with a Hamming distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more base variations. Alternative distance metrics to Hamming distance can also be used to define the minimum required distance between designed nucleic acids.
[0131] Some sequencing methods and instruments require the input nucleic acid to contain specific sequences, such as adapter sequences or primer binding sites. These sequences can be referred to as "method-specific sequences." A typical preliminary workflow for the sequencing instruments and methods involves assembling method-specific sequences into a nucleic acid library. However, if it is known in advance that a synthetic nucleic acid library (e.g., an identifier library) will be sequenced by a particular instrument or method, these method-specific sequences can be designed into the nucleic acids (e.g., components) that comprise the library (e.g., an identifier library). For example, sequencing adapters can be assembled onto members of a synthetic nucleic acid library in the same reaction step as the members themselves are assembled from individual nucleic acid components.
[0132] Nucleic acid can be designed to avoid sequences that can easily cause DNA damage.For example, sequences that contain sites for site-specific nucleases can be avoided.As another example, UVB (ultraviolet-B) light can cause adjacent thymines to form pyrimidine dimers, which then inhibits sequencing and PCR.Therefore, if synthetic nucleic acid library is intended to be stored in an environment that is exposed to UVB, it can be beneficial to design its nucleic acid sequence to avoid adjacent thymines (i.e., TT).
[0133] Method for computing with identifiers It may be possible to perform computational calculations on the data encoded in the identifier library using chemical manipulations. This may be advantageous because such manipulations can be performed in a parallelized manner on any subset of the entire archive or on the entire archive. Second, computational calculations can be performed in vitro without decoding the data, thus ensuring confidentiality while enabling computational calculations. In one embodiment, computational calculations, including Boolean logic operations such as AND, OR, NOT, NAND, etc., can be performed on the encoded bitstream using an identifier representing each bit position, where the presence of the identifier encodes the bit value as "1" and the absence of the identifier encodes the bit value as "0."
[0134] In one embodiment, all identifiers are constructed as single-stranded nucleic acid molecules (or initially as double-stranded nucleic acid molecules and then isolated into single-stranded form). For any single-stranded identifier x, we define an identifier that is the reverse complement of x as x. * For any set S of single-stranded identifiers, we denote by S the set of reverse complements of each identifier in S. * We denote by U the set of all possible single-stranded identifiers in the library, and U *We denote these sets by the universe and the universe * It is called U s and U s * By this, we have * A second pair of sets is displayed, and each identifier in these sets is extended with additional nucleic acid sequences known as search regions that can be targeted or selected by chemical methods.
[0135] Computational computation on a given library of identifiers can be performed by a series of chemical operations involving hybridization and cleavage. Abstractions of these operations are described below. Each operation takes a pool of identifiers as input, performs an operation, and returns a pool of identifiers as output.
[0136] The operation single(X) takes a pool of identifiers (double-stranded and / or single-stranded) and returns only single-stranded nucleic acid identifiers (removes all double-stranded identifiers). The operation double(X) takes a pool of identifiers (double-stranded and / or single-stranded) and returns only double-stranded identifiers (removes all single-stranded identifiers). The operations make-single(X) and make-single *(X) converts all double-stranded nucleic acid identifiers to their single-stranded form (the starred version returns the minus strand, while the non-starred version returns the plus strand). The operation get(X,q) returns a pool of all identifiers that match the query q. If q="all", the query matches and operates on all identifiers. The operation delete(X,q) deletes all identifiers (double-stranded or single-stranded) that satisfy the query q. Queries can be implemented with random access as previously described. The operation combine(P,Q) returns a pool containing all identifiers in P or Q. We define the operation assign(X,Y), which assigns the result of Y to a variable name X. Briefly, we also denote this operation in the form: X=Y. We assume that the assignment operation runs under ideal conditions, allowing variables to be reused without any "contamination" issues.
[0137] In the following, we consider that bitstreams a and b, both of length l, are written into double-stranded identifier libraries dsA and dsB, respectively, and we consider that some sub-bitstreams s=a i …a j and t=b i …b j Suppose we are interested in computing on dsA, dsB, s, t, and the results of the computation should be stored in sub-bitstream s. That is, we assume that the following operations are initially performed in a specified order, denoted by the operation initialize(dsA, dsB, s, t): [Table 1]
[0138] FIG. 9 illustrates an example setup for computing with an identifier library. This figure illustrates an example of a combinatorial space of identifiers, depicted as an abstract tree data structure (labeled 4). In this example, each level of the tree chooses between two components (labeled 2). Each path from the root of the tree corresponds to a unique identifier (as illustrated by the example in label 3) and determines its order (or rank). Label 4 indicates a single-stranded universal identifier library. Label 5 indicates a single-stranded identifier library that encodes a specific bitstream, e.g., called "a." Label 7 indicates a sub-bitstream of "a," called "s," containing 7 bits. Similarly, label 10 indicates a sub-bitstream "t" of bitstream "b" of the same length. As described in the initialization procedure for computing initialize(dsA, dsB, s, t), the sub-bitstreams to be computed are available in pools P and Q (labeled 6 and 9, respectively) and are ready for computation.
[0139] The operation and(s,t), defined as the bitwise AND of bits in bitstreams s and t, can be implemented using the sequence of operations below. [Table 2]
[0140] The operation not(s), defined as the bitwise logical negation of the bits in a bitstream s, can be implemented using the sequence of operations below. [Table 3]
[0141] The operation or(s,t), defined as the bitwise OR of bits in bitstreams s and t, can be implemented using the sequence of operations below. [Table 4]
[0142] The operation nand(s,t), defined as the bitwise logical negation of the conjunction of bits in bitstreams s and t, can be implemented using the sequence of operations below. [Table 5]
[0143] In one embodiment, the operation single(X) hybridizes X to a universal identifier such that a single-stranded identifier derived from X hybridizes to a universal identifier. s or U s * It may involve first combining with either U s and U s * The universal identifier in has a special search region so that those molecules that hybridize to the universal identifier can be accessed in a targeted manner.
[0144] In one embodiment, the operation double(X) may involve treating the identifiers in X with a single-strand-specific nuclease, such as S1 nuclease, and then running the resulting pool of DNA on a gel to isolate only the uncut (and therefore fully double-stranded) identifiers.
[0145] Figure 10 illustrates an example of how logical operations can be performed on bitstreams "s" and "t" encoded by an identifier library. In this figure, we use a universal library (labeled 14) that is complementary to the pool being computed. The columns labeled AND / AND show how the product of bitstreams "s" and "t" (labeled 5 and 7, respectively) can be computed. We show that the pool can be computed using the exact universal library (U or U *) are reformatted using the . When two pools are combined, as shown (e.g., labeled 9), the complementary single-stranded identifiers hybridize to form double-stranded identifiers. The resulting collection of double-stranded identifiers in the pool (labeled 10) encodes the result of a computational calculation of AND: separation of the double-stranded products yields an identifier library representation of and(s,t). Alternatively, separation of the single-stranded products yields an identifier library representation of nand(s,t). The column labeled OR indicates how the disjunction of bitstreams "s" and "t" can be computed. When pools containing identifiers representing "s" and "t" are combined, the resulting library contains a representation of or(s,t). The column labeled NOT indicates how the negation of bitstream "s" can be computed. Thus, the single-stranded identifier library representing bitstream "s" is combined with the complementary universal identifier library (labeled 15). As a result (labeled 19), all double-stranded products formed (e.g., labeled 18) represent a "1" bit in "s" and can be discarded. The remaining single-stranded products (e.g., labeled 17) represent a "0" bit in "s" and thus correspond to a "1" bit in not(s). These single-stranded products provide an identifier library representation of not(s) and can be used for further computation.
[0146] DNA-based data randomization, encryption, and authentication methods The ability to generate and store random bit streams using DNA may have applications in computer - based computations in cryptography and combinatorial algorithms. Many encryption algorithms, e.g., DES, require the use of random bits to ensure security. Other encryption algorithms, e.g., AES, require the use of an encryption key. Typically, any systematic pattern or bias in the random bits or keys can be exploited to attack and break the encrypted message, so such random bits and keys are generated using a robust source of randomness. Further, the keys used in encryption are typically required to be archived for decryption. The strength of the security of an encryption method depends on the length of the key used in the algorithm: generally, the longer the key, the stronger the encryption. Methods such as the one - time - pad are one of the most robust encryption methods but are limited in use due to their long key requirements.
[0147] Using the methods described in this document, an extremely large collection of random keys, which can be tens, hundreds, thousands, tens of thousands or more bits in length, can be generated and archived. In one implementation, a nucleic acid library can be generated where each nucleic acid molecule meets the following design: has a length of n bases and has a variable region of k < n bases. The bases in the variable region can be randomly selected during the construction of the library. For example, n can be 100 and k can be 80; thus, a library of different molecules of size 10 50 can potentially be generated. For example, sequencing a random sample of such a library of 1000 - sized molecules can yield up to 1000 - bit random keys that can be used for encryption.
[0148] In another implementation, the nucleic acid keys (nucleic acid molecules representing the keys) described above can be attached to identifiers to obtain an ordered collection of key sets. The ordered key sets can be used to synchronize the order in which keys are used by various parties in a cryptographic context. For example, an identifier library can be constructed combinatorially using a product scheme to generate 10 12 A number of unique identifiers can be obtained. Using microfluidic methods, each identifier can be juxtaposed with a nucleic acid key and assembled to form a nucleic acid sample comprising a unique identifier and a random key. Because the identifiers in the identifier library are ordered, the keys can then be ordered, accessed, and sequenced in any specified order.
[0149] In another implementation, a key attached to the identifier can be used to instantiate a random function that maps the input identifier to a string of random bits. Such a random function can be useful in applications requiring a function whose value is easy to compute but difficult to invert from a given value, such as hashing. In such applications, a library of keys assembled with each unique identifier is used as the random function. If a value is to be hashed, it is mapped to the identifier. The identifier is then accessed from the key library using a random access method, such as hybridization capture or PCR. The identifier is attached to a key containing a sequence of random bases. This key is sequenced, translated into a string of bits, and used as the output of the random function.
[0150] Because libraries of nucleic acid molecules can be cheaply and quickly copied and can be transported covertly in small volumes, nucleic acid key sets generated as described above can be useful in situations where large numbers of encryption keys must be distributed periodically in a robust and covert manner among multiple parties that are not geographically co-located. Additionally, keys can be reliably archived for extremely long periods of time, allowing for robust storage of encrypted archived data.
[0151] 11-16 illustrate implementations of methods for creating, storing, accessing, and using random or encrypted data stored in DNA. DNA is depicted as a string of characters including gray and black bars and symbols. Each depicted DNA represents a distinct species. A "species" is defined as one or more DNA molecules of the same sequence. When "species" is used in the plural sense, it can be assumed that all species in the plurality of species have distinct sequences, although this may be clarified by writing "distinguishable species" instead of "species."
[0152] Figure 11 depicts an example of an entropy (or random data) generator using a large combinatorial space of DNA and a sequencer. The method begins with a random pool of DNA species, called seeds. The seeds ideally contain all species of a defined combinatorial set of DNA, e.g., 50 bases (4 50The seed should contain a uniform distribution of all DNA species (with members N...N). However, because the complete combinatorial space may be too large for all members to be represented in the seed, it is acceptable for the seed to contain a random subset of the combinatorial space instead of the entire combinatorial space. Seed species can be designed to have a common sequence at the ends (black and light gray bars) followed by distinct sequences in the middle (N...N). A degenerate oligonucleotide synthesis strategy can be used to produce this starting seed in a rapid and inexpensive manner. The common end sequence may allow for amplification of the seed by PCR or compatibility with certain readout (or sequencing) methods. As an alternative to degenerate oligonucleotide synthesis, combinatorial DNA assembly (multiplexed in a single reaction) can also be used to generate seeds rapidly and inexpensively. The sequencer randomly samples species from the seed and samples them in a random order. Because there is uncertainty in the species read by the sequencer at any given time, the system can be classified as an entropy generator, which can be used to generate random numbers or random streams of data, for example, as encryption keys.
[0153] Figure 12A illustrates an example schematic of a method for storing randomly generated data in DNA. The method begins with (1) a large random pool of DNA species, called seeds. The seeds ideally contain all species of a defined combinatorial set of DNA, e.g., 50 bases (4 50The seed should contain a uniform distribution of all DNA species (having 1,000 members). However, because the complete combinatorial space may be too large for all members to be represented in the seed, it is acceptable for the seed to contain a random subset of the combinatorial space. The seed itself can be generated from degenerate oligonucleotide synthesis or combinatorial DNA assembly. (2) Random data (or entropy) is generated by receiving a random subset of the species in the seed. For example, this can be achieved by receiving proportional, fractional volumes of seed solution. For example, if the seed solution consists of an estimated million species per microliter (uL), a random subset of approximately 1,000 species can be selected by receiving a 1 nanoliter (nL) aliquot from the seed solution (assuming it is well mixed). Alternatively, the subset can be selected by flowing an aliquot of the seed solution through a nanopore membrane and collecting only the species that pass through the membrane. Counting the number of species that pass through the membrane can be achieved by measuring the voltage difference across the nanopore. This process can continue until a desired number of signatures are detected (e.g., 100, 1000, 10,000, or more species signatures). As another alternative, single species can be isolated in small droplets (e.g., in an oil emulsion). Small droplets containing a single species can be detected by their fluorescent signature and sorted into a collection chamber by a series of microfluidic channels. (3) We refer to each selected species as an identifier, and further refer to a complete subset of selected species as a "random identifier library" or RIL. To stabilize the information in the RIL and protect it from degradation, the RIL can be amplified with PCR primers that bind to consensus sequences at the ends of the species. To determine the identifiers in the RIL (and therefore the data stored therein), the RIL can be sequenced. True identifiers can be defined by species in the sample that have enrichment above a defined noise threshold.(4) Once the data contained in the RIL has been determined, extra error checking and error correction seeds can be added to the RIL. For example, "integer DNA" (e.g., a checksum or parity check) can be added to the RIL that contains information about how many identifiers to expect. The integer DNA can inform how deeply to sequence the RIL to retrieve all of the information.
[0154] RILs can be barcoded with unique DNA tags. Several barcoded RILs can then be pooled together so that any given RIL can be individually accessed by hybridization assay (or PCR) for its unique DNA tag. Unique DNA tags can be combinatorially assembled or synthesized and then assembled in its corresponding RIL. Figure 12B shows an example of an RIL containing four species, each containing 100 random bases. The combinatorial space of possible species is 4 100 Therefore, RIL is log2(4 100 choose4) can contain ≈725 bits of information. Figure 12C also shows an example of a RIL containing four seeds, each containing 100 random bases. 4 100 As an alternative to storing the information in specific unordered combinations of four species chosen from the combinatorial space (as in Figure 12B), we can reserve the final 90 random bases of each species and store them in log2(4 90 ) = 180 bits of information can be stored, while the first 10 random bases can be reserved to establish a relative order between the information stored in each of the four seeds. The relative order can be defined by a lexicographic ordering of the 10-base string based on a defined ordering of the 4 bases (similar to how words in English are ordered according to the order of letters in the alphabet). This method for assigning information to RILs can be computationally faster than the method described in Figure 12B for mapping to binary strings.
[0155] In the previous figures (Figures 12A-12C), we discuss a strategy for barcoding multiple RILs and pooling them together. To do this, an input-output mapping is created where the input corresponds to a barcode hybridization probe (for accessing individual RILs) and the output corresponds to a random data string (encoded by the targeted RILs). In this method, predefined barcodes are assembled into random data for retrieval from the combined pool, while Figure 13A demonstrates a different method for creating an input-output mapping between nucleic acid probes and random data strings, whereby barcodes (for accessing data) are randomly generated along with the random data itself. For example, a barcode can be a pair of short sequences of DNA that can appear at both ends of one or more species. In this implementation, the combinatorial space of possible barcodes may be small compared to the total number of all possible species in the pool, so that each barcode is, by chance, associated with one or more species. For example, if the barcode is three bases at each end of a random DNA sequence (flanked by a consensus sequence) in a species, then 4 6 = 4096 possible barcodes, so 4 can be constructed to access these 6 There are = 4096 primer pairs (corresponding to a 12-bit input). If the pool of DNA is selected to have approximately 400K species, then each barcode can be associated with approximately 100 species on average. In this implementation, the RILs are defined by the subset of species associated with each barcode. Following the preceding example, if each species contains 25 random bases (or random sequences) apart from the bases (or sequences) used for barcoding, then the barcodes associated with the RILs of the 100 species will be at most log2(4 25 choose100) can contain ≒ 4475 bits of information.
[0156] Figure 13B demonstrates an implementation of a scheme for accessing and reading stored random data from a pool of barcoded RILs. The sequencer (or reader) can further include functions to manipulate the sequence data before returning the output. For example, a hash function could perform a reverse chemical query using the output data string to make it difficult to find the input. This functionality could be useful, for example, if the input is a key or credential used for authentication.
[0157] Methods for generating and storing queryable (or accessible) random strings of data can be particularly useful for generating and archiving encryption keys (generated from random data strings). Each input can be used to access a different encryption key. For example, each input can correspond to a specific user, time range, and / or project in a private archive database. The encrypted data in a private archive database (potentially amounting to very large amounts of data) can be stored on conventional media by the archive service provider, while the encryption key can be stored in DNA by the owner. Furthermore, the potential latency and sophistication required to perform chemical access protocols for specific inputs can increase the encryption method's security barrier to hacking.
[0158] FIG. 14 illustrates an example of a system for securing and authenticating access to an artifact. The system requires a physical key containing a specific combination of DNA species taken from a large pool of possible species. The target combination of species, also referred to as an "identifier key," can be generated automatically, for example, by a combinatorial microfluidic channel, electrowetting, or printing device, or manually by pipetting. A reader or sequencer with a built-in lock verifies the matching identifier key and enables access to the artifact. Alternatively, the reader can behave as a credential-token system, whereby instead of directly unlocking access to the artifact, it returns a token that can be used to access the artifact. The token can be generated, for example, by a built-in hashing function within the reader, whereby the hashing function is electronically applied to the read or sequence data originating from the reader. For example, the reader includes a processor configured to execute steps of a program on a processor-readable medium, whereby the steps involve taking the read or sequence data, applying one or more mathematical or logical operations to the data, and outputting a hashed value or hashed token.
[0159] Rather than applying a hashing function electronically after sequencing the identifiers within a reader or otherwise, the hashing function can be applied chemically via one or more reactions applied to a library of identifiers to generate a hashed library, after which the now-hashed identifiers are sequenced or read. This approach is advantageous because it represents an air-gapped approach for greater security of the information encoded by the identifiers, since sequencing or reading only the hashed identifiers does not reveal the sequence data of the original library of identifiers. Figure 15 shows a flowchart describing a method 1500 for preparing a library of nucleic acid molecules for use in security and authentication. Method 1500 involves steps 1502 and 1504. Step 1502 involves obtaining a library of nucleic acid molecules representing security tokens. Step 1504 involves applying a chemical operation to the library representing the security tokens to obtain a hashed library of nucleic acid molecules representing the hashed tokens.
[0160] The chemical operations can be designed to result in one or more Boolean functions on the security tokens. For example, the Boolean functions described above with respect to Figures 9 and 10 can be applied to the library and, thereby, the tokens it represents. Such Boolean functions can constitute hash functions that are chemically applied to the library to obtain a hashed library representing the hashed tokens. The hashed library can be a subset of the original library, where the subset is determined by selecting a portion of the nucleic acid molecules in the library.
[0161] In some implementations, method 1500 further includes sequencing at least a portion of the nucleic acid molecules in the hashed library to obtain a sequencing read. Additionally, method 1500 may involve comparing the sequencing read to a database or lookup table to determine the presence or absence of a matching sequence. Based on the presence or absence of a matching sequence in the sequencing read, access to a secured asset or location can be granted or denied. Suitable types of sequencing include Sanger sequencing, high-throughput sequencing, shotgun sequencing, and nanopore sequencing.
[0162] Rather than sequencing the hashed library, a verification function can be applied to authenticate the hashed token without having to sequence the entire library. If the hashed token matches a reference sequence, the verification function is performed by one or more additional chemical operations on the hashed library to produce an output molecule. The chemical operations can have the effect of performing Boolean logic on the hashed token, such as those described above with respect to Figures 9 and 10. An assay is then used to determine the presence or absence of the output molecule. The verification function chemical operations can involve nested PCR, PCR with target-specific primers, application of a set of probes (e.g., affinity-tagged probes or degradation-targeted probes), or application of an enzyme or protein that interacts with the nucleic acids of the hashed library. For example, the verification function chemical operation can have the effect of comparing the hashed token to a reference pattern / sequence by applying primers to the hashed library, where the primers are designed to hybridize only to nucleic acid molecules having a sequence that matches the reference pattern. Another example involves comparing or evaluating hashed tokens by using CRISPR-associated proteins, such as zinc finger nucleases, transcription activator-like effector nucleases (Talen), or Cas9, to target nucleic acid molecules with sequences corresponding to the reference pattern. These proteins can cleave the targeted nucleic acid molecules to create fragments. Cas9, in particular, can use guide RNAs that are complementary to the target nucleic acid. The output molecule can be a small molecule, a nucleic acid molecule, a nucleic acid molecule with a specific sequence, a nucleic acid fragment from one of the nucleic acids in the library, a protein, an enzyme, a functionalized protein, a tagged molecule, or a molecule configured to decay over a short period of time. For example, the output molecule can be RNA (e.g., RNA from a library) that degrades by methylation of uracil to thymine or oxidative degradation of uracil, a process that modifies the sequence of the RNA and confers a limited lifetime of sequence fidelity on the RNA.
[0163] For example, PCR, reverse transcription PCR (RT-PCR), qPCR, affinity tagging, fluorimetry, or electrophoresis can be used as assays to complete the verification function. Fluorimetry can be particularly useful when the output molecule is or is tagged with a fluorophore. RT-PCR, which produces complementary DNA (cDNA), which is more chemically stable than RNA, is useful for assaying RNA as the output molecule. Assays can be used together or instead to verify the chemical identity of the output molecule. The method can further involve granting or denying access to a secured asset or location based on the assay results. The method can further involve determining the authenticity of library-related artifacts based on the assay results.
[0164] In some implementations, the library includes a unique molecular barcode. The library may be lyophilized for stabilized storage. The security token may be unique to the user of the token. The security token may encode a message, a code word, a randomized code word / key / string, an identity, or a monetary value. The token may be part of a two-factor authentication system where a user logs into the system by entering a password, presents the library, which is hashed and verified to confirm or deny access to the system. The library may be configured to decay after a period of time. For example, the library may be RNA (e.g., library-derived RNA) that degrades by methylation of uracil to thymine or oxidative degradation of uracil, processes that modify the sequence of the RNA and give the RNA a limited lifetime of sequence fidelity.
[0165] In some implementations, the library is collocated with the artifact, and the security token is unique to the artifact. For example, the artifact is a container configured to enclose the library, such as a well, a droplet, a spot, a sealed container, a gel, a suspension, or a solid matrix. Other suitable artifacts include fluids (e.g., liquids, gases, oils, inks, compressed gases, or drugs), biological objects, currency, or documents. When the artifact is a document, an ink or stamp containing the library is imprinted on the document.
[0166] The library can encode at least about 1 kilobit of information. The security token can include a plurality of symbols, each represented by a distinct sequence of nucleic acid molecules in the library. In some implementations, the library is randomly generated, such as any of the random libraries described with respect to Figures 11-13. In some implementations, the security token is represented by a library of nucleic acid molecules in an encoding scheme, where the tokens are mapped to a plurality of symbols having one of two possible symbol values, where a symbol in the plurality of symbols is represented by the presence of a distinct nucleic acid molecule in the library if the symbol has a first of the two possible symbol values, and where a symbol in the plurality of symbols is represented by the absence of a distinct nucleic acid molecule if the symbol has a second of the two possible symbol values.
[0167] How to tag artifacts and track entities with DNA Identifier libraries dissolved in a solvent can be sprayed, diffused, dispensed, or injected into or onto a physical artifact to tag it with information. Identifier libraries in solid form (e.g., lyophilized) can be deposited, electrostatically attached, chemically bonded, or aerosolized and sprayed into or onto a physical artifact to tag it with information. For example, unique identifier libraries can be used to tag distinguishable instances of a certain type of artifact. Identifier library tags on artifacts can act as unique barcodes or values, or can contain more sophisticated information such as product numbers, manufacturing or shipping dates, location of origin, or any other information related to the artifact's history, e.g., a transaction list of previous owners. The main advantages of using identifiers to tag artifacts are that the identifiers are undetectable, durable, and well-suited for individually tagging large numbers of artifact instances.
[0168] Physical objects can be marked or painted with uniquely identifiable synthetic DNA samples. Even gases (e.g., compressed air) and liquids (e.g., ink or oil) can be tagged, which is not possible with conventional methods. If ink, such as that in a print cartridge or pen, is tagged with a unique DNA library and used to print or write on a document, the authenticity of the document can be verified by swabbing the DNA from the document and sequencing it. Furthermore, covert messages can be contained within the ink that either complement or verify material in the document. Tags are discreet and can be used, for example, to identify whether an object has moved through a particular physical space or interacted with another object. Tags are also quantitative and can therefore be used to verify whether a particular object has been tampered with or diluted (in the case of a liquid or gas). For example, if a liquid is tagged with 1,000 copies per mL but later recovered at 100 copies per mL, it can be inferred that the liquid has been diluted. Tags and barcodes can be easily created and placed. They can contain up to several kilobits or more of information. They can be created by accepting a subset of identifiers from a pre-made combinatorial space of possible identifiers.
[0169] Identifier libraries can be easily generated and used as tokens to gain access to secure assets. Tokens can be small, for example, encoding 1 kilobit of information while still being secure. The identifier library representing a token can be created by receiving a subset of identifiers from a pre-created combinatorial space of possible identifiers. For example, a token can be given to an owner after deposit and accepted after the asset is withdrawn. Alternatively, a token can be created by the owner, like a physical key. Because of its physical nature, the token will be impervious to electronic theft or tampering. Similarly, because of its confidential nature, the token will be difficult to counterfeit. Chemical methods can be used to hash or verify tokens to prevent them from ever entering electronic or readable form. Chemical operations can be used to perform hash or verification functions, such as the Boolean logic gates described above with respect to Figures 9 and 10. For example, chemical logic gates such as AND, OR, NOT, and NAND can be configured together to form a hash function that makes it difficult to deduce the original token by sequencing the hashed token. The value of the hashed token can be matched to a database to determine authorization for the asset. Due to the irreversible nature of the hash function, the database can be viewed by unauthorized parties, yet still not compromise the security of the asset and the authorized parties' ability to access it. Additionally or alternatively, the chemical logic gate can include a verification function for the token that can verify the token without requiring sequencing of the DNA molecule that contains it. For example, the verification function can be used to produce a specific output identifier if and only if the token matches a precise pattern. For example, the presence of the identifier can be determined by an assay such as real-time PCR (qPCR), fluorimetry, or gel electrophoresis.
[0170] 16 shows a flowchart describing a method 1600 for tagging a fluid for tracking or authentication. Method 1600 includes steps 1602 and 1604. Step 1602 involves obtaining a library of nucleic acid molecules representing information. Step 1604 involves combining the fluid with a tag containing the library to obtain a tagged fluid for tracking or authentication. For example, the tags containing the library of nucleic acid molecules are dispersed approximately uniformly throughout the tagged fluid.
[0171] In some implementations, method 1600 further includes sampling the library of nucleic acid molecules from the tagged fluid to obtain a sample. Sampling may involve swabbing the tags or tagged fluid, extracting at least a portion of the library from the tagged fluid (e.g., by pipetting or removing a volume from the fluid), or removing the tags from the tagged fluid (e.g., by a separation process such as filtration). In some implementations, the tags further include magnetic beads, and sampling involves applying a magnet to the fluid to extract the tags via the magnetic beads. Method 1600 may further include sequencing the sample of nucleic acid molecules to obtain a sequencing readout. Any of the sequencing methods described above may be used for this step. The sequencing readout may be communicated to a computer system, such as computer network 802 described in FIG. 8. According to the methods described herein, the sequencing readout may be hashed using a hashing function to obtain hashed data for information security.
[0172] The library can encode at least about 1 kilobit of information. The amount of information can be scaled based on the size of the library and / or fluid. In some implementations, the tag includes a molecular barcode specific to the tag or fluid. The information encoded by the library of nucleic acid molecules can be a message, such as an encrypted message. The information can represent a monetary value. The tag can be part of a two-factor authentication system.
[0173] In some implementations, the fluid is a liquid, gas, oil, ink, compressed gas, or drug. Method 1600 may involve measuring the concentration of the tag in the tagged fluid to determine the amount of dilution. In some implementations, the tag is configured to decay or dilute within a period of time. For example, this period of time begins when the tag or fluid is accessed or sampled. For example, the nucleic acid of the tag is RNA, which degrades by methylation of uracil to thymine or oxidative degradation of uracil, processes that modify the sequence of the RNA and give the RNA a limited lifetime of sequence fidelity. Alternatively, the fluid is contained in a locked container, and when the locked container is broken, a reagent is released into the fluid to react with the tag.
[0174] A library of nucleic acid molecules can be an identifier library that encodes information, as described above. The information can include or be mapped to a plurality of symbols, each represented by a distinct sequence of nucleic acid molecules in the library. In some implementations, the library is a subset of a larger library. In some implementations, the library is randomly generated, as described above with respect to Figures 11-13. In some implementations, the information is represented by a library of nucleic acid molecules in an encoding scheme, where the information is mapped to a plurality of symbols having one of two possible symbol values, where if a symbol has a first of the two possible symbol values, a symbol of the plurality of symbols is represented by the presence of a distinct nucleic acid molecule in the library, and if a symbol has a second of the two possible symbol values, the symbol is represented by the absence of a distinct nucleic acid molecule. For example, the two possible symbol values are 0 and 1, and a nucleic acid molecule corresponding to a symbol having a value of 0 is absent from the tag, and a nucleic acid molecule corresponding to a symbol having a value of 1 is present in the tag.
[0175] In another implementation, one or more physical locations can each be tagged with a unique identifier from an identifier library. For example, physical sites A, B, and C can be universally tagged with an identifier library. An entity, such as a vehicle, person, or any other object, that visits or comes into contact with site A can, intentionally or unintentionally, pick up a sample of the identifier library. Later, after the entity's access, a sample can be collected from the entity, chemically processed, and decoded to identify which sites were visited by the entity. An entity can visit more than one site and pick up more than one sample. If the identifier libraries are disjoint, a similar process can be used to identify some or all of the sites visited by an entity. Such a scheme can have application in covert tracking of entities. Some advantages of using this scheme are that the identifiers are undetectable unless specifically sought, can be designed to be biologically inert, and can be used to uniquely tag a vast number of sites or entities.
[0176] In another implementation, the identifier library can tag entities. Entities can leave samples of injected identifiers at sites they visit. Such samples can be collected, processed, and decoded to identify which entities have visited the sites. [Example]
[0177] Example 1: Encoding, Writing, and Reading a Single Poem in a DNA Molecule The data to be encoded is a text file containing the poem. To construct the identifier using a production scheme implemented with overlap-extension PCR, the data is manually encoded by mixing together DNA components from two layers of 96 components using a pipette. The first layer, X, contains 96 total DNA components. The second layer, Y, also contains 96 total components. Before writing to DNA, the data is mapped to binary and then re-encoded into a uniform weight format in which every contiguous (adjacent, separate) string of 61 bits in the original data is translated into a 96-bit string with exactly 17 bit values of 1. This uniform weight format may have inherent error-checking qualities. The data is then hashed into a 96 x 96 table to form a reference map.
[0178] The center panel of Figure 17A shows a two-dimensional reference map of a 96x96 table in which the poem is encoded into multiple identifiers. Black dots correspond to "1" bit values and white dots correspond to "0" bit values. The data is encoded into identifiers using two layers of 96 components. A component is assigned to each X and Y value in the table, and for each (X,Y) coordinate with a "1" value, overlap extension PCR is used to assemble the X and Y components into an identifier. The data was read back (e.g., decoded) by sequencing the identifier library to determine the presence or absence of each possible (X,Y) assembly.
[0179] The right panel of Figure 17A shows a two-dimensional heat map of the abundance of sequences present in the identifier library, as determined by sequencing. Each pixel represents a molecule containing the corresponding X and Y components, and the grayscale intensity at that pixel represents the relative abundance of that molecule compared to other molecules. Identifiers are taken as the top 17 most abundant (X,Y) assemblies in each row (uniform weight encoding ensures that each consecutive column of 96 bits can have exactly 17 "1" values, and therefore 17 corresponding identifiers).
[0180] (Example 2: Encoding a 62824-bit text file) The data to be encoded are three text files totaling 62,824 bits. To construct identifiers using a production scheme implemented with overlap-extension PCR, the data is mixed and encoded using a Labcyte Echo® liquid handler, mixing DNA components from two layers of 384 components together. The first layer, X, contains 384 total DNA components. The second layer, Y, also contains 384 total components. Before writing to DNA, the data is mapped to binary and then re-encoded to reduce the weight (the number of "1" bits) and include a checksum. The checksum is established so that for every consecutive string of 192 bits of data, there is an identifier corresponding to the checksum. The weight of the re-encoded data is approximately 10,100, which corresponds to the number of identifiers to be constructed. The data can then be hashed into a 384 x 384 table to form a reference map.
[0181] The center panel of Figure 17B shows a two-dimensional reference map of a 384 x 384 table in which a text file is encoded into multiple identifiers. Each coordinate (X, Y) corresponds to a bit of data at position X + (Y - 1) * 192. Black dots correspond to a bit value of "1" and white dots correspond to a bit value of "0." The black dots on the right side of the figure are checksums, and the pattern of black dots at the top of the figure is a codebook (e.g., a dictionary for decoding the data). Components can be assigned to each X and Y value in the table, and for each (X, Y) coordinate with a "1" value, overlap extension PCR can be used to assemble the X and Y components into an identifier. The data was read back (e.g., decoded) by sequencing the identifier library to determine the presence or absence of each possible (X, Y) assembly.
[0182] The right panel of Figure 17B shows a two-dimensional heat map of the abundance of sequences present in the identifier library, as determined by sequencing. Each pixel represents a molecule containing the corresponding X and Y components, and the grayscale intensity at that pixel represents the relative abundance of that molecule compared to other molecules. The identifier is taken as the top S most abundant (X, Y) assembly in each row, where S in each row can be a checksum value.
[0183] While preferred implementations of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such implementations are provided by way of example only. The present invention is not limited by the specific examples provided herein. While the present invention has been described with reference to the above specification, the descriptions and illustrations of implementations herein are not intended to be construed in a limiting sense. Numerous modifications, changes, and substitutions will readily occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the implementation of the invention described herein can be used in practicing the invention. Accordingly, it is intended that the present invention encompass any and all such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.
Claims
1. A method for preparing a library of nucleic acid molecules for use in artifact security and authentication, comprising: obtaining a library of nucleic acid molecules encoding security tokens, the library being collocated with an artifact, the security tokens being unique to the artifact, and the security tokens encoding a message, a code word, a randomized code word / key / character string, an identity, or a monetary value; applying a chemical operation to the library encoding the security tokens to hash the security tokens to obtain a hashed library of nucleic acid molecules encoding hashed tokens; A method comprising:
2. The method of claim 1 , wherein the chemical manipulation results in one or more Boolean functions on the security token.
3. The method of claim 2 , wherein the one or more Boolean functions apply a hash function to the security token to obtain the hashed token represented by the hashed library.
4. Sequencing at least some of the nucleic acid molecules of the hashed library to obtain sequencing reads. The method of claim 1 further comprising:
5. comparing the sequencing reads to a database or look-up table to determine the presence or absence of matching sequences. The method of claim 4 further comprising:
6. granting or denying access to a secured asset or location based on said determined presence or absence, respectively, of said matching sequence. The method of claim 5 further comprising:
7. 5. The method of claim 4, wherein the sequencing comprises any one of high-throughput sequencing, shotgun sequencing, or nanopore sequencing.
8. if the hashed token matches a reference sequence, applying additional chemical operations to the hashed library to produce an output molecule; determining the presence or absence of said output molecule by assay; The method of claim 1 further comprising:
9. 9. The method of claim 8, wherein the assay is one of polymerase chain reaction (PCR), real-time PCR, reverse transcription PCR (RT-PCR), fluorimetry, and gel electrophoresis.
10. 9. The method of claim 8, wherein the output molecules are distinguishable nucleic acid molecules of the hashed library.
11. granting or denying access to the secured asset or location based on the presence of the output molecule; The method of claim 8 further comprising:
12. 10. The method of claim 1, wherein the library comprises unique molecular barcodes.
13. The method of claim 1 , wherein the security token comprises a randomly generated key.
14. The method of claim 1 , wherein the library is lyophilized.
15. The method of claim 1 , wherein the artifact is a fluid.
16. 16. The method of claim 15, wherein the fluid is one of oil, ink, compressed gas, or a drug.
17. 16. The method of claim 15, further comprising measuring the concentration of the library in the fluid to determine the amount of dilution.
18. The method of claim 1 , wherein the artifact is a living organism.
19. The method of claim 1 , wherein the artifact is a document.
20. 10. The method of claim 1, wherein the library is contained in any one of a well, a droplet, a spot, a sealed container, a gel, a suspension, or a solid matrix.
21. The method of claim 1 , wherein the library is generated by selecting a subset of nucleic acid molecules from a pool of nucleic acid molecules.
22. The method of claim 1 , wherein the security token is part of a two-factor authentication system.
23. 10. The method of claim 1, wherein the security token comprises a plurality of symbols, each symbol represented by a distinct sequence of a nucleic acid molecule in the library.
24. The method of claim 1 , wherein the library is randomly generated.
25. The security tokens are represented by the library of nucleic acid molecules in an encoding scheme, the security tokens are mapped to a plurality of symbols having one of two possible symbol values, and the symbol has a first symbol value of the two possible symbol values.
2. The method of claim 1, wherein a symbol of the plurality of symbols is represented by the presence of a distinguishable nucleic acid molecule in the library if the symbol has a second of the two possible symbol values, and the symbol is represented by the absence of the distinguishable nucleic acid molecule if the symbol has a second of the two possible symbol values.
26. The method of claim 1 , wherein the security token includes at least 1 kilobit of information.
27. The method of claim 1 , wherein the security token is unique to a user.
28. The method of claim 1 , wherein the hashed library is a subset of the library.
Citation Information
Patent Citations
DNA authentication system
JP2010020524A
Nucleic acid sequence assembly process and system
JP2017526046A
Nucleic acid-based data storage
WO2018094108A1
Encryption method, decryption method, encryption system and decryption system
WO2019066007A1
Methods and reagents for assessing the presence or absence of replication competent virus
WO2019152747A1