DNA-based vector data structure using parallel computing
By employing combinatorial strategies and parallel operations, the method addresses the inefficiencies of base-by-base nucleic acid synthesis, achieving cost-effective and rapid data storage and retrieval in nucleic acid sequences.
Patent Information
- Application Number
- JP2025536950
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-21
- Publication Date
- 2026-01-21
AI Technical Summary
Current methods for encoding digital information into nucleic acid sequences rely on costly and time-consuming base-by-base synthesis, leading to errors and high costs in sequencing and retrieval.
Encoding digital information into nucleic acid sequences without base-by-base synthesis by using combinatorial strategies and unique nucleic acid sequences, enabling parallel operations through chemical methods like hybridization, ligation, and amplification for efficient data storage and retrieval.
Reduces the cost and time of encoding and sequencing by utilizing parallel processes, allowing for faster and more reliable data storage and retrieval in nucleic acid molecules.
Smart Images

Figure 2026502172000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 434,154, entitled "DNA-based Vector Data Structure with Parallel Operations," filed December 21, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002] background
[0002] Nucleic acid digital data storage is a stable method for encoding and storing information for long periods of time, storing data at higher densities than magnetic tape or hard drive storage systems. Furthermore, digital data stored in nucleic acid molecules stored in low-temperature, dry conditions can be retrieved for periods of time as long as 60,000 years or more.
[0003]
[0003] To access the digital data stored in nucleic acid molecules, the nucleic acid molecules can be sequenced. Thus, nucleic acid digital data storage can be an ideal method for storing data that is not accessed frequently, but for long-term storage or archiving of large amounts of information.
[0004]
[0004] Current methods rely on encoding digital information (e.g., binary code) into nucleic acid sequences base by base, such that the relationships between bases in the sequence are directly translated into digital information (e.g., binary code). Because the cost of de novo nucleic acid synthesis base by base can be high, sequencing digital data stored in base-by-base sequences that can be read into bitstreams or bytes of digitally encoded information can be prone to errors and expensive to encode. Opportunities for new methods of implementing nucleic acid digital data storage may provide approaches to encoding and retrieving data that are lower cost and easier to commercially implement. Summary of the Invention [Means for solving the problem]
[0005] overview
[0005] Described herein are methods and systems for encoding digital information into nucleic acid (e.g., deoxyribonucleic acid, DNA) molecules without base-by-base synthesis by encoding bit-value information into the presence or absence of unique nucleic acid sequences in a pool, including designating each bit location in a bit stream with a unique nucleic acid sequence and designating the bit value at that location by the presence or absence of a corresponding unique nucleic acid sequence in the pool. However, more generally, designating unique bytes in a byte stream with a unique subset of nucleic acid sequences. Also described are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzyme-based editing of nucleic acid sequences).
[0006]
[0006] Also described herein is a technology including a vector data structure for representing data objects using identifiers, applied to searches for data objects similar or identical to a given query data object and vector arithmetic. The technology includes a DNA data structure and a computational architecture that provides parallel operations on individual bits of each data object. The described architecture utilizes chemical methods for DNA selection, hybridization, ligation, or amplification and uses them to implement computational instructions that provide parallel searches and arithmetic.
[0007] In one aspect, the present disclosure provides a method for encoding digital information into a nucleic acid sequence. The method includes encoding the digital information into a symbol sequence, converting the symbol sequence into a codeword, and parsing the codeword into a coded symbol sequence. The method includes mapping the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences. The method includes enumerating an identifier library, each symbol of the coded symbol sequence being encoded by one or more identifiers. Each identifier of the plurality of identifiers comprises a plurality of components from one or more layers, each layer of the one or more layers comprising a distinct set of components. Each identifier of the identifier library comprises one or more data layers encoding symbols and one or more key layers encoding keys.
[0008] In one aspect, the present disclosure provides a nucleic acid-based integrated storage system. The system includes a data encoding unit configured to write digital information to one or more nucleic acid sequences, the data encoding unit writing the digital information to the one or more nucleic acid sequences in the absence of base-by-base nucleic acid synthesis. The system includes a storage unit configured to store the one or more nucleic acid sequences encoding the digital information. The system includes a read unit configured to access and read the digital information encoded in the one or more nucleic acid sequences. The system includes one or more computer processors operably coupled to the data encoding unit, the storage unit, and the read unit. The one or more computer processors are individually or collectively programmed to (i) direct the data encoding unit to encode the digital information into the one or more nucleic acid sequences, (ii) direct the storage unit to store the digital information encoded in the one or more nucleic acid sequences, and (iii) direct the read unit to access and decode the digital information stored in the one or more nucleic acid sequences. Encoding includes encoding the digital information into a symbol sequence, converting the symbol sequence into a codeword, and parsing the codeword into a coded symbol sequence. Encoding includes mapping the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences. Encoding includes enumerating an identifier library, each symbol of the coded symbol sequence being encoded by one or more identifiers. Each identifier of the plurality of identifiers comprises multiple components from one or more layers, each layer of the one or more layers comprising a distinct set of components. Each identifier of an identifier library comprises one or more data layers encoding symbols and one or more key layers encoding keys.
[0009]
[0009] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description should be regarded as illustrative in nature and not as restrictive.
[0010] Incorporation by Reference
[0010] All published patents, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual published patent, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that published patents, patents, or patent applications incorporated by reference conflict with the disclosure contained herein, it is intended that the present specification supersede and / or take precedence over any such conflicting matter.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Brief explanation of the drawings]
[0012] [Figure 1]
[0012] A schematic overview of the process for encoding, writing, accessing, reading and decoding digital information stored in nucleic acid sequences is presented. [Figure 2A]
[0013] 1 illustrates, in simplified form, how objects or identifiers (e.g., nucleic acid molecules) can be used to encode digital data, referred to as "data at an address," showing the combination of a rank object (or address object) with a byte value object (or data object) to create an identifier. [Figure 2B]
[0013] This paper outlines a method for using objects or identifiers (e.g., nucleic acid molecules) to encode digital data, referred to as "data-at-addresses," and illustrates one embodiment of the data-at-address method in which rank objects and byte value objects are themselves combinatorial concatenations of other objects. [Figure 3A]
[0014] FIG. 1 is a schematic diagram illustrating an example method for encoding digital information using objects or identifiers (e.g., nucleic acid sequences), showing the encoding of digital information using rank objects as identifiers. [Figure 3B]
[0014] A diagram illustrating an example method for encoding digital information using objects or identifiers (e.g., nucleic acid sequences), showing an embodiment of the encoding method in which the address object is itself a combinatorial concatenation of other objects. [Figure 4]
[0015] 1 shows a schematic overview of a method for writing information into a nucleic acid sequence (e.g., deoxyribonucleic acid). [Figure 5]
[0016] 1 illustrates a schematic representation of an example combinatorial space of identifiers organized as an m-level n-variable tree. [Figure 6]
[0017] 1 illustrates an example of a method for minimizing the number of identifiers structured to write a bitstream; [Figure 7]
[0018] 1 illustrates a schematic diagram of an example method for remapping words to codewords to ensure uniform weight codewords for error detection. [Figure 8]
[0019] 1 illustrates a schematic of an example method for minimizing write time by generating a minimal set of reactions. [Figure 9]
[0020] 1 illustrates schematically an isometric mapping of addresses to identifiers and a dual encoding of data. [Figure 10]
[0021] 1 illustrates a schematic diagram of an example of a method for masking encoding and decoding to protect against unauthorized decoding. [Figure 11]
[0022] 1 shows an example component carousel. [Figure 12]
[0023] 1 illustrates schematically how electrowetting can be used for component calculations. [Figure 13]
[0024] 1 illustrates an example print-based method for dispensing components. [Figure 14]
[0025] 1 shows an example of microfluidic injection of components. [Figure 15]
[0026] An example of selective condensation of component mists is shown. [Figure 16]
[0027] 1 shows a schematic diagram of an example of a method for producing an identifier by weaving or knitting. [Figure 17]
[0028] 1 illustrates generally one example of a method for generating an identifier from a set of components. [Figure 18]
[0029] 1 shows a schematic diagram of an example of a method for producing an identifier from a separate film or thread. [Figure 19]
[0030] 1 illustrates a schematic diagram of an example of a method for using subtraction to write information. [Figure 20]
[0031] An example of a method for reading by hybridization is shown schematically. [Figure 21]
[0032] 10 illustrates a schematic of an approach for selecting identifiers that match a query in a subset of layers. [Figure 22]
[0033] 10 illustrates a schematic of a technique for modifying selected identifiers for computational purposes. [Figure 23]
[0034] 1 shows a schematic diagram of an example of a method for reading by nanopore sequencing. [Figure 24]
[0035] 1 illustrates a computerized control system programmed or otherwise configured to carry out the methods provided herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description
[0036] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein are available.
[0014]
[0037] The term "digital message," as used herein, generally refers to a symbol sequence provided for encoding into a nucleic acid molecule. A digital message can be original text written into a nucleic acid molecule.
[0015]
[0038] The term "symbol," as used herein, generally refers to a representation of a unit of digital information. Digital information can be divided or converted into strings of symbols. In one example, a symbol can be a bit, and a bit can have a value of "0" or "1."
[0016]
[0039] The terms "distinct" or "unique," as used herein, generally refer to an entity that can be distinguished from other entities in a group. For example, a distinct or unique nucleic acid sequence can be a nucleic acid sequence that does not have the same sequence as any other nucleic acid sequence. A distinct or unique nucleic acid molecule does not have the same sequence as any other nucleic acid molecule. A distinct or unique nucleic acid sequence or molecule can share regions of similarity with another nucleic acid sequence or molecule.
[0017]
[0040] The term "component," as used herein, generally refers to a nucleic acid sequence. A component can be a separate sequence. A component can be linked or assembled with one or more other components to produce another nucleic acid sequence or molecule.
[0018]
[0041] The term "stratum," as used herein, generally refers to a group or pool of components. Each stratum may contain a distinct set of components such that the components in one stratum differ from the components in another stratum. Components from one or more stratums may be assembled to generate one or more identifiers.
[0019]
[0042] The term "identifier," as used herein, generally refers to a nucleic acid molecule or sequence that represents the position and value of a bit string within a larger bit string. More generally, an identifier may refer to any object that represents or corresponds to a symbol in a symbol string. In some embodiments, an identifier may include one or more concatenated components.
[0020]
[0043] The term "combinatorial space," as used herein, generally refers to the set of all possible distinct identifiers that can be generated from a starting set of objects, such as components, and an allowable set of rules for how to modify those objects to form identifiers. The size of the combinatorial space of identifiers created by assembling or linking components can depend on the number of layers of components, the number of components in each layer, and the particular assembly method used to generate the identifiers.
[0021]
[0044] The term "identifier rank," as used herein, generally refers to a relationship that defines the ordering of identifiers within a set.
[0022]
[0045] The term "identifier library," as used herein, generally refers to a collection of identifiers that correspond to symbols in a string representing digital information. In some embodiments, the absence of a given identifier in an identifier library may indicate a symbol value at a particular location. One or more identifier libraries may be combined in a pool, group, or set of identifiers. Each identifier library may include a unique barcode that identifies the identifier library.
[0023]
[0046] The term "universal library," as used herein, generally refers to a collection of identifiers corresponding to the set of all possible distinct identifiers that can be generated from a starting set of objects, such as components, and a set of allowable rules for how to modify those objects to form the identifiers.
[0024]
[0047] The term "word," as used herein, generally refers to a block of a string of symbols. The length of the block may or may not be constant. A string of symbols may be divided into one or more words that are L symbols long. In one example, a string of length 16 symbols may be divided into four words, each four symbols long.
[0025]
[0048] The term "codeword," as used herein, generally refers to a string of symbols that encodes a word. The length of the string may or may not be constant. A source bitstream may be parsed into words, which are subsequently converted into codewords using a codebook. The codebook may correlate words to codewords. Codewords may be selected to reduce writing time, minimize identifier construction, or detect writing errors.
[0026]
[0049] The term "nucleic acid," as used herein, generally refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or variants thereof. Nucleic acids can include one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Nucleotides can include A, C, G, T, U, or variants thereof. Nucleotides can include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits can be specific to A, C, G, T, U, or one or more complementary A, C, G, T, or U, or any other subunit that is complementary to a purine (i.e., A or G or variants thereof) or pyrimidine (i.e., C, T, or U or variants thereof). In some examples, nucleic acids can be single-stranded or double-stranded, and in some cases, the nucleic acid molecule is circular.
[0027]
[0050] The terms "nucleic acid molecule" or "nucleic acid sequence," as used herein, generally refer to a polymeric form of nucleotides or polynucleotides, which may be of various lengths, either deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs thereof. Oligonucleotide, as used herein, generally refers to a single-stranded nucleic acid sequence, typically composed of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), guanine (G), and thymine (T) (though uracil (U) replaces thymine (T) when the polynucleotide is RNA). The term "nucleic acid sequence" can refer to the alphabetical representation of a polynucleotide molecule; alternatively, the term can apply to the physical polynucleotide itself. This alphabetical representation can be entered into a database in a computer having a central processing unit, which can map the nucleic acid sequence or molecule to symbols or bits and use them to encode digital information. A nucleic acid sequence or oligonucleotide may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0028]
[0051] Examples of modified nucleotides include, but are not limited to, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylketone, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5- Examples include methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylqueosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid(v), wybutoxosin, pseudouracil, queosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine. Nucleic acid molecules may have modified base moieties (e.g., one or more atoms normally available to form hydrogen bonds with a complementary nucleotide and / or one or more atoms normally incapable of forming hydrogen bonds with a complementary nucleotide), modified sugar moieties, or modified phosphate backbones. Nucleic acid molecules may also contain amine-modified groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexylacrylamide-dCTP (aha-dCTP), to allow covalent attachment of amine-reactive moieties, such as N-hydroxysuccinimide ester (NHS).
[0029]
[0052] The term "primer," as used herein, generally refers to a nucleic acid strand that serves as the starting point for nucleic acid synthesis, such as in a polymerase chain reaction (PCR). In one example, during replication of a DNA sample, an enzyme that catalyzes replication initiates replication at the 3' end of a primer bound to the DNA sample and copies the opposite strand.
[0030]
[0053] The term "polymerase" or "polymerase enzyme," as used herein, generally refers to any enzyme capable of catalyzing a polymerase reaction. Examples of polymerases include, but are not limited to, nucleic acid polymerases. Polymerases can be naturally occurring or synthetic. An example of a polymerase is Φ29 polymerase or a derivative thereof. In some cases, a transcriptase or a ligase (i.e., an enzyme that catalyzes the formation of a bond) is used in conjunction with or as a substitute for a polymerase to construct new nucleic acid sequences. Examples of polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, and bacteriophage T4. Examples of DNA polymerases include Φ29 (Phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, platinum Taq polymerase, Tbr polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3' to 5' exonuclease activity, and variants, modified products, and derivatives thereof.
[0031]
[0054] Digital information, such as computer data in the form of binary code, may include a sequence or string of symbols. Binary code may, for example, encode or represent text or computer processor instructions using a binary system with two binary symbols called bits, typically 0 and 1. Digital information may be represented in the form of a non-binary code, which may include a sequence of non-binary symbols. Each encoded symbol may be reassigned to a unique bit string (or "byte"), and the unique bit string or byte may be arranged into a byte string or byte stream. The bit value for a given bit may be one of two symbols (e.g., 0 or 1). A byte may include a string of N bits, for a total of 2 N For example, a byte containing 8 bits can have a total of 2 8 Or 256 possible unique byte values can result, where each of the 256 bytes can correspond to one of the 256 possible distinct symbols, characters, or instructions that can be encoded in the byte. Raw data (e.g., text files and computer instructions) can be represented as a string of bytes or a byte stream. Zip files or compressed data files containing raw data can also be stored as a byte stream; these files can be stored in a compressed format as a byte stream and then restored to raw data before being read by a computer.
[0032]
[0055] The disclosed methods and systems may be used to encode computer data or information into multiple identifiers, each of which may represent one or more bits of the original information. In some examples, the disclosed methods and systems encode the data or information using identifiers, each of which represents two bits of the original information.
[0033]
[0056] Traditional methods of encoding digital information into nucleic acids rely on base-by-base synthesis of nucleic acids, which can be costly and time-consuming. Alternative methods may improve the commercial viability of digital information storage by reducing the reliance on base-by-base nucleic acid synthesis to encode digital information, eliminating the need for de novo synthesis of separate nucleic acid sequences for every new information storage requirement.
[0034]
[0057] Instead of relying on base-by-base or de novo nucleic acid synthesis (e.g., phosphoramidite synthesis), the new methods can encode digital information (e.g., binary code) into multiple identifiers or nucleic acid sequences that comprise combinatorial arrangements of components. Thus, the new strategies can produce a first set of distinct nucleic acid sequences (or components) for a first request for information storage, and then reuse the same nucleic acid sequences (or components) for subsequent information storage requests. These approaches can significantly lower the cost of DNA-based information storage by reducing the role of de novo synthesis of nucleic acid sequences in the process of encoding and writing information into DNA. Furthermore, unlike implementations of base-by-base synthesis, such as phosphoramidite chemistry or template-free polymerase-based nucleic acid extension, which require cyclic delivery of each base to each extending nucleic acid, the new methods of writing information into DNA, using identifier construction from components, are highly parallelizable processes that cannot use cyclic nucleic acid extension. Therefore, the new methods can increase the speed at which digital information can be written into DNA compared to older methods.
[0035] Methods for encoding and writing information into nucleic acid sequences
[0058] In one aspect, the present disclosure provides a method for encoding a symbol sequence for writing into a nucleic acid sequence. The method for encoding a symbol sequence for writing into a nucleic acid sequence may include (a) converting the symbol sequence into a codeword using one or more codebooks, (b) parsing the codeword into a coded symbol sequence, (c) mapping the coded symbol sequence to a plurality of identifiers, (d) generating an identifier library, and (e) attaching a description of the one or more codebooks and the plurality of identifiers to the identifier library. Each symbol in the coded symbol sequence may be encoded by one or more identifiers.
[0036]
[0059] FIG. 1 illustrates the overall process of encoding information into a nucleic acid sequence, writing information into a nucleic acid sequence, reading the information written into the nucleic acid sequence, and decoding the read information. Digital information or data can be translated into one or more strings of symbols. In one example, the symbols are bits, each having a value of either "0" or "1." Each symbol can be mapped or encoded into an object (e.g., an identifier) that represents that symbol. Each symbol can be represented by a separate identifier. The separate identifiers can be nucleic acid molecules composed of components. The components can be nucleic acid sequences. Digital information can be written into a nucleic acid sequence by generating an identifier library corresponding to the information. The identifier library can be physically created by physically constructing an identifier corresponding to each symbol in the digital information. Every portion of the digital information can be accessed at one time. In one example, a subset of identifiers is accessed from the identifier library. The subset of identifiers can be read by sequencing and identifying the identifiers. To decode the digital data, the identified identifiers can be associated with corresponding symbols. FIG. 1 illustrates the overall process of encoding information into a nucleic acid sequence, writing information into a nucleic acid sequence, reading the information written into the nucleic acid sequence, and decoding the read information without using base-by-base synthesis. Digital information or data can be translated into one or more symbol strings. In one example, the symbols are bits, each bit having a value of either "0" or "1." Each symbol can be mapped or encoded to a physical object (e.g., an identifier) that represents that symbol. Each symbol can be represented by a separate identifier. The separate identifier can be a nucleic acid molecule composed of components. The components can be nucleic acid sequences. Digital information can be written into a nucleic acid sequence by generating an identifier library corresponding to the information. The identifier library can be generated by assembling identifiers corresponding to each symbol of the digital information. All or a portion of the digital information can be accessed at one time. In one example, a subset of identifiers is removed from the identifier library. The subset of identifiers can be read by identifying the identifiers. The identified identifiers can be associated with corresponding symbols to decode the digital data.
[0037]
[0060] A method for encoding and reading information using the approach of Figure 1 can include, for example, receiving a bitstream. This can include using the identifier rank to map each bit (a bit with a bit value of "1") in the bitstream to a distinct nucleic acid identifier. Constructing a nucleic acid sample pool or identifier library that includes copies of identifiers corresponding to bit values of 1 (and excludes identifiers with a bit value of 0). Reading the sample can include using molecular biology methods (e.g., sequencing, hybridization, PCR, etc.), determining which identifiers are represented in the identifier library, and assigning a bit value of "1" to the bits corresponding to those identifiers and a bit value of "0" elsewhere (again, referring to the identifier rank to identify the bit in the original bitstream to which each identifier corresponds), thus decoding the information into the encoded original bitstream.
[0038]
[0061] Encoding a string of N distinct bits allows an equal number of unique nucleic acid sequences to be used as possible identifiers. This approach to encoding information may use de novo synthesis of an identifier for each new item of information (string of N bits) to be stored. In other cases, the cost of synthesizing identifiers (N or fewer in number) de novo may be reduced by performing de novo synthesis once and continually maintaining all possible identifiers, such that encoding a new item of information may involve mechanically selecting pre-synthesized (or pre-made) identifiers and mixing them together to form an identifier library. In other cases, synthesizing and maintaining several (fewer than N, and in some cases much less than N) nucleic acid sequences and then modifying these sequences through enzymatic reactions to generate up to N identifiers for each new item to be stored may reduce both or any combination of (1) the cost of de novo synthesizing up to N identifiers for each new item of information to be stored, and (2) the cost of maintaining and selecting from N possible identifiers for each new item of information to be stored.
[0039]
[0062] Identifiers may be rationally designed and selected to facilitate read, write, access, copy, and delete operations. Identifiers may be designed and selected to minimize write errors, mutations, decomposition, and read errors.
[0040]
[0063] 2A and 2B show an example of a method called "data-at-address" for encoding digital data into objects or identifiers (e.g., nucleic acid molecules). FIG. 2A shows encoding a bitstream into an identifier library, where each identifier is constructed by concatenating a single component that specifies an identifier rank with a single component that specifies a byte value. In general, the data-at-address method uses identifiers to modularly encode information by including two objects: one object is a "byte-value object" (or "data object") that identifies a byte value, and one object is a "rank object" (or "address object") that identifies the identifier rank (or relative position of the byte within the original bitstream). FIG. 2B shows an example of the data-at-address method, where each rank object is constructed combinatorially from a set of components, and each byte-value object can be constructed combinatorially from a set of components. Such a combinatorial structure of rank and byte-value objects allows more information to be written into an identifier than would be possible if the objects were made from only a single component (e.g., FIG. 2A).
[0041]
[0064] 3A and 3B schematically illustrate another example of a method for encoding digital information into an object or identifier (e.g., a nucleic acid sequence). FIG. 3A illustrates encoding a bitstream into an identifier library, where the identifier is constructed from a single component that specifies an identifier rank. The presence of the identifier at a particular rank (or address) specifies a bit value of "1," and the absence of the identifier at a particular rank (or address) specifies a bit value of "0." This type of encoding uses identifiers that encode only their rank (the relative position of the bit in the original bitstream), and the presence or absence of those identifiers in the identifier library can be used to encode a bit value of "1" or "0," respectively. Reading and decoding the information can involve identifying identifiers present in the identifier library and assigning a bit value of "1" to their corresponding ranks and a bit value of "0" for identifiers not present in the identifier library. FIG. 3B illustrates an example of an encoding method in which each identifier can be combinatorially constructed from a set of components, with each possible combinatorial construction specifying a rank. Such a combinatorial structure allows more information to be written into an identifier than if the identifier were made from only a single component (e.g., FIG. 3A). For example, a component set may include five distinct components. The five distinct components may be assembled to generate ten distinct identifiers, each containing two of the five components. The ten distinct identifiers may each have a rank (or address) corresponding to the position of a bit in the bit stream. The identifier library may include a subset of those ten possible identifiers that correspond to positions of bit value "1" within a bit stream of length 10, and may exclude a subset of those ten possible identifiers that correspond to positions of bit value "0."
[0042]
[0065] Figure 4 shows an overall method for writing information into a nucleic acid sequence. Before writing information, the information may be translated into a symbol string and encoded into multiple identifiers. Writing information may include preparing reactions to produce possible identifiers. Reactions may be prepared by depositing inputs into compartments. The inputs may include nucleic acids, components, enzymes, or chemical reagents. Compartments may be wells, tubes, locations on a surface, chambers in a microfluidic device, or droplets in an emulsion. Multiple reactions may be prepared in multiple compartments. In one example, one or more reactions may be prepared to generate a universal library. Reactions may proceed to generate identifiers through programmed temperature incubation or cycling. Reactions may be selectively or universally removed (e.g., deleted). Reactions may also be selectively or universally interrupted, combined, and refined to collect identifiers into a single pool. Identifiers from multiple identifier libraries may be collected into the same pool. Individual identifiers may include a barcode or tag to identify the identifier library to which they belong. Alternatively or additionally, the barcode may contain metadata for the encoded information. The supplemental nucleic acids or identifiers may be included in an identifier pool along with an identifier library. The supplemental nucleic acids or identifiers may contain metadata for the encoded information or may function to obscure the encoded information.
[0043]
[0066] The identifier rank may include a method for determining the ordering of identifiers. The method may include a lookup table with all identifiers and their corresponding ranks. The method may also include a lookup table with the ranks of all components that make up the identifier and a function for determining the ordering of any identifier that includes a combination of those components. Such a method may be called lexicographic ordering and may be similar to the way words in a dictionary are ordered alphabetically. In a data-at-address encoding method, the identifier rank (encoded by the identifier's rank object) may be used to determine the location of a byte (encoded by the identifier's byte value object) in a bitstream. In one example encoding method, the identifier rank of the current identifier (encoded by the entire identifier itself) may be used to determine the location of a bit value "1" in a bitstream.
[0044]
[0067] Identifiers can be constructed by combinatorially assembling component nucleic acid sequences. For example, information can be encoded by taking a set of nucleic acid molecules (e.g., identifiers) from a defined group of molecules (e.g., combinatorial space). Each possible identifier for a defined group of molecules can be an assembly of nucleic acid sequences (e.g., components) from a set of pre-made components that can be separated into layers. Each individual identifier can be constructed by linking one component from every layer in a fixed order. For example, if there are M layers, each layer having n components, then the maximum C=n M A unique identifier can be constructed and up to 2 C different items or C bits can be encoded and stored. For example, storing 1 megabit of information requires 1 x 10 6 distinct identifiers or size C = 1 x 10 6 In this example, the identifier can be assembled from a variety of components organized in different ways. The assembly can be made from M=2 prefabricated layers, each layer having n=1×10 3 Alternatively, the assembly can be made of M=3 layers, each layer containing n=1×10 2components. As this example shows, it may be possible to use more layers to encode the same amount of information, resulting in a smaller total number of components. Using fewer components overall may be advantageous in terms of write costs.
[0045]
[0068] In one example, one can start with two layers, X and Y, each layer having x and y nucleic acid sequences (e.g., components), respectively. Each nucleic acid sequence from X can be assembled with each nucleic acid sequence from Y. The total number of nucleic acid sequences maintained in the two sets can be the sum of x and y, while the total number of nucleic acid molecules, and therefore possible identifiers, that can be generated can be the product of x and y. If sequences from X can be assembled with sequences of Y in any order, even more nucleic acid sequences (e.g., identifiers) can be generated. For example, if the assembly order can be programmable, the number of nucleic acid sequences (e.g., identifiers) generated can be twice the product of x and y. This set of all possible nucleic acid sequences that can be generated can be referred to as XY. The order of assembled units of unique nucleic acid sequences within XY can be controlled using nucleic acids with distinct 5' and 3' ends, and restriction digestion, ligation, polymerase chain reaction (PCR), and sequencing can be performed on the distinct 5' and 3' ends of the sequences. Such an approach can reduce the total number of nucleic acid sequences (e.g., components) used to encode N distinct bits by encoding the information in an assembly product combination and order. For example, to encode 100 bits of information, two layers of 10 distinct nucleic acid molecules (e.g., components) can be assembled in a fixed order to encode 10 bits of information. * 10, or 100 distinct nucleic acid molecules (e.g., identifiers) can be generated, or one layer of 5 distinct nucleic acid molecules (e.g., components) and another layer of 10 distinct nucleic acid molecules (e.g., components) can be assembled in any order to generate 100 distinct nucleic acid molecules (e.g., identifiers).
[0046]
[0069] The nucleic acid sequences (e.g., components) within each layer may include a central unique (or distinct) sequence or barcode, a common hybridization region at one end, and another common hybridization region at the other end. The barcode may include a sufficient number of nucleotides to uniquely identify every sequence within the layer. For example, there are typically four possible nucleotides for each base position in the barcode. Thus, a three-base barcode may be a four-base barcode. 3 = 64 nucleic acid sequences. Barcodes can be designed to be randomly generated. Alternatively, barcodes can be designed to avoid sequences that may introduce complexity into the chemistry of the identifier or sequencing structure. Additionally, barcodes can be designed so that each barcode has a minimum Hamming distance from other barcodes, thereby reducing the likelihood that variations in base resolution or reading errors can interfere with proper identification of the barcode.
[0047]
[0070] The hybridization region at one end of a nucleic acid sequence (e.g., a component) can be different within each layer but the same for each member within a layer. Adjacent layers have complementary hybridization regions for their components and can interact with each other. For example, any component from layer X can have complementary hybridization regions and thus be able to attach to any component from layer Y. The hybridization region at the opposite end can serve the same purpose as the hybridization region at the first end. For example, any component from layer Y can attach to any component of layer X at one end and to any component of layer Z at the opposite end.
[0048]
[0071] Combinatorial assembly of two or more components, each from a different layer (e.g., X, Y, or Z) to construct an identifier can be achieved using polymerase chain reaction (PCR), ligation, or recombination. Generally, any method of linking two or more distinct nucleic acid sequences can be used to construct identifiers in an identifier library. In some cases, all or a portion of the combinatorial space of possible identifiers can be constructed before digital information can be encoded or written, in which case the writing process can involve mechanically selecting and pooling identifiers (encoding targeted information) from an already existing set. In other cases, identifiers can be constructed after one or more steps of the data encoding or writing process have occurred (i.e., as information is written). Methods for constructing identifiers include, but are not limited to, overlap extension polymerase chain reaction (PCR) (or polymerase cycling assembly), sticky end ligation, recombinase assembly, template-directed ligation (or bridge strand ligation), BioBrick assembly, Golden Gate assembly, Gibson assembly, and ligase cycling reaction assembly. Methods for constructing an identifier can also include deleting nucleic acid sequences (e.g., components) from or inserting nucleic acid sequences (e.g., components) into a parent nucleic acid sequence (or parent identifier). In one example, an identifier can be generated from a parent identifier that is composed of multiple components. Components can be cleaved from or inserted into a parent identifier to generate a unique identifier. Enzymes that modify the parent identifier can include double-strand-specific nucleases, single-strand-specific nucleases, and Cas9.
[0049]
[0072] Enzymatic reactions can be used to assemble components from different layers. Because components of each layer have specific hybridization or attachment regions for components of adjacent layers, assembly can occur in a one-pot reaction. For example, nucleic acid sequence (e.g., component) X1 from layer X, nucleic acid sequence Y1 from set Y, and nucleic acid sequence Z1 from set Z can form an assembled nucleic acid molecule (e.g., identifier) X1Y1Z1. In addition, multiple nucleic acid molecules (e.g., identifiers) can be assembled in a single reaction by including multiple nucleic acid sequences from each layer. For example, including both Y1 and Y2 in the one-pot reaction of the previous example can result in two assembled products (e.g., identifiers), namely, X1Y1Z1 and X1Y2Z1. This reaction multiplexing can be used to accelerate writing times if multiple identifiers can be physically constructed. Assembly of nucleic acid sequences can be performed within a time period of about 1 day, 12 hours, 10 hours, 9 hours, 8 hours, 7 hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours, or 1 hour or less. The accuracy of the encoded data may be at least about 90%, 95%, 96%, 97%, 98%, 99% or higher.
[0050]
[0073] Writing information into a nucleic acid sequence may include parsing the information into a symbol string, mapping the symbol string to a unique identifier, and generating an identifier library containing identifiers corresponding to the symbol string. The identifier library may include an identifier for each identifier rank or may exclude identifiers for an identifier rank if they correspond to a selected symbol value (e.g., 0 or 1). The information may include a symbol string. In one example, the symbol string includes symbols taken from a fixed, finite alphabet of symbols. The string may be converted into a second symbol sequence. The second symbol sequence may include a formal data structure. The second symbol sequence may be parsed into words. The words may be converted into codewords using a codebook. The codebook may be an explicit codebook or an implicit codebook. The codewords may be parsed into a third symbol sequence. Each symbol in the third symbol string may be mapped to a unique identifier. A set of identifiers (e.g., an identifier library) may be defined such that each symbol can be encoded into one or more identifiers. A set of identifiers (eg, an identifier library) may include or be accompanied by information relating to one or more codebooks, data structures, and combinatorial spaces.
[0051]
[0074] The formal data structure may include a tree, a trie, a table, a set, a key-value dictionary, or a set of multidimensional vectors. The formal data structure may be queriable by one or more query types, including, but not limited to, a range query, a rank query, a count query, a membership query, a nearest neighbor query, a match query, a selection query, or any combination thereof. The second symbol sequence having the formal data structure may be parsed into a word sequence to minimize the number of identifiers used to encode the bitstream. Each bit of the source bitstream may be associated with an identifier in the combination space.
[0052]
[0075] The identifier combination space may include unique identifiers that can be generated by one or more construction algorithms from a library of T total components. In one embodiment, the construction algorithm may generate identifiers using a Cartesian product scheme with M layers, where the i-th layer includes Ni components. The number of identifiers in the combination space may depend on the number of layers, the number of components in each layer, and the method used to assemble the identifiers. FIG. 5 shows an example of an identifier combination space using a product scheme with M layers and N components in each layer. In this example, M=4 and N=2. Items 501-504 in FIG. 5 represent layers in this example. Items 511 and 512 represent two components in layer 1 in this example. Similarly, items 509-510, 507-508, and 505-506 represent components belonging to layers 2, 3, and 4. The components are placed in a repeating pattern to illustrate the 16 distinct identifier combination spaces that result from this scheme. The steps in one example of a combinatorial algorithm to generate each identifier in the combinatorial space can be shown as a tree diagram, shown in item 513. The tree diagram can be divided into M layers. Each layer contains nodes representing the available choices for ingredients in that layer. For example, in layer 1, two arrows emanating from the node labeled "a" represent the two ingredient choices in layer 1, shown by items 511 and 512. In layer 2, an arrow emanating from node b represents the ingredient choices in layer 2, shown as elements 509 and 510, conditional on ingredient 511 being selected in layer 1. The left and right arrows emanating from each node correspond to the ingredient pattern shown in the layer in item 515. The arrows emanating from each node are ordered according to the ingredient ranking defined in the production scheme. Each path down the tree diagram, starting from the top node labeled "a" to any node below, corresponds to a distinct identifier. One such path is shown by item 514. The combinatorial space of all 16 identifiers in this example is shown by item 518. Item 517 shows one bit value in an example bitstream that may be encoded using this combination space, with each bit in the bitstream corresponding to a distinct identifier shown below that bit.In one embodiment, the value of a bit is represented by including or excluding its identifier from a constructed identifier library. To encode a bitstream, all identifiers corresponding to bits having a value of "1" can be constructed and pooled, while those corresponding to bits having a value of "0" can be excluded. The excluded identifiers are marked using a dark overlay, and item 519 indicates one such excluded identifier corresponding to the 10th bit having a value of "0".
[0053]
[0076] Information can be encoded into identifiers using data in an addressing scheme abbreviated as the DAA scheme. The source bitstream can be divided into words of fixed length L. The bitstream can then be interpreted as a symbol stream of L-bit symbols (e.g., each symbol containing L bits). Unique identifiers can be constructed for each symbol in the symbol stream (i.e., each symbol containing L bits) and pooled or grouped together. In one embodiment, the identifiers can be constructed using a production method that includes M layers with N components per layer. Each identifier can be contained in two parts (or objects). The first part can include up to k < M layers and can provide information regarding the address of the symbol. The second part of the unique identifier can include components from the M - k layers and can provide information regarding the value of the symbol. Alternatively or in addition, the source bitstream can be divided into a stream of words of length L bits. Using a codebook, the words can be mapped to codewords over a nucleic acid alphabet containing the four bases A, T, C, and G. Each codeword can be constructed of the four bases. The identifier for each L-bit word can be constructed by assembling or concatenating the corresponding synthesized codeword with components that specify the address of that codeword.
[0054]
[0077] Before writing the source bitstream to the identifier library, the source bitstream may be encoded into an intermediate bitstream. The source bitstream may be divided into words. Another codeword may be chosen to replace the word. The length of the codeword may be longer, equal to, or shorter than the length of the corresponding word. In one embodiment, each word X containing some number N(X) of Y symbols may be replaced with a codeword containing fewer or more Y symbols. For example, a word containing N(X) "1" symbols may be replaced with a codeword containing fewer than N(X) "1" symbols. In one example encoding method, this may reduce the size of the identifier library used to encode a given piece of digital information. By minimizing the number of identifiers physically assembled, the time to write information to the identifiers and the time to read the information encoded in the identifiers may be reduced. Figure 6 schematically illustrates an example of a method for minimizing the number of identifiers constructed for writing to a bitstream using extended codewords. The bitstream may be divided into words, and in this example, each word may be of a fixed length of 2 bits. A list of words containing two bits includes "00," "01," "10," and "11." Each word can appear zero or more times in the bit stream. For example, the bit stream "0110101010011101" can be divided into the two-bit words {01, 10, 10, 10, 10, 01, 11, 01}, where the "00" word appears zero times, the "01" word appears three times, the "10" word appears four times, and the "11" word appears once. The total number of "1" symbols in this word sequence is nine, indicating that the encoding method may require the assembly of nine separate identifiers to represent the bit stream. However, the words can be reordered so that fewer identifiers can be used to encode a given bit stream.
[0055]
[0078] Digital information encoded into nucleic acids may first be converted into a symbol sequence and then reorganized into a formal data structure suitable for one or more query types. This data structure may then be serialized into a second symbol string. This second symbol string may be coded using one or more codebooks for one or more purposes, including error protection, encryption, write speed optimization, or identifier library size minimization. FIG. 6 illustrates an example of a method for minimizing the size of an identifier library. Item 620 illustrates a tree diagram representation of the combinatorial space, the notation of which was explained in FIG. 5. In this example, item 621 indicates bit values from a bit stream of 16-bit values. Item 622 indicates a set of identifiers corresponding to bit values in the bit stream having a value of "1." Thus, in its current state, coding may require assembling nine separate identifiers corresponding to the nine bits having a value of "1." However, the size of this identifier library may be reduced by re-encoding the bit stream using a code book that maps 2-bit words to 3-bit code words, such that the new 3-bit code words have fewer "1" symbols, resulting in a smaller identifier library.
[0056]
[0079] In one example of this re-encoding method, the bitstream may be divided into eight consecutive 2-bit words and the number of occurrences of each 2-bit word may be recorded. In this example, these counts are shown under the count column in table 623. All possible 3-bit codewords are listed as columns to form a matrix, with cell (i,j) containing the cost of mapping 2-bit word i to distinct 3-bit codeword j. This cost may be calculated by multiplying the number of "1" symbols in the codeword by the number of occurrences of the word in the original bitstream and calculating the number of identifiers that can be used to construct using this word-to-codeword permutation. For example, the word "01" occurs three times in the original bitstream. If it is mapped to the codeword "111," the number of "1" symbols in the re-encoded bitstream resulting from this permutation may increase from 3 to 12. These costs are calculated for all such possible permutations. The resulting matrix, indicated by item 623, can be translated into a weighted bipartite graph, and a minimum-weight perfect matching can be obtained using an algorithm such as the Kuhn-Munkers algorithm. A minimum perfect matching can be equivalent to choosing exactly one cell in each row and each column in matrix 623 such that the sum of all the chosen cells is minimized. The cost of each cell in one such minimum re-encoding is shown in table 623 using shaded cells. In this minimum re-encoding, word "00" maps to codeword "011," "01" maps to "001," "10" maps to "000," and "11" maps to "010." The new bitstream thus coded has a total of four "1" symbols. Therefore, the cost can be reduced from 9 in the original bitstream to 4 in the newly re-encoded bitstream. The new bitstream contains 3-bit codewords, indicated in the tree diagram by item 624. Each 3-bit codeword uniquely maps to a 2-bit codeword from the original set of 2-bit codewords indicated by item 625 .Item 626 shows the new identifier library to be assembled.
[0057]
[0080] The choice of symbols used to encode digital information can enable the detection and / or correction of coding errors. Re-encoding a symbol stream to include error protection symbols calculated from symbols in the original string can enable the detection or correction of errors that occurred during the process of writing the symbol stream using nucleic acids. In one embodiment, a symbol stream can be divided into fixed-length words, and one or more error protection symbol strings can be calculated from each such word and appended to the word to obtain a recoded string. For example, the number of identifiers constructed in a fixed-length block of K identifiers can be counted. If this count is even, an additional identifier can be added to the block; if the count is odd, such additional identifiers need not be added. The combination space can be chosen to accommodate these additional identifiers. When such an identifier block is read, any write errors in which an identifier is mistakenly omitted or an additional identifier is mistakenly added can be detected, since such an event would negate the required property that each block has an odd number of identifiers. In another embodiment, the number of identifiers in any fixed-length block of K identifiers is counted and K minus the count is calculated. This value, called the error protection value, can be attached to the block and coded. The combination space can be chosen to accommodate the identifiers corresponding to these error protection values. In this case, when the block and error protection value are read, any errors in which the identifier was accidentally omitted can be detected. If the omitted identifier could be in the original block, this can be reflected by a mismatch in the error protection value. If the omitted identifier is in the error protection value, a lower value can indicate that an error may exist in the error protection value. If there is an error in both the block and the value, the mismatch can lead to the detection of an error. In another embodiment, the symbol stream can be divided into fixed-length words of W symbols. Each word can then be remapped to a codeword such that each codeword leads to the construction of an identifier of fixed length V. Figure 7 illustrates this uniform-weight codeword error detection scheme. Item 727 indicates an identifier library that can be constructed to encode the bitstream shown in the tree diagram in Figure 7.In the original bitstream, for any constant word length W, the number of identifiers is not constant: for example, if W=2, there may be one identifier in each of the first six words and two identifiers in the second word. Table 727 shows an example re-encoded codebook that maps words of length W=2 to codewords of length V=4. The example codebook maps the words "00," "01," "10," and "11" to the codewords "0011," "0101," "0110," and "1001," respectively. Because every codeword has exactly two "1" symbols and the word and codeword lengths are constant, the resulting bitstream has exactly two "1" symbols for any codeword length of four symbols. This is illustrated in the example tree diagram for the re-encoded bitstream shown in 730. Item 729 indicates which words map to distinct codewords, such as those shown by item 728. A certain rate and number of identifiers are expected in the identifier library, so that any missing identifier errors can be detected during decoding.
[0058]
[0081] Write time can be minimized by interpreting the input bitstream to be a multi-valued Boolean function. In one embodiment, the input bitstream can be divided into blocks of a fixed length L before undergoing write time minimization. The input bitstream can be subjected to a heuristic logic minimization algorithm, such as espresso-mv or mvsis, to obtain a multi-valued algebraic expression representing the source bitstream. In one embodiment, the input bitstream can be encoded using an M-layer product scheme to construct identifiers. In this embodiment, the input bitstream can be interpreted as an M-input multi-valued Boolean function with a single Boolean output. For a Boolean function, a set of functions can be defined as the set of all inputs of the function that output the value "1". Using techniques from logic minimization, the Boolean function can be converted into an algebraic expression, including a sum-of-products expression. The obtained expression includes every identifier in the set of source bitstreams. Each term in the expression can be converted into a set of identifiers (constructed in a multiplicative manner) that can be executed within a single reaction compartment (e.g., a partition or reaction vessel). The obtained formula can be used to minimize the number of reaction compartments used and maximize the number of identifiers assembled in a single compartment. The formula can also be used to minimize the total time used to prepare the identifier assembly reaction, for example, if the write time can be proportional to the number of reaction compartments. Similar methods can also be used to prepare reactions used to query a subset of bits from a source bitstream.
[0059]
[0082] Figure 8 shows the output of an example of a reaction set minimization scheme. Consider a multiplication scheme with a bit stream of length L and M layers, where layer i has Ni components such that the product of all Ni is at least L. Each component in a layer can be labeled with an integer ranging from 0 to Ni-1. The bit stream of length L can be interpreted as a Boolean function F of M variables, where each variable V can take one of Ni values from 0 to Ni-1. All combinations of these variable values can be represented as an M-dimensional vector, and the values of variable V can be represented as integers in the i-th dimension of the vector. The Boolean function F can be defined using these vectors as inputs and each bit value in the bit stream as an output. If the multiplication scheme has a combinatorial space of size greater than L, the output of F for those additional input vectors can be defined to be distinct "don't-care" values.
[0060]
[0083] FIG. 8 illustrates an example in which information represented by item 831, which can be represented as a 64-bit long bitstream as shown by item 832, is encoded through a product scheme including two layers with 13 and 5 components in each layer. The defined Boolean function F has 65 possible input vectors, and each vector can be two-dimensional. Dimensional variables V1 and V2 have 13 and 5 values, respectively, with V1 ranging from 0 to 12 and V2 ranging from 0 to 4. The set of all possible variable value combinations can be represented as a tree diagram. The output of function F as defined above, which is "1," can also be represented as a tree diagram with a subset of arrows. This tree diagram is shown at the top of FIG. 8. The set of variable value combinations for which F takes the value "1" corresponds to the set of identifiers that need to be constructed to encode the bitstream. Thus, the paths from the root of the tree diagram to the individual values shown in the tree diagram correspond to the set of reactions required to assemble each identifier. In this example, the arrows represented by items 833 and 834 indicate one set of paths corresponding to three bits in the bitstream to be encoded. These three paths also correspond to the three identifiers that need to be assembled to encode those three bits. The vectors describing these "1" values of F differ in their second dimension, so that taking values 0, 3, and 4 results in their corresponding identifiers being different in the second layer, taking the 0th, 3rd, and 4th components in the second layer. All three identifiers have the same component in their first layer, corresponding to value V1=10. Therefore, all three identifiers can be assembled in a single reaction using component V1=10 and component V2 from the set {0,3,4}. The resulting set of combinations (10,0), (10,3), and (10,4) corresponds to the correct set of identifiers to be constructed. From the tree diagram, 13 such reaction sets are required to encode a given bitstream. However, tree diagrams may be included in a set of tree diagrams using a heuristic guided search so that all identifiers in each factor tree can be assembled in a single reaction.For example, a greedy heuristic may be used, in which all values of V1 for some value V2 = v are grouped together so that all assembled identifiers correspond to a "1" value in F. Item 835 shows the set of values where V2 = 0 and V1 = {3, 4, 5}. In another embodiment, multiple heuristics may be combined to obtain a minimal set that covers all "1" values in F. In another embodiment, heuristic techniques from logic minimization [Brayton et al. Logic Minimization Algorithms for VLSI Synthesis Kluwer Academic Publishers, incorporated herein by reference in its entirety] may be used to minimize the number of reaction sets. The five tree diagrams shown under the label "Heuristic Search-Guided Optimized Solutions" together cover all "1" values in F. As a result, five reaction sets may be used to prepare five separate partitions, rather than the 13 partitions in the original tree diagram.
[0061]
[0084] Each symbol (e.g., a bit in a bit stream) can be mapped to one or more unique identifiers in the combinatorial space. The set of identifiers may be determined and enumerated in computer memory or may be generated by combinatorially assembling a set of identifiers into an identifier library. When presented with digital information to be encoded into an identifier library, in one embodiment, each symbol in the digital information can be mapped to a distinct identifier in the combinatorial space. There can be a vast number of ways to map a given bit stream into a combinatorial space generated from a combinatorial scheme and containing some selected number of components (e.g., product schemes, permutation schemes, or some other scheme). Some of these mappings can be beneficial in reducing the number of queries when the encoded data is later queried. Specifically, mappings that preserve the locality of symbols in the original symbol stream after mapping them into the combinatorial space can be useful in reducing the number of accesses used to answer a query. An access can be a request to select a set of identifiers from an identifier library or pool of identifiers, described by a single nucleic acid sequence called an access sequence. In one embodiment, when identifiers are assembled from components, a single access can access the set of all identifiers containing a particular component. The nucleic acid sequences of the components can be the access sequences in this example. A family of mappings that preserves the localization of the original symbols is called isometric mapping. Furthermore, a single digital message can be mapped into two orthogonal combination spaces, each with its own component library, to generate two orthogonal identifier libraries that represent the same digital message. The two mappings can be beneficial in reducing the number of accesses to the two sets of queries. This type of encoding using multiple mappings can be called multi-encoding, and if the number of mappings can be fixed at two mappings, it can be called dual-encoding.
[0062]
[0085] FIG. 9 illustrates a schematic of an isometric mapping of addresses to identifiers and dual encoding of data. The process of encoding a digital message may involve converting information into a symbol sequence and converting the symbol sequence into a second symbol sequence with a formal data structure suitable for one or more query types. FIG. 9 illustrates an example in which the digital information to be encoded may be a two-dimensional image, shown in item 936. Item 937 illustrates a schematic diagram of the image, with the shaded circle indicating the lower right quadrant of the image. The original sequence of symbols, in this case bit values, may be encoded in the order presented. This order is shown in item 938, and the tree diagram generated for the product scheme is shown in item 939. If one reads the lower right quadrant of the image, this may result in querying the shaded circle in the encoded bitstream. In the combinatorial space, this may translate into a query for four identifiers. In this example, assuming each layer in the product scheme has two components each, two queries may be used: a query for all identifiers starting with component 101* and a query for all identifiers starting with component 111*. where * indicates that identifiers with any component in that layer can be returned as an answer to the query. Two queries are used because the lower quadrant of the image can be mapped into the combination space such that nearby regions of the image are mapped to identifiers that are not nearby in the combination space.
[0063]
[0086] Item 941 shows an alternative mapping in which neighboring regions of an image are mapped to neighboring identifiers. This can be called an isometric (i.e., distance-preserving) mapping. In this case, a single query can be used: all identifiers starting with 11** are sufficient to answer the query. This can be generalized to multidimensional data structures, including multi-column tables, tries, trees, sets, and vectors. More generally, the product scheme uniquely encodes data multidimensionally, thereby optimizing and parallelizing queries on many types of data. Item 945 shows a multidimensional data set with four dimensions, X, Y, Z, and W. In this example, X, Y, and Z each take two values, and the fourth dimension, W, takes four values. Each four-dimensional vector corresponds to a single bit value in this example. In general, this can be extended to integer values. Item 946 shows a tree diagram for encoding this 32-bit bitstream using a four-layer product scheme. Specifically, the product-based structure preserves the dimensionality of the original data structure: dimensions X, Y, and Z can be mapped to binary layers, and the four-valued dimension W can be mapped to a layer with four components. Furthermore, items 947 and 948 show two mappings of a dataset to the same combinatorial space. The two mappings differ in which regions of the data structure are mapped to adjacent regions of identifiers in the combinatorial space. In the mapping of item 947, the data regions corresponding to X=0, Y=0 and X=1, Y=1 are mapped to non-adjacent identifiers, while in the mapping of item 948, they are mapped to adjacent identifiers. Item 949 shows a possible query for unshaded bit values. Item 952 shows the component access sequence used to retrieve these bit values using the mapping shown in item 947. In this example, the query can be answered using a single access to component 0 in layer W. Item 50 shows a more complex query that can be answered by two parallel accesses to components W=0 and Y=1 followed by a serial access to component X=1. This answers the query for all unshaded values in item 950. Item 951 shows a more complex query.Using the mapping in item 947, this query might require five or more accesses. However, using the mapping in item 948, this query can be answered using one access followed by a single decomposition step. The decomposition step removes all identifiers that contain a particular pattern. In this example, the pattern is component 1 from layer W. In this way, mapping a data structure to a combinatorial space can reduce the complexity of answering data queries. In some embodiments, multiple mappings of the same data structure can be encoded into a single pool of identifiers using orthogonal or distinguishable sets of components. This is shown in the mappings shown in items 947 and 948: two identifier libraries can encode the data structure shown in item 945, and the query can be answered using either mapping depending on the number of accesses used by each mapping.
[0064]
[0087] Digital information submitted for encoding into an identifier library may contain information that can be protected from unauthorized decoding. The methods of writing information into DNA described herein may provide an additional level of protection against unauthorized decoding of the encoded information. Biochemical methods of encryption, authentication, obfuscation, and destruction may be used to protect the encoded information. In one embodiment, information may be encoded and obfuscated by including decoy identifiers in the identifier library. Decoy identifiers may be identifiers included to not encode any information that was part of the original digital information submitted for encoding, but to make the decoding process prohibitively expensive and / or unwieldy without possessing the decoy key. The decoy key may be a set of sequences of components such that selecting identifiers containing the components can isolate some or all of the identifiers that make up the original identifier library, or conversely, deleting all identifiers containing the components can remove some or all of the decoy identifiers.
[0065]
[0088] FIG. 10 illustrates an example of a method for masking encoding and decoding to protect against unauthorized decoding. A bitstream can be encoded into a unique identifier, and an identifier library can be assembled. Additional nucleic acid sequences can be added to the identifier library. The additional or supplemental nucleic acid sequences can be of similar length to the unique identifier and indistinguishable from the unique identifier without a key to decode the information. Decoding the information can include subjecting the pool of identifiers to one or more selections and / or decompositions of target nucleic acid sequences until a unique identifier is extracted from the supplemental nucleic acid sequence. Item 1056 illustrates a tree diagram showing encoding of a bitstream using a five-layer stacking scheme, where each layer contains two components. The original bitstream, represented by item 1057, contains 16 bits, represented as circles circumscribing the values. However, this bitstream is encoded with a larger combinatorial space than that used to encode 16 bits; for example, as represented by item 1058, the remaining undefined symbols are represented as empty circles. The illustrated five-layer binary scheme allows for a combinatorial space of 32 distinct identifiers. Some of the identifiers corresponding to "1" bit values in the original bit sequence are shown in item 1060. Some of the remaining identifiers that do not correspond to any bit values in the original bitstream are shown by item 1059 and are labeled "potential decoy identifiers." These identifiers are shaded to indicate the minimum number of components sufficient to distinguish them from identifiers corresponding to bit values in the original bitstream. These identifiers are referred to as potential decoy identifiers. The selection of which identifiers are chosen as decoy identifiers and which identifiers are chosen to correspond to bit values in the original bitstream may be arbitrary in this example, but may be governed by the data structure of the bitstream, the query constraints, and the strength of the obfuscation or concealment used. From the set of potential decoy identifiers, some decoy identifiers are selected for inclusion in the identifier library encoding the original bitstream, as shown in item 1062, and are labeled "selected decoy identifiers."A bitstream may be encoded into a pool of identifiers that includes both identifiers corresponding to bit values and decoy identifiers that do not correspond to any bit values in the original bitstream. Thus, any unauthorized decryption of the pool may not be able to correctly decode the original bitstream without information about the set of selected decoy identifiers. A set of component sequences that describes a chosen set of decoy identifiers may be referred to as a decoy key. The decoy key in this example is shown by item 1064 and includes two component sequences: components 1, 0, 1, 1 from layers 0-4 and components 0, 1, 1 from layers 0-3. The decoy key may be interpreted as follows: Each component in the component sequence in the decoy key corresponds to an access query. All identifiers that match that component query are accessed from the current pool. The last component in the component sequence may not be accessed; instead, it may be used to remove all identifiers that match that component from the current pool. Table 1063 shows the steps required to implement the decoy key shown in 1064. Starting from the pool of all identifiers shown in tree diagram 1061, a series of accesses followed by deletions results in the exact identifier library corresponding to the original bitstream surviving: all decoy identifiers are removed. The surviving identifiers are shown in the shaded cells of table 1063.
[0066] Systems for encoding information into nucleic acid sequences and decoding information from nucleic acid sequences
[0089] Systems for encoding digital information into nucleic acids (e.g., DNA) can include systems, methods, and devices that convert files and data (e.g., raw data, compressed zip files, integer data, and other forms of data) into bytes and encode the bytes into nucleic acids, typically DNA segments, sequences, or combinations thereof.
[0067]
[0090] In one aspect, the present disclosure provides a system for writing information into a nucleic acid sequence. The system for writing information into a nucleic acid may include an assembly unit and one or more computer processors. The assembly unit may be configured to generate an identifier library that encodes a symbol sequence. The identifier library may include at least a subset of a plurality of identifiers. The one or more computer processors may be operably coupled to the assembly unit. The computer processors may be individually or collectively programmed to: (i) convert the symbol sequence into a code word using one or more code books; (ii) parse the code word into a coded symbol sequence; (iii) map the coded symbol sequence to a plurality of identifiers; (iv) instruct the assembly unit to generate the identifier library; and (v) instruct the assembly unit to attach a description of the one or more code books and a plurality of identifiers to the identifier library. Each symbol in the coded symbol sequence may be encoded by one or more identifiers.
[0068]
[0091] In another aspect, the present disclosure provides an integrated system for nucleic acid-based data storage. The integrated system for nucleic acid-based data storage may include a data encoding unit, a storage unit, a reading unit, and one or more computer processors. The data encoding unit may be configured to write digital information to a nucleic acid sequence. The storage unit may be configured to store a nucleic acid sequence encoding the digital information. The reading unit may be configured to access and read the digital information encoded in the nucleic acid sequence. One or more computer processors may be coupled to the data encoding unit, the storage unit, and the reading unit. The one or more computer processors may be individually or collectively programmed to (i) instruct the data encoding unit to encode the digital information into the nucleic acid sequence, (ii) instruct the storage unit to store the digital information encoded in the nucleic acid sequence, and (iii) instruct the reading unit to access and decode the digital information stored in the nucleic acid sequence. Digital information may be encoded in the nucleic acid sequence without base-by-base nucleic acid synthesis.
[0069]
[0092] The system may include one or more computer processors and a human-machine interface (HMI) for controlling and programming the computer processors. The system may encode and record digital information using any method as described elsewhere herein. The system may generate a list of identifiers that make up the identifier library. Alternatively or additionally, an external computer processing unit may generate a list of identifier sequences that make up the identifier library. The system may have an interface for receiving the list of identifier sequences. The interface unit may convert the list of identifier sequences into instructions for downstream units or modules of the system to generate and pool identifiers.
[0070]
[0093] The system can include an assembly module. The assembly module can be configured to receive a plurality of substrates (e.g., components) and reactants (e.g., enzymes) and output a plurality of reactions to produce identifiers that constitute one or more identifier libraries. One or more identifiers can be produced in a given reaction. One or more identifiers can be produced in multiple reactions. The multiple reactions can be about 1, 2, 4, 6, 8, 10, 20, 30, 50, 75, 100, 150, 200, 300, 400, 500, 750, 1000, 10000, 1×10 5 , 1×10 6 , 1×10 7 , 1×10 8 , 1×10 9 The plurality of reactions may include about 1×10 or more, or even more. 9 , 1×10 8 , 1×10 7 , 1×10 6 , 1×10 5 , 10,000, 1,000, 750, 500, 400, 300, 200, 150, 100, 75, 50, 30, 20, 10, 8, 6, 4, 2, or fewer reactions. One or more reactions may be performed simultaneously or sequentially. One or more reactions may be combined to generate an identifier library. An assembly unit may selectively remove one or more of the reactions that do not produce a selected identifier. An assembly unit may include one or more sections, containers, or partitions. An assembly unit may include multiple sections, containers, or partitions. Each section, container, or partition may generate, store, maintain, facilitate, or terminate one or more assembly reactions.
[0071]
[0094] The assembly unit may include a reaction module. The reaction module may collect reagents, one or more nucleic acid sequences, one or more components, one or more templates, or any combination thereof. The reaction module may be configured to incubate or agitate an assembly reaction to produce one or more identifiers. The reaction module may further include a detection unit. The detection unit may monitor the assembly of the identifiers. The reaction module may include multiple partitions. Each of the multiple partitions may contain one or more assembly reactions. The multiple partitions may be wells or droplets on a chemically modified surface.
[0072]
[0095] The substrate or input may include one or more M layers. Each layer may include one or more components. The components in each layer may be distinct from the components in the other layers. The substrate may include assembly templates, primers, probes, and any other elements for directing and facilitating the identifier assembly reaction. The reagents may include enzymes, buffers, nucleic acid sequences, cofactors, or any combination thereof. The enzymes may be produced by overexpression of the corresponding recombinant gene in living cells. The reagents may be combined in individual assembly reactions or combined as a master mix before being added to the assembly reactions.
[0073]
[0096] The system may further include a storage unit (e.g., a database). The assembly unit may output one or more identifier libraries. The one or more identifier libraries may be received by the storage unit. The storage unit may include one or more pools, containers, or partitions. The storage unit may combine individual identifier libraries with one or more additional identifier libraries to form one or more pools of identifier libraries. Each individual identifier library may include a barcode or tag that identifies and distinguishes identifiers from each library from one another. The storage unit may provide conditions for long-term storage of the identifier library (e.g., conditions to reduce identifier degradation). The identifier library may be stored in powder, liquid, or solid form. The database may provide protection from ultraviolet light, low temperature (e.g., refrigeration or freezing), and chemical and enzyme degradation. Before transfer to the database, the identifier library may be lyophilized or frozen. The identifier library may include ethylenediaminetetraacetic acid (EDTA), other metal chelators, or other reaction blocking agents to inactivate nucleases and / or buffers and maintain the stability of the nucleic acid molecules.
[0074]
[0097] The system may further include a selection unit. The selection unit may be configured to select one or more identifiers from the identifier library or from a group of identifier libraries. The assembly unit may prepare all possible reactions to generate a combination space, and the selection unit may selectively remove reactions that do not produce the target identifier and retain reactions that do. The selection unit may include an optical or mechanical ablation module that removes reactions, a dispenser that delivers degradative enzymes to non-target reactions, or a dispenser that delivers primers or affinity-tagged probes to target reactions. The selection unit may facilitate evaluation of stored data. Access to information stored in nucleic acid molecules (e.g., identifiers) may be performed by selectively removing a portion of the identifier library or an identifier library from a group or pool of combined identifier libraries. Access to data may be performed by selectively capturing or amplifying identifiers corresponding to the accessed data and / or removing identifiers that do not correspond to the accessed data. Methods for selecting identifiers may include using polymerase chain reaction, affinity-tagged probes, and degradation-tagged probes. A pool of identifiers (e.g., an identifier library) may include identifiers with a common sequence at each end, identifiers with a variable sequence at each end, or identifiers with either a common sequence or a variable sequence at each end. Identifiers may include the same common sequence at each end or different common sequences at each end. An identifier library may include distinct common sequences in the library that allow selective access to a single library from a pool or group of two or more identifier libraries. The common sequence or variable sequence may be a primer binding site. One or more primers may bind to the common region of the identifiers. The primed identifiers may be amplified by PCR. The amplified identifiers may far outnumber the unamplified identifiers.
[0075]
[0098] The common sequence of the identifier may share complementarity with one or more probes. The one or more probes may bind or hybridize to the accessed identifier. The probe may include an affinity tag. The affinity tag may bind to a bead, generating a complex including the bead, at least one probe, and at least one identifier. The beads may be magnetized, and the selection unit may include one or more magnetic or electronic areas. The beads may collect and extract the accessed identifier. Alternatively or additionally, the beads may collect the non-accessed identifier. The identifier may be removed from the bead under denaturing conditions before reading. The affinity tag may be bound to a column, and the selection unit may include one or more affinity columns. The accessed identifier may bind to the column, and the accessed identifier may flow through the column, while the non-accessed identifier may bind to the column. The column-bound accessed identifier may be debound or denatured from the column before reading. Accessing the identifier may include applying one or more probes to the identifier library simultaneously or sequentially applying one or more probes to the identifier library / group of identifier libraries. In one example, one or more identifier libraries are combined, each containing one or more distinct consensus sequences. One set of probes can be applied to the library to extract a first subset of identifiers. Subsequently, a second set of probes can be applied to the library to extract a second subset of identifiers. This operation can be repeated until all identifiers have been extracted.
[0076]
[0099] The common sequence of the identifier may share a complementary relationship with one or more probes. The probe may bind or hybridize to the common sequence of the identifier. The probe may be a target for a denaturing enzyme. In one example, one or more identifier libraries may be combined. A set of probes may hybridize to one of the identifier libraries. The set of probes may include RNA, and the RNA may induce a Cas9 enzyme. The Cas9 enzyme may be introduced into one or more identifier libraries. Identifiers hybridized to the probes may be degraded by the Cas9 enzyme. Accessed identifiers may not be degraded by the degradative enzyme. In another example, the identifiers may be single-stranded, and the identifier library may be combined with a single-strand-specific endonuclease that selectively degrades non-accessed identifiers. Accessed identifiers may be hybridized to a complementary set of identifiers to protect them from degradation by the single-strand-specific endonuclease. Accessed identifiers may be separated from degradation products by size selection, such as size-selective chromatography (e.g., agarose gel electrophoresis). The selection unit may be capable of performing one or more size selection techniques. Alternatively or additionally, undegraded identifiers may be selectively amplified (for example using PCR) so that degradation products are not amplified. Undegraded identifiers may be amplified using primers that hybridize to each end of undegraded identifiers and therefore do not hybridize to each end of degraded or cleaved identifiers.
[0077]
[0100] Individual nucleic acid sequences (e.g., components and templates) that make up or assist in the construction of the identifiers may be synthesized by the system or may be synthesized and amplified externally. The system may further include a nucleic acid synthesis module. The nucleic acid synthesis module may perform base-by-base construction of components and templates. Nucleic acid sequences (e.g., components and templates) may be constructed using phosphoramidite chemistry. Components may first be constructed using phosphoramidite chemistry, and then the original phosphoramidite templates may be replicated using PCR. Components may first be constructed using phosphoramidite chemistry, and then copies of the templates may be produced by cloning the components into one or more high-copy vectors. The vector may be transformed into a living cell, where the vector, along with the embedded nucleic acid sequence, may be replicated during cell growth. The vector may be isolated from the cell culture, and the components may be separated from the vector using restriction digestion. Double-stranded nucleic acid sequences may be converted to single-stranded nucleic acid sequences by using affinity-tagged probes that share complementarity with one of the two nucleic acid strands.
[0078]
[0101] The system may use techniques to minimize the number of reactions used to generate the identifier library, and thus the write time. One or more of the techniques may include heuristic techniques. The heuristic techniques may minimize the set of compartmentalized sets of reactions used to construct a given set of identifiers from components. The heuristic techniques may include on-set covering heuristics. The physical distance traveled by the writer may also be minimized to reduce the write time. Figure 8 shows an example of a method for minimizing write time through minimal reaction set generation.
[0079]
[0102] The system may transfer fluids (e.g., reagents, components, templates) using pressure, vacuum, or suction. The assembly unit may combine one or more nucleic acid sequences with one or more reagent mixtures. The assembly unit may combine substrates (e.g., enzymes, components, and templates) with reactions using one or more of electrowetting, misting, printing, laser ablation, weaving or knitting of materials coated with nucleic acid sequences, slip technology, stamping, laser printing, or droplet microfluidics. A co-located set of biomolecules may generate an identifier. For example, instead of linking the components to each other, by assembling separate components from each layer onto a shared substrate such as a bead. Various techniques may be used to co-locate a set of biomolecules. As one example, instead of constructing an identifier by linking a set of separate components to each other, the identifier may be constructed by associating the components with a shared substrate such as a bead. As another example, instead of constructing an identifier by linking a set of separate components to each other, the identifier may be constructed by assembling each component into a barcode sequence that identifies the association of the components.
[0080]
[0103] A component carousel can be used to co-locate a set of biomolecules. FIG. 11 shows a top-down view 1108 of an example component carousel and a cross-sectional view 1109 of the component carousel along line 1110 in the top view 1108. In this example, the component carousel includes multiple inlets and multiple outlets. The inlets can be on the outer periphery of the carousel, and the outlets can be on the inner periphery of the carousel. Each inlet can selectively introduce a single input (typically a component, but possibly a nucleic acid, enzyme, or reaction mixture) into a reaction chamber connected to the outlet. After introducing one input, the carousel can shift one position to selectively introduce an adjacent input into the reaction chamber. This process can be repeated until a selected number of inputs can be combined.
[0081]
[0104] The component carousel may be comprised of two substrates 1101 and 1102, with flat surfaces facing each other. In the embodiment shown in FIG. 11, the two surfaces are configured to rotate relative to each other. In some cases, it may be advantageous to introduce oil or another lubricant between the two surfaces to reduce sliding friction. While any lubricating fluid can be used, a fluorinated oil may be used to minimize the transfer of biological materials into the oil or between the chambers. In this example, the inlet 1103 and outlet 1104 consist of paired through-holes in one of the substrates 1101. The second substrate 1102 has one chamber 1105 for each pair of through-holes. When the surfaces of the two substrates face each other and contact each other, the chambers 1105 in the second substrate 1102 align with grooves or channels 1106 in the first substrate to complete a flow path between the pairs of through-holes. The two substrates are designed to slide relative to each other, such that each channel sequentially connects to every chamber as the two surfaces slide past each other through a complete rotation. In this way, all inputs can be selectively added to each chamber. For example, in one embodiment, the first substrate has 72 pairs of through-holes, and the second substrate has 72 chambers. The system is configured to selectively introduce different components into the chambers every five degrees of rotation of the surfaces. At the end of a complete rotation, outlet 1107 allows the reaction mixture to exit the chamber as a bolus 1111. After the reactants are cleared from the chamber, the chamber can be reused for subsequent reactions. Typically, one pass is used to remove the reaction bolus 1111, and a subsequent pass is used to clean the reaction chamber. The introduction of a master mix into the reaction chamber can optionally have separate passes, or the master mix may be introduced with each input. In this example, the remaining 70 passes allow 70 unique inputs to be sequentially introduced into a given reaction chamber. If the input is a set of 22 layers of 3 components and 1 layer of 4 components, the combinatorial space of the multiplication method is 4 * 3 22 =1.2×10 1115 unique identifiers. Increasing the number of channels slightly to facilitate 96 components allows the 96 components to be arranged in 32 layers with 3 components per layer, generating up to 1.8e15 unique identifiers. In some embodiments, the chambers are filled with oil or gas before the first input is introduced. In some embodiments, oil or gas is used to evacuate reactants from the reaction chambers after the last input and reaction master mix are introduced. There is no limit to the number of chambers and the number of inputs that can be introduced. In some embodiments, 10 or fewer chambers are used, in some embodiments, 10 to 100 chambers are used, and in other embodiments, 100 to 1000 chambers are used. In other embodiments, more than 1000 chambers are used. There is no limit to the amount of biological material that can be introduced into the chambers. In some cases, the inputs can be factors for amino acid or peptide synthesis; in other cases, the inputs can be reagents for synthesizing small molecules; and in other cases, the inputs can include cells, bacteria, viruses, droplets or other particles, lysis buffers, or reagents for tagging, amplifying, binding, or identifying biological material within cell lysates or on the surfaces of cells, bacteria, viruses, or other particles. In some cases, the chambers index between port pairs at a rate of several times per hour or several times per minute. However, this indexing frequency can be any number of times and may be selected to be fast. In some cases, it may be once per second, 10 times per second, 100 times per second, 1000 times per second, 10,000 times per second, or more. External fluid control may be used to selectively introduce inputs into the chambers on demand.
[0082]
[0105] Electrowetting can be used to co-locate a set of biomolecules. Figure 12 shows an electrowetting method for input operations. Inputs (e.g., nucleic acids, components, templates, enzymes, or reaction mixtures) can be introduced through separate ports 1201. Each port 1201 can introduce one input or a mixture of multiple inputs. Droplets can be generated using electrowetting and combined to assemble selected inputs into an identifier. Droplets are created, combined, mixed, or split by selectively applying voltages to electrode patches 1202. In some embodiments, the electrode patches are arranged in a rectangular array. The patches are typically configured to be separated from the droplets by an insulating coating with low electrical conductivity. The electrowetting device can be open or closed at the top. The electrowetting chamber can contain an insulating fluid such as oil. Any oil can be used, such as silicone oil, mineral oil, or hydrocarbon oil. In one example, a fluorinated oil is used. Surfactant mixtures of other additives may be utilized to improve device performance by modifying the surface energy at the droplet-oil interface or the interface with the chamber wall.
[0083]
[0106] Electrowetting techniques can be used to create and manipulate small volumes of fluid, ranging from the subpicoliter to nanoliter range. For example, FIG. 12 shows an electrowetting device configured to selectively combine inputs in a programmable manner. Systems can be easily configured to simultaneously process tens, hundreds, thousands, thousands, millions, or more droplets using electrowetting techniques. In some embodiments, it can be advantageous to combine droplets and then split the combined droplet into two mixed droplets. In some cases, mixing can be enhanced by combining and splitting in a generally orthogonal direction. Each of the split droplets can receive a different subsequent input. The process can be repeated until all inputs required to construct the identifier are introduced into the droplets. For example, component C 1,1Droplet 1203 containing (component 1 of layer 1) and component C 2,1 Droplets 1204 containing (component 1 of layer 2) are combined to form mixed droplets 1205C 1,1 C 2,1 where the mixed droplet has both components. The mixed droplet can then be split into two daughter droplets 1206, both of which have the same mixed composition. Component C from the third layer 3,1 1207 and C 3,2 Additional droplets containing 1208 can be introduced into mixed droplet 1206 to form droplets 1209 and 1210 containing components from the first three layers. This process of combining, mixing, and splitting droplets can be repeated until the components used to construct the appropriate identifier are complete. In some cases, a master mix for assembling or constructing the identifier can be introduced into a separate input droplet, either together with the nucleic acid input. In a multiplication scheme, at least one component from each layer can be introduced into a droplet to assemble the complete identifier. In a multiplex reaction, multiple components from one or more layers can be introduced into a given droplet. In embodiments utilizing droplet splitting, it can be advantageous to have components at different initial concentrations to facilitate balancing the concentrations of each component. Due to the parallelism in which droplets can be processed at different locations on the same electrode array, it may be possible to process droplets at arbitrarily high speeds, preparing reaction conditions for thousands, millions, or even billions of droplets every second.
[0084]
[0107] Print-based methods can be used to co-locate biomolecules. Figure 13 shows an example of a print-based method for dispensing inputs. Inputs (e.g., nucleic acids, components, templates, enzymes, or reaction mixtures) can be brought together into stationary reaction regions by dispensing or printing directly into those regions. Reaction regions can be discrete locations on a substrate 1301. Component inputs 1306 can be assembled into identifiers in discrete regions. Surfaces can be patterned using chemical modifications to create regions of varying hydrophobicity. Regions of varying hydrophobicity can be useful for preventing inputs from migrating from one region to a neighboring region. The regions can have dimensions of about 0.1 micrometers (μm) or more, about 0.5 μm or more, about 1 μm or more, about 2 μm or more, about 4 μm or more, about 6 μm or more, about 8 μm or more, about 10 μm or more, about 20 μm or more, about 40 μm or more, about 60 μm or more, about 80 μm or more, about 100 μm or more, or more. The regions can have dimensions of about 100 μm or less, about 80 μm or less, about 60 μm or less, about 40 μm or less, about 20 μm or less, about 10 μm or less, about 8 μm or less, about 6 μm or less, about 4 μm or less, about 2 μm or less, about 1 μm or less, about 0.5 μm or less, about 0.1 μm or less, or less. The reaction regions can be separated by physical barriers, such as walls. The walls can be formed by lithography in an otherwise flat surface to create microwells. Alternatively or additionally, microwells can be molded or embossed into a plastic substrate. Microwell volumes can be about 0.1 picoliters (pL) or more, about 1 pL or more, about 10 pL or more, about 100 pL or more, about 1 nanoliter (nL) or more, about 10 nL or more, or greater. Microwell volumes can be about 10 nL or less, about 1 nL or less, about 100 pL or less, about 10 pL or less, about 1 pL or less, about 0.1 pL or less, or less. The substrate can comprise glass, paper, or a plastic film. The substrate can optionally be patterned using one or more methods, such as hydrophobicity, embossed wells, etched wells, molded features, deposited features, etc.In a reel-to-reel system 1302, rollers may be used to pattern indentations directly on the substrate prior to dispensing. The substrate may be translated under a stationary printhead, or optionally the printhead may be translated over the surface of the substrate. Dispensing may utilize a wide range of commercially available printing techniques. The printhead may contain 1, 10, 100, 1,000, 10,000, or more nozzles. Each nozzle of a printhead may dispense the same input, or one or more nozzles may dispense separate inputs. In some embodiments, a sufficient number of printheads are utilized so that a given nozzle can dispense a single input. For example, if each printhead dispenses four inputs, a collection of 50 printheads can dispense 200 inputs. Such an arrangement with printheads aligned to dispense swaths can optionally be combined with reel-to-reel operation, in which the substrate passes under all printheads to dispense all inputs to all reaction regions. Each nozzle in a printhead can dispense at a rate of 10, 100, 1,000, 20,000, 50,000, 100,000, or more per second. Nozzles can be configured to operate in parallel, such that a printhead with 1,000 nozzles operating at 50,000 dispenses per second can dispense up to 50 million times per second. Printer driver circuits can enable higher and lower frequencies and drop-on-demand operation, any of which can be utilized to dispense the inputs. These systems include, but are not limited to, inkjet, bubble jet, and piezoelectric arrays. In some cases, electrostatic charges and electric fields are used to direct or control droplet placement. In other cases, electrostatically neutral droplets are dispensed.
[0085]
[0108] Similar to printhead operation, laser forward transfer is an optical technique for selectively transferring material, including input 1303, from one substrate 1304 to a receiving surface 1305. Precise positioning of the laser pulse controls the transfer of material. By controlling laser focus, pulse width, power, and location, the amount of material transferred can be controlled to pattern a given input onto the substrate. Sequential transfer of each input provides a robust mechanism and time-efficient method for preparing a reaction set. In some embodiments, optically detectable markers, such as fluorescent or absorbing dyes, can be introduced into the input fluid to enhance imaging-based inspection and confirm that the input is distributed among the reactants as intended.
[0086]
[0109] A 1.0×10 bin is generated by (1) recoding the string into a uniform weight form in which every contiguous (i.e., adjacent and disjoint) span of 250 bits has exactly 75 bit values of "1", (2) encoding the recoded bit stream into an identifier library (excluding identifiers corresponding to bit values of "0" from the library) using an example encoding method, and (3) constructing an identifier whose components are divided into 8 layers using a product method. 12 Encoding and writing bit strings. In this example protocol, sequential words of length 216 bits from the original information string may be encoded using code words containing a subset of exactly 75 identifiers from each sequential set of 250 possible identifiers. This 250-choose-75 uniform encoding scheme can be used to encode 216-bit words to a 1 terabit (1 x 10 12 If expressed as a string of bits, it is at least (250 / 216) * 1.0×10 12 =1.15×10 12 A combinatorial space of distinct identifiers can be used. In this example, we use 7 layers with 20 components in each layer and an 8th layer with 1000 components. The identifiers available in this example are then 1000 * 20 7 =1.28×10 12 , which is 1.15×10 12exceeding the minimum required number of 1.0 × 10 12 This may be sufficient to uniquely represent one component from each of the first seven layers and 75 from the eighth layer. * Multiplexed assembly reactions can be constructed by assembling components representing 4 code words, dispensing 4 = 300 components into each reaction, for a single multiplexed reaction volume. Seven components from the first seven layers are assembled with 300 components from the eighth layer to create the original 1.0 × 10 12 Bitstream unique 4 * Generate 300 unique identifiers representing 216 = 864-bit portions. 1.0 x 10 12 The identifier library that represents the entire bit string is 1.0 × 10 12 It can be assembled using 1 / 864 = 1.16e9 reactions, each with one component from each of the first seven layers and 300 components from the eighth layer (i.e., a total of 307 components across all layers). In this example, using a 100 μm separation between reactions, it is roughly 12.8 square meters (m 2 ) area can be covered with reactions. A single printhead operating at 5000 dispenses per second, with 160 nozzles per component, can cover an area of 1.16 × 10 9 An assembly with 10 printheads dispensing four components, each using 160 nozzles per component and operating at 5000 dispenses per second, can accommodate all 1140 components in 1.16 × 10 9 All reactions can be distributed in approximately 12.6 hours using continuous dispensing.
[0087]
[0110] Microfluidic injection can be used to co-locate biomolecules. Figure 14 shows an example of microfluidic injection of inputs. Microfluidic devices can be constructed by any method, such as injection molding, embossing a plastic substrate, etching a glass channel, or cross-linking a polymer. Fluids are introduced into the microfluidic device through ports and can be driven by any method, such as electroosmotic flow, external pressure or vacuum, or a positive displacement pump. In one embodiment, a stream of master mix 1401 is introduced into a stream of carrier oil 1402, and droplets of master mix 1403 form the oil stream. In some embodiments, the master mix droplets can be 1 nL or greater; in other embodiments, they are less than 100 pL, 50 pL, 10 pL, 5 pL, or 1 pL in volume. The master mix droplets can contact the channel wall, or a layer of carrier oil can separate them from the channel wall. The carrier oil can be any oil, such as a hydrocarbon oil, a fluorocarbon oil, a mineral oil, or any combination of oils. In one example, the oil is a fluorocarbon oil. In some embodiments, the oil may further include a surfactant or other additive. The master mix may include an aqueous fluid. Inputs are introduced into the microfluidic device through ports and multiple input streams 1405 that intersect with the main channel 1404. The inputs (e.g., components or templates such as nucleic acids, enzymes, or reagents) can be selectively added to droplets as they pass through one or more injection orifices. Injection can be controlled through the selective application of an electric field through the application of a voltage to an electrode 1406 positioned near the main channel. The electrode can be separated from the channel by an insulating layer. In one embodiment, all possible distinct identifier-producing reaction droplets can be generated, and a targeted subpopulation of droplets can be collected using a sorting branch in the channel. Sorting can be achieved by any method, including, but not limited to, an electric field gradient, a laser pulse, a bubble, a piezoelectric actuator, an external valve, an acoustic wave, or any other sorting mechanism. In another embodiment, droplets containing target identifier-producing reactants are generated. Reactants can proceed to completion either on the microfluidic device in which they are fabricated or elsewhere.The droplets may be collected in a reaction reservoir 1407 either on the microfluidic device or elsewhere.
[0088]
[0111] Each identifier can be constructed using a multiplication method by assembling components, with at least one component from each layer being introduced into the same droplet. Each pico-injector includes a component stream 1405 and a method for applying an external electric field 1406. The components are assembled into the identifier using enzymes. In some embodiments, the component fluid 1405 further includes an enzyme or master mix. As an example, a microfluidic device includes 10 sets of 10 pico-injectors configured so that a set of 100 pico-injectors can be used to introduce any combination of components from 10 layers of 10 components each into a flowing droplet. This example system illustrates 10 layers constructed using a multiplication method. 10 It may be possible to generate unique identifiers. For M layers with N pico-injectors (e.g., component inputs) at each layer, N×M pico-injectors may be generated. M This can be easily generalized to allow construction of ×N identifiers. More generally, if one layer is designed as a multilayer with ×N pico-injectors, construction of ×N identifiers can be multiplexed in each droplet. The advantage of having one layer with more components than other layers is that that layer can be used as a multilayer to assemble multiple identifiers in the same droplet, thus reducing the total number of droplets required to write information. Each droplet receives one component from each layer except the multilayer, and may receive up to all components from the multilayer; ×N identifiers are constructed in each droplet.
[0089]
[0112] There can be flexibility in how components can be divided into layers to assemble identifiers using a multiplication scheme. For example, an input in a given set of 200 pico-injectors can be divided into 11 layers of components, each with 10 components (and pico-injectors for dispensing the components), and multiple layers with 100 components. In that case, the combinatorial space of identifiers can be expanded to 10 10 ×100=1012 Alternatively, we can use the same 200 pico-injectors and divide them into 40 layers of 4 components and multiple layers of 40 components. In that case, the size of the combinatorial space is 4 40 ×40=4.8×10 25 Typically, the more layers, the longer the DNA identifier can be.
[0090]
[0113] In one example droplet microfluidic system, the identifier is assembled from 12 layers of 16 components using a stacking method. In this example, the microfluidic device is configured with 16 pico-injectors in each layer (16 x 12 = 192 pico-injectors). 12 =2.8×10 24 It may be possible to assemble 10 unique identifiers. An alternative organization of 11 layers of 10 components and 1 layer of 100 components (11 x 10 + 100 = 210 pico-injectors) would result in 10 11 ×100=10 13 This creates a combinatorial space of unique identifiers. Using uniform weight encoding with codewords containing a subset of 18 identifiers from every block of 100 identifiers, we can encode 64-bit long words from the original compressed bitstream. It takes 1.56 x 10 10 droplets can be used. At a rate of 180,845 droplets / sec or 1,809 droplets / sec on 100 parallel devices, 1.0e12 bit strings can be written to DNA in 24 hours. If an initial droplet volume of 100 pL is used and 10 pL is added for each pico-injector used, then 100 pL + 100 pL (first 10 layers) + 180 pL (multiple layers) = 380 pL / droplet. Total droplet volume used: 380 x 10 -12 x1.5 x 10 10 Droplets = 5.7 L. After enzymatic assembly of the identifiers in the droplets, the contents of each droplet can be combined and condensed or lyophilized in preparation for storage.
[0091]
[0114] Selective condensation of component mist can be used to co-locate biomolecules. Figure 15 shows an example of selective condensation of component mist for co-locating biomolecules. A mist nozzle 1501 generates a mist or cloud of micrometer- or sub-micrometer-sized droplets 1502. The droplets can contain one or more inputs (e.g., nucleic acid sequences, enzymes, or reagents, such as components or templates). The mist can be generated using a vibrating membrane, electrospray, nebulizer, or any other method. The mist can be directed onto a thin-film transistor array 1503. The thin-film transistors can utilize individual electrodes 1504 to condense the mist or electrode pairs 1505, such as in an in-plane switching configuration, to selectively condense mist droplets in specific regions of the transistor array. Inputs can be introduced to the array 1503 one at a time or in groups of multiple inputs. The array can be dried between sequential introductions of the inputs. After the inputs are directed to the array, a master mix can be introduced to every reaction spot in the array to construct an identifier.
[0092]
[0115] Other methods, such as slip technology, microfluidic devices with elastomeric valves, and contact stamping, may be used to generate and select libraries of identifiers. Slip technology can involve parallel input streams for introducing components into multiple chambers or partitions in parallel. The chambers can slide to allow access to different compartments. In one example, components can be introduced into the chambers through elastomeric valves. In another example, microfluidic channels can be arranged along the periphery of tandemly arranged barrels, such that each barrel's chamber can be used to add one layer of components. The barrels can be rotated relative to each other by one channel diameter increment.
[0093]
[0116] Various methods can be used to generate all possible identifiers from the combinatorial space. Figure 16 shows a schematic diagram of an example of a method for generating identifiers by weaving or knitting. A flexible material can be coated with specific components in specific areas. The material can be plastic, metal, thread, or natural material. The flexible material can be woven, knitted, pinched, or intertwined together to co-locate the components to be assembled. Segments of components can be joined at knitted or woven joints and separated into reaction volumes. After all identifiers are constructed, any subset of identifiers can be deleted, including those that are inconsistent with the bitstream to be encoded. The family of methods that can encode information by deleting identifiers from constructed identifiers or established sets of identifier-producing reactions, or by deleting co-located components that assemble into identifiers, is referred to as the family of subtractive writing methods. In one embodiment, the components can be placed on a thread or film. Items 1601-1604 show an example in which four threads or four sheets of film are marked with components in a specific pattern. For example, a length of yarn or film labeled 1601 is divided into two regions: Region 0 is loaded with component 0 from layer 0, as indicated by level 1611, and Region 1 is loaded with component 1 from layer 0, as indicated by level 1612. A length of film, yarn, or fiber labeled 1602 is similarly divided into four regions: Region 0 is loaded with component 0 of layer 1 (labeled 1609), Region 1 is loaded with component 1 of layer 1 (labeled 1610), Region 2 is loaded with component 0 of layer 1, and Region 3 is loaded with component 1 of layer 1. In general, a film, yarn, or fiber corresponding to the ith layer containing Ni components is labeled Ni-1 *The substrate is divided into Ni regions, each loaded with one of the Ni components in the i-th layer, cycling through the list of Ni components in order, repeatedly. This method of organizing component regions on a substrate and loading them with components is called combinatorial marking. Other patterns, sequences, and schemes may be used to organize components into films and yarns. In one embodiment, each yarn, film, or fiber may be loaded with a single component. A set of such single-component yarns, fibers, or films may be woven into a lattice, as shown at 1613 and 1614. In this example, each intersection between a weft and warp yarn co-locates two components, as shown at 1615. In another embodiment, many yarns may intersect at a single location, thereby co-locating multiple components. These intersections may be used to construct an identifier, or sets of such co-located components may be extracted from these sites and used to assemble identifiers elsewhere. In one embodiment, each yarn may have a specific pattern of regions and components, as described above. These yarns may be woven together to form a mesh, as shown at 1617. Regions of this woven mesh may co-locate all of the components used to construct the identifier, as shown at 1616. These regions of the woven mesh may be used as reaction sites, or the set of components so located in these regions may be extracted from these sites and used to assemble an identifier elsewhere. In another embodiment, a product scheme may be prepared in which the number of components in each layer Ni is relatively prime to the number of components in all other layers. That is, for any Ni and Nj, denoting the number of components in layers i and j, where i is not equal to j, Ni is not a divisor of Nj, and vice versa. An example is shown at 1618, where two yarns, films, or fibers are shown, with Yarn 0 containing two components, labeled 5 and 6, and Yarn 1 containing five components, labeled 7, 8, 9, A, and B. The numbers of components in these layers, 2 and 5, are relatively prime because 2 is not a divisor of 5, and vice versa. The components are loaded onto the yarn and repeated in a cyclical sequence.Thus, Yarn 0 has a repeating sequence of two components 5, 6, 5, 6, ... as shown, and Yarn 1 has a repeating sequence of five components 7, 8, 9, A, B, 7, 8, 9, A, B, ... as shown. In one embodiment, these yarns can be pinched, twinned, or co-located together so that each component-loaded region on one yarn can align with a corresponding component-loaded region on another yarn. Because the number of components on each yarn is relatively prime, all possible combinations of components are generated at the pinched or twinned sites. The components thus located at these sites may be used as reaction sites to construct an identifier from these components or may be used to extract the co-located components to construct an identifier elsewhere. In another embodiment, a similar scheme using relatively prime component numbers may be used to generate a knitted yarn mesh. The weft yarn is shown at 1621. The weft yarn may be repeated as many times as the product of the number of components in the warp yarns.
[0094]
[0117] Figure 17 shows schematically an alternative method for generating an identifier from a set of components. The components are first stored in separate reservoirs, as shown at 1723. The reservoirs may also store assembly reagents and other equipment. The components may be placed together in a set of reaction compartments, an example of which is shown at 1724. Using a transfer method such as printing or fluid manipulation, each combination of components is placed together into individual components, as shown at 1726. These compartments may now be used as sites for assembling the identifier using multiple biochemical processes.
[0095]
[0118] FIG. 18 shows a schematic diagram of an example of a method for generating an identifier from separate films or yarns. 1832 shows a device called a collocator that takes as input a wound set of yarns, films, fibers, or substrates, each of which can be marked using a combinatorial marking scheme or some other marking scheme, and collects the components in each corresponding region on the individual yarns, films, or substrates. The collected components are co-located on an output film, yarn, or fiber, shown at 1833. As each region on each yarn or fiber passes through the collocator, new combinations of components can be generated in new regions on the output film or yarn, as shown by 1834. Item 1835 shows a schematic diagram of co-located components that can be used as reaction sites for assembling an identifier. Item 1836 shows a closer view of one embodiment of a collocator. Item 1837 shows one embodiment of a method for collecting components. In this example, the collocator punches holes in the passing fiber, yarn, or film, collecting the punched pieces or fragments and outputting them to an output substrate. In another embodiment, the collocator may use rubbing, suction, or other mechanical, electrical, optical, magnetic, knitting, weaving, pinching, or stamping mechanisms to place all components from all films or yarns together on the output film, yarn, or substrate.
[0096]
[0119] A subtractive writing method may be one in which a given digital message is encoded by subtracting identifiers from a previously constructed identifier library or an established library of identifier-producing reactions, or by subtracting co-located components that are prepared to be assembled into an identifier. In one embodiment, this library contains all possible identifiers in the combinatorial space. A subtractive writing method may be advantageous because it may eliminate the complexity of constructing a specific, given set of identifiers on demand. Rather, the construction of identifiers may be independent of the specific digital message being encoded and may be performed in advance of any encoding request. Furthermore, the encoding process may require simpler subtraction operations at the time of writing, rather than biochemical assembly or construction of identifiers. In one embodiment, a subtractive writing method requires a method for generating all possible identifiers. In one embodiment, when encoding is used in conjunction with a product method, all possible identifiers may be generated by preloading a simple sequence of components for each layer and then combining the preloaded streams of components. The preloaded sequence of components may be such that, when the component streams are combined, all possible component combinations are generated. This can be achieved using printing, yarn processing, knitting, weaving, twinning, pinching, stamping, and other methods.
[0097]
[0120] FIG. 19 shows an example of how subtraction can be used to write information. The subtraction target identifier can be removed enzymatically (e.g., using a CRISPR / Cas system) or by cutting, optical, thermal, electronic, electrostatic or electron discharge or other charged particle beam, sorting, liquid jet, acoustic, mechanical scraping, or perforation methods. In certain embodiments where components are co-located to form an identifier but have not yet reacted, the components at each location can be assembled after subtracting out unwanted identifier-producing reaction preparations. Item 1927 shows a tree diagram of a given bitstream encoded using a product scheme involving four binary layers. In this example, the combinatorial space contains 16 different identifiers. All 16 identifiers can first be placed together in individual compartments, as shown in 1925. These identifiers can then be mapped to individual symbols in the encoded information—bit values in this example—following the considerations outlined in FIG. 9. Once the correspondence between bits and identifiers is fixed, each compartment containing the set of components used to construct the identifier can be mapped to a bit value in the bitstream. For each partition mapped to a bit with a value of "0," the components in that partition may be destroyed, deleted, or otherwise manipulated so that an identifier is not assembled in that partition (item 1930). For each partition mapped to a bit with a bit value of "1," the components in that partition are provided with all reagents used to assemble the identifier and are not deleted or destroyed (item 1931). In another embodiment, all identifiers are assembled, and identifiers corresponding to bit values of "0" are deleted or destroyed after assembly. Finally, all surviving identifiers are pooled together to encode and store a given bit stream in a compact format.
[0098]
[0121] The system may include a unit that reads the generated identifier library. In one example, decoding of the nucleic acid-encoded data may be achieved by base-by-base sequencing of nucleic acid strands, such as Illumina® sequencing, or by utilizing sequencing techniques that indicate the presence or absence of specific nucleic acid sequences, such as fragment analysis by capillary electrophoresis. Sequencing may employ the use of reversible terminators. Sequencing may employ the use of natural or unnatural (e.g., recombinant) nucleotides or nucleotide analogs. Alternatively or additionally, decoding of the nucleic acid sequence may be performed using a wide variety of analytical techniques, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal. A wide variety of sequencing techniques may be used, including but not limited to polymerase chain reaction (PCR), digital PCR, Sanger sequencing, high-throughput sequencing, sequencing-by-synthesis, single-molecule sequencing, sequencing-by-ligation, RNA-Seq (Illumina), next-generation sequencing, digital gene expression (Helicos), clonal single microarrays (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, or massively parallel sequencing.
[0099]
[0122] Various readout methods can be used to extract information from the encoded nucleic acids. In one example, microarrays (or any type of fluorescent hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can also be used to read out the encoded sequences and digitally encoded data by extension. A subset of data (e.g., data belonging to a specific barcode) can be accessed from the pool by PCR using one primer that binds to the 5' barcode in the forward direction and one primer that binds to the common 3' sequence in the reverse direction.
[0100]
[0123] The accessed data may be read in the same device, or the accessed data may be transferred to another device. The reading device may include a detection unit that detects and identifies the identifier. The detection unit may be part of a sequencer, hybridization array, or other unit for identifying the presence or absence of the identifier. The sequencing platform may be specifically designed to decode and read information encoded in a nucleic acid sequence. The sequencing platform may be specialized for sequencing single-stranded or double-stranded nucleic acid molecules. The sequencing platform may decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of the entire nucleic acid sequence embedded within the nucleic acid molecule. Alternatively, the sequencing platform may be a system such as Illumina® sequencing or capillary electrophoretic fragment analysis. Alternatively or additionally, decoding of the nucleic acid sequence may be performed using a wide variety of analytical techniques implemented by the device, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal.
[0101] Vector data structures and parallel operations Combination writing
[0124] The techniques described herein utilize, for example, a combinatorial writing scheme such as that described above in this specification. In one example scheme, one starts with n non-empty sets S1, ..., Sn of oligonucleotides (oligos), where every oligo has a different sequence from every other oligo. As described above, these oligos are components, and the sets are layers. The set of all components within sets S1, ..., Sn constitutes a component library. Components within each set are assigned a fixed linear order: the first component, the second component, and so on. Using the chemistry described above, one component from each layer is concatenated in layer order to form a longer DNA molecule c1-c2-...-cn, where each c1 is a component from layer S1. As described above, these concatenated molecules constitute an identifier, and the position of each component within the identifier is also referred to as a layer: component c1 in the above example is said to occur in layer i. Every different selection of components from each layer leads to the generation of a different identifier: the set of all possible identifiers is called the combinatorial space generated by the layer. The fixed linear order of the components in each set is extended lexicographically across all identifiers that can be generated from the layer Si: thus, the combinatorial space can now be ordered in a fixed linear order and can now be used as a DNA memory. Each identifier can be used to encode a bit value corresponding to the binary choice of creating that identifier from the components.
[0102] Layer Allocation
[0125] Layers of an identifier may be assigned to serve a specific purpose. In some embodiments, some layers may be assigned to encode values or symbols (or both) and are referred to as codeword layers or data layers. Some layers may be assigned to encode addresses or keys and are referred to as address layers or key layers. Layers may be assigned to serve other purposes, multiple purposes, or no purpose. In one example implementation of the technology described herein, a set of consecutive layers, starting with layer 1 or layer n, are assigned to encode values and are therefore data layers. The remaining layers are assigned to encode keys. In some embodiments, the key layers and data layers may not be the same for all identifiers. There may be multiple sets of key layers and data layers in the same identifier library context.
[0103] Codeword Vector of Identifiers (CVI) data structure
[0126] The set of all possible identifiers in the combination space forms an identifier vector. When the identifiers are assembled, they can be used to indicate a bit value, e.g., a 1, at that location in the space. Any subset of this combination space also forms an identifier vector. A special subset of identifiers whose set of identifiers all have identical components in the key layer is said to form an identifier codeword vector (CVI). This is denoted a CVI data structure, or CVI for short. Each unique key corresponds to a unique instance of the CVI data structure. Each unique combination in the data layer is used to encode a unique data value. Each CVI instance can be used to store a data object. In some implementations of the technology described herein, a CVI can be large enough to encode an entire data object or a portion of a data object, where a data object is defined as a bit sequence that encodes an independent unit of information, as defined by the type and use of the data set. For example, a word in a sentence, an image in an image set, a hash in a hash set, a record in a table in a database, or a snippet of audio from an audio file can each be considered a data object. In some embodiments, data objects may be larger than these examples—e.g., paragraphs of text instead of words—or smaller—e.g., patches of pixels from an image. In some embodiments, a sequence of data objects may be mapped to a sequence of data object combinations. For example, a word sequence may first be converted to a sequence of word triplets. An image may first be converted to a sequence of overlapping patches of pixels and then stored in CVI. In some embodiments, a dataset may be converted to multiple resolutions using special data structures or transforms (e.g., like wavelet trees) and then stored in CVI.
[0104] Universe
[0127] In some types of datasets, each data object may be an instance of an element from a finite universe of possible data objects. For example, in the case of a text dataset, a dictionary of possible words may serve as the finite universe, and each word in a particular archive may be an instance of a word from the dictionary. This finite universe is denoted the universe of the dataset. In some implementations, the universe may not be completely known at the time of writing. In this case, a sufficiently large universe may be assumed.
[0105] Vector representation of data objects
[0128] Without wishing to be bound by theory, it is deemed sufficient to discuss, without loss of generality, elements of the universe from which data objects are taken, rather than specific instances of data objects. These elements are referred to as data values. Data values can be represented using bit sequences. However, in some implementations, data values may need to be appropriately converted to other representations to make them suitable for computation. Similarly, when data values are written to a CVI data structure, in some implementations, they may need to be converted to a bit vector suitable for CVI operations. In some implementations, the data value may be encoded using its original bit representation, where one or more bits of the original data object representation are mapped to one or more identifiers in CVI. In some implementations, the data value may first be transformed to generate a new bit vector more suitable for CVI data structures and their operations. This bit vector is referred to as a vector representation of the data value. For example, a data value may be mapped to a sparser or denser bit vector and then encoded into CVI. In some implementations, each unique data value may be mapped to a separate vector representation that differs from its original representation: for example, each unique word in a language may be mapped to a separate bit vector. In some implementations, vector representations of data values may be chosen uniformly at random from a set of all possible vectors. In some implementations, vector representations may be generated by a procedure that embeds data values into a linear algebraic vector space so as to preserve a distance metric across the data values. For example, similar patches of pixels may be mapped to similar (but not necessarily identical) vector representations, and similar, synonymous, or homophoneous words may be mapped to similar vector representations. For example, vector representations of two similar pixel patches or two synonymous words may differ only slightly in bit positions, but the bit positions may be largely identical.
[0106]
[0129] The vector representation of the data value is ultimately written to the CVI. Thus, each bit of the vector representation is mapped to one or more identifiers in the CVI. A key layer in the CVI is used to distinguish between different instances of the CVI. Thus, only the component combination in the data layer encodes each bit of the vector representation. Furthermore, the same data value written to two different CVI instances will have the same component combination in the data layer.
[0107] CVI length and weight
[0130] The length of the CVI and the vector representation of the data value may be the same, constant, or variable for all data objects or data values in a dataset. Discussed herein is an example in which the lengths of the vector representation and the CVI are the same. In one embodiment, the number of bits having the value 1 in the vector representation of a data object may be constant for all data objects in the dataset. For example, the vector representation may be 1e6 bits long, of which a set of 20 bits has the value 1. This number is called the weight of the vector representation. In some embodiments, the weight may be 1, some other constant, or a percentage of the length of the bit vector representation, such as 10%, 25%, or 50%. In some embodiments or applications, the weight may be variable: some CVIs may have a different number of 1 bits than others. In some applications, the CVI may only encode vector representations of integers or floating-point numbers and therefore may have any weight. The weight can be an important parameter of this data structure: the smaller the weight, the fewer the total number of identifiers written or read, while the larger the weight, the more signal redundancy is useful for calculations. In some embodiments, the length and weight are chosen as a function of the size of the universe. For example, if the universe has 1000 items, a CVI length and weight of 1000 and 1, respectively, can be chosen. In another example, a CVI length and weight of 20000 and 10, respectively, can be chosen. This choice can be made to fit a universe of unknown size or to improve signal redundancy during calculations.
[0108]
[0131] Described above is a CVI data structure implemented using the combined writing method, which can operate in several ways, as described below.
[0109] Parallel CVI Queries
[0132] Recall that a data value is assigned a vector representation, which is then written to an instance of a CVI. Here, each bit in the vector representation maps to one or more identifiers in the CVI. In one example embodiment, each bit maps to exactly one identifier in the CVI. A set of CVIs that all use the same universe and vector representation can be queried in parallel. Given a target query data value, its vector representation is first obtained from a list of all vector representations of all data values in the universe. This vector representation has some number of bits with value 1, e.g., w bits (equal to the weight of the vector representation). These w bits correspond to w identifiers in each CVI that contain the target data value. Each bit is written to the identifier such that only its identifier data layer encodes the bit value: the key layer is the same for all identifiers in each instance of the CVI. For example, consider w=1, and a single bit in the vector representation of the target data value is mapped to an identifier in a CVI with a data layer containing components {c1, c2, ..., c}. In that case, selecting every identifier that has this particular combination of components in the data layer selects all occurrences of this data value from all CVI instances. This extends to the case where w>1: multiple identifiers can be selected, corresponding to multiple bits. Using this method, given a dataset written to a set of CVIs, and given a query data value, all CVIs containing the target data value can be found in a number of steps that is independent of the number of CVIs or the number of occurrences of the data value in the dataset. In some implementations, the vector representations of two or more data values may share the same value in some bit positions. For example, two data values X and Y may be mapped to vector representations (1,0,1,0,1) and (1,1,0,1,0). Using the above query method, a search for data value X matches three bits of the vector representation of X, but also matches one bit—bit number 1—of the vector representation of Y. This issue is addressed by calculating the 1-norm of the CVI after the query, as described below.
[0110]
[0133] One benefit of CVI queries is that the queries are executed in parallel on all CVIs, so that the resources required are independent of the number of CVIs and, therefore, the number of data objects in the dataset. The CVI query techniques described herein can be used with hundreds of millions to trillions of CVIs. The queries narrow the dataset down to a small number of data objects, reducing the read load to a fraction of the cost of reading the entire dataset.
[0111]
[0134] Figure 20 illustrates several approaches to selecting identifiers that contain a particular query in a subset of layers. An example library includes a set of identifiers with four components. The "target query layer" represents a particular query component in the first and second target layers. In the figure, the top two identifiers have both the first and second target components. The third identifier has a non-query component in the first layer, and the fourth identifier has a non-query component in the second layer.
[0112]
[0135] In Method #1, oligonucleotides (e.g., primers) specific to the query component can be used in PCR to selectively amplify only identifiers containing the query component. This can be done recursively, reducing the number of targets with each amplification round until only identifiers that perfectly match the query remain. Several methods can be employed to remove background non-target identifiers. First, the average number of copies per identifier used as PCR input can be tightly controlled to avoid adding excessive background while still preserving all target identifiers. Furthermore, the number of cycles used in the PCR reaction can be optimized to be as low as possible to avoid amplifying non-targets while still sufficiently amplifying the target. Finally, within a round of PCR selection, reactions can be subsampled. This can be done by removing a portion of the PCR reaction after just a few cycles and using that portion as input for a new PCR reaction. The subsampled input will likely contain copies of the amplified target identifiers, but will stochastically not contain all of the non-target identifiers that were not amplified. Multiple iterations of this subsampling ensure optimal removal of non-target background.
[0113]
[0136] In Method #2, oligonucleotides specific to the query component can be used as a guide to direct an endonuclease to a target identifier. If a match is found, the endonuclease cleaves the identifier to shorten it. These shortened identifiers can be selected based on size. This process can be done recursively, reducing the number of targets with each round until only identifiers that perfectly match the query remain.
[0114]
[0137] In Method #3, oligonucleotides specific to the query component can be attached to beads (e.g., magnetic beads) or a surface and then mixed with a single-stranded version of the library (alternatively, DNA can be mixed and hybridized, and then oligonucleotides can be attached). Only identifiers containing the query component will hybridize, and thus only identifiers attached to the beads / surface can be pulled from solution, while other identifiers can be removed by washing. This process can be done recursively, reducing the number of targets with each round until only identifiers that exactly match the query remain.
[0115]
[0138] In Method #4, the single-stranded version of the library can be hybridized with oligonucleotides corresponding to all components in the unknown layer, and with nucleotides corresponding only to the query component in a subset of layers encoding bit positions. After hybridization, an endonuclease targeting single-stranded DNA can be added, thereby cleaving all identifiers present in single-stranded DNA, while all fully hybridized target identifiers become double-stranded and are therefore protected from nuclease cleavage. Full-length identifiers can then be selected based on size. In some embodiments, the single-stranded version of the library can be immobilized, for example, on beads or a surface, so that after hybridization with the query component, the newly formed strand can be denatured and released from the surface.
[0116] Regular sampling
[0139] In some embodiments, instead of querying specific data values, statistics on the vector representation may be desired. For example, one may want to estimate the relative proportion of data objects in which a bit from a particular subset is set to 1 in the vector representation. In some subsets, the bit positions of interest may be expressed as a set of query operations, as described in the previous query section. For example, bit positions such as the first half of the vector representation or alternative bit positions in the vector representation, i.e., every third and seventh bit position, may be translated into a union of a small number of query operations. These query operations may be combined to estimate the Fourier spectrum of a Boolean function. In one embodiment, the Boolean function is a vector representation of the data values of interest. In one embodiment, quantitative PCR of samples generated from the union of the query operations may estimate the relative proportion of data objects having a 1 bit in the subset of positions.
[0117] Reserved Bits and Incremental Writes
[0140] In one embodiment, certain bit positions in the vector representation may be reserved for a particular purpose at any point in the life of the archive. For example, the codeword length may be 20,000 identifiers, the weight may be 10, with all 10 bits occurring in the first 10,000 identifier positions of the CVI, and the remaining 10,000 identifiers being reserved. The reserved identifiers may be used to incrementally write additional bits of information after the initial data values are written. For example, an initial data set may first be written using the scheme described above. Later, one may wish to modify the written archive in some way. These modifications may be encoded in the reserved bits. In one application, the desired modification may be to encode a relationship between a set of data objects in the archive that was unknown at the time of the initial write. In this application, a particular reserved bit in the vector representation may be set to 1 by writing the identifier corresponding to that bit position to the CVI for the set of related data objects. In some embodiments, more than one bit may be used to encode the relationship. In some embodiments, the strength of the relationship may be encoded by the number of copies of the identifier in the CVI representing the bit encoding the relationship. For example, a weaker relationship between data objects may be encoded by writing fewer copies of the identifier representing that bit in the vector representation compared to a reference copy number, while a stronger relationship may be encoded by writing more copies of the identifier. This copy number modulation can be achieved at write time by increasing or decreasing the number of reaction copies created to assemble a particular reserved bit identifier. When the data objects are read, stronger relationships are more likely to be read immediately than weaker relationships. This copy number modulation can be used to encode the importance or other aspects of the data objects. In some embodiments, other aspects such as the version number of the data object or its deletion status can be encoded in the reserved bit.
[0118] Parallel CVI Editing
[0141] Data values written to a CVI instance can be modified. As described above, data values are encoded into a vector representation of data, with bits in the vector representation mapping to component combinations in the data layer of an identifier in the CVI. For example, consider w=1, where a single bit in the vector representation of a target data value maps to an identifier in a CVI with a data layer containing components {c1, c2, ..., ck}. Changing component ck to the next component in that set then has the effect of moving the identifier to the next identifier in the CVI. This has the effect of setting the corresponding bit in the vector representation to 0 and setting the bit in the next position to 1. This can be extended to multiple bits if w>1. This edit can be performed on the original bits, and new bits can be written as described above. In another embodiment, specific identifiers can be targeted and deleted, and new bits can be written as described above, thereby achieving the effect of editing the CVI.
[0119]
[0142] Described herein are several methods for achieving CVI editing. Figure 21 illustrates several approaches for changing the identifier. Directed editing uses guide nucleotides to target a programmable nuclease-recombinase enzyme complex to the component to be changed and to target a new template oligonucleotide for the recombinase to use to replace the component of interest. PCR-based editing uses an oligonucleotide that only partially matches the identifier as one of the PCR primers. The remaining portion of the oligonucleotide contains the sequence of the altered adjacent component. During rounds of PCR amplification, the mismatched tail sequence is incorporated into the newly generated identifier, thereby replacing the old component sequence. Adding a new layer or sequence tag can be achieved by using an oligonucleotide that partially matches the identifier as one of the PCR primers. The remaining portion of the oligonucleotide contains the sequence of the new layer or tag. During rounds of PCR amplification, the mismatched tail sequence is added to the newly generated identifier, thereby adding a new layer or other sequence to the end of the identifier.
[0120]
[0143] An important benefit of CVI is that it allows us to write identifiers corresponding to reserved bits without modifying existing identifiers, simply by combining them with the previous pool of identifiers. Similarly, bit value editing can be done in parallel in several steps with increasing optimization, independent of the data size, e.g., from hundreds of millions of CVI instances to trillions.
[0121] Parallel CVI 1-norm calculation
[0144] The 1-norm of a vector is the sum of the absolute values that comprise the vector: for example, the 1-norm of the vector (1,-1,2) is 4. Given a set of CVI instances, we can convert them to their individual 1-norms in parallel. First, we define the concept of a sub-identifier. Given an identifier, a sub-identifier is a DNA fragment that contains a subset of the contiguous layers from the identifier. For example, in a CVI, the key layer of an identifier creates a sub-identifier when it is separated from the identifier. We refer to this new operation, where the key layer of an identifier is separated from the data layer, as key extraction. In a CVI, all identifiers are identical in the key layer and only differ in the data layer. Therefore, when the key layer is extracted from a CVI, the sub-identifiers generated from each identifier in the CVI are identical. If a CVI has weight w, and each of w identifiers appears in c molecular copies, then approximately w * There is a single sub-identifier subset with c copies. If the average number of copies c is known from measurements or from the design of the identifier library, then the w * The c copies encode the 1-norm of the CVI in the form of a copy count of the key sub-identifier.
[0122]
[0145] This 1-norm encoding can be used in subsequent computation or reading steps via sequencing. In one embodiment, after a parallel CVI query for all w identifiers of the target data value, a key extraction operation is performed on the generated identifiers. This generates a sub-identifier, i.e., a key, from every CVI containing a data value that matches the query data value. This sub-identifier is in the wc copy and signals the CVI instance containing the queried data value. In comparison, a CVI in which only v bits out of w bits match the vector representation of the query data value will have a sub-identifier in the vc copy. Careful design of the vector representation of each data value in the universe can ensure that v is much smaller than w. Thus, a CVI that matches the query will have a sub-identifier in a much higher copy count than a CVI that only partially matches the query.
[0123]
[0146] In one embodiment, an identifier query is performed. The generated library contains all identifiers that match the queried component in the data layer. Thus, all data objects in the dataset with the queried bit value set to 1 are identified and their queried bit is retrieved. The retrieved identifiers are assumed to occur at some standard average copy count of c copies per identifier. This copy count can be multiplied by a given scalar as follows: If the given scalar is s, and s=2^0 * b_0+2^1 * b_1+…+2^(k-2) * b_(k-2)+2^(k-1) * b_(k-1) Consider a positive k-bit binary representation given by: Take the library generated from the query and amplify it k-fold using polymerase chain reaction (PCR). Then split into k aliquots, labeling each aliquot with a number between 1 and k-1. Next, take the i-th aliquot and look up the bit number b_i in the binary representation of a given scalar s. If bit b_i is 0, destroy that aliquot. If bit b_i is 1, run i cycles of PCR on that aliquot. After i PCR cycles, each identifier in the aliquot will be approximately 2^i * c copies. All aliquots on which PCR was performed are then pooled. In this final sample, each identifier will be approximately s * In this way, we can multiply the number of copies of an identifier by any known scalar and obtain the result in unary form.
[0124]
[0147] In one implementation, this operation can be used to convert a binary representation of a number to its unary representation. For example, consider that CVI encodes a non-negative integer value encoded in binary form. Consider that this integer uses w bits. Using the query operation described above, an identifier representing the i-th bit of each data object's vector representation is extracted. Their copy counts are multiplied by the scalar 2^i. Key extraction is performed on the resulting libraries, and then the w libraries are pooled together. The result of this series of operations is a data value written in binary form in CVI, e.g., v=2^i. * b_0+2^1 * b_1+…+2^(w-2) * b_(w-2)+2^(w-1) * b_(w-1) is now converted to its unary form and encoded into the copy count of the subidentifier.
[0125]
[0148] In another embodiment, a known vector v = (v_0, v_1, ..., v_(w-1)) containing w non-negative values is multiplied by each CVI to calculate the dot product of the CVI and v. This can be done using the same technique as above, except that instead of the scalar 2̂i, a given scalar v_i is used in each multiplication.
[0126] Parallel addition and subtraction of CVI pairs in unary form.
[0149] In some embodiments, two data sets can be written to two sets of CVIs. Key extraction can be performed on some subset of the CVIs in both data sets to convert data values to unary form. When two identifier libraries are pooled, the resulting library contains CVIs that encode data values that are the sum of the data values encoded by the CVIs in the initial data sets.
[0127]
[0150] In some embodiments, two data sets may be written into two CVI sets, A and B. Key extraction may be performed on some subset of the CVIs in both data sets to convert the data values to unary form. An approximate subtraction, AB, may then be performed in unary form as follows: Each sub-identifier has two strands, referred to herein as the plus and minus strands. In one embodiment, the sub-identifiers in A undergo a chemical procedure to melt the sub-identifiers into their constituent strands, preserving only the plus strand of each sub-identifier. Similarly, the sub-identifiers in B undergo a chemical procedure to preserve only the minus strand of the sub-identifier. Importantly, the chemical procedure used to melt the sub-identifiers and preserve the strands does not change the relative copy counts of the sub-identifiers in A and B. The plus strands of the sub-identifiers in A are then pooled with the minus strands of the sub-identifiers in B. The pooled sample is held at a temperature that promotes re-annealing of the plus and minus strands if they are perfectly complementary. For each identifier present in A and B, the copy number of the double-stranded subidentifier is min(c,c'), where c is the copy number of the subidentifier in A and c' is the copy number of the subidentifier in B. The copy number of the plus single strand of each subidentifier is c-c' if c>c'. The copy number of the minus single strand of each subidentifier is c'-c if c'>c. If c=c', there are no single-stranded copies of that subidentifier. After pooling and annealing, the plus single stranded subidentifiers, minus single stranded subidentifiers, or double-stranded subidentifiers may be separated. The plus single stranded subidentifiers and minus single stranded subidentifiers may then be converted to double-stranded forms in separate reactions. The copy number of the double-stranded subidentifier obtained from the plus single stranded product is represented as strand c-c' if c>c'. Similarly, the copy number of the double-stranded subidentifier obtained from the minus single stranded product is represented as c'-c if c'>c. The copy number of the double-stranded sub-identifier obtained directly from pooling and reannealing represents the minimum data value encoded between two sub-identifiers with the same key. In this way, data values encoded by CVI can be subtracted in parallel.
[0128]
[0151] Described herein are some examples of chemical methods for performing the subtraction described above. In one embodiment, one of the strands of each sub-identifier can be attached to a bead or a protein molecule, such as biotin, and this bead or protein "handle" can be used to separate and remove all single-stranded sub-identifiers. In another embodiment, methods such as high performance liquid chromatography can be used to separate single-stranded and double-stranded sub-identifiers. Single-stranded sub-identifiers from either the plus or minus strand can be converted to double-stranded form using PCR.
[0129]
[0152] In some embodiments, the addition and subtraction can be performed using a hybrid approach, where some steps are performed using chemistry and others are performed using conventional computers after sequencing. In one embodiment, two CVI libraries can be subjected to key extraction and sequencing. The sequencing output can be analyzed and converted into read counts for each sub-identifier in each library. These counts can then be calculated using a conventional computer, for example, multiplied by a weight and added or subtracted from each other.
[0130] Transverse Striped Writing
[0153] In some embodiments, the CVI encoding a data object may be written to several separate samples. For example, if the CVI length is N identifiers and there are M data objects, the data is encoded into DNA by writing N separate DNA libraries, each with M identifiers. This is referred to herein as horizontal stripe writing (TSW). (This approach is similar to the vertical approach, in which a key layer of identifiers encodes the identity of the data object to be encoded, and one or more data layers encode portions of the data.) This approach differs from approaches in which a single DNA library containing N identifiers encodes a data set, with all N identifiers having distinct sequences. In the TSW approach, all identifiers in each of the N libraries have distinct sequences, but do not individually encode complete data objects. Each of the M identifiers encodes portions of a different data object. However, a set of N identifiers encoding a single data object resides in separate DNA libraries that all have the same sequence but cannot be pooled without losing some information. Thus, each of the N libraries is a "stripe" from the complete data object, encoding only a portion.
[0131]
[0154] An advantage of TSW is that it prepares the identifiers in a state that facilitates parallel addition and subtraction, as described above: all identifiers that encode a single data object have the same sequence and can therefore be pooled or reannealed with the intended result, which would not be possible if the identifiers that encoded the data object had different sequences.
[0132]
[0155] In another embodiment, a non-striped DNA library can be converted to striped form by performing selection using the methods described above. For example, the i-th identifier corresponding to the i-th bit in the vector representation of the data object can first be selected using the method described in the CVI query section. The selected identifier or sub-identifier can then undergo key extraction. This can be performed for all w identifiers in the CVI. The resulting w separate libraries all have the same key layer and therefore have the same sequence, which is compatible with the addition and subtraction methods described above. If the CVI weights are not constant across all CVIs, different numbers and positions of bits can be converted to striped form.
[0133] example
[0156] In one application of the technology described herein, text data can be written to a CVI as follows: First, the text data is scanned to list all unique words present in the text. These words serve as the universe. In one particular implementation, a text dataset containing 208,183 words taken from a universe of 13,051 (unique) words was used. Each word in the universe is assigned a bit vector representation that is n bits long, with only w bits set to 1. For example, in a particular implementation, n=18,750 and w=10. The text is scanned, and sequences of words are converted into sequences of bit vectors. The bit vectors are written to DNA using the combinatorial writing scheme for identifiers described above in this specification. In one implementation, this results in 2,081,830 distinct identifiers. Each identifier is 14 layers long, with six layers (L0-L5) being data layers and eight layers (L6-L13) being key layers. Next, a query word is selected from the universe, and its vector representation, taken from the list of vector representations assigned to that word, is assigned to a word in the universe. From the vector representation and layer allocation scheme, a specific six-component combination is obtained corresponding to each of the 10 bits of the vector representation of the query word. Using recursion technique #1 shown in FIG. 20, the identifier library is queried with each six-component combination, recursively, one component at a time, starting with L0 and progressing through L5. The resulting libraries are pooled, and key extraction is performed to generate a library of sub-identifiers comprising layers L6-L13. The resulting libraries are sequenced with a sequencing load that is a fraction of the load of reading the entire text dataset. In some implementations, depending on the frequency of the query word in the text archive, this load may be as little as 1% of the load of reading the entire library.
[0134]
[0157] In another example embodiment, the CVI may encode a particular data object that needs to be searched based on a score given by the model. The model may be a linear model, where the score is given by a linear function of each bit in the vector representation of the data object. In one example embodiment, a match may be defined as a data object whose score exceeds a certain threshold. For example, the data object may be an image encoded into a single CVI, and the linear model may be a linear function of various pixel locations in the vector representation. The linear model may be of the form a_0 * x_0+a_1 * x_1+…+a_(n-1) * x_(n-1)>b, where a_i is a coefficient, x_i is a position in the vector representation of the image, and b is a match threshold. Each identifier corresponding to a pixel location can be retrieved using the CVI query method described above. Alternatively, the image can be written into n separate stripes using TSW. Each pixel location x_i can be multiplied in parallel by its coefficient a_i using the multiplication method described above. The results can be added using the pooling method for addition described above, pooling all n libraries into a single library. The result can be checked against threshold b using the subtraction method described above. In this case, a library encoding the value b in each possible key is first constructed and then used as an operand in the subtraction, as described above. Alternatively, the comparison to the threshold can be performed on a conventional computer, as described above. In another embodiment, the result of this linear model calculation can serve as input to another linear model calculation, resulting in a network of linear models similar to a multilayer perceptron.
[0135]
[0158] Identifying identifiers in an identifier library can be performed using any identification or sequencing method. Figure 22 shows an example of a method for reading encoded information by hybridization. The readout unit can include one or more hybridization arrays. The hybridization array can include identifiers 2201 bound to a surface or support 2202. The identifiers can be spatially oriented to enable single-molecule or molecular population resolution using optical detection. Probe sequences 2203 that share complementarity with one or more components of the identifiers can be introduced into the array. The probe sequences can include one or more fluorophores 2204. In one example, the probe includes a fluorophore and a quencher 2205. The quencher can be another dye, fluorophore, or quenchbody. Hybridization of the probe to the identifier can separate the fluorophore and quencher, creating a detectable signal. In other embodiments, the probe includes an array of fluorophores that can be detected as an optical signature indicative of a particular probe or a particular probe set. Individual components can be detected by optical imaging of the area, such as confocal techniques, or scanning of the area. Sequential introduction of probes, imaging of probes, and removal of probes can be used to identify some or all of the components on a given identifier. There may be no limit to the number of components that can be identified at one time. Probes to different components may have different optical signatures or the same optical signature.
[0136]
[0159] Another method for detecting identifier sequences can include nanopore sequencing. Figure 23 shows an example of a nanopore sequencing readout method. Molecules can have a unique impedance signature when they move through a pore or channel to which a voltage is applied. Several existing nucleic acid sequencing platforms use this property to identify the sequence of base pairs in nucleic acid molecules. These platforms have the advantage of being able to sequence longer molecules of nucleic acid and detect the presence or absence of non-natural nucleotides and chemical moieties that can be used to modify both natural and non-natural nucleotides. In one example, identifier sequence 2303 is combined with probe 2304 that hybridizes to a component of the identifier sequence. The probe can include a molecule that generates a unique impedance signal while moving through pore 2301. The probe or channel can be microfabricated to the nanometer scale in substrate 2302, which can include a biological membrane or a crystalline material. Alternatively or additionally, each component within each layer can include a unique molecule that generates a unique impedance signature. The unique molecules may include sequence-based nucleotide / protein / hybrid tags, chemical modifications of nucleotides, fluorescent probes, or any combination thereof. In some embodiments, the signal may be a current through the probe or channel, while in other embodiments, the detectable signal is detected by an impedance detector adjacent to the probe or channel. The burst of signal 2305 provides a signature indicative of the individual identifier.
[0137]
[0160] Systems for encoding, writing, and reading data stored in nucleic acid molecules may be automated or non-automated. The systems may be networked to allow cloud-based access to the data, or they may not be networked. The systems may be capable of operating in zero- or low-gravity environments and / or at high or low atmospheric pressure or vacuum. The systems may be shielded from electromagnetic and other radiation to avoid degradation of the identifiers and other internal electronics, chemicals, and enzymes. The systems may use or include an external power source. The systems may include a method for generating electricity. One or more of the units of the system may be modules and mobile devices. The modules or mobile devices may be installed in or embedded in third-party vehicles. One or more of the units or modules of the system may physically or digitally interact with an external machine. For example, the system may receive physical or digital input from an external machine, or the system may output physical material or digital information to an external machine.
[0138]
[0161] Storing information in nucleic acid molecules can have a variety of applications, including, but not limited to, long-term information storage, confidential information storage, and medical information storage. In one example, a person's medical information (e.g., medical history and medical records) can be stored in a nucleic acid molecule and given to the person. The information can be stored outside the body (e.g., in a wearable device) or inside the body (e.g., in a subcutaneous capsule). When the patient is brought to a clinic or hospital, a sample can be taken from the device or capsule, and the information can be decoded using a nucleic acid sequencer. Personal storage of medical records in nucleic acid molecules can provide an alternative to computer and cloud-based storage systems. Personal storage of medical records in nucleic acid molecules can reduce the incidence or prevalence of medical record hacking. Nucleic acid molecules used in capsule-based storage of medical records can be derived from human genome sequences. The use of human genome sequences can reduce the immunogenicity of the nucleic acid sequences in the event of capsule failure or leakage.
[0139]
[0162] The combinatorial assembly method described herein can be used to create DNA libraries for encoding amino acid chains. The amino acid chains can be peptides or proteins. The DNA components can form junctions along functionally or structurally inactive codons that can be common to all members of the combinatorial library. The DNA components can form junctions along introns so that the processed peptide or protein has no gaps between the variable amino acid chains. Each combinatorial DNA molecule can be assembled in a separate reaction chamber. In vivo expression assays can be performed to detect expression. Each combinatorial DNA molecule can be pooled together, and individual in vitro expression assays can be performed by encapsulating the molecules in droplets. In vivo expression assays can be performed by transforming the molecules into cells. The DNA can act as a barcode so that cells and droplets containing specific amino acid chain variants can be identified. The assay can have a fluorescent output, allowing cells / droplets to be sorted into bins by fluorescence intensity and sequenced to correlate each combinatorial DNA sequence with a specific output. The combinatorial DNA molecule can encode RNA. If the output itself is RNA abundance (e.g., RNA aptamer screening and testing), the pooled assay can be performed outside of droplets or cells. Combinatorial DNA can encode combinations of CRISPR gRNAs or microRNAs that up- or down-regulate genes inside cells. To test how combinatorial gene regulation affects cellular properties during cellular perturbations, combinatorial DNA libraries can be transformed into cells. Combinatorial DNA libraries can encode combinations of genes in pathways. Each DNA component can contain a gene expression construct, and the DNA components can form junctions along inactive DNA sequences between genes. DNA sequences can be transformed into cells to examine how different combinations of gene overexpression affect cellular properties during different cellular perturbations.
[0140] Computer Control System
[0163] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 24 shows a computer system 2401 programmed or otherwise configured to encode digital information into a nucleic acid sequence and / or read (e.g., decode) information derived from a nucleic acid sequence. The computer system 2401 can coordinate various aspects of the encoding and decoding procedures of the present disclosure, such as, for example, bit value and bit location information for a given bit or byte from an encoded bitstream or bytestream.
[0141]
[0164] The computer system 2401 includes one central processing unit (CPU, also referred to herein as "processor" and "computer processor") 2405, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 2401 also includes memory or memory locations 2410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 2415 (e.g., a hard disk), a communication interface 2420 (e.g., a network adapter) for communicating with one or more other systems, and peripherals 2425, such as cache, other memory, data storage, and / or electronic display adapters. The memory 2410, storage unit 2415, interface 2420, and peripherals 2425 communicate with the CPU 2405 through a communication bus (solid lines) such as a motherboard. The storage unit 2415 may be a data storage unit (or data repository) for storing data. The computer system 2401 may be operably coupled to a computer network ("network") 2430 using the communication interface 2420. Network 2430 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 2430 is a telecommunications network and / or a data network. Network 2430 may include one or more computer servers that may enable distributed computing, such as cloud computing. In some cases, network 2430 may implement a peer-to-peer network using computer system 2401, which may enable devices coupled to computer system 2401 to act as clients or servers.
[0142]
[0165] The CPU 2405 can execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 2410. The instructions may be directed to the CPU 2405, which may then be programmed or otherwise configured to perform the methods of the present disclosure. Examples of operations performed by the CPU 2405 may include fetch, decode, execute, and writeback.
[0143]
[0166] The CPU 2405 may be part of a circuit, such as an integrated circuit. One or more other components of the system 2401 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0144]
[0167] Storage unit 2415 can store files such as drivers, libraries, and saved programs. Storage unit 2415 can store user data, such as user preferences and user programs. Computer system 2401 may, in some cases, include one or more additional data storage units external to computer system 2401, such as located on a remote server that communicates with computer system 2401 through an intranet or the Internet.
[0145]
[0168] Computer system 2401 can communicate with one or more remote computer systems through network 2430. For example, computer system 2401 can communicate with a user's remote computer system and / or with a machine that can be used by the user in the process of analyzing data encoded or decoded in a nucleic acid sequence (e.g., a sequencing or other system for chemically identifying the order of nitrogenous bases in a nucleic acid sequence). Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 2401 via network 2430.
[0146]
[0169] The methods described herein may be implemented by machine (e.g., a computer processor) executable code stored in an electronic storage location of computer system 2401, such as memory 2410 or electronic storage unit 2415. The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by processor 2405. In some cases, the code may be retrieved from storage unit 2415 and stored in memory 2410 for easy access by processor 2405. In some situations, electronic storage unit 2415 may be omitted, and machine-executable instructions may be stored in memory 2410.
[0147]
[0170] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or run-time compiled manner.
[0148]
[0171] Aspects of the systems and methods provided herein, such as computer system 2401, can be embodied in programming. Various aspects of the present technology can be thought of as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried or embodied on some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage"-type media can include any tangible memory of a computer, processor, etc., or associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage for software programming at any time. All or portions of the software may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable software to be loaded from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Accordingly, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices through wired and optical landline networks and via various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0149]
[0172] Thus, a machine-readable medium, such as a computer-executable code, may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer, such as may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs and EPROMs, Flash EPROMs, any other memory chips or cartridges, carrier waves transporting data or instructions, cables or links transporting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0150]
[0173] The computer system 2401 may include or communicate with an electronic display 2435 including a user interface (UI) 2440 to provide sequence output data including, for example, sequences and bits, bytes, or bitstreams that are encoded or read by a machine or computer system encoding or decoding, for example, nucleic acids to be encoded or decoded into DNA storage data, raw data, files, and compressed or decompressed zip files. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0151]
[0174] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented in software when executed by the central processing unit 2405. The algorithms can be used in conjunction with, for example, a DNA index and raw data or zip file compressed or decompressed data to determine a customized method of coding digital information from the raw data or zip file compressed data prior to encoding the digital information.
[0152]
[0175] While preferred embodiments of the present invention have been illustrated and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. While the present invention has been described with reference to the above specification, the descriptions and illustrations of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it is to be understood that all aspects of the present invention are not limited to the specific illustrations, configurations, or relative parts set forth herein, which depend upon a variety of conditions and variables. It is to be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the present invention shall encompass any and all such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of the claims and their equivalents be covered thereby. [Example]
[0153] Example
[0176] In one aspect, the disclosure provides a method for encoding digital information into a nucleic acid sequence, the method including: (a) encoding the digital information into a symbol sequence and converting the symbol sequence into a codeword using one or more codebooks; (b) parsing the codeword into a coded symbol sequence; (c) mapping the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; (d) enumerating an identifier library, each symbol of the coded symbol sequence being encoded by one or more identifiers; and (e) attaching a description of the one or more codebooks and the plurality of identifiers to the identifier library.
[0154]
[0177] In some embodiments, the coded symbol sequence includes symbols taken from a fixed alphabet of symbols. In some embodiments, the method further includes converting the coded sequence to a second symbol sequence. In some embodiments, the second symbol sequence includes a formal data structure. In some embodiments, the formal data structure includes one or more members selected from the group consisting of a tree structure, a trie structure, a table structure, a key-value dictionary structure, and a set. In some embodiments, the formal data structure is queriable by a range query, a rank query, a count query, a membership query, a nearest neighbor query, a match query, a select query, or any combination thereof.
[0155]
[0178] In some embodiments, the method further includes parsing the second symbol sequence into a word sequence. In some embodiments, the method further includes converting the word sequence into a codeword sequence using one or more codebooks. In some embodiments, the method further includes converting the codeword sequence into a third symbol sequence. In some embodiments, converting the word sequence into a codeword sequence minimizes the number of one or more types of symbols in the third symbol sequence.
[0156]
[0179] In some embodiments, the coded symbol sequence includes one or more symbol blocks. In some embodiments, converting the word sequence to the codeword sequence generates a fixed number of one or more types of symbols in each symbol block of the one or more symbol blocks in the third symbol sequence. In some embodiments, the codebook attaches one or more error protection symbols to individual codewords of the codeword sequence. In some embodiments, the one or more error protection symbols are calculated from one or more words of the word sequence.
[0157]
[0180] In some embodiments, the plurality of identifiers is selected from a combinatorial space of identifiers. In some embodiments, each identifier of the plurality of identifiers comprises one or more components. In some embodiments, each component of the one or more components comprises a nucleic acid sequence. In some embodiments, the nucleic acid sequences are distinct sequences.
[0158]
[0181] In some embodiments, each symbol in the symbol string is one of two possible symbol values. In some embodiments, one symbol value at each position in the symbol string can be represented by the absence of a distinct identifier in the identifier library. In some embodiments, the two possible symbol values are bit values of 0 or 1, and an individual symbol in the symbol string having a bit value of 0 can be represented by the absence of the distinct identifier in the identifier library, and an individual symbol in the symbol string having a bit value of 1 can be represented by the presence of the distinct identifier in the identifier library, or vice versa. In some embodiments, the presence of an individual identifier in the identifier library corresponds to a first symbol value in the binary string, and the absence of an individual identifier from the identifier library corresponds to a second symbol value in the binary string. In some embodiments, the first symbol value is "1" and the second symbol value is "0." In some embodiments, the first symbol value is "0" and the second symbol value is "1." In some embodiments, the identifier library includes a supplemental nucleic acid sequence. In some embodiments, the supplemental nucleic acid sequence includes metadata about or an encoding of the first symbol sequence. In some embodiments, the supplemental nucleic acid sequence does not correspond to the digital information, and the supplemental nucleic acid sequence hides the digital information encoded in the identifier library.
[0159]
[0182] In some embodiments, the one or more identifiers are generated by combinatorial assembly of one or more components. In some embodiments, the method further includes constructing a universal identifier library. In some embodiments, the identifier library is constructed from the universal identifier library by decomposing or excluding individual identifiers not present in the identifier library. In some embodiments, constructing the universal identifier library includes using one or more reactions. In some embodiments, one or more reactions corresponding to individual identifiers not present in the identifier library are removed, eliminated, decomposed, or inhibited. In some embodiments, the one or more reactions include components, templates, and / or reagents, and the components, templates, and / or reagents are loaded onto a film, yarn, fiber, or other substrate. In some embodiments, the components, templates, and / or reagents are positioned adjacent to one another by stamping, intertwining, knitting, pinching, or weaving the film, yarn, fiber, or other substrate.
[0160]
[0183] In another aspect, the present disclosure provides a system for encoding digital information into nucleic acid sequences, the system including: an assembly unit configured to generate an identifier library encoding a symbol sequence, the identifier library including at least a subset of a plurality of identifiers; and one or more computer processors operably coupled to the assembly unit, the one or more computer processors individually or collectively programmed to: (i) encode the digital information into a symbol sequence and convert the symbol sequence into a codeword using one or more codebooks; (ii) parse the codeword into a coded symbol sequence; (iii) map the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; (iv) instruct the assembly unit to generate the identifier library, wherein each symbol of the coded symbol sequence is encoded by one or more identifiers; and (v) instruct the assembly unit to append a description of the one or more codebooks and the plurality of identifiers to the identifier library.
[0161]
[0184] In some embodiments, one or more identifiers are assembled in one or more assembly reactions, hi some embodiments, one or more products of the one or more assembly reactions are combined to generate an identifier library.
[0162]
[0185] In some embodiments, the assembly unit includes one or more containers. In some embodiments, the one or more containers are partitions. In some embodiments, the assembly unit includes reagents, one or more layers of components, one or more templates, or any combination thereof. In some embodiments, the assembly unit is configured to receive reagents, one or more layers of components, one or more templates, or any combination thereof. In some embodiments, the assembly unit is configured to output an identifier library.
[0163]
[0186] In some embodiments, the assembly unit includes a reaction module. In some embodiments, the reaction module is configured to collect reagents, one or more layers, one or more templates, or any combination thereof. In some embodiments, the reagents include an enzyme, one or more nucleic acid sequences, a buffer, a cofactor, or any combination thereof. In some embodiments, the reagents are combined into a master mix before entering the reaction module. In some embodiments, the reaction module is configured to incubate or agitate an assembly reaction, wherein the assembly reaction produces one or more identifiers. In some embodiments, the reaction module includes a detector unit, wherein the detector unit monitors the assembly of the one or more identifiers.
[0164]
[0187] In some embodiments, the system further includes a storage unit, and the assembly unit transfers the generated identifier library to the storage unit. In some embodiments, the storage unit includes one or more pools, receptacles, or partitions. In some embodiments, the storage unit combines one or more identifier libraries into one or more pools, one or more receptacles, or one or more partitions.
[0165]
[0188] In some embodiments, the system further comprises a selection unit configured to select one or more identifiers, hi some embodiments, the selection unit comprises a size selection module, an affinity capture module, a nuclease cleavage module, or any combination thereof.
[0166]
[0189] In some embodiments, the system further comprises a nucleic acid synthesis unit configured to synthesize one or more nucleic acid sequences, hi some embodiments, the one or more nucleic acid sequences are constructed using base-by-base synthesis.
[0167]
[0190] In some embodiments, the assembly unit generates a plurality of reactions for assembling one or more identifiers, hi some embodiments, the assembly unit selectively removes individual reactions from the plurality of reactions that do not generate at least a subset of the plurality of identifiers in the identifier library.
[0168]
[0191] In some embodiments, the assembly unit generates the identifier library using one or more of electrowetting, misting, printing, laser ablation, weaving or knitting or intertwining of materials coated with nucleic acid sequences, slip technology, stamping, laser printing, or droplet microfluidics.
[0169]
[0192] In some embodiments, the one or more computer processors are individually or collectively programmed to use heuristic techniques to minimize the number of reactions for generating the identifier library or to minimize the time it takes to prepare the number of reactions for generating the identifier library, hi some embodiments, the heuristic techniques include on-set covering heuristics or heuristics that minimize the travel path of the device.
[0170]
[0193] In another aspect, the present disclosure provides a nucleic acid-based integrated storage system, the system comprising: a data encoding unit configured to write digital information to one or more nucleic acid sequences, the data encoding unit writing the digital information to the one or more nucleic acid sequences in the absence of base-by-base nucleic acid synthesis; a storage unit configured to store one or more nucleic acid sequences encoding the digital information; a reading unit configured to access and read the digital information encoded in the one or more nucleic acid sequences; and one or more computer processors operably coupled to the data encoding unit, the storage unit, and the reading unit, wherein the one or more computer processors are individually or collectively programmed to: (i) direct the data encoding unit to encode the digital information into the one or more nucleic acid sequences; (ii) direct the storage unit to store the digital information encoded in the one or more nucleic acid sequences; and (iii) direct the reading unit to access and decode the digital information stored in the one or more nucleic acid sequences.
[0171]
[0194] In some embodiments, one or more computer processors parse the digital information into a plurality of symbols. In some embodiments, the plurality of symbols are mapped to a plurality of identifiers. In some embodiments, each symbol of the plurality of symbols corresponds to one or more identifiers of the plurality of identifiers. In some embodiments, the plurality of identifiers comprises a plurality of components. In some embodiments, each component of the plurality of components comprises a distinct nucleic acid sequence.
[0172]
[0195] In some embodiments, the data encoding unit generates one or more identifier libraries that include one or more sets of identifiers corresponding to the digital information, and in some embodiments, reading the digital information includes identifying one or more sets of identifiers in the one or more identifier libraries.
[0173]
[0196] In some embodiments, the system is automated. In some embodiments, the system is networked. In some embodiments, the system is configured to operate in a zero or low gravity environment. In some embodiments, the system is configured to operate at sub-atmospheric pressure, vacuum, or super-atmospheric pressure. In some embodiments, the system includes a power source or a method for generating electricity. In some embodiments, the system includes a radiation shield.
[0174]
[0197] In some embodiments, the generated identifier library is a universal library. In some embodiments, the system further comprises multiple modules. In some embodiments, a first module generates the identifier library. In some embodiments, a second module performs deletion of individual identifiers or identifier reactions. In some embodiments, a third module separates individual identifiers present in the identifier library from individual identifiers not present in the identifier library. In some embodiments, a fourth module groups or pools the identifier library into one or more partitions. In some embodiments, one or more partitions are stored separately from the system. In some embodiments, one or more reaction compartments, vessels, partitions, or substrates are mounted or stored on a disk, plate, film, fiber, tape, or thread separate from the system before, after, or both before and after generation of the identifier library or universal library.
[0175]
[0198] Item 1. A method for encoding digital information into a nucleic acid sequence, the method comprising: (a) encoding the digital information into a symbol sequence and converting the symbol sequence into a codeword; (b) parsing the codeword into a coded symbol sequence; (c) mapping the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; and (d) enumerating an identifier library, each symbol of the coded symbol sequence being encoded by one or more identifiers, each identifier of the plurality of identifiers comprising multiple components from one or more layers, each layer of the one or more layers comprising a distinct set of components, each identifier of the identifier library comprising one or more data layers encoding symbols and one or more key layers encoding keys.
[0176]
[0199] Item 2. The method of item 1, wherein a set of identifiers having identical components in the key layer form a codeword vector (CVI) of identifiers.
[0177]
[0200] Item 3. The method of items 1 or 2, wherein the digital information includes a data value represented by a bit sequence, and wherein encoding includes converting the bit sequence into a bit vector, each bit of the bit vector being mapped to one or more identifiers.
[0178]
[0201] Item 4. The method of item 3, wherein the data value is encoded using its original bit representation.
[0179]
[0202] Item 5. The method of item 3, wherein the data values are encoded using sparser or denser bit vectors.
[0180]
[0203] Item 6. The method according to any one of Items 3 to 5, wherein the data layer encodes a bit vector.
[0181]
[0204] Item 7. The method of any one of items 1 to 6, comprising performing a data query operation on an identifier library.
[0182]
[0205] Item 8. The method of item 7, wherein the query operation is performed on one or more query components from a subset of layers.
[0183]
[0206] Item 9. The method of item 8, comprising performing a selective polymerase chain reaction (PCR) using a query oligonucleotide to amplify only identifiers that include one or more query components.
[0184]
[0207] Item 10. The method of Item 9, comprising performing multiple cycles of PCR using a different query oligonucleotide in each cycle.
[0185]
[0208] Item 11. The method of items 9 or 10, comprising performing a size selection process to isolate a targeted subset of identifiers.
[0186]
[0209] Item 12. The method of any one of Items 8 to 11, comprising performing directed digestion to digest one or more query components in an identifier that includes one or more query components using a nuclease guided by a query oligonucleotide.
[0187]
[0210] Item 13. The method of item 12, comprising performing multiple cycles of directed digestion.
[0188]
[0211] Item 14. The method of items 12 or 13, comprising performing a size selection process to isolate a targeted subset of identifiers.
[0189]
[0212] Item 15. The method of any one of Items 8 to 14, comprising performing selective hybridization to hybridize a modified oligonucleotide to one or more query components in a single-stranded identifier comprising one or more query components.
[0190]
[0213] Item 16. The method of Item 15, comprising performing multiple cycles of selective hybridization.
[0191]
[0214] Item 17. The method of items 15 or 16, comprising performing a removal of hybridized identifiers to extract a targeted subset of identifiers.
[0192]
[0215] Item 18. The method according to any one of Items 15 to 17, wherein each of the modified oligonucleotides is attached to a bead or a surface.
[0193]
[0216] Item 19. The method of any one of Items 8 to 18, comprising performing selective hybridization to hybridize a query oligonucleotide to one or more query components in the single-stranded identifiers and hybridizing oligonucleotides corresponding to all components in the layer other than the layer of query components, and performing digestion of the unhybridized single-stranded components to isolate the target subset of identifiers.
[0194]
[0217] Item 20. The method according to any one of Items 7 to 19, comprising performing a sampling operation.
[0195]
[0218] Item 21. The method according to any one of Items 1 to 20, comprising performing multiple cycles of steps (a) to (d).
[0196]
[0219] Item 22. The method of any one of items 1 to 21, comprising modulating the copy number of a subset of identifiers to encode a relationship between a set of data objects.
[0197]
[0220] Item 23. The method of any one of items 2 to 22, comprising performing one or more editing processes on the CVI.
[0198]
[0221] Item 24. The method of item 23, wherein the editing process involves using a guide oligonucleotide to direct a programmable nuclease-recombinase enzyme complex to the target component and using a recombinase to replace the target component with a new template oligonucleotide.
[0199]
[0222] Item 25. The method according to Item 23 or 24, wherein the editing process comprises PCR using, as one PCR primer, an oligonucleotide that (a) only partially matches the identifier and (b) includes a sequence for the modified flanking component.
[0200]
[0223] Item 26. The method of any one of Items 23 to 25, wherein the editing process includes adding one or more components by PCR using oligonucleotides containing sequences for the flanking components as one PCR primer.
[0201]
[0224] Item 27. The method of any one of items 1 to 26, including performing a key extraction process, wherein the key layer for each identifier is separated from the data layer to generate a set of sub-identifiers.
[0202]
[0225] Item 28. The method of item 27, wherein the set of sub-identifiers includes a key layer.
[0203]
[0226] Item 29. The method of item 28, comprising performing a sub-identifier counting operation.
[0204]
[0227] Item 30. The method of items 28 or 29, including performing a conversion operation, wherein the data object is converted from binary form to unary form.
[0205]
[0228] Item 31. The method of Item 30, comprising performing PCR on an aliquot of a subset of the identifiers, the number of PCR cycles corresponding to the bit position of at least one identifier.
[0206]
[0229] Item 32. The method of any one of Items 26 to 31, comprising performing an addition operation or a subtraction operation.
[0207]
[0230] Item 33. The method of Item 32, wherein the addition operation includes pooling two or more identifier libraries.
[0208]
[0231] Item 34. The method of item 32, wherein the subtraction operation AB includes (a) converting the binary representation of A to unary form by a first CVI and converting the binary representation of B to unary form by a second CVI; (b) performing a chemical procedure to melt the subidentifiers into their constituent strands, retaining only the positive strand of each subidentifier of A and only the negative strand of each subidentifier of B; (c) pooling the retained subidentifiers; (d) heating the pooled subidentifiers to a temperature that promotes reannealing of the positive and negative strands; (e) discarding any duplexes resulting from process (e); and (f) separating and counting the remaining subidentifiers that represent A or B.
[0209]
[0232] Item 35. The method of Item 34, comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and / or B into double-stranded sub-identifiers.
[0210]
[0233] Item 36. The method of any one of items 2 to 35, comprising writing a CVI encoding a data object into a plurality of separate samples.
[0211]
[0234] Item 37. The method of any one of items 1 to 36, wherein one or more data layers encode symbol values, symbol positions, or both.
[0212]
[0235] Item 38. The method of any one of Items 1 to 37, wherein the components are assembled in linear order.
[0213]
[0236] Item 39. The method of any one of items 1 to 38, wherein the coded sequence of symbols includes symbols taken from a fixed alphabet of symbols.
[0214]
[0237] Item 40. The method of any one of Items 1 to 38, further comprising converting the coded sequence into a second symbol sequence.
[0215]
[0238] Item 41. The method of Item 40, wherein the second sequence of symbols includes a formal data structure.
[0216]
[0239] Item 42. The method of Item 41, wherein the formal data structure includes one or more members selected from the group consisting of a tree structure, a trie structure, a table structure, a key-value dictionary structure, and a set.
[0240]
[0217]
[0241] Item 43. The method of items 41 or 42, wherein the formal data structure is queriable by a range query, a rank query, a count query, a membership query, a nearest neighbor query, a match query, a selection query, or any combination thereof.
[0218]
[0242] Item 44. The method of any one of Items 41 to 43, further comprising parsing the second symbol sequence into a word sequence.
[0219]
[0243] Item 45. The method of Item 44, further comprising converting the word sequence to the codeword sequence using one or more codebooks.
[0220]
[0244] Item 46. The method of Item 45, further comprising converting the codeword sequence into a third symbol sequence.
[0221]
[0245] Item 47. The method of Item 46, wherein converting the word sequence into the code word sequence minimizes the number of symbols of one or more types in the third symbol sequence.
[0222]
[0246] Item 48. The method of any one of items 1 to 47, wherein the coded sequence of symbols comprises one or more symbol blocks.
[0223]
[0247] Item 49. The method of Item 48, wherein converting the word sequence into the code word sequence generates a fixed number of one or more types of symbols in each symbol block of the one or more symbol blocks in the third symbol sequence.
[0224]
[0248] Item 50. The method of any one of Items 1 to 49, wherein the codebook attaches one or more error protection symbols to each codeword of the codeword sequence.
[0225]
[0249] Item 51. The method of Item 50, wherein the one or more error protection symbols are calculated from one or more words of the word sequence.
[0226]
[0250] Item 52. The method according to any one of Items 1 to 51, wherein the plurality of identifiers are selected from a combination space of identifiers.
[0227]
[0251] Item 53. The method of any one of Items 1 to 52, wherein each component of the plurality of components comprises a nucleic acid sequence.
[0228]
[0252] Item 54. The method of Item 53, wherein the nucleic acid sequences are distinct sequences.
[0229]
[0253] Item 55. The method of any one of Items 1 to 54, wherein the presence of the individual identifier in the identifier library corresponds to a first symbol value and the absence of the individual identifier in the identifier library corresponds to a second symbol value.
[0230]
[0254] Item 56. The method of Item 55, wherein the first symbol value is "1" and the second symbol value is "0."
[0231]
[0255] Item 57. The method of Item 55, wherein the first symbol value is "0" and the second symbol value is "1."
[0232]
[0256] Item 58. The method of any one of Items 1 to 57, wherein the identifier library comprises complementary nucleic acid sequences.
[0233]
[0257] Item 59. The method of Item 58, wherein the supplemental nucleic acid sequence includes metadata about the first symbol sequence and an encoding of the first symbol sequence.
[0234]
[0258] Item 60. The method of Item 58, wherein the supplemental nucleic acid sequence does not correspond to digital information, and the supplemental nucleic acid sequence conceals the digital information encoded in the identifier library.
[0235]
[0259] Item 61. The method of any one of items 1 to 60, wherein the one or more identifiers are generated by combinatorial assembly of one or more components.
[0236]
[0260] Item 62. The method of any one of Items 1 to 61, further comprising constructing a universal identifier library.
[0237]
[0261] Item 63. The method of Item 62, wherein the identifier library is constructed from the universal identifier library by ignoring or excluding the individual identifiers that are not present in the identifier library.
[0238]
[0262] Item 64. The method according to Item 62, wherein constructing the universal identifier library includes using one or more reactions.
[0239]
[0263] Item 65. The method according to Item 64, wherein the one or more reactions corresponding to the individual identifiers not present in the identifier library are removed, deleted, ignored, or blocked.
[0240]
[0264] Item 66. The method of Item 64, wherein the one or more reactions include components, templates, and / or reagents, and the components, templates, and / or reagents are loaded onto a film, thread, fiber, or other substrate.
[0241]
[0265] Item 67. The method of item 66, wherein the components, the template, and / or the reagents are dispensed or co-located adjacent to one another by stamping, entangling, knitting, pinching, or weaving the film, the thread, the fiber, or the other substrate.
[0242]
[0266] Item 68. A nucleic acid-based integrated storage system, comprising: a data encoding unit configured to write digital information to one or more nucleic acid sequences, wherein the data encoding unit writes the digital information to the one or more nucleic acid sequences in the absence of base-by-base nucleic acid synthesis; a storage unit configured to store the one or more nucleic acid sequences encoding the digital information; a read unit configured to access and read the digital information encoded in the one or more nucleic acid sequences; and one or more computer processors operably coupled to the data encoding unit, the storage unit, and the read unit, wherein the one or more computer processors individually or collectively (i) instruct the data encoding unit to encode the digital information into the one or more nucleic acid sequences; (ii) instruct the storage unit to store the digital information encoded in the one or more nucleic acid sequences; a coding unit configured to code the digital information stored in the one or more nucleic acid sequences; and a coding unit configured to code the digital information stored in the one or more nucleic acid sequences; wherein the coding unit is programmed to: (a) encode the digital information into a symbol sequence and convert the symbol sequence into a codeword; (b) parse the codeword into a coded symbol sequence; (c) map the coded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; and (d) enumerate an identifier library, wherein each symbol of the coded symbol sequence is encoded by one or more identifiers, each identifier of the plurality of identifiers comprising multiple components from one or more layers, each layer of the one or more layers comprising a distinct set of components, and each identifier of the identifier library comprising one or more data layers encoding symbols and one or more key layers encoding keys.
[0243]
[0267] Item 69. The system of Item 68, wherein a set of identifiers having identical components in the key layer form a codeword vector (CVI) of identifiers.
[0244]
[0268] Item 70. The system described in Item 68 or 69, wherein the digital information includes a data value represented by a bit sequence, and encoding includes converting the bit sequence into a bit vector, each bit of the bit vector being mapped to one or more identifiers.
[0245]
[0269] Item 71. The system of item 70, wherein data values are encoded using their original bit representation.
[0246]
[0270] Item 72. The system of Item 70, wherein the data values are encoded using sparser or denser bit vectors.
[0247]
[0271] Item 73. A system described in any one of Items 70 to 72, wherein the data layer encodes a bit vector.
[0248]
[0272] Item 74. The system of any one of Items 68 to 73, wherein the encoding or decoding includes performing a data query operation on an identifier library.
[0249]
[0273] Item 75. The system of Item 74, wherein the query operation is performed on one or more query components from a subset of layers.
[0250]
[0274] Item 76. The system of Item 75, wherein the encoding or decoding comprises performing a selective polymerase chain reaction (PCR) using the query oligonucleotide to amplify only identifiers that include one or more query components.
[0251]
[0275] Item 77. The system of Item 76, wherein the encoding or decoding comprises performing multiple cycles of PCR using a different query oligonucleotide in each cycle.
[0252]
[0276] Item 78. The system of Item 76 or 77, wherein the encoding or decoding includes performing a size selection process to isolate a target subset of identifiers.
[0253]
[0277] Item 79. The system of any one of Items 75 to 78, wherein the encoding or decoding comprises performing directed digestion to digest one or more query components in the identifier comprising one or more query components using a nuclease guided by the query oligonucleotide.
[0254]
[0278] Item 80. The system of Item 79, wherein the encoding or decoding comprises performing multiple cycles of directed digestion.
[0255]
[0279] Item 81. The system of items 79 or 80, wherein the encoding or decoding includes performing a size selection process to isolate a target subset of identifiers.
[0256]
[0280] Item 82. The system of any one of Items 75 to 81, wherein the encoding or decoding comprises performing selective hybridization to hybridize modified oligonucleotides to one or more query components in a single-stranded identifier comprising one or more query components.
[0257]
[0281] Item 83. The system of Item 82, wherein the encoding or decoding comprises performing multiple cycles of selective hybridization.
[0258]
[0282] Item 84. The system of items 82 or 83, wherein the encoding or decoding includes performing a removal of hybridized identifiers to extract a targeted subset of identifiers.
[0259]
[0283] Item 85. The system of any one of items 82 to 84, wherein each of the modified oligonucleotides is attached to a bead or surface.
[0260]
[0284] Item 86. The system of any one of Items 75 to 85, wherein the encoding or decoding comprises performing selective hybridization to hybridize a query oligonucleotide to one or more query components in the single-stranded identifiers and hybridizing oligonucleotides corresponding to all components in the layer other than the layer of query components, and performing digestion of the unhybridized single-stranded components to isolate the target subset of identifiers.
[0261]
[0285] Item 87. The system of any one of Items 75 to 86, wherein the encoding or decoding includes performing a sampling operation.
[0262]
[0286] Item 88. The system of any one of Items 68 to 87, wherein encoding includes performing multiple cycles of steps (a) to (d).
[0263]
[0287] Item 89. The system of any one of items 68 to 88, wherein encoding includes modulating the number of copies of a subset of identifiers to encode a relationship between a set of data objects.
[0264]
[0288] Item 90. The system of any one of items 69 to 89, wherein encoding includes performing one or more editing processes on the CVI.
[0265]
[0289] Item 91. The system of Item 90, wherein the editing process includes using a guide oligonucleotide to direct a programmable nuclease-recombinase enzyme complex to the target component and using a recombinase to replace the target component with a new template oligonucleotide.
[0266]
[0290] Item 92. The system of Item 90 or 91, wherein the editing process includes PCR using, as one PCR primer, an oligonucleotide that (a) only partially matches the identifier and (b) includes a sequence for the modified flanking component.
[0267]
[0291] Item 93. The system of any one of items 90 to 92, wherein the editing process includes adding one or more components by PCR using oligonucleotides containing sequences for the flanking components as one PCR primer.
[0268]
[0292] Item 94. The system of any one of items 68 to 93, including performing a key extraction process, wherein the key layer for each identifier is separated from the data layer to generate a set of sub-identifiers.
[0269]
[0293] Item 95. The system of item 94, wherein the set of sub-identifiers includes a key layer.
[0270]
[0294] Item 96. The system of Item 95, wherein the encoding or decoding includes performing a counting operation on the sub-identifiers.
[0271]
[0295] Item 97. The system of items 95 or 96, wherein encoding or decoding includes performing a conversion operation, wherein the data object is converted from binary form to unary form.
[0272]
[0296] Item 98. The system of Item 97, wherein the encoding or decoding includes performing PCR on an aliquot of a subset of the identifiers, the number of PCR cycles corresponding to a bit position of at least one identifier.
[0273]
[0297] Item 99. The system of any one of Items 93 to 98, wherein encoding or decoding includes performing an addition or subtraction operation.
[0274]
[0298] Item 100. The system of Item 99, wherein the addition operation includes pooling two or more identifier libraries.
[0275]
[0299] 101. The subtraction operation AB is (a) converting the binary representation of A to monadic form by a first CVI and the binary representation of B to monadic form by a second CVI; (b) performing a chemical procedure to fuse the sub-identifiers into their constituent strands, preserving only the positive strand of each sub-identifier of A and only the negative strand of each sub-identifier of B; (c) pooling the retained sub-identifiers; and (d) heating the pooled sub-identifiers to a temperature that promotes reannealing of the plus and minus strands; (e) discarding any double strands resulting from process (e); and (f) separating and counting the remaining sub-identifiers representing A or B; Item 99. The system of item 99, comprising:
[0276]
[0300] Item 102. The system of Item 101, comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and / or B into double-stranded sub-identifiers.
[0277]
[0301] Item 103. The system of any one of items 69 to 102, wherein encoding or decoding includes writing a CVI that encodes the data object into multiple separate samples.
[0278]
[0302] Item 104. The system of any one of items 68 to 103, wherein one or more data layers encode symbol values, symbol positions, or both.
[0279]
[0303] Item 105. The system of any one of Items 68 to 104, wherein the multiple components are assembled in a linear order.
[0280]
[0304] Item 106. The system according to any one of Items 68 to 105, wherein the system is automated.
[0281]
[0305] Item 107. A system described in any one of Items 68 to 106, wherein the system is network-connected.
[0282]
[0306] Item 108. A system described in any one of Items 68 to 107, wherein the generated identifier library is a universal library.
[0283]
[0307] Item 109. The system described in any one of Items 68 to 108, further comprising a plurality of modules.
[0284]
[0308] Item 110. The system of Item 109, wherein the first module generates an identifier library.
[0285]
[0309] Item 111. The system according to Item 109 or 110, wherein the second module performs deletion of the individual identifiers or identifier responses.
[0286]
[0310] Item 112. The system of any one of Items 109 to 111, wherein the third module separates the individual identifiers present in the identifier library from the individual identifiers not present in the identifier library.
[0287]
[0311] Item 113. The system of any one of Items 109 to 112, wherein the fourth module groups or pools the identifier library into one or more partitions.
[0288]
[0312] Item 114. The system according to any one of items 68 to 113, wherein one or more reaction compartments, vessels, partitions, or substrates are mounted or stored on a disc, plate, film, fiber, tape, or thread separate from the system before, after, or both before and after generation of the identifier library or universal library.
[0289]
[0313] Item 115. The method of Item 19, wherein the single-stranded identifier is immobilized on a bead or a surface.
[0290]
[0314] Item 116. The system according to Item 86, wherein the single-stranded identifier is immobilized on a bead or a surface.
Claims
1. 1. A method for encoding digital information into a nucleic acid sequence, comprising: (a) encoding the digital information into a sequence of symbols and converting the sequence of symbols into a codeword; (b) parsing said codeword into a coded symbol sequence; (c) mapping the encoded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; (d) enumerating an identifier library, wherein each symbol of the coded symbol sequence is encoded by one or more identifiers, each identifier of the plurality of identifiers comprising a plurality of components from one or more layers, each layer of the one or more layers comprising a distinct set of components, and each identifier of the identifier library comprising: one or more data layers that encode symbols; and One or more key layers that encode the keys and A method comprising:
2. The method of claim 1 , wherein a set of identifiers having identical components in the key layer form a codeword vector of identifiers (CVI).
3. 3. The method of claim 1, wherein the digital information comprises a data value represented by a bit sequence, and wherein encoding comprises converting the bit sequence into a bit vector, each bit of the bit vector being mapped to one or more identifiers.
4. The method of claim 3 , wherein a data value is encoded using its original bit representation.
5. The method of claim 3 , wherein the data values are encoded using sparser or denser bit vectors.
6. The method of any one of claims 3 to 5, wherein the data layer encodes the bit vector.
7. The method of any one of claims 1 to 6, comprising performing a data query operation on the identifier library.
8. The method of claim 7 , wherein the query operation is performed on one or more query components from a subset of layers.
9. 9. The method of claim 8, comprising performing a selective polymerase chain reaction (PCR) using a query oligonucleotide to amplify only identifiers that include the one or more query components.
10. 10. The method of claim 9, comprising performing multiple cycles of PCR with a different query oligonucleotide in each cycle.
11. 11. The method of claim 9 or 10, comprising performing a size selection process to isolate a targeted subset of identifiers.
12. 12. The method of any one of claims 8 to 11, comprising performing a directed digestion to digest the one or more query components in an identifier that comprises the one or more query components using a nuclease guided by a query oligonucleotide.
13. 13. The method of claim 12, comprising performing multiple cycles of directed digestion.
14. 14. The method of claim 12 or 13, comprising performing a size selection process to isolate a targeted subset of identifiers.
15. 15. The method of any one of claims 8 to 14, comprising performing selective hybridization to hybridize modified oligonucleotides to one or more query components in the single-stranded identifier comprising said one or more query components.
16. 16. The method of claim 15, comprising performing multiple cycles of selective hybridization.
17. 17. The method of claim 15 or 16, comprising performing a removal of hybridized identifiers to extract a targeted subset of identifiers.
18. The method of any one of claims 15 to 17, wherein each of the modified oligonucleotides is attached to a bead or a surface.
19. 19. The method of any one of claims 8 to 18, comprising performing selective hybridization to hybridize a query oligonucleotide to one or more query components in the single-stranded identifiers and hybridizing oligonucleotides corresponding to all components in a layer other than the layer of the query component, and performing digestion of the unhybridized single-stranded components to isolate a target subset of identifiers.
20. A method according to any one of claims 7 to 19, comprising performing a sampling operation.
21. 21. The method of any one of claims 1 to 20, comprising carrying out multiple cycles of steps (a) to (d).
22. A method according to any preceding claim, comprising modulating the number of copies of a subset of identifiers to encode a relationship between a set of data objects.
23. A method according to any one of claims 2 to 22, comprising performing one or more editing processes on the CVI.
24. 24. The method of claim 23, wherein the editing process comprises using a guide oligonucleotide to target a programmable nuclease-recombinase enzyme complex to a target component and using the recombinase to replace the target component with a new template oligonucleotide.
25. 25. The method of claim 23 or 24, wherein the editing process comprises PCR using, as one PCR primer, an oligonucleotide that (a) only partially matches the identifier and (b) includes a sequence for a modified flanking component.
26. 26. The method of any one of claims 23 to 25, wherein the editing process comprises adding one or more components by PCR using oligonucleotides comprising sequences for the flanking components as one PCR primer.
27. A method according to any preceding claim, comprising performing a key extraction process, wherein the key layer for each identifier is separated from the data layer to generate a set of sub-identifiers.
28. 28. The method of claim 27, wherein the set of sub-identifiers comprises the key layer.
29. 30. The method of claim 28, comprising performing a counting operation on the sub-identifiers.
30. 30. A method according to claim 28 or 29, comprising performing a conversion operation, wherein the data object is converted from binary form to unary form.
31. 31. The method of claim 30, comprising performing PCR on an aliquot of the subset of identifiers, the number of PCR cycles corresponding to a bit position of at least one identifier.
32. A method according to any one of claims 26 to 31, comprising performing an addition or subtraction operation.
33. 33. The method of claim 32, wherein the addition operation comprises pooling two or more identifier libraries.
34. The subtraction operation A-B is (a) converting the binary representation of A to monadic form by a first CVI and the binary representation of B to monadic form by a second CVI; (b) performing a chemical procedure to fuse the sub-identifiers into their constituent strands, preserving only the positive strand of each sub-identifier of A and only the negative strand of each sub-identifier of B; (c) pooling the retained sub-identifiers; and (d) heating the pooled sub-identifiers to a temperature that promotes reannealing of the plus strand and the minus strand; (e) discarding any double strands resulting from process (e); (f) separating and counting the remaining sub-identifiers representing A or B; 33. The method of claim 32, comprising:
35. 35. The method of claim 34, comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and / or B into double-stranded sub-identifiers.
36. A method according to any one of claims 2 to 35, comprising writing a CVI encoding a data object into a plurality of separate samples.
37. The method of any preceding claim, wherein the one or more data layers encode symbol values, symbol positions, or both.
38. The method of any one of claims 1 to 37, wherein the components are assembled in a linear order.
39. A method according to any preceding claim, wherein the coded sequence of symbols comprises symbols taken from a fixed alphabet of symbols.
40. The method of any one of claims 1 to 38, further comprising converting the coded sequence into a second sequence of symbols.
41. 41. The method of claim 40, wherein the second sequence of symbols comprises a formal data structure.
42. 42. The method of claim 41, wherein the formal data structure comprises one or more members selected from the group consisting of a tree structure, a trie structure, a table structure, a key-value dictionary structure, and a set.
43. 43. The method of claim 41 or 42, wherein the formal data structure is queriable by a range query, a rank query, a count query, a membership query, a nearest neighbor query, a match query, a selection query, or any combination thereof.
44. 44. The method of any one of claims 41 to 43, further comprising parsing the second sequence of symbols into a sequence of words.
45. 45. The method of claim 44, further comprising converting the word sequence to the codeword sequence using one or more codebooks.
46. 46. The method of claim 45, further comprising converting the codeword sequence to a third symbol sequence.
47. 47. The method of claim 46, wherein converting the word sequence to the codeword sequence minimizes a number of one or more types of symbols in the third symbol sequence.
48. A method according to any preceding claim, wherein the coded sequence of symbols comprises one or more symbol blocks.
49. 49. The method of claim 48, wherein converting the word sequence to the codeword sequence produces a fixed number of one or more types of symbols in each symbol block of the one or more symbol blocks in the third symbol sequence.
50. A method according to any preceding claim, wherein a codebook attaches one or more error protection symbols to each codeword of the codeword sequence.
51. 51. The method of claim 50, wherein the one or more error protection symbols are computed from one or more words of the word sequence.
52. A method according to any preceding claim, wherein the plurality of identifiers are selected from a combinatorial space of identifiers.
53. 53. The method of any one of claims 1 to 52, wherein each component of the plurality of components comprises a nucleic acid sequence.
54. 54. The method of claim 53, wherein the nucleic acid sequences are distinct sequences.
55. 55. The method of any one of claims 1 to 54, wherein the presence of the individual identifier in the identifier library corresponds to a first symbol value and the absence of the individual identifier in the identifier library corresponds to a second symbol value.
56. 56. The method of claim 55, wherein the first symbol value is a "1" and the second symbol value is a "0."
57. 56. The method of claim 55, wherein the first symbol value is "0" and the second symbol value is "1."
58. 58. The method of any one of claims 1 to 57, wherein the identifier library comprises complementary nucleic acid sequences.
59. 59. The method of claim 58, wherein the supplemental nucleic acid sequence comprises metadata about or an encoding of the first symbol sequence.
60. 59. The method of claim 58, wherein the complementary nucleic acid sequences do not correspond to digital information, and the complementary nucleic acid sequences conceal the digital information encoded in the identifier library.
61. 61. The method of any one of claims 1 to 60, wherein the one or more identifiers are generated by combinatorial assembly of one or more components.
62. 62. The method of any one of claims 1 to 61, further comprising constructing a universal identifier library.
63. 63. The method of claim 62, wherein the identifier library is constructed from the universal identifier library by ignoring or excluding the individual identifiers that are not present in the identifier library.
64. 63. The method of claim 62, wherein constructing the universal identifier library comprises using one or more reactions.
65. 65. The method of claim 64, wherein the one or more reactions corresponding to the individual identifiers not present in the identifier library are removed, deleted, ignored, or blocked.
66. 65. The method of claim 64, wherein the one or more reactions include components, templates, and / or reagents, and the components, templates, and / or reagents are loaded onto a film, thread, fiber, or other substrate.
67. 67. The method of claim 66, wherein the components, the template, and / or the reagents are dispensed or co-located adjacent to one another by stamping, intertwining, knitting, pinching, or weaving the film, thread, fiber, or other substrate.
68. 1. A nucleic acid-based integrated storage system comprising: a data encoding unit configured to write digital information to one or more nucleic acid sequences, the data encoding unit writing the digital information to the one or more nucleic acid sequences in the absence of base-by-base nucleic acid synthesis; a storage unit configured to store the one or more nucleic acid sequences encoding the digital information; a reading unit configured to access and read the digital information encoded in the one or more nucleic acid sequences; one or more computer processors operatively coupled to the data encoding unit, the storage unit, and the reading unit; Equipped with The one or more computer processors individually or collectively: (i) instructing the data encoding unit to encode the digital information into the one or more nucleic acid sequences; (ii) instructing the storage unit to store the digital information encoded in the one or more nucleic acid sequences; (iii) instructing the reading unit to access and decode the digital information stored in the one or more nucleic acid sequences; It is programmed to To encode, (a) encoding the digital information into a sequence of symbols and converting the sequence of symbols into a codeword; (b) parsing said codeword into a coded symbol sequence; (c) mapping the encoded symbol sequence to a plurality of identifiers, each identifier of the plurality of identifiers comprising one or more nucleic acid sequences; (d) enumerating an identifier library, wherein each symbol of the coded symbol sequence is encoded by one or more identifiers, each identifier of the plurality of identifiers comprising a plurality of components from one or more layers, each layer of the one or more layers comprising a distinct set of components, and each identifier of the identifier library comprising: one or more data layers that encode symbols; and One or more key layers that encode the keys and Including, the system.
69. 69. The system of claim 68, wherein a set of identifiers having identical components in the key layer form a codeword vector (CVI) of identifiers.
70. 70. The system of claim 68 or 69, wherein the digital information comprises a data value represented by a bit sequence, and wherein encoding comprises converting the bit sequence into a bit vector, each bit of the bit vector being mapped to one or more identifiers.
71. 71. The system of claim 70, wherein a data value is encoded using its original bit representation.
72. 71. The system of claim 70, wherein data values are encoded using sparser or denser bit vectors.
73. 73. The system of any one of claims 70 to 72, wherein the data layer encodes the bit vector.
74. The system of any one of claims 68 to 73, wherein encoding or decoding comprises performing a data query operation on the identifier library.
75. 75. The system of claim 74, wherein the query operation is performed on one or more query components from a subset of layers.
76. 76. The system of claim 75, wherein encoding or decoding comprises performing a selective polymerase chain reaction (PCR) using a query oligonucleotide to amplify only identifiers that include the one or more query components.
77. 77. The system of claim 76, wherein the encoding or decoding comprises performing multiple cycles of PCR with a different query oligonucleotide in each cycle.
78. 78. The system of claim 76 or 77, wherein the encoding or decoding comprises performing a size selection process to isolate a target subset of identifiers.
79. 79. The system of any one of claims 75 to 78, wherein encoding or decoding comprises performing directed digestion to digest the one or more query components in identifiers that comprise the one or more query components using a nuclease guided by a query oligonucleotide.
80. 80. The system of claim 79, wherein the encoding or decoding comprises performing multiple cycles of directed digestion.
81. 81. The system of claim 79 or 80, wherein the encoding or decoding comprises performing a size selection process to isolate a target subset of identifiers.
82. 82. The system of any one of claims 75-81, wherein the encoding or decoding comprises performing selective hybridization to hybridize modified oligonucleotides to one or more query components in the single-stranded identifier comprising the one or more query components.
83. 83. The system of claim 82, wherein the encoding or decoding comprises performing multiple cycles of selective hybridization.
84. 84. The system of claim 82 or 83, wherein the encoding or decoding comprises performing a removal of hybridized identifiers to extract a targeted subset of identifiers.
85. 85. The system of any one of claims 82 to 84, wherein each of the modified oligonucleotides is attached to a bead or surface.
86. 86. The system of any one of claims 75 to 85, wherein the encoding or decoding comprises performing selective hybridization to hybridize a query oligonucleotide to one or more query components in the single-stranded identifiers and hybridizing oligonucleotides corresponding to all components in a layer other than the layer of query components, and performing digestion of unhybridized single-stranded components to isolate a target subset of identifiers.
87. A system according to any one of claims 75 to 86, wherein encoding or decoding comprises performing a sampling operation.
88. The system of any one of claims 68 to 87, wherein the encoding comprises performing multiple cycles of steps (a) to (d).
89. A system according to any one of claims 68 to 88, wherein the encoding comprises modulating the number of copies of a subset of the identifiers to encode a relationship between the set of data objects.
90. The system of any one of claims 69 to 89, wherein encoding comprises performing one or more editing processes on the CVI.
91. 91. The system of claim 90, wherein the editing process comprises using a guide oligonucleotide to target a programmable nuclease-recombinase enzyme complex to a target component and using the recombinase to replace the target component with a new template oligonucleotide.
92. 92. The system of claim 90 or 91, wherein the editing process comprises PCR using an oligonucleotide as one PCR primer that (a) only partially matches the identifier and (b) includes a sequence for a modified flanking component.
93. 93. The system of any one of claims 90 to 92, wherein the editing process comprises adding one or more components by PCR using oligonucleotides comprising sequences for the adjacent components as one PCR primer.
94. A system according to any one of claims 68 to 93, comprising performing a key extraction process, wherein the key layer for each identifier is separated from the data layer to generate a set of sub-identifiers.
95. 95. The system of claim 94, wherein the set of sub-identifiers comprises the key layer.
96. 96. The system of claim 95, wherein encoding or decoding comprises performing a counting operation on the sub-identifiers.
97. 97. A system according to claim 95 or 96, wherein the encoding or decoding comprises performing a conversion operation, wherein the data object is converted from binary form to unary form.
98. 98. The system of claim 97, wherein encoding or decoding comprises performing PCR on an aliquot of the subset of identifiers, the number of PCR cycles corresponding to a bit position of at least one identifier.
99. A system according to any one of claims 93 to 98, wherein encoding or decoding comprises performing an addition or subtraction operation.
100. 100. The system of claim 99, wherein the addition operation comprises pooling two or more identifier libraries.
101. The subtraction operation A-B is (a) converting the binary representation of A to monadic form by a first CVI and the binary representation of B to monadic form by a second CVI; (b) performing a chemical procedure to fuse the sub-identifiers into their constituent strands, preserving only the positive strand of each sub-identifier of A and only the negative strand of each sub-identifier of B; (c) pooling the retained sub-identifiers; and (d) heating the pooled sub-identifiers to a temperature that promotes reannealing of the plus strand and the minus strand; (e) discarding any double strands resulting from process (e); (f) separating and counting the remaining sub-identifiers representing A or B; 100. The system of claim 99, comprising:
102. 102. The system of claim 101, comprising, after step (e), converting the remaining single-stranded sub-identifiers of A and / or B into double-stranded sub-identifiers.
103. A system according to any one of claims 69 to 102, wherein encoding or decoding comprises writing a CVI encoding the data object into a plurality of separate samples.
104. The system of any one of claims 68 to 103, wherein the one or more data layers encode symbol values, symbol positions, or both.
105. The system of any one of claims 68 to 104, wherein the plurality of components are assembled in a linear order.
106. The system of any one of claims 68 to 105, wherein the system is automated.
107. A system according to any one of claims 68 to 106, wherein the system is networked.
108. The system of any one of claims 68 to 107, wherein the generated identifier library is a universal library.
109. The system of any one of claims 68 to 108, further comprising a plurality of modules.
110. 110. The system of claim 109, wherein the first module creates an identifier library.
111. 111. The system of claim 109 or 110, wherein the second module performs deletion of the individual identifiers or identifier responses.
112. A system according to any one of claims 109 to 111, wherein a third module separates the individual identifiers that are present in the identifier library from the individual identifiers that are not present in the identifier library.
113. The system of any one of claims 109 to 112, wherein a fourth module groups or pools the identifier library into one or more partitions.
114. The system of any one of claims 68 to 113, wherein one or more reaction compartments, vessels, partitions, or substrates are mounted on or stored on a disk, plate, film, fiber, tape, or thread separate from the system before, after, or both before and after generation of the identifier library or universal library.
115. 20. The method of claim 19, wherein the single-stranded identifier is immobilized on a bead or a surface.
116. 87. The system of claim 86, wherein the single-stranded identifier is immobilized on a bead or a surface.