Systems for nucleic acid-based data storage

The combinatorial genome strategy for encoding digital information in nucleic acid molecules addresses the inefficiencies of base-by-base synthesis, reducing costs and time, thereby improving the commercial viability and efficiency of nucleic acid digital data storage.

KR102995741B1Inactive Publication Date: 2026-07-29CATALOG TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
CATALOG TECHNOLOGIES INC
Filing Date
2017-11-16
Publication Date
2026-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for encoding digital information in nucleic acid sequences are costly and time-consuming due to the need for base-by-base synthesis of nucleic acids, making nucleic acid digital data storage inefficient and commercially unviable.

Method used

A method and system for encoding digital information in nucleic acid molecules without base-to-base synthesis by using a combinatorial genome strategy, where bit values are represented by the presence or absence of unique nucleic acid sequences, and identifiers are generated through a combinatorial assembly of components, reducing the need for de-novo synthesis.

Benefits of technology

This approach significantly reduces the cost and time required for encoding and retrieving data, enhancing the commercial viability and efficiency of nucleic acid digital data storage by allowing reuse of pre-synthesized nucleic acid sequences and enabling parallelizable processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112023039755705-PAT00001_ABST
    Figure 112023039755705-PAT00001_ABST
Patent Text Reader

Abstract

A method and system for encoding digital information in a nucleic acid (e.g., deoxyribonucleic acid) molecule without base-by-base synthesis by encoding bit value information in the presence or absence of a unique nucleic acid sequence in a pool, comprising the steps of specifying each bit position in a bit-stream having a unique nucleic acid sequence, and specifying a bit value at that position by the presence or absence of a corresponding unique nucleic acid sequence in a pool, but more generally, specifying a unique byte of a byte stream by a unique subset of nucleic acid sequences. Additionally, a method for generating a unique nucleic acid sequence without base-by-base synthesis using a combinatorial genome strategy (e.g., assembly of multiple nucleic acid sequences or enzyme-based editing of a nucleic acid sequence) is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference

[0002] This application claims priority based on U.S. Provisional Application No. 62 / 423,058 filed November 16, 2016, U.S. Provisional Application No. 62 / 457,074 filed February 9, 2017, and U.S. Provisional Application No. 62 / 466,304 filed March 2, 2017, the contents of each of which are incorporated herein by reference in their entirety. Background Technology

[0003] Nucleic acid digital data storage is a stable approach that encodes and stores information for long periods with data stored at a higher density than magnetic tape or hard drive storage systems. Furthermore, digital data stored in nucleic acid molecules in cold and dry conditions can be retrieved for 60,000 years or more.

[0004] To access digital data stored in nucleic acid molecules, the nucleic acid molecules can be sequenced. As such, the storage of nucleic acid digital data can be an ideal method for storing data containing large amounts of information that is not accessed frequently but is stored or retained for a long time.

[0005] Current methods encode digital information (e.g., binary code) into base-level nucleic acid sequences, allowing the relationships between bases within the sequence to be directly converted into digital information (e.g., binary code). Sequencing digital data stored as bit streams, bytes of digitally encoded information, or base-level sequences readable as bit streams can be error-prone and costly to encode due to the high cost of synthesizing new base-level nucleic acids. The opportunity for new methods to perform nucleic acid digital data storage can provide a low-cost and commercially feasible approach to encoding and retrieving data. means of solving the problem

[0006] A method and system for encoding digital information in a nucleic acid (e.g., deoxyribonucleic acid, DNA) molecule without base-to-base synthesis by encoding bit value information in the presence or absence of a unique nucleic acid sequence in a pool comprises the step of specifying each bit position in a bit-stream having a unique nucleic acid sequence and specifying a bit value at that position by the presence or absence of a corresponding unique nucleic acid sequence in said pool. However, more generally, a unique byte in a byte stream is specified by a unique subset of nucleic acid sequences. Additionally, a method for generating a unique nucleic acid sequence without base-to-base synthesis using a combinatorial genome strategy (e.g., assembly of multiple nucleic acid sequences or enzyme-based editing of a nucleic acid sequence) is disclosed.

[0007] In one embodiment, the present invention comprises: (a) coding digital information into a series of symbols and converting a sequence of said symbols into code words using one or more codebooks; (b) parsing said code words into a coded sequence of symbols; (c) mapping said coded sequence of symbols to a plurality of identifiers, wherein each identifier among said plurality of identifiers comprises one or more nucleic acid sequences; (d) enumerating an identifier library such that each symbol of said sequence of coded symbols is encoded by one or more identifier(s); and (e) adding said one or more codebooks and descriptions of said plurality of identifiers to said identifier library.

[0008] In some embodiments, the coded sequence of symbols comprises symbols taken from fixed alphabetic symbols. In some embodiments, the method further comprises the step of converting the coded sequence into a second sequence of symbols. In some embodiments, the second sequence of symbols comprises a formal data structure. In some embodiments, the formal data structure comprises one or more members selected from a group consisting of a tree structure, a trie structure, a table structure, a key-value dictionary structure, and a set. In some embodiments, the formal data structure is queried by a range query, a rank query, a count query, a membership query, a nearest neighbor query, a match query, a select query, or any combination thereof.

[0009] In some embodiments, the method further includes the step of parsing a second sequence of symbols into a word sequence. In some embodiments, the method further includes the step of converting the word sequence into a sequence of code words using the one or more codebooks. In some embodiments, the method further includes the step of converting the sequence of code words into a sequence of third symbols. In some embodiments, converting the word sequence into a sequence of code words minimizes the number of one or more types of symbols in the third sequence of symbols.

[0010] In some embodiments, the coded sequence of symbols includes one or more blocks of symbols. In some embodiments, converting a sequence of words into a sequence of code words generates a fixed number of symbols of types in each block of one or more symbols within the third sequence of symbols. In some embodiments, the codebook adds one or more error-prevention symbols to individual code words of the sequence of code words. In some embodiments, one or more error-prevention symbols are calculated from one or more words of the word sequence.

[0011] In some embodiments, a plurality of identifiers are selected from a combination space of identifiers. In some embodiments, individual identifiers of the plurality of identifiers include one or more components. In some embodiments, individual components of one or more components include a nucleic acid sequence. In some embodiments, the nucleic acid sequence is a distinct sequence.

[0012] In some embodiments, each symbol in the symbol string is one of two possible symbol values. In some embodiments, one symbol value at each position in the symbol string may be represented by the absence of a unique identifier in the identifier library. In some embodiments, the two possible symbol values ​​are bit values ​​of 0 and 1, and an individual symbol having a bit value of 0 in the string of symbols may be represented by the absence of an individual identifier in the identifier library, and an individual symbol having a bit value of 1 in the string of symbols may be represented by the presence of an individual identifier in the identifier library, and vice versa. In some embodiments, the presence of an individual identifier in the identifier library corresponds to a first symbol value in the binary string, and the absence of an individual identifier from the identifier library corresponds to a second symbol value in the binary string. In some embodiments, the first symbol value is '1' and the second symbol value is '0'. In some embodiments, the first symbol value is '0' and the second symbol value is '1'. In some embodiments, the identifier library includes a supplementary nucleic acid sequence. In some specific embodiments, the supplementary nucleic acid sequence includes metadata regarding the sequence of the first symbol or the encoding of the sequence of the first symbol. In some specific embodiments, the supplementary nucleic acid sequence does not correspond to digital information, and the supplementary nucleic acid sequence conceals digital information encoded in the identifier library.

[0013] In some embodiments, one or more identifier(s) are generated by a combinatorial assembly of one or more components. In some embodiments, the method further includes the step of constructing a universal identifier library. In some embodiments, the identifier library is constructed from the universal identifier library by decomposing or excluding individual identifiers that are not present in the identifier library. In some embodiments, constructing the universal identifier library involves using one or more reactions. In some embodiments, one or more reactions corresponding to individual identifiers that are not present in the identifier library are removed, deleted, degraded, or prohibited. In some specific embodiments, one or more reactions include components, templates, and / or reagents, and the components, templates, and / or reagents are loaded onto a film, thread, fiber, or other substrate. In some embodiments, the components, templates, and / or reagents are placed adjacent to each other by stamping, intertwining, braiding, pinching, or weaving the film, thread, fiber, or other substrate.

[0014] In another aspect, the present invention provides a system for coding digital information into nucleic acid sequence(s), comprising: an assembly unit configured to generate an identifier library that encodes a sequence of symbols, wherein the identifier library comprises at least a subset of a plurality of identifiers; and one or more computer processors operably coupled to the assembly unit, wherein the one or more computer processors are programmed individually or collectively to (i) code the digital information into a sequence of symbols and convert the sequence of symbols into a codeword using a codebook; (ii) parse the codeword into a coded sequence of symbols; (iii) map the coded sequence of symbols to the plurality of identifiers, wherein each of the plurality of identifiers comprises one or more nucleic acid sequences; (iv) instruct the assembly unit to generate an identifier library, wherein each symbol of the coded sequence of symbols is encoded by one or more identifiers; and (v) instruct the assembly unit to attach descriptions of the one or more codebooks and the plurality of identifiers to the identifier library.

[0015] In some embodiments, one or more identifier(s) are assembled into one or more assembly reactions. In some embodiments, one or more products of one or more assembly reactions are combined to create an identifier library.

[0016] In some embodiments, the assembly unit comprises one or more containers. In some embodiments, one or more containers are partitions. In some specific embodiments, the assembly unit comprises a reagent, one or more layers of components, one or more templates, or any combination thereof. In some embodiments, the assembly unit is configured to accommodate a reagent, one or more layers of components, one or more templates, or any combination thereof. In some embodiments, the assembly unit is configured to output an identifier library.

[0017] In some embodiments, the assembly unit includes a reaction module. In some specific embodiments, the reaction module is configured to collect a reagent, one or more layers, one or more templates, or any combination thereof. In some specific embodiments, the reagent includes an enzyme, one or more nucleic acid sequences, a buffer, a cofactor, or any combination thereof. In some specific embodiments, the reagent is combined into a master mix before entering the reaction module. In some specific embodiments, the reaction module is configured to incubate or agitate the assembly reaction, and the assembly reaction generates one or more identifiers. In some embodiments, the reaction module includes a detector unit, and the detector unit monitors the assembly of one or more identifier(s).

[0018] In some embodiments, the system further includes a storage unit, and the assembly unit transfers the generated identifier library to the storage unit. In some embodiments, the storage unit includes one or more pools, containers, or partitions. In some embodiments, the storage unit combines one or more identifier libraries into one or more pools, one or more containers, or one or more partitions.

[0019] In some embodiments, the system further includes a selection unit configured to select one or more identifier(s). In some embodiments, the selection unit includes a size selection module, an affinity capture module, a nuclease cleavage module, or any combination thereof.

[0020] In some embodiments, the system further comprises a nucleic acid synthesis unit configured to synthesize one or more nucleic acid sequences. In some embodiments, one or more nucleic acid sequences are produced by base-by-base synthesis.

[0021] In some embodiments, the assembly unit generates a plurality of responses for assembling one or more identifier(s). In some embodiments, the assembly unit selectively removes individual responses from a plurality of responses that do not generate at least a subset of the plurality of identifiers in the identifier library.

[0022] In some embodiments, the assembly unit may generate an identifier library using one or more of nucleic acid sequence-coated material, slip technology, stamping, laser printing, or electro-wetting, misting, printing, laser ablation, weaving or braiding, or intertwining of droplet microfluid.

[0023] In some embodiments, the one or more computer processors are programmed individually or collectively to use heuristic techniques to minimize the number of responses for generating the identifier library or to minimize the time taken to set up multiple responses for generating the identifier library. In some embodiments, the heuristic techniques include a heuristic or on-set covering heuristic that minimizes the movement path of the device.

[0024] In another aspect, the present invention comprises: (a) a data encoding unit configured to record digital information in one or more nucleic acid sequences, wherein the data encoding unit records the digital information in the one or more nucleic acid sequences in the absence of base-wise nucleic acid synthesis; (b) a storage unit configured to store the one or more nucleic acid sequences encoding the digital information; (c) a reading unit configured to access and read the digital information encoded in the one or more nucleic acid sequences; and (d) one or more computer processors operably coupled to the data encoding unit, the storage unit, and the reading unit, wherein the one or more computer processors are individually or collectively programmed to (i) instruct the data encoding unit to encode the digital information in the one or more nucleic acid sequences, (ii) programmed to instruct the storage unit to store the digital information encoded in the one or more nucleic acid sequences, and (iii) programmed to instruct the reading unit to access and decode the digital information stored in the one or more nucleic acid sequences.

[0025] In some embodiments, one or more computer processors parse digital information into a plurality of symbols. In some embodiments, the plurality of symbols are mapped to a plurality of identifiers. In some embodiments, individual symbols of the plurality of symbols correspond to one or more identifiers of the plurality of identifiers. In some embodiments, the plurality of identifiers includes a plurality of components. In some specific examples, individual components of the plurality of components include distinct nucleic acid sequences.

[0026] In some embodiments, the data encoding unit generates one or more identifier libraries comprising one or more sets of identifiers corresponding to digital information. In some embodiments, the step of reading digital information includes the step of identifying one or more sets of identifiers in one or more identifier libraries.

[0027] In some embodiments, the system is automated. In some embodiments, the system is networked. In some embodiments, the system is configured to operate in a zero or low gravity environment. In some embodiments, the system is configured to operate at sub-atmospheric pressure, under vacuum, or at pressure higher than atmospheric pressure. In some embodiments, the system includes a power source or a method for generating power. In some embodiments, the system includes a radiation shield.

[0028] In some embodiments, the generated identifier library is a general-purpose library. In some embodiments, the system further includes a plurality of modules. In some embodiments, a first module generates the identifier library. In some embodiments, a second module implements the deletion of individual identifiers or identifier reactions. In some embodiments, a third module separates individual identifiers present in the identifier library from individual identifiers not present in the identifier library. In some embodiments, a fourth module groups or pools the identifier library into one or more partitions. In some embodiments, one or more partitions are stored separately from the system. In some embodiments, one or more reaction compartments, containers, partitions, or substrates are mounted or stored on a disk, plate, film, fiber, tape, or thread separated from the system before, after, or both before and after the generation of the identifier library or general-purpose library.

[0029] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, in which exemplary embodiments of the present disclosure are illustrated and described. As can be realized, the present disclosure may be modified in various obvious respects without departing from the content of the disclosure. Accordingly, the drawings and description should be considered illustrative in nature and are not limiting.

[0030] Integration by reference

[0031] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is specifically and individually directed to be incorporated by reference. Where any publication or patent or patent application included in this specification contradicts the disclosures included in this specification, this specification is intended to replace and / or prioritize such contradictory material. Brief explanation of the drawing

[0032] Novel features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description describing exemplary embodiments in which the principles of the present invention are utilized and the accompanying drawings (also "Drawings" and "Drawings"). Among these: Figure 1 schematically illustrates an overview of the process of encoding, recording, accessing, reading, and decoding digital information stored in a nucleic acid sequence. FIGS. 2A and 2B schematically illustrate a method for encoding digital data referred to as "address data" using an object or identifier (e.g., a nucleic acid molecule). FIG. 2A illustrates combining a rank object (or address object) with a byte-value object (or data object) to generate an identifier. FIG. 2B illustrates an example of data in an address method in which the rank object and the byte-value object are themselves a combination link of other objects. FIGS. 3a and 3b schematically illustrate exemplary methods for encoding digital information using objects or identifiers (e.g., nucleic acid sequences). FIG. 3a illustrates encoding digital information using a ranking object as an identifier. FIG. 3b illustrates an embodiment of an encoding method in which an address object is itself a combination link of other objects. Figure 4 schematically illustrates an overview of a method for recording information in a nucleic acid sequence (e.g., deoxyribonucleic acid). Figure 5 schematically illustrates an exemplary combination space of identifiers organized as an m-level n-ary tree. FIG. 6 schematically illustrates an exemplary method for minimizing the number of identifiers configured to record a bit stream. FIG. 7 schematically illustrates an exemplary method for remapping words to codewords to ensure uniform weight codewords for error detection. Figure 8 schematically illustrates an exemplary method for minimizing recording time by generating a minimum set of reactions. Figure 9 schematically illustrates the dual mapping of addresses to identifiers and the dual encoding of data. FIG. 10 schematically illustrates an exemplary method for masking encoding and decoding to protect against unauthorized decoding. FIG. 11 illustrates an exemplary component carousel. FIG. 12 schematically illustrates a method of using electrowetting for component operation. Figure 13 illustrates an example of a print-based method for distributing components. Figure 14 illustrates an example of microfluidic injection of a component. Figure 15 illustrates an example of selective condensation of component mist. FIG. 16 schematically illustrates an exemplary method of generating an identifier by weaving or braiding. FIG. 17 schematically illustrates an exemplary method for generating an identifier from a set of components. FIG. 18 schematically illustrates an exemplary method of generating an identifier from individual films or threads. Figure 19 schematically illustrates an exemplary method of recording information using subtraction. FIG. 20 schematically illustrates an exemplary method of reading by hybridization. Fig. 21 schematically illustrates an exemplary method of reading by nanopore sequencing. FIG. 22 illustrates a computer control system programmed or otherwise configured to implement the method provided in this specification. Specific details for implementing the invention

[0033] Although various embodiments of the present invention have been illustrated and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. Various modifications, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be adopted.

[0034] As used herein, the term "digital message" generally refers to a series of symbols provided to be encoded into a nucleic acid molecule. A digital message may be the original text recorded on a nucleic acid molecule.

[0035] As used herein, the term "symbol" generally refers to a representation of a unit of digital information. Digital information can be divided into several symbols. For example, a symbol can be a bit, and a bit can have a value of '0' or '1'.

[0036] As used herein, the terms “distinctive” or “unique” generally refer to an object that is distinguishable from other objects within a group. For example, a distinct or unique nucleic acid sequence may be a nucleic acid sequence that does not have the same sequence as any other nucleic acid sequence. A unique or unique nucleic acid molecule cannot have the same sequence as another nucleic acid molecule. A distinct or unique nucleic acid sequence or molecule may share similar regions with other nucleic acid sequences or molecules.

[0037] As used herein, the term "component" generally refers to a nucleic acid sequence. A component may be a distinct nucleic acid sequence. A component may be linked or assembled with one or more other components to form a different nucleic acid sequence.

[0038] As used herein, the term “layer” generally refers to a group or pool of components. Each layer may include a distinct set of components such that the components of one layer differ from the components of another layer. Components of one or more layers may be assembled to generate one or more identifiers.

[0039] As used herein, the term “identifier” generally refers to a nucleic acid molecule or nucleic acid sequence representing the position and value of a bit-string within a larger bit sequence. More generally, an identifier may represent a symbol in a sequence of symbols or any entity corresponding to a symbol. In some embodiments, an identifier may comprise one or more combined components.

[0040] As used herein, the term “combination space” generally refers to all possible sets of individual identifiers that can be generated from a starting set of objects, such as components, and a set of acceptable sets of methods for modifying said objects to form identifiers. The size of the combination space of identifiers created by combining or linking components may vary depending on the number of component layers, the number of components in each layer, and the specific assembly method used to generate the identifiers.

[0041] The term "identifier rank" as used in this specification generally refers to a relationship that defines the order of identifiers within a set.

[0042] As used herein, the term “identifier library” generally refers to a set of identifiers corresponding to symbols in a symbol string representing digital information. In some embodiments, if no identifier is given in the identifier library, a symbol value may be represented at a specific location. One or more identifier libraries may be combined into a pool, group, or set of identifiers. Each identifier library may include a unique barcode that identifies the identifier library.

[0043] As used herein, the term “universal library” generally refers to a set of identifiers corresponding to a set of all possible individual identifiers that can be generated from a starting set of objects, such as components, and an acceptable set of rules for methods of modifying these objects for identifier formation.

[0044] As used herein, the term “word” generally refers to a block of a series of symbols. The length of the block may or may not be fixed. A string of symbols may be divided into one or more words comprising a length of L symbols. In one example, a string of symbols may be divided into four words of length 16 symbols, each of which has a length of 4 symbols.

[0045] As used herein, the term “codeword” generally refers to a symbol string that codes a word. The length of the string may be fixed or not fixed. A source bit stream may be parsed into words that are subsequently converted into codewords using a codebook. The codebook may correlate words with codewords. Codewords may be selected to reduce write time, minimize identifier construction, or detect write errors.

[0046] As used herein, the term “nucleic acid” generally refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or variants thereof. Nucleic acids may comprise one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Nucleotides may comprise A, C, G, T, or U, or variants thereof. Nucleotides may comprise any subunit that can be incorporated into a growing nucleic acid strand. Such subunits may be A, C, G, T, or U, or any other subunit specific to one of the more complementary A, C, G, T, or U, or complementary to purines (i.e., A or G, or variants thereof), or pyrimidines (i.e., C, T, or U, or variants thereof). In some examples, nucleic acids may be single-stranded or double-stranded, and in some cases, nucleic acid molecules are circular.

[0047] As used herein, the terms “nucleic acid molecule” or “nucleic acid sequence” generally refer to a polymeric form of nucleotide or polynucleotide that may have various lengths and is either deoxyribonucleotide (DNA) or ribonucleotide (RNA), or include analogs thereof. As used herein, oligonucleotide generally refers to a single-stranded nucleic acid sequence and generally consists of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (uracil (U) for thymine (T) when the polynucleotide is RNA). The term “nucleic acid sequence” may refer to an alphabetic representation of a polynucleotide molecule; alternatively, the term may be applied to the physical polynucleotide itself. This alphabetic notation may be entered into a database of a computer with a central processing unit and is used to map the nucleic acid sequence or nucleic acid molecule to symbols or bits that encode digital information. The nucleic acid sequence or oligonucleotide may include one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.

[0048] Examples of modified nucleotides include diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl, 2-methylguanine, 2-methylguanine, 3-methylcytosine, 2-methylguanine, 2-methylguanine, 2-methylguanine, 5-methylguanine, 5-methylguanine, 5-methylguanine, 5-methylguanine, beta-D-mannosylcytosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, and 2-methylthio-D46-iso Pentenyl adenine, uracil-6-methylcytosine, N-methylcytosine, N6-adenine,(v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyllacuracil, uracil-5-methyl oxyacetate ester, uracil-5-oxyacetate(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil,(acp3)w, 2,6-diaminopurine, etc. Nucleic acid molecules may also be modified in the base portion (e.g., with a complementary nucleotide and / or with one or more atoms typically available to form hydrogen bonds, or with a complementary nucleotide), sugar residues, or phosphate backbones. Nucleic acid molecules may also contain amine modification groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexylacrylamide-dCTP (aha-dCTP), to allow for the covalent attachment of amine-reactive residues, such as N-hydroxysuccinimide ester (NHS).

[0049] As used herein, the term "primer" generally refers to a nucleic acid strand that acts as a starting point for nucleic acid synthesis, such as in the polymerase chain reaction (PCR). For example, during the replication of a DNA sample, the enzyme catalyzing replication starts replication at the 3' end of a primer attached to the DNA sample and replicates the opposite strand.

[0050] As used herein, the terms “polymerase,” “polymerase,” or “polymerase enzyme” generally refer to any enzyme capable of catalyzing a polymerase reaction. Examples of polymerases include, without limitation, nucleic acid polymerases. Such polymerases may occur naturally or be synthesized. An exemplary polymerase is 129 polymerase or a derivative thereof. In some cases, transcription enzymes or ligases (i.e., enzymes that catalyze bond formation) are used in conjunction with polymerases or as alternatives to polymerases for constructing new nucleic acid sequences. Examples of polymerases include DNA polymerase, RNA polymerase, heat-stable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, 129 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Pfu polymerase, Pwo polymerase, Vent polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tma polymerase, Tf polymerase, Tfi polymerase, Platinum Taq polymerase, Tbr polymerase, Tfl polymerase, Pfutubo polymerase, and pyrost polymerase. Includes KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase variants having 3'~5' exonuclease activity, modified products and derivatives thereof.

[0051] Digital information, such as computer data, in the form of binary code may contain sequences or strings of symbols. Binary code may encode or represent text or computer processor instructions using a binary number system having two binary symbols, typically 0 and 1, referred to as bits, for example. Digital information may be represented in the form of non-binary code, which may contain sequences of non-binary symbols. Each encoded symbol may be reassigned to a unique bit string (or "byte"), and a unique bit string or byte may be arranged as a string of bytes or byte streams. The bit value for a given bit may be one of two symbols (e.g., 0 or 1). A byte containing an N-bit string may have a total of 2N unique byte values. For example, a byte containing 8 bits may generate a total of 28 or 256 possible unique byte values, and each of the 256 bytes may correspond to one of 256 possible distinct symbols, characters, or instructions that can be encoded into a byte. Raw data (e.g., text files and computer instructions) can be represented as bytes or strings of byte streams. Zip files or compressed data files containing raw data can be stored as byte streams, and these files can be stored as byte streams in a compressed format and then decompressed into raw data before being read by a computer.

[0052] The method and system of the present invention may be used to encode computer data or information into a plurality of identifiers, each of which may represent one or more bits of the original information. In some examples, the method and system of the present disclosure encode data or information using identifiers, each of which represents two bits of the original information.

[0053] Previous methods for encoding digital information into nucleic acids relied on the base-by-base synthesis of nucleic acids, which can be expensive and time-consuming. An alternative method can improve efficiency, enhance the commercial viability of digital information storage by reducing dependence on base-by-base nucleic acid synthesis for encoding digital information, and eliminate the need for novel synthesis of specific nucleic acid sequences for all new information storage requirements.

[0054] The new method can encode digital information (e.g., binary code) using multiple identifiers or nucleic acid sequences comprising combinational arrangements of components, instead of relying on base-by-base or novel nucleic acid synthesis (e.g., phosphoamidite synthesis). As such, the new strategy can generate a first set of different nucleic acid sequences (or components) for the initial request for information storage and reuse the same nucleic acid sequences (or components) for subsequent information storage requests. This approach can significantly reduce the cost of DNA-based information storage by reducing the role of de-novo synthesis of nucleic acid base sequences in the information-DNA encoding and writing process. Furthermore, unlike implementations of base-by-base synthesis such as phosphoramidite chemistry or template-free polymerase-based nucleic acid elongation, which require the periodic delivery of each base to each nucleic acid, the use of identifier construction from components in the new information-DNA writing method is a highly parallelizable process that does not utilize cyclic nucleic acid elongation. Therefore, the new method can increase the speed of writing digital information to DNA compared to previous methods.

[0055] Method of encoding and recording information in nucleic acid sequence(s)

[0056] In one embodiment, the present disclosure provides a method for coding a symbol sequence for recording as nucleic acid sequence(s). A method for coding a symbol sequence for recording as nucleic acid sequence(s) may include: (a) converting the symbol sequence into a code word using one or more codebooks; (b) parsing the code word into a coded sequence of symbols; (c) mapping the coded sequence of symbols to a plurality of identifiers; (d) creating an identifier library; and (e) adding descriptions of the one or more codebooks and the plurality of identifiers to the identifier library. Each symbol of the coded symbol sequence may be encoded by one or more identifier(s).

[0057] FIG. 1 illustrates an overview process of encrypting information into a nucleic acid sequence, writing information to the nucleic acid sequence, reading the information written to the nucleic acid sequence, and decoding the read information. Digital information or data may be converted into one or more strings of symbols. For example, a symbol is a bit, and each bit has a value of '0' or '1'. Each symbol may be mapped to or encoded by an object representing that symbol (e.g., an identifier). Each symbol may be represented by a unique identifier. The distinct identifier may be a nucleic acid molecule composed of components. The components may be nucleic acid sequences. Digital information may be written to the nucleic acid sequence by creating an identifier library corresponding to the information. The identifier library may be physically created by physically configuring identifiers corresponding to each symbol of the digital information. All or part of the digital information may be accessed at once. For example, a subset of identifiers is accessed from the identifier library. A subset of identifiers may be read by sequencing and identifying the identifiers. The identified identifier may be associated with a corresponding symbol to decode the digital data. FIG. 1 illustrates an overview process of encoding information into a nucleic acid sequence, writing information to the nucleic acid sequence, reading the information written to the nucleic acid sequence, and decoding the read information without using base-based synthesis. Digital information or data may be converted into one or more strings of symbols. For example, a symbol is a bit, and each bit has a value of '0' or '1'. Each symbol may be mapped to or encoded by a physical object (e.g., an identifier) ​​representing the symbol. Each symbol may be represented by a unique identifier. The distinct identifier may be a nucleic acid molecule composed of components. The components may be a nucleic acid sequence. Digital information may be written to the nucleic acid sequence by creating a library of identifiers corresponding to the information.An identifier library can be generated by combining identifiers corresponding to each symbol of digital information. All or part of the digital information can be accessed at once. For example, a subset of identifiers is removed from the identifier library. A subset of identifiers can be read by identifying the identifiers. The identified identifiers can be associated with corresponding symbols to decode the digital data.

[0058] A method for encoding and reading information using the approach of FIG. 1 may include, for example, receiving a bit stream. This may include mapping each 1-bit (a bit having a bit value '1') within the bit stream to a distinct nucleic acid identifier using an identifier rank. Construct a nucleic acid sample pool or identifier library containing copies of identifiers corresponding to bit values ​​of 1 (and identifiers for bit values ​​of 0). Reading the sample uses a molecular biological method (e.g., sequencing, hybridization, PCR, etc.) to determine which identifiers appear in the identifier library and assigns a bit value of '1' to the bit corresponding to that identifier (in other words, referring to the identifier rank for identifying the bit in the original bit stream to which the identifier corresponds), thereby decoding the information into the encoded original bit stream.

[0059] Encoding a string of N distinct bits may use a number of unique nucleic acid sequences equivalent to the number of possible identifiers. This approach to information encoding may use the synthesis of a new identifier for each new information item (a string of N bits) to be stored. In other cases, the cost of synthesizing a new identifier for each new item of information to be stored (greater than or equal to N) may be reduced by a one-time new synthesis and subsequent maintenance of all possible identifiers. Encoding a new information item may involve mechanically selecting and mixing pre-synthesized (or pre-made) identifiers to form an identifier library. In other cases, the cost may be reduced by (1) synthesizing N identifiers as new until the new information item is stored, (2) maintaining and selecting N identifiers when the new information item is stored, or by synthesizing and maintaining (less than N, in some cases less than N) nucleic acid sequences and then modifying these sequences through enzymatic reactions to generate N identifiers for each new information item to be stored.

[0060] Identifiers can be reasonably designed and selected to facilitate read, write, access, copy, and delete operations. Identifiers can be designed and selected to minimize write errors, mutations, performance degradation, and read errors.

[0061] FIGS. 2A and 2B schematically illustrate an exemplary method referred to as "data in address" for encoding digital data in an object or identifier (e.g., a nucleic acid molecule). FIG. 2A illustrates encoding a bit stream into an identifier library, where individual identifiers are constructed by concatenating a single component specifying an identifier rank with a single component specifying a byte value. Generally, data in address methods utilizes identifiers that encode information in a modular manner by including two objects: one object is a "byte value object" (or "data object") identifying a byte value, and the other is a "rank object" (or "address object") identifying an identifier rank. FIG. 2B illustrates an example of data in address methods where each rank object can be combined from a set of components and each byte-value object can be combined from a set of components. This combined construction of rank and byte-value objects allows more information to be written to the identifier than in an object made of only a single component (e.g., FIG. 2A).

[0062] FIGS. 3A and 3B schematically illustrate other exemplary methods for encoding digital information in an object or identifier (e.g., a nucleic acid sequence). FIG. 3A illustrates encoding an identifier string into an identifier library, wherein the identifier consists of a single component that specifies an identifier rank. The presence of the identifier at a specific rank (or address) specifies a bit-value of '1', and the absence of the identifier at a specific rank (or address) specifies a bit-value of '0'. This type of encoding may use identifiers that encode only the rank (the relative position of a bit in the original bit stream) and use the presence or absence of these identifiers in the identifier library to encode bit-values ​​of '1' or '0'. The steps of reading and decoding information may include identifying the identifier present in the identifier library, assigning a bit-value of '1' to a corresponding rank, and assigning a bit-value of '0' elsewhere. FIG. 3B illustrates an exemplary encoding method in which each identifier may be combinatorially constructed from a set of components such that each possible combination configuration specifies a rank. This combination configuration allows more identifiers to be written to the identifier than an identifier made from only a single component (e.g., FIG. 3a). For example, a set of components may include five distinct components. The five distinct components may be assembled to generate ten distinct identifiers, each containing two of the five components. Each of the ten distinct identifiers may have a rank (or address) corresponding to a bit position within a bit stream. The identifier library may include a subset of ten possible identifiers corresponding to a position of bit-value '1' and exclude a subset of ten possible identifiers corresponding to a position of bit-value '0' within a bit stream of length ten.

[0063] FIG. 4 illustrates an overview method for recording information in a nucleic acid sequence. Before writing the information, the information may be translated into a string of symbols and encoded into multiple identifiers. Writing the information may involve setting up a reaction to generate possible identifiers. Inputs may be introduced into compartments to set up the reaction. Inputs may include nucleic acids, components, enzymes, or chemical reagents. A compartment may be a well, a tube, a location above a surface, a chamber within a microfluidic device, or a droplet within a fluid. Multiple reactions may be set up in multiple compartments. For example, one or more reactions may be set up to generate a universal library. Reactions may proceed to produce identifiers through programmed temperature incubation or cycling. Reactions may be optionally or ubiquitously removed (e.g., deleted). Reactions may also be optionally or ubiquitously interrupted, integrated, and purified to collect identifiers in a single pool. Identifiers from a multi-identifier library may be collected in the same pool. Individual identifiers may include a barcode or tag identifying the identifier library to which they belong. Alternatively or additionally, the barcode may contain metadata about the encoded information. Supplementary nucleic acids or identifiers may also be included in an identifier pool along with an identifier library. Supplementary nucleic acids or identifiers may serve to contain metadata about the encoded information or to obfuscate the encoded information.

[0064] An identifier rank may include a method for determining the order of identifiers. This method may include a lookup table having all identifiers and their corresponding ranks. The method may also include a function for determining the order of any identifier, including a lookup table having the ranks of all components constituting the identifier and combinations of these components. Such a method may be called lexicographical ordering and may be similar to the way words in a dictionary are sorted in lexicographical order. In data in an address encoding method, an identifier rank (encoded by the identifier's rank object) may be used to determine the position of a byte within a bit stream (encoded by the identifier's byte-value object). In an exemplary encoding method, the identifier rank for the current identifier (encoded by the entire identifier itself) may be used to determine the position of a bit-value of '1' within the bit stream.

[0065] Identifiers can be constructed by assembling and combining component nucleic acid sequences. For example, information can be encoded by taking a set of nucleic acid molecules (e.g., identifiers) from a defined group of molecules (e.g., combination space). Each possible identifier of a defined group of molecules may be an assembly of nucleic acid sequences (e.g., components) from components of an assembled set that can be divided into layers. Each individual identifier can be constructed by connecting components of all layers in a fixed order. For example, if there are M layers and each layer has n components, then at most C = n M Can be configured with 2 unique identifiers, up to a maximum of 2 C 10⁶ different information items or C bits can be encoded and stored. For example, to store megabits of information, 1 x 10⁶ unique identifiers or C = 1 x 10⁶ 6A combination space of size can be used. The identifier in this example can be assembled into various components configured in various ways. Each assembly is n = 1 x 10 3 It can be made with M = 2 prefabricated layers containing components. Alternatively, the assembly can be n = 1 x 10 respectively. 2 It can be made into M = 3 layers containing components. As can be seen from this example, encoding the same amount of information using a large number of layers can reduce the total number of components. Using fewer total components can be advantageous in terms of write costs.

[0066] In one example, one may start with two layers, X and Y, each having nucleic acid sequences (e.g., components) x and y, respectively. Each nucleic acid sequence from X can be assembled into each nucleic acid sequence from Y. The total number of nucleic acid sequences maintained in the two sets may be the sum of x and y, the total number of nucleic acid molecules, and, if possible, the identifiers that can be generated may be the result of x and y. If sequences from X can be assembled into sequences of Y in any order, more nucleic acid sequences (e.g., identifiers) may be generated. For example, if the assembly order can be programmable, the number of generated nucleic acid sequences (e.g., identifiers) may be twice the product of x and y. The set of all possible nucleic acid sequences that can be generated may be referred to as XY. The order of assembly units of unique nucleic acid sequences in XY can be controlled using nucleic acids with distinct 5' and 3' ends, and restriction cleavage digestion, ligation, polymerase chain reaction (PCR), and sequencing analysis involve distinct sequences with 5' and 3' ends. This approach can reduce the total number of nucleic acid sequences (e.g., components) used to encode N individual bits by encoding information as a combination and order of assembly products. For example, to encode 100 bits of information, two layers of 10 distinct nucleic acid molecules (e.g., components) can be assembled in a fixed order to produce 10 * 10 or 100 distinct nucleic acid molecules (e.g., identifiers), or another layer consisting of 5 distinct nucleic acid molecules (e.g., components) and 10 distinct nucleic acid molecules (e.g., components) can be assembled in any order to produce 100 distinct nucleic acid molecules (e.g., identifiers) in a single layer.

[0067] The nucleic acid sequences (e.g., components) within each layer may include a unique (or distinct) sequence or barcode in the middle, a common hybridization region at one end, and another common hybridization region on the other end. The barcode may contain a sufficient number of nucleotides to uniquely identify all sequences within the layer. For example, generally, there are four possible nucleotides for each base position within the barcode. Therefore, three base barcodes are 4 3 = 64 Nucleic acid sequences can be uniquely identified. Barcodes can be designed to be generated randomly. Alternatively, barcodes can be designed to avoid sequences that may cause problems with the identifier's composition chemistry or sequencing. Additionally, barcodes are designed to have a minimum Hamming distance from other barcodes to reduce the likelihood that basic resolution variations or reading errors will interfere with the correct identification of the barcode.

[0068] The hybridization region on one end of a nucleic acid sequence (e.g., a component) may differ in each layer, but the hybridization region may be the same for each member within the layer. Adjacent layers are layers that have hybridization regions complementary to the components so that they can interact with each other. For example, all components of layer X may have complementary hybridization regions and thus can attach to components of layer Y. The hybridization region on the opposite end can serve the same purpose as the hybridization region at the first end. For example, all components of layer Y can attach to the X layer component at one end and the Z layer component at the opposite end.

[0069] To construct an identifier, the assembly of two or more components from different layers (e.g., X, Y, or Z) can be achieved using polymerase chain reaction (PCR), ligation, or recombination. Generally, any method for linking two or more different nucleic acid sequences may be used to construct identifiers from an identifier library. In some cases, all or part of the possible identifier combination space may be constructed before digital information is encoded or written, and the writing process may then involve mechanically selecting and pooling identifiers from an already created set of information. In other examples, identifiers may be constructed after one or more steps of the data encoding or writing process have occurred (i.e., as information is written). Methods for constructing identifiers include, but are not limited to, linking components such as nested extension PCR (or polymerase cyclic assembly), sticky end-linking, recombinase assembly, template-induced linking (or cross-linked strand linking), biobrick assembly, Golden Gate assembly, Gibson assembly, and ligation cycling reaction assembly. A method for constructing an identifier may also include deleting a nucleic acid sequence (e.g., a component) from a parent nucleic acid sequence (or parent identifier) ​​or inserting a nucleic acid sequence (e.g., a component) into the parent sequence. For example, an identifier may be generated from a parent identifier composed of multiple components. Components may be detached from or inserted into the parent identifier to generate a unique identifier. Enzymes that modify the parent identifier may include double-strand specific nucleases, single-strand specific nucleases, and Cas9.

[0070] Enzyme reactions can be used to assemble components from different layers. Since the components of each layer have specific hybridization or attachment regions to the components of adjacent layers, assembly can occur as a single-pot reaction. For example, a nucleic acid sequence (e.g., component) X1 from layer X, a nucleic acid sequence Y1 from set Y, and a nucleic acid sequence Z1 from set Z can form an assembled nucleic acid molecule (e.g., identifier) ​​X1Y1Z1. Additionally, multiple nucleic acid molecules (e.g., identifiers) can be assembled in a single reaction by including multiple nucleic acid sequences from each layer. For example, if both Y1 and Y2 are included in the single-pot reaction of the previous example, two assembled products (e.g., identifiers), X1Y1Z1 and X1Y2Z1, can be produced. This reaction multiplexing can be used to reduce write time if multiple identifiers can be physically configured. The assembly of nucleic acid sequences can be performed at time intervals of about 1 day, 12 hours, 10 hours, 9 hours, 8 hours, 7 hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours, or less than 1 hour. The accuracy of the encoded data can be at least approximately 90%, 95%, 96%, 97%, 98%, 99%, or higher.

[0071] The step of recording information into a nucleic acid sequence may include parsing the information into a symbol string, mapping the symbol string to a unique identifier, and generating an identifier library containing identifiers corresponding to the symbol string. The identifier library may include identifiers for each identifier rank, or exclude identifiers for an identifier rank if they correspond to a selected symbol value (e.g., 0 or 1). The information may include a series of symbols. For example, a string of symbols may include symbols taken from a fixed, limited, finite set of symbols. The string may be converted into a second symbol sequence. The second sequence of symbols may include a formal data structure. The second sequence of symbols may be parsed into words. The words may be converted into codewords using a codebook. The codebook may be an explicit codebook or an implicit codebook. The codewords may be parsed into a third symbol string. Each symbol of the third symbol string may be mapped to a unique identifier. A set of identifiers (e.g., an identifier library) may be enumerated or defined so that each symbol can be encoded into one or more identifiers. A set of identifiers (e.g., an identifier library) may include or add information related to one or more codebooks, data structures, and combination spaces.

[0072] The formal data structure may include a tree, trie, table, set, key-value dictionary, or set of multidimensional vectors. The formal data structure may be queried by one or more query types, including range queries, rank queries, count queries, membership queries, nearest neighbor queries, match queries, select queries, or any combination thereof. A second sequence of symbols containing the formal data structure may be parsed into a sequence of words to minimize the number of identifiers used to encode the bit stream. Each bit of the source bit stream may be associated with an identifier in the combination space.

[0073] The combination space of identifiers may include unique identifiers that can be generated by one or more construction algorithms from a library of T total components. In one embodiment, the construction algorithm may generate identifiers using a Cartesian product scheme comprising M layers in which the i-th layer contains Ni components. The number of identifiers in the combination space may vary depending on the number of layers, the number of components in each layer, and the method used to assemble the identifiers. FIG. 5 illustrates an example of a combination space of identifiers using a product scheme comprising M layers and N components in each layer. In this example, M = 4 and N = 2. Items (501-504) 5 in FIG. 5 show the layers of this example. Items 511 and 512 show two components of Layer 1 in this example. Similarly, items 509-510, 507-508, and 505-506 show components belonging to Layers 2, 3, and 4. The components are arranged in a repetitive pattern to show a combination space of 16 unique identifiers generated in this scheme. The steps in the example of the combination algorithm for generating each identifier in the combination space can be illustrated by the tree diagram shown in Item 513. The tree diagram can be divided into M layers. Each layer has a node representing the selections available for the components of that layer. For example, two arrows emanating from the node labeled "a" in Layer 1 show the selection of two components labeled Items 511 and 512 in Items 1 and 2. The arrows emanating from node b in Layer 2 represent the elements (509 and 510) conditioned according to the selection of component (511) in component layer 1, and the left and right arrows emanating from each node correspond to the pattern of components shown in the layer of item (515).Nodes are arranged according to the component ranking defined for the product scheme. Each path under the tree diagram corresponds to a unique identifier, starting from the top node labeled "a" to one of the bottom nodes. One such path is exemplified by item (514). The combination space of all identifiers (a total of 16 in this example) is illustrated by item (518). Item (517) represents a single bit value within an exemplary bit stream that can be encoded using this combination space. Each bit of the bit stream corresponds to a unique identifier indicated below the bit. In one embodiment, the value of the bit is represented by the inclusion or exclusion of identifiers from a configured identifier library. To encode the bit stream, all identifiers corresponding to bits having a value of "1" can be configured and pooled, while identifiers corresponding to bits having a value of "0" can be excluded. Excluded identifiers are indicated by a dark overlay. Item 519 indicates one of the excluded identifiers corresponding to the 10th bit with a value of "0".

[0074] Information can be encoded into identifiers abbreviated as data (DAA scheme) in an address scheme. A source bit stream can be divided into words of fixed length L. The bit stream can be interpreted as a symbol stream of L-bit symbols (e.g., each symbol contains L-bits). In the symbol stream (i.e., for each symbol containing L-bits), a unique identifier can be constructed for each symbol and can be pooled or grouped together. In one embodiment, the identifier can be constructed using a product scheme comprising M layers, each layer having N components. Each identifier can be decomposed into two parts (or objects). The first part is at most k <M 개의 층을 포함할 수 있으며, 심볼의 어드레스에 관한 정보를 제공할 수 있다. 고유 식별자의 제 2 부분은 M-k 층들로부터의 구성요소들을 포함할 수 있고 심볼의 값에 관한 정보를 제공할 수 있다. 대안으로, 또는 이에 부가하여, 소스 비트 스트림은 길이가 L 비트의 워드들의 스트림으로 분할될 수 있다. 코드북은 4 개의 염기 A, T, C 및 G를 포함하는 핵산 알파벳 상으로 워드를 코드 워드로 매핑하는데 사용될 수 있다. 각각의 코드 워드는 4 개의 염기로 구성될 수 있다. 각 L- 비트 워드에 대한 식별자는 대응하는 합성된 코드 워드를 그 코드 워드의 어드레스를 지정하는 구성 요소들의 어셈블리에 어셈블 링하거나 연결함으로써 구성될 수 있다.

[0075] Before writing the source bit stream to the identifier library, the source bit stream may be encoded into an intermediate bit stream. The source bit stream may be divided into words. Different code words may be selected to replace a word. The length of the code word may be greater than, equal to, or less than the length of the corresponding word. In one embodiment, each word X containing a number N(X) of Y symbols may be replaced with a code word containing a smaller or larger number of Y symbols. For example, a word containing N(X) "1" symbols may be replaced with a code word containing fewer than N(X) "1" symbols. In an exemplary encoding method, this may result in a reduction in the size of the identifier library used to encode the given digital information. Minimizing the number of physically assembled identifiers can reduce the time required to write information to the identifiers and to read the information encoded by the identifiers. FIG. 6 schematically illustrates an exemplary method for minimizing the number of identifiers to be configured to write the bit stream using extended code words. A bit stream can be divided into words, and in this example, each word can be of a fixed length of 2 bits. A list of words containing 2 bits includes '00', '01', '10', and '11'. Each word may appear zero or more times in the bit stream. For example, the bit stream '0110101010011101' is a 2-bit word {01,10,10,10,10,01,11,01} where the word '00' appears zero times, the word '01' appears three times, the word '10' appears four times, and the word '11' appears once. In this series of words, the total number of "1" symbols is 9, which indicates that the encoding method may require 9 individual identifiers to represent the bit stream. However, words can be recoded so that fewer identifiers can be used to encode a given bit stream.

[0076] Digital information to be encoded into nucleic acid can first be converted into a series of symbols and then reconstructed into a formal data structure dependent on one or more query types. This data structure can be serialized into a second string of symbols. This second string of symbols can be coded using one or more codebooks for one or more purposes, including error protection, encryption, optimization of write speed, or minimization of the identifier library size. FIG. 6 illustrates an exemplary method for minimizing the identifier library size. Item (620) illustrates a tree diagram representation of the combination space, the notation of which is described in FIG. 5. In this example, Item (621) represents a bit value from a bit stream of 16-bit values. Item 622 illustrates a set of identifiers corresponding to bit values ​​of the bit stream having the value "1". Thus, this encoding may require a combination of nine individual identifiers corresponding to nine bits having the value "1". However, the size of this identifier library can be reduced by re-encoding the bit stream using a codebook that maps 2-bit words to 3-bit codewords, so that the new 3-bit codewords have fewer "1" symbols and appear as a smaller identifier library.

[0077] In this example of a re-encoding method, the bit stream can be divided into eight consecutive 2-bit words, and the occurrence count of each 2-bit word can be recorded. In this example, this count is displayed under the Count column in Table 623. All possible 3-bit codewords are listed as columns to form a matrix, where cell(i, j) contains the cost of mapping 2-bit word i to a distinct 3-bit codeword j. This cost can be calculated by multiplying the number of "1" symbols in the codeword by the number of word occurrences in the original bit stream to determine the number of identifiers that can be generated using this word-to-codeword substitution. For example, the word "01" occurs 3 times in the original bit stream. If it is mapped to the codeword "111", the number of "1" symbols resulting from this substitution in the re-coded bit stream can increase from 3 to 12. This cost is calculated for all such possible substitutions. As indicated by item 623, the matrix thus obtained can be converted into a weighted two-part graph, and the minimum weighted perfect match can be obtained using an algorithm such as the Kuhn-Munkres algorithm. The minimum perfect match can be equivalent to selecting exactly one cell in each row and column of the matrix (623) such that the sum of all selected cells is minimized. The cost of each cell in one such minimum re-encoding is shown in Table 623 with shaded cells. In this minimum re-encoding, the word "00" is mapped to the codewords "011", "01" through "001", "10" through "000", and "11" through "010". The new bit stream thus coded has a total of four "1" symbols. Thus, the cost can be reduced from nine in the original bit stream to four in the newly encoded bit stream. The new bit stream contains a three-bit codeword as illustrated in the tree structure diagram by item (624).Each 3-bit codeword uniquely maps to a 2-bit codeword from the original 2-bit codeword set described by item (625). Item (626) shows a new identifier library to be assembled.

[0078] The selection of symbols encoding digital information can enable the detection and / or correction of encoding errors. By re-encoding a symbol stream to include error-prevention symbols calculated from the symbols of the original string, errors encountered during the process of recording the symbol stream using nucleic acids can be detected or corrected. In one embodiment, the symbol stream may be divided into fixed-length words, and one or more error-prevention symbol strings may be calculated from each of these words and attached to the words to obtain a recorded string. For example, the number of identifiers to be configured within a fixed-length block of K identifiers may be counted. If this count is even, an additional identifier may be added to the block, and if the count is odd, an additional identifier may not be added. A combination identifier space may be selected to accommodate such additional identifiers. When such an identifier block is read, any recording error in which an identifier is incorrectly omitted or an additional identifier is incorrectly added may be detected, because such events may invalidate the required attribute that each block has an odd number of identifiers. In another embodiment, the number of identifiers within any fixed-length block of K identifiers is counted, and a count minus is calculated from the count. This value, called an error protection value, can be added to the block and encoded. A combination space can be selected to accommodate identifiers corresponding to this error protection value. In this case, when the block and the error protection value are read, any error in which an identifier is incorrectly omitted can be detected. If the omitted identifier may be present in the original block, this can be reflected in the mismatch error protection value. If the omitted identifier is present in the error protection value, a lower value may indicate that the error may be present in the error protection value. If there is an error in both the block and the value, the mismatch error can be detected. In another embodiment, the symbol stream can be divided into fixed-length words of W symbols.Then, each word is remapped back to a codeword so that each codeword constitutes a fixed number of identifiers V. FIG. 7 schematically illustrates this uniformly weighted codeword error detection method. Item (727) illustrates an identifier library that can be configured to encode the bit stream illustrated in the tree structure diagram of FIG. 7. In the original bit stream, for any fixed word length W, the number of identifiers is not constant: for example, for W = 2, there may be one identifier for each of the first six words and two identifiers for the second word. Table 727 shows a re-encoding example codebook that maps words of length W = 2 to codewords of length V = 4. The example codebook maps the words "00", "01", "10", and "11" to the codewords "0011", "0101", "0110", and "1001", respectively. All codewords have exactly two "1" symbols, and since the word and codeword lengths are fixed, the resulting bit stream has exactly two "1" symbols in all codewords of length 4 symbols. This is illustrated by an exemplary tree diagram for a re-encoded bit stream shown in 730. Entry (729) represents a word mapped to a distinct codeword as illustrated by entry (728). In a fixed rate and identifier library, missing identifier errors can be detected during decoding.

[0079] Write time can be minimized by interpreting the input bit stream as a multi-valued Boolean function. In one embodiment, the input bit stream can be divided into blocks of fixed length L before minimizing write time. The input bit stream can be subjected to a heuristic logic minimization algorithm, such as Espresso-mv or mvsis, to obtain a multi-valued algebraic representation representing the source bit stream. In one embodiment, the input bit stream can be encoded using an M-layer multiplication method to construct identifiers. In this embodiment, the input bit stream can be interpreted as an M-input multi-valued Boolean function having a single Boolean output. For a Boolean function, a set of functions can be defined as the set of all inputs for a function where the function outputs a value of "1". Using the logic minimization technique, the algebraic representation of the Boolean function includes a sum-of-product formula. The obtained representation includes all identifiers within a set of the source bit stream. Each term in the representation can be converted into a series of identifiers (constructed in a multi-layered manner) in a single reaction compartment (e.g., compartment or reaction vessel). The obtained expression can be used to minimize the number of response compartments used and maximize the number of identifiers assembled in a single compartment. The expression can also be used to minimize the total time used to set up the identifier assembly response. For example, the write time may be proportional to the number of response compartments set at the top. A similar method can be used to set up the response used to query a subset of bits from a source bit stream.

[0080] Figure 8 schematically illustrates the output of an exemplary scheme for minimizing a set of reactions. Consider a product scheme with a bit stream of length L and M layers. Each layer i has a component of Ni, and thus the product of all Ni is at least L. Each component of a layer can be interpreted as a Boolean function F of M variables, where each variable Vi can take one of the values ​​of Ni between 0 and Ni-1. All combinations of these variable values ​​can be represented as M-dimensional vectors, where the value of variable Vi can be represented as an integer in the i-th dimension of the vector. A Boolean function F can be defined using these vectors as inputs and each bit value of the bit stream as an output. If the product scheme has a combination space of size greater than L, the output of F on those additional input vectors can be defined as distinct "don't-care" values.

[0081] FIG. 8 illustrates an example in which the information described in Item 831 can be represented as a bit stream of length 64 bits as shown in Item 832 and encoded through a Product Scheme comprising two layers having 13 and 5 components in each layer. The Boolean function F includes 65 possible input vectors, each vector being two-dimensional. Each dimensional variable V1 and V2 takes values ​​of 13 and 5, respectively, where V1 takes values ​​in the range of 0 to 12 and V2 takes values ​​in the range of 0 to 4. A set of all possible variable-value combinations can be depicted as a tree diagram. When the output of the aforementioned function F is "1", it may also be depicted as a tree diagram containing a subset of arrows. This tree diagram is shown at the top of FIG. 8. The set of variable value combinations in which F takes the value "1" corresponds to a set of identifiers that must be constructed to encode the bit stream. Therefore, the paths from the root of the tree diagram to the individual values ​​shown in the tree diagram correspond to the series of reactions required to combine each identifier. In this example, the arrows labeled 833 and 834 represent a set of paths corresponding to 3 bits in the bit stream to be encoded. These three paths also correspond to three identifiers that must be assembled to encode the three bits. Because the vectors describing these "1" values ​​of F differ from each other in the second dimension, having values ​​0, 3, and 4, the corresponding identifiers differ in the second layer, and the 0th, 3rd, and 4th components are taken from the second layer. All three identifiers have the same component corresponding to the value V1 = 10 in the first layer. Consequently, all three identifiers can be combined in a single reaction using the components V1 = 10 and V2 of the set {0, 3, 4}. The resulting sets of combinations (10, 0), (10, 3), and (10, 4) correspond to the correct sets of identifiers to be generated.To encode a given bit stream in a tree diagram, 13 sets of reactions are required. However, since the tree diagram can be decomposed into sets of tree diagrams using heuristic-guided search, all identifiers in each element tree can be assembled into a single reaction. For example, a greedy heuristic can be used where all values ​​of V1 for some value V2 = v are grouped together so that all identifiers correspond to the "1" values ​​of F. Item 835 is V2 = 0 and V1 = {3, 4, 5}. In another embodiment, multiple heuristics can be combined to obtain a minimum set covering the "1" values ​​of F. In another embodiment, a heuristic technique from logic minimization [Brayton et al. Logic Minimization Algorithms for VLSI Synthesis Kluwer Academic Publishers (the full text incorporated herein by reference) can be used to minimize the number of response sets. The five tree diagrams shown under the label "Heuristic search guided optimized solution" include all "1" values ​​of F. Consequently, five response sets can be used to be set in five individual sections rather than 13 sections within the original tree diagram.

[0082] Each symbol (e.g., a bit within a bit stream) may be mapped to one or more unique identifiers within a combination space. A set of identifiers may be determined and enumerated in computer memory or may be generated by combinatorially assembling a set of identifiers into an identifier library. When digital information is presented to be encoded into an identifier library, in one embodiment, each symbol within the digital information may be mapped to a distinct identifier within the combination space. There may be a vast number of methods for mapping a given bit stream into a combination space that is generated from a combination method (e.g., a product method, a permutation method, or some other method) and includes any selected number of components. Some of these mappings may be useful for reducing the number of queries when querying the encoded data later. In particular, a mapping that preserves the locality of a symbol from the original symbol stream after mapping the symbol into the combination space may be useful for reducing the number of accesses used to respond to a query. An access may be a request to select a set of identifiers from an identifier library or an identifier pool described by a single nucleic acid sequence referred to as an access sequence. In one embodiment, when an identifier is assembled from components, a set of all identifiers containing a specific component can be accessed with a single access. The nucleic acid sequence of the component may be the access sequence in this embodiment. Mapping families that preserve the locality of the original symbol are called isometric mappings. Furthermore, a single digital message can be mapped to two orthogonal combination spaces, each having its own component library and becoming two orthogonal identifier libraries representing the same digital message. The two mappings can be useful for reducing the number of accesses for two sets of queries.This type of encoding using multiple mappings can be referred to as multi-encoding, and if the number of mappings can be fixed to two mappings, it can be called double encoding.

[0083] FIG. 9 schematically illustrates the isometric mapping of identifiers of addresses and the double encoding of data. The process of encoding a digital message may include the step of converting information into a sequence of symbols and converting the sequence of symbols into a second sequence of symbols having a formal data structure according to one or more query types. FIG. 9 illustrates an example where the digital information to be encoded may be the two-dimensional image shown in Item 936. Item 937 illustrates a schematic diagram of the image, and the diagonally marked circle represents the lower right quadrant of the image. The original sequence of symbols, in this case the bit values, may be encoded in the presented order. This order is shown in Item 938, and the resulting tree diagram of the product scheme is shown in Item 939. Reading the lower right quadrant of the image allows querying the shaded circle of the encoded bit stream. In combinational space, this can be interpreted as a query for four identifiers. In this example, assuming there are two components in each layer of the Product Scheme, queries for all identifiers starting with component 101* and queries for all identifiers starting with component 111* can be used. * indicates that all components of that layer can be returned as answers to the query. Two queries are used because sub-regions of an image can be mapped to the join space in such a way that adjacent regions of the image are mapped to identifiers that are not adjacent in the join space.

[0084] Item (941) describes an alternative mapping where neighboring regions of an image are mapped to nearby identifiers. This can be called an equidistant (i.e., distance-preserving) mapping. In this case, a single query can be used. The query can be responded to with any identifier starting with 11**. This can be generalized to multidimensional data structures including multi-column tables, trials, trees, sets, and vectors. More generally, since product schemes encode data in a unique multidimensional manner, they can query various types of data to optimize and parallelize processing. Item 945 shows a multidimensional data set consisting of four dimensions: X, Y, Z, and W. In this example, X, Y, and Z each take two values, and the fourth dimension W takes four values. Each of the four dimensions corresponds to a single bit value in this example. Generally, this can be extended to integer values. Item (946) illustrates a tree diagram for encoding this 32-bit bitstream using a four-layer product scheme. In particular, the Product scheme structure preserves the dimensions of the original data structure. Dimensions X, Y, and Z can be mapped to a binary hierarchy, and dimension W, which takes four values, can be mapped to a hierarchy with four components. Additionally, items (947, 948) represent two mappings of the dataset to the same combinatorial space. The two mappings differ in how regions of the data structure are mapped to the proximate regions of identifiers in the combinatorial space. In the mapping of item 947, data regions corresponding to X = 0, Y = 0 and X = 1, Y = 1 are mapped to non-adjacent identifiers, whereas in the mapping of item 948, they are mapped to adjacent identifiers. Item 949 shows possible queries for unshaded bit values. Item 952 shows the component access order used to retrieve these bit values ​​using the mappings shown in item 947.In this example, the hierarchy can respond to a query using a single access to component 0 of the W hierarchy. Item 50 is a more complex query, which can be responded to by two parallel accesses to components W = 0 and Y = 1, followed by a serial access to component X = 1. This responds to a lookup for all unshaded values ​​in item 950. Item 951 represents a more complex lookup. Using the mapping of item 947, this query may require more than four accesses. However, using the mapping of item (948), this query can be responded to using one access and one degradation step. The degradation step removes all identifiers that constitute a specific pattern. In this example, the pattern is component 1 of the W hierarchy. Mapping the data structure to a combination space in this way can reduce the complexity of responding to data queries. In some embodiments, multiple mappings of the same data structure can be encoded in a single pool of identifiers using an orthogonal or distinguishable set of components. This is described in the mappings shown in items 947 and 948. Two identifier libraries can encode the data structures shown in item 945 and can respond to queries using the mappings based on the number of accesses used for each mapping.

[0085] Digital information provided for encoding into an identifier library may include information that can be protected from unauthorized decoding. The method of recording information in DNA described herein may provide an additional level of protection against unauthorized decoding of the encoded information. Biochemical methods of encryption, authentication, obfuscation, and destruction may be used to protect the encrypted information. In one embodiment, information may be encoded and obfuscated by including decoy identifiers in the identifier library. A decoy identifier may be an identifier that does not encode information that is part of the original digital information provided for encoding, and is included to create a decoding process that is extremely expensive and difficult to handle without possessing a decoy key. A decoy key may be a set of sequences of components such that selecting an identifier containing a component can separate some or all of the identifiers constituting the original identifier library, or conversely, deleting all identifiers containing a component can delete some or all of the decoy identifiers.

[0086] FIG. 10 schematically illustrates an exemplary method for masking encoding and decoding to protect against unauthenticated decoding. A bit stream may be encoded with a unique identifier and an identifier library may be assembled. Additional nucleic acid sequences may be added to the identifier library. Additional or supplementary nucleic acid sequences may have a length similar to a unique identifier without a key for decoding information and cannot be distinguished from the unique identifier. Decoding information may include the step of applying the identifier pool to one or more selected and / or degraded target nucleic acid sequences until a unique identifier is extracted from the supplementary nucleic acid sequence. Item (1056) illustrates a tree diagram depicting the encoding of a bit stream using a 5-tier Product Scheme, with each tier containing two components. The original bit stream is illustrated by Item (1057), contains 16 bits, and is depicted with circles surrounding the values. However, this bit stream is encoded in a combination space larger than that used to encode 16 bits, and, as indicated by entry (1058), for example, the remaining undefined symbols are shown as empty circles. The indicated 5-layer binary scheme enables 32 unique identifier combination spaces. Some of the identifiers corresponding to the "1" bit value of the original bit stream are shown in entry 1060. Some of the remaining identifiers that do not correspond to any bit value of the original bit stream are shown by entry 1059 and are called "potential decoy identifiers." These identifiers are shaded to indicate a minimum number of components sufficient to distinguish them from the identifiers corresponding to the bit values ​​of the original bit stream. These identifiers are called potential decoy identifiers.In this example, which identifiers are selected as decoy identifiers and which identifiers are selected to correspond to bit values ​​of the original bitstream may be arbitrary, but may be governed by the data structure, query constraints, and strength of the bitstream, which may be awkward or hidden. Some useful identifiers from the set of potential deceptive identifiers are selected to be included in the identifier library encoding the original bitstream, as shown in Item 1062, and are designated as "selected decoy identifiers." The bitstream may be encoded as a pool of identifiers containing both identifiers corresponding to bit values ​​and decoy identifiers that do not correspond to any bit values ​​of the original bitstream. Therefore, any unauthorized decoding of the pool may not be able to faithfully decode the original bitstream if there is no information regarding the set of selected decoy identifiers. A set of components describing the set of selected decoy identifiers may be called a decoy key. The decoy key for this example is denoted as item 1064 and includes two sequences: component 1, 0, 1, 1 of layer 0-4 and component 0, 1, 1 of layer 0-3. The decoy key can be interpreted in the following way: Each component in the component sequence of the decoy key corresponds to an access query. All identifiers matching that component query are accessed from the current pool. The last component of the component sequence is not accessed. Instead, it can be used to delete all identifiers matching that component from the current pool. Table (1063) illustrates the steps required to execute the decoy key illustrated in 1064. Starting from the pool of all identifiers illustrated in the tree diagram (1061), deletion after a series of accesses results in the survival of the exact identifier library corresponding to the original bit stream: all decoy identifiers are removed. The remaining identifiers are indicated in the shaded cells of Table 1063.

[0087] A system that encodes information into nucleic acid sequences and decodes information from nucleic acid sequences

[0088] A system for encoding digital information into nucleic acids (e.g., DNA) may include a system, method, and apparatus for converting files and data (e.g., raw data, compressed zip files, integer data, and other data formats) into bytes and encoding the bytes into segments or sequences of nucleic acids, typically DNA or a combination thereof.

[0089] In one embodiment, the present disclosure provides a system for writing information into nucleic acid sequence(s). A system for writing information into nucleic acid sequence(s) may include an assembly unit and one or more computer processors. The assembly unit may be configured to generate an identifier library that encodes a sequence of symbols. The identifier library may include at least a subset of a plurality of identifiers. One or more computer processors may be operably coupled to the assembly unit. The computer processors include the steps of: (i) converting a sequence of symbols into a code word using one or more codebooks; (ii) parsing the code word into a coded sequence of symbols; (iii) mapping the coded sequence of symbols to a plurality of identifiers; (iv) instructing the assembly unit to generate an identifier library; and (v) instructing the assembly unit to attach the descriptions of the one or more codebooks and the plurality of identifiers to the identifier library. Each symbol of the coded sequence of symbols may be encoded by one or more identifier(s).

[0090] In another aspect, the present invention provides an integrated system for nucleic acid-based data storage. The integrated system for nucleic acid-based data storage may include a data encoding unit, a storage unit, a read unit, and one or more computer processors. The data encoding unit may be configured to write digital information to a nucleic acid sequence. The storage unit may be configured to store a nucleic acid sequence that encodes digital information. The read unit may be configured to access and read digital information encoded in a nucleic acid sequence. One or more computer processors may be connected to the data encoding unit, the storage unit, and the read unit. One or more computer processors may be programmed to (i) instruct the data encoding unit to encode digital information into a nucleic acid sequence, (ii) instruct the storage unit to store the encoded digital information in a nucleic acid sequence, and (iii) instruct the read unit to access and decode the digital information stored in the nucleic acid sequence. Digital information may be encoded into a nucleic acid sequence in the absence of base-by-base nucleic acid synthesis.

[0091] The system may include one or more computer processors and a Human Machine Interface (HMI) for controlling and programming the computer processors. The system may encode and recode digital information using the methods described herein. The system may generate a list of identifiers that constitute an identifier library. Alternatively, or additionally, an external computer processing unit may generate a list of sequences of identifiers that constitute an identifier library. The system may have an interface for receiving the list of sequences of identifiers. The interface unit may generate and pool identifiers by converting the sequences of identifiers into commands for downstream units or modules of the system.

[0092] The above system may have an assembly module. The assembly module may be configured to receive multiple substrates (e.g., components) and reactants (e.g., enzymes) and output multiple reactions to generate identifiers that constitute one or more identifier libraries. One or more identifiers may be generated from a given reaction. One or more identifier(s) may be generated from multiple reactions. The multiple reactions are 1, 2, 4, 6, 8, 10, 20, 30, 50, 75, 100, 150, 200, 300, 400, 500, 750, 1000, 10000, 1x10 5 , 1x10 6 , '1x10 7 , 1x10 8 , 1x10 9 It may include or more reactions. Many reactions are approximately 1 x 10⁻⁶ 9 , 1x10 8 , 1x10 7 , 1x10 6 , 1x10 5 It may include , 10000, 1000, 750, 500, 400, 300, 200, 150, 100, 75, 50, 30, 20, 10, 8, 6, 4, 2 or fewer reactions. One or more reactions may be performed simultaneously or sequentially. One or more or multiple reactions may be combined to generate an identifier library. The assembly unit may optionally remove one or more multiple reactions that do not generate a selected identifier. The assembly unit may include one or more sections, containers, or compartments. The assembly unit may include multiple sections, containers, or partitions. Each section, container, or partition may generate, store, maintain, facilitate, or terminate one or more assembly reactions.

[0093] The assembly unit may include a reaction module. The reaction module may collect reagents, one or more nucleic acid sequences, one or more components, one or more templates, or any combination thereof. The reaction module may be configured to isothermally or agitate the assembly reaction to generate one or more identifiers. The reaction module may additionally include a detection unit. The detection unit may monitor the assembly of identifiers. The reaction module may include a plurality of compartments. A plurality of partitions may each contain one or more assembly reactions. A plurality of partitions may be wells or droplets of a chemically modified surface.

[0094] Substrates or inputs may include one or more M layers. Each layer may include one or more components. The components of each layer may be distinguishable from the components of other layers. Substrates may also include composition templates, primers, probes, and other elements to direct and facilitate substrate assembly reactions. Reagents may include enzymes, buffers, nucleic acid sequences, cofactors, or any combination thereof. Enzymes may be generated by the overexpression of corresponding recombinant genes in living cells. Reagents may be combined in individual assembly reactions or combined into a master mix before being added to the assembly reaction.

[0095] The above system may further include a storage unit (e.g., a database). An assembly unit may output one or more identifier libraries. One or more identifier libraries may be received by the storage unit. The storage unit may include one or more pools, containers, or partitions. The storage unit may combine individual identifier libraries with one or more additional identifier libraries to form one or more pools of identifier libraries. Each individual identifier library may include a barcode or tag that enables identifiers from each library to be identified and distinguished from one another. The storage unit may provide conditions for the long-term storage of identifier libraries (e.g., conditions to reduce the degradation of identifiers). Identifier libraries may be stored in powder, liquid, or solid form. The database may provide protection against ultra-violet light, temperature reduction (e.g., freezing), and prevention of degradation by chemicals and enzymes. Identifier libraries may be freeze-dried or frozen before being transferred to the database. The identifier library may include ethylenediaminetetraacetic acid (EDTA), other metal chelating agents, or other reaction blocking reagents that inactivate nucleases and / or buffers to maintain the stability of nucleic acid molecules.

[0096] The system may further include a selection unit. The selection unit may be configured to select one or more identifiers from an identifier library or a group of identifier libraries. An assembly unit may set all possible reactions to create a combination space, and the selection unit may selectively remove reactions that do not generate a target identifier and retain reactions that generate a target identifier. The selection unit may include an optical or mechanical removal module for removing reactions, a distributor delivering a degrading enzyme to a non-targeted reaction, or a dispenser delivering a primer or an affinity-tagged probe to a targeted reaction. The selection unit may facilitate the evaluation of stored data. Accessing information stored in nucleic acid molecules (e.g., identifiers) may be performed by selectively removing an identifier library or a portion of an identifier library from a group or pool of combined identifier libraries. Accessing data may be performed by selectively capturing or amplifying identifiers corresponding to the data to be accessed and / or removing identifiers that do not correspond to the data to be accessed. Methods for selecting identifiers may include using a polymerase chain reaction, an affinity-labeled probe, and a degradation-labeled probe. A pool of identifiers (e.g., an identifier library) may include identifiers having a common sequence at each end, a pool of identifiers (e.g., an identifier library) may include identifiers having a common sequence at each end, variable sequences at each end, or identifiers containing either a common sequence or a variable sequence at each end. Identifiers may include the same common sequence at each end or different common sequences at each end. An identifier library may include a common sequence distinct from that library so that a single library can be selectively accessed from a pool or group of one or more identifier libraries. The common sequence or variable sequence may be a primer binding site.One or more primers can bind to a common region on the identifier. Identifiers having bound primers can be amplified by PCR. The number of amplified identifiers can be much greater than the number of non-amplified identifiers.

[0097] A common sequence of identifiers may share complementarity with one or more probes. One or more probes may be bound to or hybridized with the identifier to be accessed. The probes may include an affinity tag. The affinity tag may be bound to a bead and may form a complex comprising a bead, one or more probes, and at least one identifier. The bead may be magnetic, and the selection unit may include one or more magnetic or electronic regions. The bead may collect and extract the identifier to be accessed. Alternatively or additionally, the bead may collect the identifier not to be accessed. The identifier may be removed from the bead under degrading conditions prior to reading. The affinity tag may be bound to a column, and the selection unit may include one or more affinity columns. The identifier to be accessed may be bound to the column of the identifier to be accessed, and the identifier not to be accessed may be bound to the column. Accessing the identifier bound to the column may be unbound before the column or changed within the column. Accessing identifiers may involve applying one or more probes to an identifier library simultaneously or applying one or more probes sequentially to an identifier library / identifier library group. In one example, one or more identifier libraries are combined, and each identifier library contains one or more distinct common sequences. A set of probes may be applied to the library to extract a first subset of identifiers. Subsequently, a second set of probes may be applied to the library to extract a second subset of identifiers. This operation may be repeated until all identifiers are extracted.

[0098] The common sequence of identifiers may share complementarity with one or more probes. Probes may bind to or hybridize the common sequence of identifiers. Probes may be targets of degrading enzymes. For example, one or more identifier libraries may be combined. A series of probes may be hybridized with one of the identifier libraries. A set of probes may contain RNA, and the RNA may induce a Cas9 enzyme. The Cas9 enzyme may be introduced into one or more identifier libraries. Identifiers hybridized with probes may be degraded by the Cas9 enzyme. Identifiers to be accessed may not be degraded by the degrading enzyme. In another example, identifiers may be single-stranded, and identifier libraries may be combined with single-stranded specific endonuclease(s) that selectively degrade identifiers that are not accessed. Identifiers to be accessed may be hybridized with a complementary set of identifiers to protect them from degradation by a single-stranded specific endonuclease (end-1). The identifier to be accessed can be separated from the degradation products by size selection, such as size selection chromatography (e.g., agarose gel electrophoresis). The selection unit can perform one or more size selection techniques. Alternatively, or additionally, the undegraded identifier can be selectively amplified so that the degradation products are not amplified (e.g., using PCR). The undegraded identifier can be amplified using primers that hybridize with each end of the undegraded identifier, and thus not with each end of the degraded or cleaved identifier.

[0099] Individual nucleic acid sequences (e.g., components and templates) that constitute an identifier or assist in the construction of the identifier may be synthesized by the system or synthesized and amplified outside the system. The system may further include a nucleic acid synthesis module. The nucleic acid synthesis module may perform base-based construction of the components and templates. Nucleic acid sequences (e.g., components and templates) may be constructed using phosphoamidite chemistry. Components may be initially constructed using phosphoamidite chemistry, and then the initial phosphoamidite template may be replicated using PCR. Components may be initially constructed using phosphoamidite chemistry, and then the replication of the template may be constructed by cloning the components into one or more high-copy vectors. The vector may be transformed into living cells, and the vector may be replicated during cell growth along with the inserted nucleic acid sequences. The vector may be isolated from the cell culture, and the components may be isolated from the vector using restriction digests. A double-stranded nucleic acid sequence can be converted into a single-stranded nucleic acid sequence by using an affinity-labeled probe that shares complementarity with one of the two nucleic acid strands.

[0100] The system may use techniques to minimize the number of responses used to generate an identifier library, thereby reducing write time. One or more techniques may include heuristic techniques. Heuristic techniques may minimize a compartmentalized set of responses used to construct a given set of identifiers from components. Heuristic techniques may include heuristics that include onset. The physical distance traveled by the recording device may also be minimized to reduce write time. Figure 8 illustrates an exemplary method for minimizing write time by generating a minimum set of responses.

[0101] The system can deliver fluids (e.g., reagents, components, templates) using pressure, vacuum, or suction. The assembly unit can combine one or more nucleic acid sequences with one or more reagent mixtures. The assembly unit converts materials coated with nucleic acid sequences, slip technology, stamping, laser printing, or droplet microfluidics (electrowetting, misting, printing, laser cutting, and templates) into reactions. The assembly unit can place biomolecules in the same location to create multiple co-located sets of biomolecules. A set of biomolecules located together can generate an identifier. For example, instead of connecting each component to each other, distinct components are assembled into a shared substrate, such as beads of each layer. Various techniques can be used to co-locate sets of biomolecules. For example, instead of forming an identifier by connecting sets of distinct components to each other, an identifier can be formed by associating components with a shared substrate, such as beads. As another example, instead of constructing an identifier by connecting a set of separate components to each other, an identifier can be constructed by assembling each component into a barcode sequence that identifies the combination of components.

[0102] A component rotation device can be used to position a set of biomolecules together. FIG. 11 illustrates a plan view (1108) of an exemplary component rotation device and a cross-sectional view (1109) of the component rotation device along the rotation axis (1110) of the plan view (1108). In this example, the component rotation device has a plurality of inlet ports and a plurality of outlet ports. The inlet ports may be located on the outer circumference of the carousel, and the outlet ports may be located on the inner circumference of the carousel. Each inlet port may selectively introduce a single input (typically a component, but possibly a nucleic acid, enzyme, or reaction mixture) into a reaction chamber connected to the outlet port. After introducing one input, the rotation device may move one position to selectively introduce an input adjacent to the reaction chamber. This process may be repeated until a number of selected inputs are combined.

[0103] The component rotation device may be composed of two substrates (1101 and 1102) having flat surfaces configured to face each other. In the embodiment illustrated in FIG. 11, the two surfaces are configured to rotate relative to each other. In some cases, it is advantageous to introduce oil or other lubricant between the two surfaces to reduce sliding friction. While lubricating oil may be used, fluorinated oil may be used to minimize the movement of biological material between the oil or chambers. In this example, the inlet (1103) and outlet (1104) ports are configured as through holes arranged in pairs on one of the substrates (1101). The second substrate (1102) has one chamber (1105) for each pair of through holes. When the surfaces of the two substrates are placed in contact facing each other, the chamber (1105) of the second substrate (1102) aligns with the lobes or channels (1106) of the first substrate to complete the flow path between the pair of through holes. Two substrates are designed to slide relative to each other in such a way that, through a full rotation, the two surfaces slide past each other, with each flow path sequentially connected through each chamber. In this manner, all inputs can be selectively added to each chamber. For example, in one embodiment, the first substrate has 72 pairs of through holes and the second substrate has 72 chambers. The system is configured so that different components can be selectively introduced into the chambers whenever the surface is indexed by 5 degrees. At the end of a full rotation, the outlet port (1107) drives the reaction mixture out of the chamber as a bore (1111). After purging the reaction from the chamber, the reaction can be reused for a subsequent reaction. Typically, one path is used to remove the reaction bowl (1111), a subsequent flow path is used to clean the reaction chamber, and the introduction of the master mixture into the reaction chamber may optionally have a separate flow path, or the master mix may be introduced with each input.In this example, the remaining 70 flow paths allow 70 unique inputs to be introduced sequentially into the given reaction chamber. If the inputs are components distributed across 22 layers of 3 components and one layer of 4 components, the combination space of the product scheme is 4 * 3. 22 = 1.2x10 11It is sufficient to generate identifiers. By slightly increasing the number of Euros to facilitate 96 components, 96 components can be arranged in 32 layers with 3 components per layer to generate up to 1.8e15 unique identifiers. In some embodiments, the chamber is filled with oil or gas before introducing the first input. In some specific embodiments, oil or gas is used to induce a reaction from the reaction chamber after the final input and reaction master-mix are introduced. There is no limit to the number of chambers or inputs that can be introduced. In some specific embodiments, 10 or fewer chambers are used, in some embodiments, 10 to 100 chambers are used, and in other embodiments, 100 to 1000 chambers are used. In other embodiments, more than 1000 chambers are used. There is no limit to the type of biological material that can be introduced into the chamber. In some cases, the input may be a factor for amino acid or peptide synthesis; in others, the input may be a reactant for synthesizing small molecules; and in still others, the input may include cells, bacteria, viruses, droplets, or other particles. It may include a lysis buffer for tagging, amplifying, binding, or identifying biological material within the cell lysate or on the surface of cells, bacteria, viruses, or other particles. In some cases, the chamber is indexed between port pairs at a rate of several times per hour or several times per minute. However, this indexing frequency can be at any timing and can be selected at a high speed. In some cases, it may exceed 1 time per second, 10 times per second, 100 times per second, 1,000 times per second, or 10,000 times per second. External fluid control may be used to selectively introduce input into the chamber as required.

[0104] Electro-wetting can be used to position a set of biomolecules together. FIG. 12 illustrates an electro-wetting method for input operations. Inputs (e.g., nucleic acids, components, templates, enzymes, or reaction mixtures) can be introduced through separate ports (1201). Each port (1201) can introduce a single input or a mixture of inputs. Selected inputs can be combined using electro-wetting to create droplets and combine identifiers. Droplets are prepared, combined, mixed, and split by selectively applying voltage to electrode patches (1202). In some embodiments, these electrode patches are arranged in a square array. The patches are typically configured to be separated from the droplets by an insulating coating having low electrical conductivity. The electro-wetting device may be open at the top or closed at the top. The electro-wetting chamber may contain an insulating fluid such as oil. Oils such as silicone oil, mineral oil, or hydrocarbon oil may be used. For example, fluorinated oil is used. A mixture of surfactants with other additives can be used to improve device performance by altering surface energy at the droplet-oil interface or the interface with the chamber wall.

[0105] The electro-wetting approach can be used to prepare and manipulate small amounts of fluid in the sub-picoliter to nanoliter range. For example, FIG. 12 illustrates an electro-wetting device configured to selectively combine inputs in a programmable manner. Using the electro-wetting approach, the system can be easily configured to process 10, 100, 1,000, 10,000, millions, or more droplets simultaneously. In some embodiments, it may be advantageous to combine droplets and then split the combined droplets into two mixed droplets. In some cases, mixing can be enhanced by combing and splitting in approximately orthogonal directions. The split droplets may each receive different subsequent inputs. The process can be repeated until all inputs necessary for identifier construction are introduced into the droplets. For example, component C ij Droplet (1203) containing (component 1 of layer 1) and component C 2,1 The droplet (1204) containing (component 1 of layer 2) is a mixed droplet (1205 C i J C 2,1 ) component. The mixed droplet can subsequently be divided into two droplets (1206) having a similar mixed composition. Third layer (C 3,1 1207 and C 3,2Additional droplets having components from 1208) may be introduced into the mixed droplets (1206) to form droplets (1209 and 1210) containing components from the first three layers. This process of combining, mixing, and splitting droplets may be repeated until the components used to form the appropriate identifier are completed. In some cases, a master mix for assembling or forming the identifier may be introduced as a nucleic acid input or a separate input droplet. In the case of a product system, at least one component of each layer may be introduced into a liquid droplet so that a complete identifier can be assembled. In a multiple reaction, multiple components from one or more layers may be introduced into a given droplet. In embodiments utilizing droplet splitting, it may be advantageous to have components of different initial concentrations to facilitate a balanced concentration of each component. Due to the parallel nature of the droplets, which can be processed at different locations on the same electrode array, it may be possible to process droplets at any high rate with tens of millions or billions of droplet reaction conditions set per second.

[0106] Print-based methods can be used to place biomolecules at the same location. FIG. 13 illustrates an example of a print-based method for distributing inputs. Inputs (e.g., nucleic acids, components, templates, enzymes, or reaction mixes) can be gathered in a stop reaction area by distributing or printing them directly to these areas. The reaction area may be a separate location on the substrate (1301). Component inputs (1306) may be assembled as identifiers in individual areas. The surface may be patterned by chemical modification to create various hydrophobic regions. Various hydrophobic regions may be useful for inhibiting input migration from one region to an adjacent region. The regions may have dimensions of about 0.1 micrometers (μm), 0.5 pm, 1 μm, 2 μm, 4 μm, 6 μm, 8 μm, 10 pm, 20 pm, 40 μm, 60 pm, 80 pm, 100 pm, or larger. The region may have dimensions of approximately 100 pm, 80 pm, 60 pm, 40 pm, 20 pm, 10 pm, 8 pm, 6 μm, 4 μm, 2 μm, 1 μm, 0.5 pm, 0.1 pm or less. The reaction region may be separated by a physical barrier such as a wall. The wall may be lithographically formed on a flat surface to create micro-wells. Alternatively, or additionally, micro-wells may be molded or embossed onto a plastic substrate. The micro-well volume may be approximately 0.1 picoliters (pL), 1 pL, 10 pL, 100 pL, 1 nanoliter (nL), or 10 nL or more. The micro-well volume may be approximately 10 nL, 1 nL, 100 pL, 10 pL, 1 pL, 0.1 pL, or less. The substrate may include glass, paper, or a plastic film. The substrate can be selectively patterned using one or more methods such as hydrophobic, embossed wells, etched wells, molded features, and deposited features.In a reel-to-reel system (1302), a roller may be used to directly pattern the indentations within the substrate before dispensing. The substrate may be translated under a stationary print head, or optionally, the print head may be translated over the surface of the substrate. Dispensing may utilize various commercially available printing methods. The print head may include 1, 10, 100, 1,000, 10,000, or more nozzles. Each nozzle of the print head may dispense the same input, or one or more nozzles may dispense distinct inputs. In some embodiments, a sufficient number of print heads are used so that a given nozzle dispenses a single input. For example, if each print head dispenses 4 inputs, a set of 50 print heads may dispense 200 inputs. Such an array having print heads aligned for dispensing into a swath can be optionally combined with reel-to-reel operation of a substrate passing under all print heads to distribute all input areas to all response areas. Each nozzle of the print head can dispense at dispensing rates of 10, 100, 1,000, 20,000, 50,000, or more than 100,000 per second. Each nozzle can be configured to operate in parallel, so a print head with 1,000 nozzles operating at 50,000 dispensing times per second can dispense up to 50 million times per second. The print driver can allow higher and lower frequencies and drop-on-demand operation, any of which can be used for input dispensing. Such systems include, but are not limited to, inkjet, bubble jet, and piezoelectric arrays. In some cases, electrostatic charges and electric fields are used to direct and control the placement of droplets. In other cases, electrostatically neutral droplets are dispensed.

[0107] In an operation similar to a print head, laser forward transmission is an optical technique that selectively delivers material containing an input (1303) from one substrate (1304) to a receiving surface (1305). The precise position of the laser pulse selectively controls the delivery of the material. By controlling the laser focus, pulse width, power, and position, the amount of material delivered can be controlled to pattern the delivery of a given input to the substrate. Sequentially delivering each input provides a robust mechanism and a time-efficient method for preparing reaction collection. In some embodiments, an optically detectable marker, such as a fluorescent or absorbent dye, may be introduced into the input fluid to enhance imaging-based inspection to verify that the input is distributed into the reaction as intended.

[0108] (1) re-code the string into a uniformly weighted form where all adjacent (i.e., adjacent and separated) stretches of 250 bits are exactly 75 bits with a value of '1', (2) use an exemplary encoding method to encode the re-encoded bit stream into an identifier library (excluding identifiers from the library corresponding to bit values ​​of '0'), and (3) use a product technique to configure the components to have identifiers separated into 8 layers, 1.0x10 12 Bit strings can be encoded and written. In this exemplary protocol, a codeword containing a subset of exactly 75 identifiers from each sequential set of 250 possible identifiers can be used to encode sequential words of length 216 bits from the original information string. One terabit (1 x 10⁻¹⁰) 12 When using the 250-select-75 uniform encoding approach to represent a 2¹⁶ bit word in a string, at least (2⁵ / 2¹⁶) * 1.0 x 10 12 = 1.15x10 12A combination space of unique identifiers can be used. In this example, seven layers with 20 components each and an eighth layer with 1000 components are used. The available identifiers in this example are 1000 * 20 7 = 1.28x10 12 and the minimum required number is 1.15x10 12 It exceeds 1.0x10 12 It may be sufficient to uniquely represent a bit. A multiplexed assembly response can be constructed by distributing one component from each of the first seven layers and distributing 75 * 4 = 300 components from the eighth layer to each response, assembling components representing four codewords into a single multiplexed response volume. The seven components from the first seven layers are assembled into the 300 components of the eighth layer, resulting in the original 1.0 x 10 12 Generates 300 unique identifiers representing a unique 4 * 2¹⁶ = 864-bit portion of the bit stream. Total 1.0 x 10 12 The identifier library representing bit strings is 1.0 x 10, where each reaction has one component in each of the first 7 layers and 300 components in the 8th layer (or a total of 307 components across all layers). 12 / 864 = 1.16e9 can be assembled using the reaction. Using a 100-micron separation between the reactions, in this example, it is approximately 12.8 metric squares (m 2 The area of ​​) can be covered by the reaction. If 160 nozzles per component are used on a single print head operating with 5,000 dispensers per second, all 1.16 x 10 9Reactions can be processed within 30 minutes. An assembly with 10 print heads dispensing 4 components using 160 nozzles per component and operating at 5,000 dispensing cycles per second can dispense all 1,140 components in approximately 12.6 hours of continuous dispensing work. 9 It can be distributed to the reaction.

[0109] Microfluidic injection can be used to place biomolecules at the same location. FIG. 14 illustrates an example of microfluidic injection of input. A microfluidic device can be manufactured by any method, such as injection molding or embossing a plastic substrate, etching a glass channel, or cross-linking a polymer. Fluid is introduced into the microfluidic device through a port and can be driven by a method such as an electroosmotic flow, external pressure or vacuum, or a positive displacement pump. In one embodiment, a stream of master mix (1401) is introduced into a stream of carrier oil (1402), and droplets of master mix (1403) form an oil stream. In some specific embodiments, the master mix droplet may be 1 nL or more, and in other embodiments, less than 100 pL, less than 50 pL, less than 10 pL, less than 5 pL, or less than 1 pL. The master mix droplet may come into contact with the channel wall, or the carrier oil layer may separate the liquid wall from the channel wall. The carrier oil may be any oil, such as hydrocarbon, fluorocarbon, silicon, or mineral oil, or any combination of oils. For example, the oil is a fluorocarbon oil. In some specific examples, the oil may additionally contain a surfactant or other additive. The master mix may contain an aqueous fluid. Inputs are introduced into the microfluidic device through a port intersecting the main channel (1404) and a plurality of input streams (1405). Inputs (e.g., nucleic acids such as components or templates, enzymes or reagents) may be optionally added to droplets as they pass through one or more injection orifices. Injection may be controlled by the selective application of an electric field by applying voltage to an electrode (1406) located near the main channel. The electrode may be separated from the channel by an insulating layer. In one embodiment, all possible distinct identifier-generating reaction droplets may be generated, and a targeted subgroup of droplets may be collected using a sorting branch in the channel.Sorting can be achieved by any method, including but not limited to using an electric field gradient, laser pulse, gas bubble, piezoelectric actuator, external valve, acoustic wave, or any other volumetric method. In another embodiment, droplets containing a target identifier generating reaction are generated. The reaction may be completed when the microfluidic device in which they are made is turned on or off. Droplets may be collected in a reaction reservoir (1407) when the microfluidic device is turned on or off.

[0110] Each identifier can be configured into a product scheme by assembling components and combining at least one component from each layer introduced into the same droplet. Multiple identifiers can be assembled within the droplet by introducing at least two components from one or more layers. Each pico-injector includes a method of applying a component stream (1405) and an external electric field (1406). The components are enzymatically assembled into identifiers. In some embodiments, the component fluid (1405) further comprises an enzyme or master mixture. As an example, one may refer to a microfluidic device comprising 10 sets of 10 pico-injectors configured such that any combination of components from 10 layers, each consisting of 10 components, can be introduced into a flowing droplet using a set of 100 pico-injectors. This example system can generate 1010 unique identifiers configured into product schemes. To enable NxM pico-injectors to construct NM identifiers, N pico-injectors (e.g., component inputs) at each layer can be easily generalized to M layers. More generally, if a layer is designated as a multiplex layer having xN pico-injectors, the construction of xN identifiers can be multiplexed in each droplet. The advantage of having more components in one layer than in another is that the layer can be used as a multilayer for combining multiple identifiers in the same droplet, thereby reducing the total number of droplets requiring write information. Each droplet receives one component from each layer, except for multilayers capable of receiving all components. xN identifiers are constructed for each droplet.

[0111] There may be flexibility in how components can be divided into layers to assemble product design and identifiers. For example, the input of a given set of 200 pico-injectors can be divided into 11 component layers, 10 layers with 10 components each (pico-injectors for distribution), and a multiplex layer with 100 components. The combination space of identifiers is 10 10 x 100 = 10 12 It can have a size of . Alternatively, using the same 200 pico-injectors, it can be divided into 40 layers of 4 components and multi-layers of 40 components. The combination space size is 4 40 x40 = 4.8x10 25 It can be. Generally, more layers generate longer DNA identifiers.

[0112] In an exemplary droplet microfluidic system, identifiers are assembled from 12 layers of 16 components in a product schema. In this example, the microfluidic device is configured to have 16 pico-injectors for each layer (16 x 12 = 192 pico-injectors). Then, 1612 = 2.8 x 1024 unique identifiers can be combined. An alternative organization of 10 layers with 11 layers and 100 layers with 100 layers (11 x 10 + 100 = 210 pico-injectors) creates a combination space of 1011 x 100 = 1013 unique identifiers. Words of length 64 bits can be encoded from the original compressed bitstream using uniformly weighted encoding into codewords containing a subset of 18 identifiers from all blocks of 100 identifiers. To represent the original 1.0e12 bit string, 1.56 x 10 10Droplets can be used. A 1.0e12-bit string can be written to DNA within 24 hours at a rate of 180,845 drops / second or 1,809 drops / second for 100 parallel devices. With an initial droplet volume of 100 pL and an additional 10 pL from each pico-injector used, 100 pL + 100 pL (first 10 layers) + 180 pL (multiplex layer) = 380 pL per drop. 380 x 10 -12 x 1.5 x 10 10 Droplet = 5.7L of total droplet volume used. After enzymatic assembly of identifiers within the droplets, the contents of each droplet can be combined and concentrated or freeze-dried for storage.

[0113] Selective condensation of component mist can be used to place biomolecules in the same location. FIG. 15 illustrates examples of selective condensation of component mist for the colocation of biomolecules. A mist nozzle (1501) can generate a mist or cloud of micron or sub-micron sized droplets (1502). The droplets may contain one or more inputs (e.g., nucleic acid sequences such as components or templates, enzymes or reagents). The mist cloud can be generated using a vibrating membrane, electrospray, atomizer, or other methods. The mist can be directed toward a thin-film transistor array (1503). The thin-film transistor array can selectively condense mist droplets by using individual electrodes (1504) to condense mist or electrode pairs (1505), such as an in-plane-switching configuration, thereby selectively condensing mist droplets within a specific region of the transistor array. Inputs can be introduced onto the array (1503) one at a time or in groups of multiple inputs. The array can be dried when inputs are introduced sequentially. After the input is sent to the array, a master mix can be introduced to all reaction points in the array. Identifiers can be constructed.

[0114] Other methods may be used to generate a selection library of identifiers such as slip technology, microfluidic devices with elastomer valves, and contact stamping. Slip technology may include parallel input streams for introducing components into multiple chambers or partitions in parallel. The chambers may slide to allow access to different compartments. For example, components may be introduced into the chambers through elastomer valves. In another example, microfluidic channels may be locations along the circumference where the barrels are arranged in a tandem manner so that the channels of each barrel can be used to add a single layer of components. The barrels may be rotated relative to each other by an increment of one channel diameter.

[0115] Various methods may be used to generate all possible identifiers from combination methods. FIG. 16 schematically illustrates an exemplary method of generating identifiers by weaving or braiding. A flexible material may be coated with specific components in specific areas. The material may be plastic, metal, screw, or natural material. The flexible material may be woven, braided, twisted, or entangled to place the parts to be assembled. Segments of the components may come together at braid or weave intersections and may be separated into reaction volumes. Once all identifiers are constructed, all subsets of identifiers, including subsets that do not match the bit stream to be encoded, may be deleted. A collection of methods for encoding information by deleting identifiers from a constructed set of identifiers or an identifier generation reaction set, or by deleting arranged components to be combined into identifiers, is called a collection of subtractive construction methods. In one embodiment, the components may be placed on threads or films. Items (1601-1604) illustrate an example where four threads or films are displayed as a specific component pattern. For example, the length of the thread or film indicated by 1601 is such that Region 0 is loaded from Layer 0 to Component 0 as illustrated by Label 1611, and Region 1 is loaded from Layer 0 to Component 1 as illustrated by Label 1612. Region 0 includes Component 0 of Layer 1 (indicated by 1609), Region 1 having Component 1 of Layer 1 (indicated by 1610), Component 1 of Layer 1 (indicated by 1610), Region 3 having Component 2 of Layer 1, and Region 3 having Component 1 of Layer 1. Generally, the film or thread or fiber corresponding to the i-th layer containing the Ni component is divided into Ni-1 * Ni regions, each region being one of the Ni components in the i-th layer, and the list of Ni components is repeatedly cycled.This method of configuring component regions on a substrate and loading them as components is called Combinatorial Marking. Components may also be configured on films and threads using different patterns, orders, and schemes. In one embodiment, each thread, film, or fiber may be loaded as a single component. A set of such single-component threads, fibers, or films may be woven into a grid as illustrated in 1613 and 1614. In this example, each intersection between horizontal and vertical threads places two components together as illustrated in 1615. In another embodiment, multiple threads may be arranged so that multiple threads intersect at a single location, thereby placing multiple components together. These intersections may be used to form an identifier, or the set of placed components may be extracted from these sites to combine identifiers at other locations. In one embodiment, each thread may have a specific pattern of regions and components as described above. Regions of this braided network may place all components used to form an identifier together, as depicted in 1616. These regions of the twisted network can be used as reaction sites or can form a network as described in 1616. A set of components placed together in these regions can be extracted from these sites and used to assemble identifiers at other locations. In another embodiment, a Product Scheme can be established in which the number of components of each layer Ni is relatively small compared to the number of components of all other layers. That is, any pair N. i and N j When represents the number of components in layers i and j, i is not equal to j. N i is N jDivide. An example is shown in 1618 in which two threads or films or fibers are represented as Thread 0, containing two components labeled 5 and 6, and Thread 1, containing five components labeled 7, 8, 9, A, and B. Layers 2 and 5 are relatively prime numbers, and 2 cannot divide 5, and vice versa. Components are loaded onto the threads and repeated in a cyclic order. Thus, Thread 0 has a repeating sequence of two components 5, 6, 5, 6, etc., as illustrated, and Thread 1 displays five components 7, 8, 9, A, B, 7, 8, 9, A, B, etc. In one embodiment, these threads may be twisted, tangled, or placed together in such a way that each area loaded with components on one thread can be aligned with a corresponding area loaded with other components on another thread. Because the number of components in each thread is relatively prime, all possible combinations of components are generated at the tightened or twisted sites. Components placed in these locations can be used as reactive locations to form identifiers for these components or to extract components placed together to form identifiers in other locations. In other embodiments, a similar scheme having a relatively small number of components can be used to create a braided network of threads. Horizontal braided threads are described in 1621. Horizontal threads can be repeated as many times as the number of components of vertical threads.

[0116] FIG. 17 schematically illustrates another method of generating an identifier from a set of components. The components are initially stored in a separate reservoir as shown in FIG. 1723. Assembly reagents and other tools may also be stored in the reservoir. The components may be placed within a set of reaction compartments, an example of which is shown in FIG. 1724. Using a transport scheme such as printing or fluid manipulation, combinations of each component are placed within individual compartments as shown in FIG. 1726. These compartments are now used as locations for assembling the identifier using multiple biochemical processes.

[0117] FIG. 18 schematically illustrates an exemplary method for generating identifiers from separate films or threads. 1832 shows a device called a Collocator, which takes as input a rolled set of threads, films, fibers, or substrates that may be marked using a combined marking technique or other marking techniques, and collects components from each corresponding area attached to each thread, film, or substrate. The collected components are arranged on an output film, thread, or fiber, indicated by 1833. As each area of ​​each thread or fiber passes through the Collocator, a new combination of components may be generated in a new area of ​​the output film or thread. Item 1835 shows a schematic diagram of the arranged components that may be used as reaction sites for combining identifiers. Item 1836 details one embodiment of the Collocator. Item (1837) illustrates one embodiment of a method for collecting components. In this example, the Collocator punctures through the passing fiber, thread, or film and collects the punctured pieces or fragments to output to an output substrate. In another embodiment, the Collocator may scrape, suck, or use other mechanical or electrical, optical or magnetic or brazing or weaving or pinch or stamping mechanisms to arrange all components from all films or threads to the output film or thread or substrate.

[0118] A subtractive recording method may be a method by which a given digital message is encoded by deleting identifiers from a pre-configured identifier library or an established library of identifier generation reactions, or by deleting arranged components prepared to be assembled into identifiers. In one embodiment, this library contains all possible identifiers in combinatorial space. The subtractive method may be advantageous because it can eliminate the complexity of configuring a specific set of identifiers when necessary. Rather, the configuration of identifiers may be independent of the specific digital message to be encoded and may be performed prior to any encoding request. Additionally, the encoding process may require simpler deletion operations at the time of writing rather than biochemical assembly or the configuration of identifiers. In one embodiment, the subtractive recording method requires a method for generating all possible identifiers. In one embodiment, when encoding is used with a product scheme, all possible identifiers may be generated by pre-loading a simple sequence of components for each layer and then combining the pre-loaded streams of components. The pre-loaded component sequences may allow all possible component combinations to be generated when the component streams are combined. This can be achieved using printing, threading, braiding, weaving, twinning, pinching, stamping, and other methods.

[0119] FIG. 19 illustrates an exemplary method of recording information using subtraction. Subtraction-targeted identifiers can be enzymatically removed (e.g., using a CRISPR / Cas system), or by cleavage, optical, thermal, electronic, static or electrical discharge, or other charged particle beams, sorting, liquid jets, acoustics, mechanical scrap, or hole punching methods. In certain embodiments where components are arranged to form identifiers but have not yet reacted, the components at each location can be assembled after subtracting the unnecessary identifier generation reaction setup. Item 1927 shows a tree diagram for a given bit stream to be encoded using a Product scheme comprising four binary layers. In this example, the combination space consists of 16 individual identifiers. All 16 identifiers can first be arranged into individual compartments as illustrated in 1925. Then, according to the considerations schematically described in FIG. 9, these identifiers can be mapped to individual symbols of the information to be encoded, bit values ​​in this example. Once the correspondence between the bits and the identifiers is established, each compartment containing a set of components used to construct the identifier can be mapped to a bit value of the bit stream. For each compartment mapped to a bit with a value of "0", the components of that compartment may be destroyed, deleted, or otherwise manipulated so that the identifier is not assembled in that compartment (Item 1930). For each compartment mapped to a bit with a value of "1", the components of that compartment are provided with all the reagents used to assemble the identifier and are not deleted or destroyed (Item 1931). In another embodiment, all identifiers are assembled, and the identifier corresponding to the bit value of "0" is deleted or destroyed after assembly. Finally, all remaining identifiers are pooled together to encode and store the given bit stream in a compressed format.

[0120] The system may include a unit for reading the generated identifier library. In one example, decoding nucleic acid-coded data may be achieved by determining the base sequence of nucleic acid strands, such as Illumina® Sequencing, or by using sequencing techniques that indicate the presence or absence of specific nucleic acid sequences, such as fragmentation analysis by capillary electrophoresis. Sequencing may employ the use of reversible terminators. Sequencing analysis may utilize natural or non-natural (e.g., engineered) nucleotides or nucleotide analogs. Optionally or additionally, decoding of nucleic acid sequences may be performed using various analytical techniques, including but not limited to any method of generating optical, electrochemical, or chemical signals. Various sequencing approaches may be used, including polymerase chain reaction (PCR), digital PCR, Sanger sequencing, high-throughput sequencing, synthetic sequencing, single molecule sequencing, ligation sequencing, etc. Various sequencing techniques may be used, including but not limited to RNA-Seq (Illumina), next-generation sequencing, digital gene expression (Helicos), clonal single microarray (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, or large-scale parallel sequencing.

[0121] Various reading methods can be used to extract information from encoded nucleic acids. For example, microarrays (or any type of fluorescent hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can be used to read additional encoded sequences and extended digitally encoded data. A subset of data (e.g., data belonging to a specific barcode) can be accessed from the pool by PCR using one primer that binds to the 5' barcode in the forward direction and one primer that binds to the common 3' sequence in the reverse direction.

[0122] The accessed data may be read from the same device or the accessed data may be transmitted to another device. The reading device may include a detection unit for detecting and identifying identifiers. The detection unit may be part of a sequencer, a hybridization array, or another unit for determining the presence or absence of identifiers. The sequencing platform may be specifically designed to decode and read information encoded as a nucleic acid sequence. The sequencing platform may be used to sequence single or double-stranded nucleic acid molecules. The sequencing platform may decode nucleic acid-encoded data by reading individual bases (e.g., base sequencing) or by detecting the presence or absence of the entire nucleic acid sequence incorporated within the nucleic acid molecule. Alternatively, the sequencing platform may be a system such as Illumina® Sequencing or fragmentation analysis by capillary electrophoresis. Decoding of nucleic acid sequences may be performed using various analytical techniques implemented by the device, including but not limited to any method of generating optical, electrochemical, or chemical signals.

[0123] Identifying an identifier in an identifier library can be performed using any identification or sequencing method. FIG. 20 illustrates an exemplary method for reading information encoded by hybridization. A reading unit may include one or more hybridization arrays. A hybridization array may include an identifier (2001) coupled to a surface or support (2002). The identifier may be spatially oriented to enable single-molecule resolution or resolution of a group of molecules using photodetection. A probe sequence (2003) sharing complementarity with one or more identifier components may be introduced into the array. The probe sequence may include one or more fluorescent materials (2004). In one example, the probe includes a fluorescent material and a quencher (2005). The quencher may be another dye or fluorescent material or a quenching device. Hybridization of the probe to the identifier may separate the fluorescent group and the quencher to generate a detectable signal. In another specific example, the probe comprises a string of fluorescent material that can be detected as an optical property indicating a specific probe or a specific set of probes. Individual components can be detected by scanning of the area, such as optical imaging of the area or confocal techniques. Sequential introduction of the probe, imaging of the probe, and removal of the probe can be used to identify some or all components on a given identifier. There is no limit to the number of components that can be identified at one time. Probes for different components may have different optical properties or the same optical properties.

[0124] Other methods for detecting identifier sequences may include nanopore sequencing. FIG. 21 illustrates an exemplary method of reading by nanopore sequencing. Pores may have unique impedance characteristics when moving through pores or channels, and voltage is applied across the pores or channels. Various existing nucleic acid sequencing platforms use this characteristic to determine the sequence of base pairs in nucleic acid molecules. These platforms have the advantage of being able to sequence longer nucleic acid molecules and detect the presence or absence of chemical residues as well as non-natural nucleotides that can be used to decorate natural and non-natural nucleotides. In one example, an identifier sequence (2103) is combined with probes (2104) that hybridize with the components of the identifier sequence. The probes may include molecules that generate a unique impedance signal while moving through the pores (2101). The pores or channels may be microfabricated to a nanometer scale on a substrate (2102) which may include a biological membrane or a crystalline material. Alternatively, or additionally, each component within each layer may include a unique molecule that generates a unique impedance signature. The unique molecule may include a sequence-based nucleotide / protein / hybrid tag, a chemical modification of a nucleotide, a fluorescent probe, or any combination thereof. In some embodiments, the signal may be a current permeating a pore or channel in other embodiments, and the detectable signal is detected by an impedance detector adjacent to the pore or channel. A burst of the signal (2105) provides a signature representing an individual identifier.

[0125] A system for encoding, recording, and reading data stored in nucleic acid molecules may or may not be automated. The system may be connected to a network to allow cloud-based access to data, or the system may not be connected to a network. The system may operate in zero or low-pressure environments, or in high / low atmospheric pressure or vacuum conditions. The system may be shielded from electromagnetic waves and other radiation to prevent the degradation of identifiers, as well as other internal electronic devices, chemicals, and enzymes. The system may use an external power source or include a power source. The system may include a power generation method. One or more units of the system may be modular and may be mobile devices. Modules or portable devices may be installed or embedded in third-party vehicles. One or more units or modules of the system may interact physically or digitally with external machines. For example, the system may take physical or digital input from an external machine, or the system may output physical materials or digital information to an external machine.

[0126] Information storage in nucleic acid molecules can have various applications, including, but not limited to, the storage of long-term information, sensitive information, and medical information. For example, a person's medical information (e.g., medical history and records) can be stored in nucleic acid molecules and delivered to the individual. The information can be stored outside the body (e.g., in a wearable device) or inside the body (e.g., within a subcutaneous capsule). When transporting a patient to a clinic or hospital, a sample can be retrieved from the device or capsule, and the information can be decoded using a nucleic acid sequencer. Privately storing medical records in nucleic acid molecules can serve as an alternative to computer and cloud-based storage systems. Private storage of medical records in nucleic acid molecules may reduce the frequency of hacked medical records. Nucleic acid molecules used for capsule-based medical record archiving can be derived from human genome sequences. The use of human genome sequences can reduce the immunogenicity of the nucleic acid sequence in the event of capsule failure or leakage.

[0127] The combinatorial assembly method described herein can be used to generate a DNA library encoding amino acid chains. The amino acid chains may be peptides or proteins. The DNA components may form junctions along codons that are functionally or structurally inactive and may be common to all members of the combinatorial library. The DNA components may form junctions along introns so that the processed peptides or proteins do not have scars between variable amino acid chains. Each combinatorial DNA molecule may be assembled in a separate reaction chamber. In vivo expression assays may be performed to detect expression. Each combinatorial DNA molecule may be collected together, and individual in vitro expression assays may be performed by encapsulating the molecules into droplets. In vivo expression assays may be performed by transforming the molecules into cells. The DNA may act as a barcode to identify cells and droplets containing specific amino acid chain variants. Since the assays may have a fluorescent output, cells / droplets can be sorted by fluorescence intensity and sequenced to correlate each combinatorial DNA sequence with a specific output. The combinatorial DNA molecules may encode RNA. If the output itself is rich in RNA (e.g., RNA after-screening and testing), pool testing can be performed in droplets or outside the cell. Combination DNA can encode combinations of CRISPR gRNAs or microRNAs that upregulate or downregulate genes inside the cell. Combination DNA libraries can be modified into cells to test how combination gene regulation affects cellular properties during cellular perturbations. Combination DNA libraries can encode combinations of genes of a pathway. Each DNA component may contain a gene expression construct, and DNA components may form junctions along inactive DNA sequences between genes.DNA sequences can be modified into cells, and various combinations of gene overexpression can be investigated to see how they affect cell properties during different cell perturbations.

[0128] Computer control system

[0129] The present disclosure provides a computer system programmed to implement the method of the present disclosure. FIG. 22 illustrates a computer system (2201) programmed or otherwise configured to encode digital information into a nucleic acid sequence and / or to read (e.g., decode) information derived from the nucleic acid sequence. The computer system (2201) may control various aspects of the encoding and decoding procedure of the present disclosure, such as bit value and bit position information for a given bit or byte from an encoded bit stream or byte stream, for example.

[0130] The computer system (2201) includes a central processing unit (CPU, where "processor" and "computer processor") (2205) which may be a single-core or multi-core processor, or a plurality of processors for parallel processing. The computer system (2201) may also include memory or memory location (2210) for communication (e.g., random access memory, read-only memory, flash memory), electronic storage unit (2215) (e.g., hard disk), one or more other systems for communication interfaces, and peripheral devices (2225) such as a cache, other memory, data storage device, and / or electronic display adapter. The memory (2210), storage unit (2215), interface (2220), and peripheral devices (2225) communicate with the CPU (2205) via a communication bus (solid line), such as a motherboard. The storage unit (2215) may be a data storage unit (or data storage) for storing data. A computer system (2201) can be operably connected to a computer network ("network") (2230) with the help of a communication interface (2220). The network (2230) is communicating with the Internet, the Internet and / or extranet, or an intranet and / or Internet. In some cases, the network (2230) is a remote communication and / or data network. The network (2230) may include one or more computer servers capable of enabling distributed computing, such as cloud computing. In some cases, with the help of the computer system (2201), the network (2230) may implement a peer-to-peer network that enables a device coupled to the computer system (2201) to operate as a client or a server.

[0131] The CPU (2205) may execute a series of machine-readable instructions that may be implemented as a program or software. The instructions may be stored in a memory location such as memory (2210). The instructions may be instructed to the CPU (2205), and the CPU (2205) may subsequently be programmed or configured to implement the method of the present disclosure. Examples of operations performed by the CPU (2205) may include patching, decoding, executing, and writing back.

[0132] The CPU (2205) may be part of a circuit such as an integrated circuit. One or more other components of the system (2201) may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0133] The storage unit (2215) can store files such as drivers, libraries, and stored programs. The storage unit (2215) can store user data such as user preferences and user programs. The computer system (2201) may include one or more additional data storage units located outside the computer system (2201), such as being located on a remote server that communicates with the computer system (2201) via an intranet or the internet.

[0134] The computer system (2201) may communicate with one or more remote computer systems through a network (2230). For example, the computer system (2201) may communicate with a remote computer system of a user or other device and / or machine, and is provided to the user in the process of analyzing data encoded or decoded in (e.g., a sequencer or other system for chemically determining the order of nitrogenous bases of a nucleic acid sequence). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad, Samsung® Galaxy Tab), a telephone, a smartphone (e.g., Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal information terminal. The user may access the computer system (2201) through the network (2230).

[0135] The method described herein may be implemented by machine-executable code (e.g., computer processor) stored in an electronic storage location of a computer system (2201), such as memory (2210) or an electronic storage unit (2215). Machine-executable or machine-readable code may be provided in the form of software. In some cases, the code may be retrieved from the storage unit (2215) and stored in memory (2210) for readiness to be accessed by the processor (2205). In some situations, the electronic storage unit (2215) may be excluded, and machine-executable instructions are stored in memory (2210).

[0136] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or it can be compiled at runtime. The code can be provided in a programming language that allows the user to choose whether to execute the code in a pre-compiled or compiled manner.

[0137] Aspects of the systems and methods provided herein, such as the computer system (2201), may be embodied in programming. Various aspects of the technology may be conceived as “products” or “manufactured items” in the form of machine (or processor) executable code and / or related data, typically performed or implemented on a type of machine-readable medium. Machine-executable code may be stored on electronic storage units such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. The medium of the “storage” type may include memory or related modules of the type such as computers, processors, etc., e.g., various semiconductor memory, tape drives, disk drives, etc., which are non-temporary storage devices capable of software programming at any time. All or part of the software may be transmitted from time to time via the Internet or various other communication networks. For example, such communication may load software from one computer or processor to another computer or processor, ranging from a management server or host computer to a computer platform of an application server. Accordingly, another type of medium capable of possessing software elements includes optical, electric, and electromagnetic waves, such as those used through physical interfaces between local devices via wired and optical wired networks and various wireless links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., may also be considered as media loaded with software. As used herein, terms such as "readable medium" of a computer or machine refer to any medium that participates in providing instructions to a processor for execution, unless limited to non-temporary, tangible "storage media."

[0138] Accordingly, machine-readable media, such as computer-executable code, may take many forms, including but not limited to physical storage media, carrier media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices of any computer(s) that may be used to implement, for example, the database illustrated in the drawing. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Types of transmission media include coaxial cables; optical fibers, including copper wires and wires that constitute a bus within a computer system. Carrier transmission media may take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communication. Accordingly, common forms of computer-readable media may be floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch card paper tapes, any other physical storage media having a hole pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, data or instructions carrying such a carrier, carrier waves carrying cables or links, or any other media in which a computer can read programming code and / or data. Such forms of computer-readable media may be associated with delivering one or more sequences of one or more instructions to a processor for execution.

[0139] A computer system (2201) is encoded or read by a machine or computer system that encodes or reads raw data, files, and compressed or decompressed zip files, and is encoded or decoded into DNA storage data, including, for example, a chromatograph, sequence and sequence output data including a user interface (UI) (2240) for providing sequence output data including bits, bytes, or bit streams. Examples of the UI include, but are not limited to, a graphical user interface (GUI) and a web-based user interface.

[0140] The method and system of the present disclosure may be implemented by one or more algorithms. The algorithm may be implemented by software when executed by a central processing unit (2205). The algorithm may be used, for example, with a DNA index and raw data or ZIP file compressed or decompressed data to determine a customized method for digital coding. Information on raw data or ZIP file compressed data may be obtained before encoding digital information.

[0141] Although preferred embodiments of the present invention have been illustrated and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. The present invention is not limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing specification, the descriptions and examples of embodiments in this specification are not to be interpreted in a limiting sense. Various modifications, changes, and substitutions will be made to those skilled in the art without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions described herein under various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used to practice the present invention. Accordingly, the present invention should include any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention and are intended to be covered by methods, structures, and equivalents within the scope of the claims.