Fixed-point number representation and calculation circuits

KR103003733B1Active Publication Date: 2026-08-12CATALOG TECHNOLOGIES INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2026-08-12

Smart Images

  • Figure 112023116830525-PCT00097_ABST
    Figure 112023116830525-PCT00097_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method for storing digital information in nucleic acid molecules in various ways. The digital information may be received as a symbol string, and each symbol in the symbol string has a symbol value and a symbol position within the symbol string. A first identifier nucleic acid molecule may be formed by storing M selected component nucleic acid molecules in a single compartment—the M selected component nucleic acid molecules are selected from a set of individual component nucleic acid molecules separated into M different layers—and physically assembling the M selected component nucleic acid molecules. Multiple identifier nucleic acid molecules may be formed, each corresponding to a respective symbol position. The identifier nucleic acid molecules may be formed in a powder, liquid, or solid form.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims priority and interest to U.S. Provisional Application No. 63 / 165,507, filed on March 24, 2021, titled "FIXED POINT NUMBER REPRESENTATION AND COMPUTATION CIRCUITS". The entire contents of the aforementioned application are incorporated herein by reference. Background Technology

[0003] Nucleic acid digital data storage is a reliable approach for encoding and storing information for long periods by storing data at a higher density than magnetic tape or hard drive storage systems. Additionally, digital data stored in nucleic acid molecules under cold and dry conditions can be retrieved for over 60,000 years.

[0004] Nucleic acid molecules can be sequenced to access the digital data stored within them. As such, nucleic acid digital data storage can be an ideal method for storing large amounts of information that is not frequently accessed but needs to be stored or preserved for a long period.

[0005] Current methods rely on encoding digital information (e.g., binary code) into base-by-base nucleic acid sequences, where the base-to-base relationships of the sequence are directly converted into digital information (e.g., binary code). Because the cost of de novo base-by-base nucleic acid synthesis can be high, sequencing digital data stored in base-by-base sequences that can be read as bitstreams or bytes of digitally encoded information can be error-prone and costly to encode. Opportunities for new methods to perform nucleic acid digital data storage may provide a less costly and commercially feasible approach to data encoding and retrieval. means of solving the problem

[0006] The present disclosure provides a system and method for storing digital information in nucleic acid molecules in various ways to improve the efficiency of retrieval and access of digital information. For example, component nucleic acid molecules (e.g., components) are selected and concatenated to form identifier nucleic acid molecules (e.g., identifiers), each corresponding to a specific symbol (e.g., a bit or a series of bits) or the position of that symbol (e.g., a rank or address) in a symbol string (e.g., a bitstream). These components may be configured in a structural manner to provide an efficient way of representing digital data. For example, the structure of the components may cause the component molecules to self-assemble, or alternatively, multiple component molecules may be stored in the same compartment or distributed and then self-aligned in a predetermined order.

[0007] In one embodiment, the present disclosure provides a method for recording information in a nucleic acid sequence. The method comprises the step of obtaining a first fixed-point number. The method comprises the step of obtaining a library of component nucleic acid sequences that defines a combination space of identifier nucleic acid sequences, each comprising an aligned subset of component nucleic acid sequences. The method comprises the step of identifying a first subset of identifier nucleic acid sequences in the combination space as a first codeword having a codeword size corresponding to the number of identifier nucleic acid sequences of the first subset. The method comprises the step of forming a first set of one or more identifier nucleic acid molecules having individual identifier nucleic acid sequences of the first subset—wherein the ratio of the number of individual identifier nucleic acid sequences represented in the first set to the codeword size approximates the first fixed-point number.

[0008] In some embodiments, the library of component nucleic acid sequences comprises a plurality of layers, each layer comprising a subset of component nucleic acid sequences. Each identifier nucleic acid sequence may comprise one component nucleic acid sequence from each layer.

[0009] In some embodiments, the first fixed-point number has a value x, the codeword size is w, and k identifier nucleic acid molecules are formed in the first set, so that the ratio is k / w and is approximately equal to x. In some embodiments, k / w is within ±20% of x. In some embodiments, the codeword size is at least 8. In some embodiments, the codeword size is at least 256. In some embodiments, the codeword size is at least 512. In some embodiments, the codeword size is at least 1024.

[0010] In some embodiments, the method comprises the steps of obtaining a second fixed-point number; identifying a second subset of identifier nucleic acid sequences in a combination space as a second codeword having the codeword size of a first codeword and corresponding to the number of identifier nucleic acid sequences in the second subset; and forming a second set of one or more identifier nucleic acid molecules having individual identifier nucleic acid sequences of the second subset. The ratio of the number of individual identifier nucleic acid sequences of the second set to the codeword size may approximate the second fixed-point number.

[0011] In some embodiments, the method comprises the step of pooling a first set and a second set to obtain a sum pool, and adding the first fixed-point number and the second fixed-point number by diluting the pooled set to obtain a scaled sum pool.

[0012] In some embodiments, the method comprises the step of pooling a first set and a second set to obtain a factor pool, and multiplying the first fixed-point number and the second fixed-point number by applying a chemical AND operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool.

[0013] In some embodiments, the chemical AND operation includes converting an identifier nucleic acid molecule into a single-strand identifier nucleic acid molecule, hybridizing a complementary identifier nucleic acid molecule, and selecting a fully hybridized double-strand nucleic acid molecule to obtain a product pool.

[0014] In some embodiments, the choice includes using at least one of an enzyme that selectively degrades single-stranded nucleic acid molecules or an enzyme that selectively degrades double-stranded nucleic acid molecules having sequence mismatch.

[0015] In some embodiments, the method comprises pooling a first set and a second set to obtain a factor pool, and applying a chemical OR operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool. In some embodiments, the method comprises the step of mixing the first set and the second set.

[0016] In some embodiments, the method comprises pooling a first set and a second set to obtain a factor pool, and applying a chemical NIMPLY operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool. In some embodiments, the chemical NIMPLY operation comprises converting the identifier nucleic acid molecules into single-strand identifier nucleic acid molecules—the single-strand identifier nucleic acid molecules of the second set contain an affinity tag—providing a molar excess of the single-strand identifier nucleic acid molecules of the second set, hybridizing the complementary identifier nucleic acid molecules, and obtaining the product pool using a specific capture mechanism for the affinity tag by selecting a fully hybridized double-strand nucleic acid molecule.

[0017] In some embodiments, the method comprises pooling a first set and a second set to obtain a factor pool, and applying a chemical NOT operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool. In some embodiments, the chemical NOT operation comprises converting the identifier nucleic acid molecules into single-strand identifier nucleic acid molecules—the single-strand identifier nucleic acid molecules of the first set contain an affinity tag—providing a molar excess of the single-strand identifier nucleic acid molecules of the first set, hybridizing the complementary identifier nucleic acid molecules, and obtaining the product pool using a specific capture mechanism for the affinity tag by selecting a fully hybridized double-strand nucleic acid molecule.

[0018] In some embodiments, the method comprises pooling a first set and a second set to obtain a factor pool, and applying a chemical XOR operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool. In some embodiments, the chemical XOR operation comprises performing two NIMPLY operations followed by an OR operation.

[0019] Inclusion by reference

[0020] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent that each individual publication, patent, or patent application is specifically and individually indicated for inclusion by reference. If any publication, patent, or patent application incorporated by reference contradicts the disclosures included in this specification, this specification is intended to replace or take precedence over such contradictory material. Brief explanation of the drawing

[0021] Novel features of the present invention are specifically described in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description describing exemplary embodiments in which the principles of the present invention are utilized and the accompanying drawings (also “Drawings” and “Figures”). Figure 1 schematically illustrates an overview of the process of encoding, recording, accessing, querying, reading, and decoding digital information stored in a nucleic acid sequence. FIGS. 2a and 2b schematically illustrate an exemplary method of encoding digital data referred to as "data at address" using an object or identifier (e.g., a nucleic acid molecule). FIG. 2a illustrates combining a rank object (or address object) with a byte-value object (or data object) to generate an identifier. FIG. 2b illustrates an example of data in an addressing method in which the rank object and the byte-value object themselves are a combinatorial connection of other objects. FIGS. 3a and 3b schematically illustrate exemplary methods for encoding digital information using objects or identifiers (e.g., nucleic acid sequences). FIG. 3a illustrates encoding digital information using a rank object as an identifier. FIG. 3b illustrates an example of an encoding method in which the address object itself is a combinatorial connection of other objects. Figure 4 shows a contour plot in logarithmic space of the relationship between the combination space of possible identifiers (C, x-axis) that can be configured to store information of a given size (contour) and the average number of identifiers (k, y-axis). Figure 5 schematically illustrates an overview of a method for recording information in a nucleic acid sequence (e.g., deoxyribonucleic acid). FIGS. 6a and 6b illustrate an exemplary method called a “multiplication method” for constructing an identifier (e.g., a nucleic acid molecule) by combinatorially assembling individual components (e.g., a nucleic acid sequence). FIG. 6a illustrates the architecture of an identifier constructed using the multiplication method. FIG. 6b illustrates an example of a combination space of identifiers that can be constructed using the multiplication method. FIG. 7 schematically illustrates the use of a nested extension polymerase chain reaction to construct an identifier (e.g., nucleic acid molecule) from a component (e.g., nucleic acid sequence). FIG. 8 schematically illustrates the use of adhesive end ligations to form an identifier (e.g., nucleic acid molecule) from a component (e.g., nucleic acid sequence). FIG. 9 schematically illustrates the use of recombinant enzyme assembly to construct an identifier (e.g., nucleic acid molecule) from a component (e.g., nucleic acid sequence). Figures 10a and 10b illustrate template-directed ligation. Figure 10a schematically illustrates the use of template-directed ligation to construct an identifier (e.g., nucleic acid molecule) from a component (e.g., nucleic acid sequence). Figure 10b shows a histogram of the copy number (abundance) of 256 individual nucleic acid sequences combinatorially assembled from each of six nucleic acid sequences (e.g., components) in a single pooled template-directed ligation reaction. Figures 11a–11g schematically illustrate an exemplary method referred to as a “permutation method” for constructing an identifier (e.g., a nucleic acid molecule) with permuted components (e.g., a nucleic acid sequence). Figure 11a illustrates the architecture of an identifier constructed using a permutation method. Figure 11b illustrates an example of a combination space of identifiers that can be constructed using a permutation method. Figure 11c shows an exemplary implementation of a permutation method using template-directed ligation. Figure 11d illustrates an example of how the implementation method of Figure 11c can be modified to construct an identifier with permuted and repeated components. Figure 11e shows how the exemplary implementation of Figure 11d can result in unwanted byproducts that can be eliminated by nucleic acid size selection. Figure 11f shows another example of a method using template-directed ligation and size selection to construct identifiers with permutations and repeated components. Figure 11g shows an example of a case where size selection may fail to separate a specific identifier from unwanted byproducts. Figures 12a–12d show a larger number M An arbitrary number among the possible components ofk An exemplary method referred to as the "MchooseK" method for constructing an identifier (e.g., a nucleic acid molecule) having assembled components (e.g., a nucleic acid sequence) is schematically illustrated. Fig. 12a illustrates the architecture of an identifier constructed using the MchooseK method. Fig. 12b illustrates an example of a combination space of identifiers that can be constructed using the MchooseK method. Fig. 12c shows an exemplary implementation of the MchooseK method using template instruction ligation. Fig. 12d shows how the exemplary implementation of Fig. 12c can result in unwanted byproducts that can be eliminated by nucleic acid size selection. FIGS. 13a and 13b schematically illustrate an exemplary method referred to as a "partitioning method" for configuring an identifier into partitioned components. FIG. 13a shows an example of a combination space of identifiers that can be configured using the partitioning method. FIG. 13b shows an example of an implementation of the partitioning method using template reference ligation. FIGS. 14a and 14b schematically illustrate an exemplary method referred to as the "unconstrained string" (or USS) method for constructing an identifier composed of any string of components from a plurality of possible components. FIG. 14a illustrates an example of a combination space of identifiers that can be constructed using the USS method. FIG. 14b shows an exemplary implementation of the USS method using template reference ligation. FIGS. 15a and 15b schematically illustrate an exemplary method called "component deletion" for constructing an identifier by removing components from a parent identifier. FIG. 15a illustrates an example of a combination space of identifiers that can be constructed using the component deletion method. FIG. 15b shows an exemplary implementation of the component deletion method using double-strand targeted cutting and repair. FIG. 16 schematically illustrates a parent identifier having a recombinant enzyme recognition site, into which additional identifiers can be formed by applying a recombinant enzyme to the parent identifier. Figures 17a–17c schematically illustrate an overview of exemplary methods for accessing parts of information stored in a nucleic acid sequence by accessing multiple specific identifiers among a larger number of identifiers. Figure 17a shows an exemplary method using a polymerase chain reaction, an affinity-tagged probe, and a degradation-targeting probe to access identifiers containing specific components. Figure 17b shows an example of a method using a polymerase chain reaction to perform an 'OR' or 'AND' operation to access identifiers containing multiple specific components. FIG. 17c illustrates an exemplary method of using affinity tags to perform an 'OR' or 'AND' operation to access an identifier containing a number of specified components. Figures 18a and 18b show examples of encoding, writing, and reading data encoded in nucleic acid molecules. Figure 18a shows an example of encoding, writing, and reading 5,856 bits of data. Figure 18b shows an example of encoding, writing, and reading 62,824 bits of data. FIG. 19 shows a computer system programmed or otherwise configured to implement the method provided in this specification. FIG. 20 illustrates an exemplary method of assembling any two selected double-strand components from a single set of double-strand components. Figure 21 shows a possible adhesive terminal component structure made of two oligos X and Y. Figure 22 shows an example of constructing an identifier from a component having multiple functional parts. Figures 23a-23b show examples of the effect of identifier ranking on PCR-based random access. Figures 24a-24b show exemplary effects of an identifier architecture with a non-uniform component distribution for PCR-based random access. Figure 25 illustrates the exemplary effect of layer increase in an identifier architecture for PCR-based random access. Figure 26 shows an example of a multi-empty position encoding scheme for an alphabet of 9 symbols. Figure 27 shows an example of a multi-bin identifier distribution encoding scheme with an identifier library of two identifiers and a set of three bins that enables the encoding of any of the nine possible messages of a 4-bit string. FIG. 28 illustrates an example of a multi-bin identifier distribution encoding scheme that utilizes identifier reuse with a library of two identifiers and a set of three bins to enable the encoding of any of the 64 possible messages of a 6-bit string. Figure 29 shows an example of encoding DNA information using integer partitioning. FIG. 30 shows an example of an encoding pipeline including an algorithm module for preparing and converting a source bitstream into a build program specification to be interpreted by the writer. FIG. 31 illustrates one embodiment of a data structure for representing an identifier library in a serialized format. Figure 32 shows an example of two source bitstreams and a universal identifier library prepared to compute using operations defined in the identifier pool. FIG. 33 shows an identifier library in a test tube ( in vitro It shows the inputs and results for three examples of logical operations performed on a pool of identifiers that can be used as a platform for computation. Figures 34a–34g show examples of saving image files and reading them at various resolutions. FIG. 35 illustrates an exemplary method for generating entropy that can be used to generate a random bit string. Figures 36a–36c show exemplary methods for generating and storing entropy (random bit strings). Figures 37a-37b show examples of how to construct and access random bit strings using inputs. Figure 38 shows an exemplary method for protecting and authenticating access to an artifact using a physical DNA key. FIG. 39 schematically illustrates an exemplary method for encoding data into DNA in FPN format using a product scheme and an example of operations on such data. Figure 40 schematically illustrates an exemplary mechanism for an AND gate using nuclease protection of dsDNA. Figure 41a schematically illustrates an overview of an exemplary mechanism for an OR gate using dsDNA. Figure 41b schematically illustrates an overview of an exemplary mechanism for an OR gate using ssDNA. FIG. 42a schematically illustrates an exemplary mechanism for a NIMPLY gate using an affinity tag. A molar excess of B is provided along with a biotin tag. Any identifier in A that matches B will be hybridized and removed by B. The identifier that survives in A is part of the A-NIMPLY-B return product. FIG. 42b schematically illustrates an exemplary mechanism for a NIMPLY gate using a nuclease. FIG. 43a schematically illustrates an exemplary mechanism for a NOT gate using an affinity tag. A molar excess of A is provided along with a biotin tag. Any identifier in B that matches A will be hybridized and removed by A. The identifier that survives in B is part of the NOT-A return product. FIG. 43b schematically illustrates an exemplary mechanism for a NOT gate using a nuclease. A molar excess of A is provided. Any identifier in B that matches A will be hybridized and removed by A. The identifier that survives in B is part of the NOT-A return product. FIGS. 44a-c schematically illustrates an overview of an exemplary mechanism for an XOR gate using affinity tags. In FIG. 44a, a molar excess of B is provided with a biotin tag. Any identifier in A that matches B will be hybridized and removed by B. The identifier that survives in A is part of the A-NIMPLY-B return product. In FIG. 44b, a molar excess of A is provided with a biotin tag. Any identifier in B that matches A will be hybridized and removed by A. The identifier that survives in B is part of the B-NIMPLY-A return product. FIG. 44c illustrates a final XOR step. Specific details for implementing the invention

[0022] Although various embodiments of the present invention have been illustrated and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. Those skilled in the art may make various modifications, changes, and substitutions without departing from the invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be adopted.

[0023] As used herein, the term "symbol" generally refers to a unit of digital information. Digital information may be divided or translated into a string of symbols. For example, a symbol may be a bit, and a bit may have a value of '0' or '1'.

[0024] As used herein, the terms "distinct" or "unique" generally refer to an object that can be distinguished from other objects within a group. For example, a distinct or unique nucleic acid sequence may be a nucleic acid sequence that does not have the same sequence as any other nucleic acid sequence. A distinct or unique nucleic acid molecule may not have the same sequence as any other nucleic acid molecule. A distinct or unique nucleic acid sequence or molecule may share a region of similarity with another nucleic acid sequence or molecule.

[0025] As used herein, the term "component" generally refers to a nucleic acid sequence. A component may be an individual nucleic acid sequence. A component may be connected to or assembled with one or more other components to produce another nucleic acid sequence or molecule.

[0026] As used herein, the term “layer” generally refers to a group or pool of components. Each layer may include a set of individual components such that the components of one layer differ from the components of another layer. Components from one or more layers may be assembled to generate one or more identifiers.

[0027] As used herein, the term “identifier” generally refers to a nucleic acid molecule or nucleic acid sequence representing the position and value of a bit string within a larger bit string. More generally, an identifier may represent a symbol within a string of symbols or refer to any corresponding object. In some embodiments, an identifier may include one or more connected components.

[0028] As used herein, the term “combination space” generally refers to the set of all possible individual identifiers that can be generated from an object, e.g., a starting set of components, and the set of acceptable rules for how to modify said object to form the identifiers. The size of the combination space of identifiers created by assembling or connecting components may vary depending on the number of layers of components, the number of components in each layer, and the specific assembly method used to generate the identifiers.

[0029] As used in this specification, the term "identifier rank" generally refers to a relationship that defines the order of identifiers within a set.

[0030] As used herein, the term “identifier library” generally refers to a collection of identifiers corresponding to symbols in a symbol string representing digital information. In some embodiments, the absence of a given identifier in the identifier library may indicate a symbol value at a specific location. One or more identifier libraries may be combined into a pool, group, or set of identifiers. Each identifier library may include a unique barcode that identifies the identifier library.

[0031] As used herein, the term “nucleic acid” generally refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or variants thereof. Nucleic acid may comprise one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) or variants thereof. Nucleotides may comprise A, C, G, T, or U or variants thereof. Nucleotides may comprise any subunit that can be incorporated into a growing nucleic acid strand. Such subunits may be A, C, G, T, or U, or specific to one of the more complementary A, C, G, T, or U, or any other subunit that may be complementary to a purine (i.e., A or G or variants thereof) or a pyrimidine (i.e., C, T, or U or variants thereof). In some examples, nucleic acid may be single-stranded or double-stranded, and in some cases, nucleic acid is circular.

[0032] As used herein, the terms “nucleic acid molecule” or “nucleic acid sequence” generally refer to a polymer form of a nucleotide or polynucleotide that may have various lengths, a deoxyribonucleotide (DNA) or a ribonucleotide (RNA), or an analogue thereof. The term “nucleic acid sequence” may mean an alphabetical representation of a polynucleotide, or alternatively, the term may be applied to the physical polynucleotide itself. This alphabetical representation may be entered into a database of a computer with a central processing unit and may be used to encode digital information by mapping the nucleic acid sequence or nucleic acid molecule to symbols or bits. A nucleic acid sequence or oligonucleotide may comprise one or more non-standard nucleotide(s), nucleotide analog(s), and / or modified nucleotides.

[0033] As used herein, "oligonucleotide" generally means a single-stranded nucleic acid sequence and generally consists of a specific sequence of the following four nucleotide bases: adenine (A), cytosine (C), guanine (G), and thymine (T), or uracil (U) if the polynucleotide is RNA.

[0034] Non-limiting examples of modified nucleotides include diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosin, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil. Examples include 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid(v), ybutoxosin, pseudouracil, cuosin, 2-thiocytosin, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, etc. Nucleic acid molecules may also be modified in base moiety (e.g., one or more atoms that can generally form hydrogen bonds with complementary nucleotides and / or one or more atoms that generally cannot form hydrogen bonds with complementary nucleotides), sugar moiety, or phosphate backbone. Nucleic acid molecules may also contain amine modification groups, e.g., aminoallyl-dUTP (aa-dUTP) and aminohexylacrylamide-dCTP (aha-dCTP), to allow covalent attachment of amine-reactive moiety, e.g., N-hydroxysuccinimide ester (NHS).

[0035] As used herein, the term "primer" generally refers to a strand of nucleic acid that serves as a starting point for nucleic acid synthesis, such as in a polymerase chain reaction (PCR). For example, during the replication of a DNA sample, the enzyme catalyzing replication initiates replication at the 3'-end of a primer attached to the DNA sample and replicates the opposite strand. For further details regarding PCR, including details on primer design, refer to Chemical Methods Section D.

[0036] As used herein, the terms “polymerase” or “polymerase enzyme” generally refer to any enzyme capable of catalyzing a polymerase reaction. Examples of polymerases include, but are not limited to, nucleic acid polymerases. Polymerases may occur naturally or be synthesized. An exemplary polymerase is Φ29 polymerase or a derivative thereof. In some cases, a transcriptionase or a ligase (i.e., an enzyme that catalyzes bond formation) is used in conjunction with or as an alternative to the polymerase to form a new nucleic acid sequence. Examples of polymerases include DNA polymerase, RNA polymerase, heat-stable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, platinum Taq polymerase. There are Tbr polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3' to 5' exonuclease activity, and their modifications, modified products, and derivatives. For details on additional polymerases available for PCR and how polymerase characteristics may affect PCR, refer to Chemical Methods Section D.

[0037] As used herein, the term “species” generally refers to one or more DNA molecule(s) of the same sequence. When “species” is used in the plural sense, it may be assumed that all species included in the plural have a distinct order, though this may sometimes be explicitly indicated by using “individual species” instead of “species.”

[0038] The terms "approximately" and "roughly" should be understood to mean within ±20% of the value following the above terms.

[0039] Digital information in the form of binary code, such as computer data, may include a sequence or string of symbols. Binary code may encode or represent text or computer processor instructions using a binary number system having two binary symbols, typically 0 and 1, called bits, for example. Digital information may be represented in the form of non-binary code, which may include a sequence of non-binary symbols. Each encoded symbol may be reassigned to a unique bit string (or "byte"), and unique bit strings or bytes may be arranged into a string of bytes or a byte stream. The bit value for a given bit may be one of two symbols (e.g., 0 or 1). A byte, which may consist of an N-bit string, is a total of 2 N It can have unique byte values. For example, a byte consisting of 8 bits has a total of 2 8It is possible to generate 256 or 10 possible unique byte values, and each of the 256 bytes can correspond to one of the 256 possible individual symbols, characters, or commands that can be encoded as bytes. Raw data (e.g., text files and computer commands) can be represented as a string of bytes or a stream of bytes. A zip file or compressed data file composed of raw data can also be stored as a stream of bytes, and these files can be stored as a stream of bytes in a compressed format and then decompressed into raw data before being read by a computer.

[0040] The method and system of the present disclosure may be used to encode computer data or information with a plurality of identifiers, each of which may represent one or more bits of the original information. In some examples, the method and system of the present disclosure encode data or information using identifiers each representing two bits of the original information.

[0041] Previous methods for encoding digital information into nucleic acids have relied on the base-by-base synthesis of nucleic acids, which can be costly and time-consuming. Alternative methods improve efficiency by reducing reliance on base-by-base nucleic acid synthesis for encoding digital information, enhance the commercial viability of digital information storage, and eliminate the de dovo synthesis of individual nucleic acid sequences for every new information storage request.

[0042] The new method can encode digital information (e.g., binary code) containing combinational arrangements of components instead of relying on multiple identifiers or base-by-base or de-novo nucleic acid synthesis (e.g., phosphoramidite synthesis) in nucleic acid sequences. Thus, the new strategy can generate a first set of individual nucleic acid sequences (or components) for the first request for information storage and reuse the same nucleic acid sequences (or components) for subsequent information storage requests. These approaches can significantly reduce the cost of DNA-based information storage by reducing the role of de-novo synthesis of nucleic acid sequences in the process of encoding and writing information to DNA. Furthermore, unlike base-by-base synthesis, such as phosphoramidite chemistry-based or template-free polymerase-based nucleic acid extension, which can periodically deliver each base to each extended nucleic acid, the new method of converting information to DNA is a highly parallelizable process that does not necessarily require periodic nucleic acid extension, as it uses the configuration of component identifiers. Therefore, the new method can increase the speed of writing digital information to DNA compared to existing methods.

[0043] Method of encoding and recording information in nucleic acid sequence(s)

[0044] In one aspect, the present disclosure provides a method for encoding information into a nucleic acid sequence. A method for encoding information into a nucleic acid sequence may include: (a) translating the information into a string of symbols; (b) mapping the string of symbols to a plurality of identifiers; and (c) configuring an identifier library comprising at least a subset of the plurality of identifiers. Among the plurality of identifiers, an individual identifier may include one or more components. Among the one or more components, an individual component may include a nucleic acid sequence. Each symbol at each position within the string of symbols may correspond to an individual identifier. An individual identifier may correspond to an individual symbol at each position within the string of symbols. Additionally, one symbol at each position within the string of symbols may correspond to no identifier. For example, in a string of binary symbols (e.g., bits) of '0' and '1', each occurrence of '0' may correspond to no identifier.

[0045] In another aspect, the present disclosure provides a nucleic acid-based computer data storage method. The nucleic acid-based computer data storage method may include: (a) receiving computer data; (b) synthesizing a nucleic acid molecule comprising a nucleic acid sequence encoding the computer data; and (c) storing the nucleic acid molecule having the nucleic acid sequence. The computer data may be encoded in a subset of at least the synthesized nucleic acid molecules, rather than in the sequence of each nucleic acid molecule.

[0046] In another aspect, the present disclosure provides a method for writing and storing information in a nucleic acid sequence. The method may include (a) receiving or encoding a virtual identifier library representing information, (b) physically configuring the identifier library, and (c) storing one or more physical copies of the identifier library in one or more separate locations. An individual identifier in the identifier library may include one or more components. An individual component among the one or more components may include a nucleic acid sequence.

[0047] In another aspect, the present disclosure provides a nucleic acid-based computer data storage method. The nucleic acid-based computer data storage method may include: (a) receiving computer data; (b) synthesizing a nucleic acid molecule comprising at least one nucleic acid sequence encoding the computer data; and (c) storing the nucleic acid molecule comprising at least one nucleic acid sequence. The synthesis of the nucleic acid molecule may not involve base-by-base nucleic acid synthesis.

[0048] In another aspect, the present disclosure provides a method for recording and storing information in a nucleic acid sequence. The method for recording and storing information in a nucleic acid sequence may include: (a) receiving or encoding a virtual identifier library representing information; (b) physically configuring the identifier library; and (c) storing one or more physical copies of the identifier library in one or more separate locations. An individual identifier in the identifier library may include one or more components. An individual component among the one or more components may include a nucleic acid sequence.

[0049] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position in the string of symbols; (b) forming a first identifier nucleic acid sequence by the following steps: (1) selecting one component nucleic acid sequence from each of the M layers from a set of individual component nucleic acid sequences separated into M different layers; (2) storing the M selected component nucleic acid sequences in one compartment; and (3) physically assembling the M selected component nucleic acid sequences of (2) such that the component nucleic acid sequences from the first and second layers correspond to the first and second terminal sequences of the identifier nucleic acid sequence, and the component nucleic acid sequence in the third layer corresponds to the third sequence of the identifier nucleic acid sequence, thereby defining the physical order of the M layers of the first identifier nucleic acid sequence, the first identifier nucleic acid sequence having the first and second terminal sequences and a third sequence located between the first terminal sequence and the second terminal sequence. A method comprising: (c) forming a plurality of additional identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has (1) a first and second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, and (2) corresponding to each symbol position, wherein the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), so that a probe can select at least two identifier nucleic acid sequences corresponding to each symbol having consecutive symbol positions within the string of symbols, and (d) collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form.

[0050] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position within the string of symbols, and the digital information comprises image data represented by a collection of vectors; (b) forming a first identifier nucleic acid sequence by depositing M selected component nucleic acid sequences in a single compartment, wherein the M selected component nucleic acid sequences are selected from a set of individual component nucleic acid sequences separated into M different layers; and (c) forming a plurality of identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has a first and a second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, corresponding to their respective symbol positions, and wherein the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), so that a single probe corresponds to at least two identifier nucleic acids corresponding to their respective symbols having a related symbol position within the string of symbols. A method comprising: enabling the selection of sequences - and (d) collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form - storing image data as nucleic acid sequences so that random neighbors of a pixel can be queried for color values ​​using a random access method.

[0051] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position within the string of symbols; (b) forming a first identifier nucleic acid sequence by storing M selected component nucleic acid sequences in a single compartment, wherein the M selected component nucleic acid sequences are selected from a set of individual component nucleic acid sequences separated into M different layers; (c) physically assembling a plurality of identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has a first and a second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, corresponding to their respective symbol positions, and wherein the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), so that a single probe can select at least two identifier nucleic acid sequences corresponding to their respective symbols having related symbol positions within the string of symbols; and (d) A method comprising the step of collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form.

[0052] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position within the string of symbols; (b) dividing the string of symbols into one or more blocks of a size not greater than a fixed length; (c) forming a first identifier nucleic acid sequence by storing M selected component nucleic acid sequences in one compartment, wherein the M selected component nucleic acid sequences are selected from a set of individual component nucleic acid sequences separated into M different layers; and (d) physically assembling a plurality of identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has a first and a second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, corresponding to their respective symbol positions, and wherein the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), so that a single probe [determines] the relevant symbol position within the string of symbols. A method comprising the step of selecting at least two identifier nucleic acid sequences corresponding to each of the symbols having, and (e) collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form.

[0053] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position within the string of symbols; (b) forming a first identifier nucleic acid sequence by storing M selected component nucleic acid sequences in a single compartment, wherein the M selected component nucleic acid sequences are selected from a set of individual component nucleic acid sequences separated into M different layers; (c) physically assembling a plurality of identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has a first and a second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, corresponding to their respective symbol positions, and wherein the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), thereby enabling a single probe to select at least two identifier nucleic acid sequences corresponding to their respective symbols having related symbol positions within the string of symbols; and (d) A method comprising the steps of: collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form; and (e) generating a new pool of nucleic acid molecules by performing a calculation on a string of symbols using the identifier nucleic acid sequence of (d), including a Boolean logical operation, e.g., AND, OR, NOT, or NAND.

[0054] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position in the string of symbols; (b) forming a first identifier nucleic acid sequence by the following: (1) selecting one component nucleic acid sequence from each of the M layers from a set of individual component nucleic acid sequences separated into M different layers; (2) storing the M selected component nucleic acid sequences in a single compartment; and (c) physically assembling a plurality of identifier nucleic acid sequences, wherein each additional identifier nucleic acid sequence has a first and a second terminal sequence and a third sequence located between the first terminal sequence and the second terminal sequence, corresponding to their respective symbol positions, and the first terminal sequence, the second terminal sequence, and the third sequence of at least one additional identifier nucleic acid sequence are identical to the target sequence of the first identifier nucleic acid sequence in (b), so that a single probe corresponds to at least two respective symbols having a related symbol position in the string of symbols. A method comprising the step of selecting identifier nucleic acid sequences, and (d) collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form.

[0055] In another aspect, the present disclosure provides a method for storing digital information in a nucleic acid sequence, the method comprising: (a) receiving digital information as a string of symbols, wherein each symbol in the string of symbols has a symbol value and a symbol position in the string of symbols; (b) forming a first identifier nucleic acid sequence by the following: (1) selecting one component nucleic acid sequence from each of the M layers from a set of individual component nucleic acid sequences separated into M different layers; (2) storing the M selected component nucleic acid sequences in a single compartment; (3) physically assembling the M selected component nucleic acid sequences of (2) to form a first identifier nucleic acid sequence comprising a specified component, wherein the specified component comprises at least one target sequence to enable access to an identifier containing the specified component; and (c) physically assembling a plurality of additional identifier nucleic acid sequences, each having a specified component, wherein the specified component comprises at least one target sequence of the first identifier nucleic acid sequence of (b), so that a probe corresponds to at least two symbols having consecutive symbol positions in the string of symbols. A method comprising the step of enabling the selection of an identifier nucleic acid sequence, and (d) collecting the identifier nucleic acid sequences of (b) and (c) in a pool having a powder, liquid, or solid form.

[0056] FIG. 1 illustrates an overview process of encoding information into a nucleic acid sequence, writing information to the nucleic acid sequence, reading the information written to the nucleic acid sequence, and decoding the read information. Digital information or data may be converted into one or more strings of symbols. In the example, a symbol is a bit, and each bit may have a value of '0' or '1'. Each symbol may be mapped to or encoded by an object representing that symbol (e.g., an identifier). Each symbol may be represented by an individual identifier. An individual identifier may be a nucleic acid molecule composed of components. The components may be nucleic acid sequences. Digital information may be written to the nucleic acid sequence by creating an identifier library corresponding to the information. The identifier library may be physically created by physically configuring identifiers corresponding to each symbol of the digital information. All or part of the digital information may be accessed at once. For example, a subset of identifiers is accessed from the identifier library. A subset of identifiers may be read by sequencing and identifying the identifiers. The identified identifiers may be associated with the corresponding symbols to decode the digital data.

[0057] A method for encoding and reading information using the approach of FIG. 1 may include, for example, receiving a bit stream and mapping each bit of the bit stream (a bit with a bit value of '1') to an individual nucleic acid identifier using an identifier rank or nucleic acid index. Construct a nucleic acid sample pool or identifier library containing copies of identifiers corresponding to bit values ​​1 (excluding identifiers for bit values ​​0). Reading the samples may include using a molecular biological method (e.g., sequencing, hybridization, PCR, etc.), determining which identifiers are represented in the identifier library, assigning a bit value of '1' to the bit corresponding to the identifier and a bit value of '0' elsewhere (referencing the identifier rank again to identify the bit of the original bit stream corresponding to each identifier) ​​to decode the information into the original encoded bit stream.

[0058] Encoding a string of N individual bits allows the same number of unique nucleic acid sequences to be used as possible identifiers. This approach to information encoding may involve the novel synthesis of identifiers (e.g., nucleic acid molecules) for each new information item to be stored (a string of N bits). In other cases, the cost of novelly synthesizing identifiers (N or fewer) for each new information to be stored is reduced through one-time novel synthesis and subsequent maintenance of all possible identifiers, so that encoding the new information may involve mechanically selecting and mixing pre-synthesized (or pre-manufactured) identifiers to form an identifier library. In other cases, the cost of (1) novelly synthesizing up to N identifiers for each new information to be stored, or (2) maintaining and selecting from N possible identifiers for each new information to be stored, or any combination thereof, may be reduced by synthesizing and maintaining a number of (less than N, and in some cases much less than N) nucleic acid sequences and then modifying these sequences through enzymatic action to generate up to N identifiers for each new information to be stored.

[0059] Identifiers may be reasonably designed and selected for ease of read, write, access, copy, and delete operations. Identifiers may be designed and selected to minimize write errors, mutations, performance degradation, and read errors. For the reasonable design of DNA sequences containing synthetic nucleic acid libraries (e.g., identifier libraries), refer to Chemical Methods Section H.

[0060] FIGS. 2a and 2b schematically illustrate an exemplary method called "data at address" for encoding digital data into an object or identifier (e.g., a nucleic acid molecule). FIG. 2a illustrates encoding a bit stream into an identifier library in which individual identifiers are constructed by connecting or assembling a single component specifying a byte-value and a single component specifying an identifier rank. Generally, the data at address method uses an identifier that encodes information modularly by including the following two objects: one object, namely, a "byte-value object" (or "data object") that identifies a byte-value, and one object, namely, a "rank object" (or "address object") that identifies an identifier rank (or the relative position of a byte within the original bit stream). FIG. 2b illustrates an example of the data at address method, wherein each rank object may be combinatorially constructed from a set of components and each byte-value object may be combinatorially constructed from a set of components. With this combination of rank and byte-value objects, more information can be written to the identifier than when the object is made from only a single component (Fig. 2a).

[0061] FIGS. 3a and 3b schematically illustrate another exemplary method for encoding digital information of an object or identifier (e.g., a nucleic acid sequence). FIG. 3a illustrates encoding a bit stream into an identifier library, wherein the identifier is composed of a single component that specifies an identifier rank. If the identifier is present at a specific rank (or address), a bit value '1' is assigned, and if the identifier is not present at a specific rank (or address), a bit value '0' is assigned. This type of encoding may use an identifier that encodes only the rank (the relative position of a bit within the original bit stream) and may encode a bit value of '1' or '0', respectively, using the presence of the corresponding identifier in the identifier library. Reading and decoding the information may include identifying the identifier present in the identifier library, assigning a bit value '1' to the corresponding rank, assigning a bit value '0' to other places, and so on. FIG. 3b illustrates an exemplary encoding method in which each identifier may be combinatorially composed from a set of components such that each possible combination configuration specifies a rank. This combination configuration allows more information to be recorded in the identifier than when the identifier is made of only a single component (e.g., FIG. 3a). For example, a set of components may include five individual components. The five individual components may be assembled to generate ten individual identifiers, each containing two of the five components. Each of the ten individual identifiers may have a rank (or address) corresponding to a position of a bit in a bitstream. The identifier library may include a subset of ten possible identifiers corresponding to a position of bit-value '1' and exclude a subset of ten possible identifiers corresponding to a position of bit-value '0' in a bitstream of length 10.

[0062] FIG. 4 shows a contour plot in logarithmic space representing the relationship between the combination space of possible identifiers (C, x-axis) and the average number of identifiers (k, y-axis) that are physically configured to store information of a given original size in bits (D, contour) using the encoding method illustrated in FIG. 3a and 3b. This plot assumes that original information of size D is recoded into a string of C bits (C may be larger than D), where the number of bits k has a bit value of '1'. Additionally, the plot assumes that information-nucleic acid encoding is performed on the recoded bit string, where identifiers are constructed for locations with a bit value of '1' and identifiers are not constructed for locations with a bit value of '0'. According to the assumptions, the combination space of possible identifiers has a size C to identify all locations of the recoded bit string, and the number of identifiers used to encode a bit string of size D is D = log 2 (C choose k) It is determined to be so, and here, Cchoose k can be a mathematical formula for the number of ways to select k unordered results from C possibilities. Therefore, as the combination space of possible identifiers increases beyond the size (in bits) of the given information, a decreasing number of physically configured identifiers can be used to store the given information.

[0063] FIG. 5 illustrates a schematic method for recording information into a nucleic acid sequence. Before recording the information, the information may be converted into a string of symbols and encoded into multiple identifiers. Recording the information may involve setting up a reaction to generate possible identifiers. The reaction may be set up by storing the input in a compartment. The input may include nucleic acids, components, templates, enzymes, or chemical reagents. The compartment may be a well, a tube, a location on a surface, a chamber within a microfluidic device, or a droplet within an emulsion. Multiple reactions may be set up in multiple compartments. The reaction may proceed through programmed temperature incubation or cycling to generate identifiers. The reaction may be optionally or ubiquitously removed (e.g., deleted). The reaction may be optionally or ubiquitously stopped, integrated, and purified to collect identifiers from a single pool. Identifiers from multiple identifier libraries may be collected in the same pool. Individual identifiers may include a barcode or tag identifying the identifier library to which they belong. Alternatively, or additionally, the barcode may include metadata regarding the encoded information. Supplementary nucleic acids or identifiers may be included in the identifier pool along with the identifier library. Supplementary nucleic acids or identifiers may serve to include metadata about encoded information, or to obfuscate or hide encoded information.

[0064] An identifier rank (e.g., nucleic acid index) may include a method or key for determining the order of identifiers. The method may include a lookup table containing all identifiers and their corresponding ranks. The method may also include a search table having ranks of all components constituting the identifier and a function for determining the order of any identifier including combinations of these components. Such a method may be described as lexicographical sorting and may be similar to sorting words in a dictionary alphabetically. In a data-at-address encoding method, an identifier rank (encoded by the identifier rank object) may be used to determine the position of a byte (encoded by the identifier byte value object) within a bit stream. Alternatively, the identifier rank for the current identifier (encoded by the entire identifier itself) may be used to determine the position of the bit value '1' within the bit stream.

[0065] A key may assign individual bytes to a unique subset of identifiers (e.g., nucleic acid molecules) within a sample. For example, in a simple form, a key may assign each bit of a byte to a unique nucleic acid sequence that specifies the position of the bit, and then assign a bit-value of 1 or 0, respectively, depending on the presence of the corresponding nucleic acid sequence within the sample. Reading encoded information from a nucleic acid sample may involve various molecular biology techniques, including sequencing, hybridization, or PCR. In some embodiments, reading an encoded dataset may involve reconstructing a portion of the dataset or reconstructing the entire encoded dataset from each nucleic acid sample. Where a sequence can be read, a nucleic acid index may be used with the presence or absence of a unique nucleic acid sequence, and the nucleic acid sample may be decoded into a bit stream (e.g., each bit string, byte, byte, or byte string).

[0066] Identifiers can be constructed by combinatorially assembling component nucleic acid sequences. For example, information can be encoded by taking a set of nucleic acid molecules (e.g., identifiers) from a defined molecular group (e.g., combination space). Each possible identifier of the defined molecular group may be an assembly of nucleic acid sequences (e.g., components) from a pre-constructed set of components that can be divided into layers. Each individual identifier can be constructed by connecting one component from all layers in a fixed order. For example, if there are M layers and each layer can have n components, maximum C = n M Can be configured with unique identifiers, and a maximum 2 C Different pieces of information or C bits can be encoded and stored. For example, to store megabit information, 1 x 10 6 individual identifiers or C = 1 x 10 6 A combination space of size can be used. The identifier in this example can be assembled from various components configured in various ways. Each assembly n = 1 x 10 3 It can be made from M = 2 prefabricated layers containing n components. Alternatively, the assembly each n = 1 x 10 2 It can be made from M = 3 layers containing components. In some embodiments, the assembly can be made of M=2, M=3, M=4, M=5, or more layers. As can be seen from this example, encoding the same amount of information using a larger number of layers can result in a smaller total number of components. Using fewer total components can be advantageous in terms of recording costs.

[0067] In one example, one may start with two sets of unique nucleic acid sequences or layers, X and Y, each having x and y components (e.g., nucleic acid sequences), respectively. Each nucleic acid sequence from X can be assembled from each nucleic acid sequence from Y. The total number of nucleic acid sequences maintained in the two sets may be the sum of x and y, but the total number of nucleic acid molecules that can be generated, and thus possible identifiers, may be the product of x and y. If sequences from X can be assembled into sequences of Y in any order, a much larger number of nucleic acid sequences (e.g., identifiers) can be generated. For example, the number of generated nucleic acid sequences (e.g., identifiers) can be twice the product of x and y if the assembly order is programmable. The set of all possible nucleic acid sequences that can be generated can be referred to as XY. The assembled unit order of the unique nucleic acid sequences of XY can be controlled using nucleic acids with individual 5' and 3' ends, and restriction digestion, ligation, polymerase chain reaction (PCR), and sequencing can occur on the individual 5' and 3' ends of the sequences. This approach can reduce the total number of nucleic acid sequences (e.g., components) used to encode N individual bits by encoding information as a combination and order of assembly products. For example, to encode 100 bits of information, two layers of 10 individual nucleic acid molecules (e.g., components) can be assembled in a fixed order to generate 10*10 or 100 individual nucleic acid molecules (e.g., identifiers), or one layer of 5 individual nucleic acid molecules (e.g., components) and another layer of 10 individual nucleic acid molecules (e.g., components) can be assembled in any order to generate 100 individual nucleic acid molecules (e.g., identifiers).

[0068] The nucleic acid sequences (e.g., components) within each layer may include a unique (or individual) sequence or barcode in the center, a common hybridization region at one end, and another common hybridization region at the other end. The barcode may contain a sufficient number of nucleotides to uniquely identify all sequences within the layer. For example, there are typically 4 possible nucleotides for each base position within a barcode. Therefore, a 3-base barcode is 4 3 = 64 nucleic acid sequences can be uniquely identified. Barcodes can be designed to be generated randomly. Alternatively, barcodes can be designed to avoid sequences that could cause complexity to the identifier's composition chemistry or sequencing. Additionally, barcodes can be designed so that each barcode has a minimum Hamming distance from other barcodes, thereby reducing the likelihood that base resolution mutations or read errors will interfere with the proper identification of the barcodes. For the reasonable design of DNA sequences, refer to Section H of the chemical methods.

[0069] A hybridization region at one end of a nucleic acid sequence (e.g., a component) may differ for each layer, but the hybridization region may be the same for each member within the layer. Adjacent layers are layers that have a hybridization region complementary to the component so that they can interact with each other. For example, all components from layer X may have a complementary hybridization region and thus can be attached to any component from layer Y. The hybridization region at the opposite end may serve the same purpose as the hybridization region at the first end. For example, any component from layer Y may be attached to any component of layer X on one end and to any component of layer Z on the opposite end.

[0070] FIGS. 6a and 6b illustrate an exemplary method called the "multiplication method" for constructing an identifier (e.g., nucleic acid molecule) by combinatorially assembling individual components (e.g., nucleic acid sequences) from each layer in a fixed order. FIG. 6a illustrates the architecture of an identifier constructed using the multiplication method. An identifier can be constructed by combining single components from each layer in a fixed order. For M layers, each containing N components... N M There are 27 possible identifiers. FIG. 6b illustrates an example of a combination space of identifiers that can be constructed using a multiplication method. For example, the combination space can be generated from three layers, each containing three individual components. The components can be combined such that one component from each layer can be combined in a fixed order. The entire combination space of this assembly method can consist of 27 possible identifiers.

[0071] FIGS. 7-10 illustrates a chemical method for implementing a multiplication method (see FIG. 6). The method illustrated in FIGS. 7-10, in conjunction with any other method for assembling two or more individual components in a fixed manner, can be used to generate an identifier library of any one or more identifiers. Identifiers may be configured at any point during the method or system disclosed herein using any of the implementation methods described in FIGS. 7-10. In some cases, all or part of the possible combination space of identifiers may be configured before digital information is encoded or written, and the writing process may then include mechanically selecting and pooling identifiers (for encoding information) from an existing set. In other cases, identifiers may be configured after one or more steps of the data encoding or writing process have occurred (i.e., while information is being written).

[0072] Enzyme reactions can be used to assemble components from different layers or sets. Since the components of each layer (e.g., nucleic acid sequences) have specific hybridization or attachment regions for the components of adjacent layers, assembly can occur as a one-pot reaction. For example, a nucleic acid sequence (e.g., component) X1 from layer X, a nucleic acid sequence Y1 from layer Y, and a nucleic acid sequence Z1 from layer Z can form an assembled nucleic acid molecule (e.g., identifier) ​​X1Y1Z1. Additionally, multiple nucleic acid molecules (e.g., identifiers) can be assembled in a single reaction by including multiple nucleic acid sequences from each layer. For example, if both Y1 and Y2 are included in the one-pot reaction of the previous example, two assembled products (e.g., identifiers) X1Y1Z1 and X1Y2Z1 can be produced. This reaction multiplexing can be used to reduce the recording time for multiple physically configured identifiers. For details on the rational design of DNA sequences related to assembly efficiency, refer to Chemical Methods Section H. The assembly of nucleic acid sequences can be performed over a period of about 1 day, 12 hours, 10 hours, 9 hours, 8 hours, 7 hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours, or less than 1 hour. The accuracy of the encoded data can be at least about 90%, 95%, 96%, 97%, 98%, 99% or higher.

[0073] The identifier may be constructed according to a multiplication method using overlap extension polymerase chain reaction (OEPCR) as exemplified in FIG. 7. Each component of each layer may include a double-stranded or single-stranded (illustrated in the figure) dissolution sequence having a common hybridization region on the sequence end that may be homologous and / or complementary to the common hybridization region on the sequence end of a component from an adjacent layer. The individual identifiers are components X1-X AOne component (e.g., a unique sequence) from layer X (or layer 1) including Y1-Y A A second component (e.g., a unique sequence) from layer Y (or layer 2) comprising, and Z1-Z B It can be constructed by linking a third component (e.g., a unique sequence) from layer Z (or layer 3) containing. The component from layer X may have a 3' end that shares complementarity with the 3' end on the component from layer Y. Thus, the single-stranded components of layers X and Y can be annealed together at the 3' end and expanded using PCR to generate a double-stranded nucleic acid molecule. The resulting double-stranded nucleic acid molecule can be melted to produce a 3' end that shares complementarity with the 3' end of the component from layer Z. The component from layer Z can be annealed with the generated nucleic acid molecule and expanded to generate a unique identifier comprising single components from layers X, Y, and Z in a fixed order. Refer to Section A of the chemical methods for OEPCR. DNA size selection (e.g., gel extraction, see Chemical Methods Section E) or polymerase chain reaction (PCR) using primers on the outermost layer side (see Chemical Methods Section D) can be implemented to separate the fully assembled identifier product from other byproducts that may be formed in the reaction. Sequential nucleic acid capture using two probes, one for each of the two outermost layers, can also be implemented to separate the fully assembled identifier product from other byproducts that may be formed in the reaction (see Chemical Methods Section F).

[0074] The identifier can be assembled according to a multiplication method using adhesive end ligation as illustrated in FIG. 8. Individual identifiers can be assembled using three layers, each containing a double-stranded component (e.g., double-stranded DNA (dsDNA)) having a single-stranded 3' overhang. For example, the identifier consists of components X1-X A One component from layer X (or layer 1) including, Y1-Y B A second component of layer Y (or layer 2) comprising, and Z1-Z C It includes a third component from layer Z (or layer 3) comprising. To combine the component from layer X with the component from layer Y, the component of layer X is of FIG. 8 a It may include a common 3' overhang labeled as such, and a component of layer Y is a common, complementary 3' overhang. a* It may include. To combine a component from layer Y with a component from layer Z, the element of layer Y is of FIG. 8 b It may include a common 3' overhang labeled as , and the element of layer Z is the common complementary 3' overhang b*It may include. The 3' overhang of a component of layer X may be complementary to the 3' end of a component of layer Y, and another 3' overhang of a component of layer Y may be complementary to the 3' end of a component of layer Z, so that the components can be hybridized and ligated. Thus, a component from layer X cannot be hybridized with another component of layer X or layer Z, and likewise, a component of layer Y cannot be hybridized with another component of layer Y. Additionally, a single component from layer Y can be ligated to a single component of layer X and a single component of layer Z while ensuring the formation of a complete identifier. For adhesive end ligation, refer to Chemical Methods Section B. DNA size selection (e.g., gel extraction, refer to Chemical Methods Section E) or polymerase chain reaction (PCR) using primers on the outermost layer side (refer to Chemical Methods Section D) may be implemented to separate the identifier product from other byproducts that may be formed in the reaction. Sequential nucleic acid capture using two probes, one for each of the two outermost layers, can also be implemented to separate the identifier product from other byproducts that may be formed in the reaction (see Chemical Methods Section F).

[0075] Sticky ends for sticky end ligation can be generated by treating the components of each layer with a restriction endonuclease (see Chemical Methods Section C for details on the restriction enzyme reaction). In some embodiments, the components of multiple layers may be generated from a single "parent" set of components. For example, there are embodiments in which a single parent set of double-stranded components may have complementary restriction sites on each end (e.g., restriction sites for BamHI and BglII). Any two components may be selected for assembly and individually digested with one or the other complementary restriction enzyme (e.g., BglII or BamHI) to generate complementary sticky ends that can be ligated together, thereby deriving an inert scar. The product nucleic acid sequence may contain restriction sites on each end (e.g., BamHI on the 5' end and BglII on the 3' end) and may be further ligated to another component from the parent set according to the same process. This process can be repeated indefinitely (Fig. 20). If the module contains N components, each cycle may be equivalent to adding an additional layer of N components to the multiplication method.

[0076] A method using ligation to construct a sequence of nucleic acids comprising elements of set X (e.g., set 1 of dsDNA) and elements of set Y (e.g., set 2 of dsDNA) may include the step of obtaining or constructing two or more pools of double-stranded sequences (e.g., set 1 of dsDNA and set 2 of dsDNA), wherein the first set (e.g., set 1 of dsDNA) has sticky ends (e.g., a ...including ) and the second set (e.g., set 2 of dsDNA) has a sticky end complementary to the sticky end of the first set (e.g., a*Includes ). Any DNA from a first set (e.g., set 1 of dsDNA) and any subset of DNA from a second set (e.g., set 2 of dsDNA) may be combined and assembled, and then ligated together to form a single double-stranded DNA having elements from the first set and elements from the second set.

[0077] The identifier can be assembled according to the multiplication method using site-specific recombination as shown in Fig. 9. The identifier can be constructed by assembling components from three different layers. The component of layer X (or layer 1) is attB on one side of the molecule. x It may include a double-stranded molecule having a recombinase site, and the component from layer Y (or layer 2) is attP on one side x It may include a double-stranded molecule having a recombinase site, and the component of layer Z (or layer 3) is attP on one side of the molecule yIt may include a recombinase site. The attB and attP sites within a pair may be recombined in the presence of the corresponding recombinase as indicated by the subscript. One component from each layer may be combined such that one component from layer X is associated with one component from layer Y, and one component from layer Y is associated with one component from layer Z. The application of one or more recombinases may recombine the components to generate a double-stranded identifier containing the aligned components. DNA size selection (e.g., gel extraction) or PCR using primers on the outermost layer side may be implemented to separate the identifier product from other byproducts that may be formed in the reaction. Generally, multiple orthogonal attB and attP pairs may be used, and each pair may be used to assemble components from additional layers. For large serine-based recombinases, up to six orthogonal attB and attP pairs may be generated per recombinase, and multiple orthogonal recombinases may also be implemented. For example, 13 layers can be assembled using 12 orthogonal attB and attP pairs, namely 6 orthogonal pairs from each of the two large serine recombinases, such as BxbI and PhiC31. The orthogonality of the attB and attP pairs ensures that one pair of attB sites does not react with another pair of attP sites. This allows the components of different layers to be assembled in a fixed order. Recombinase-mediated recombination reactions can be reversible or irreversible depending on the implemented recombinase system. For example, the large serine recombinase family catalyzes irreversible recombination reactions without requiring high-energy cofactors, whereas the tyrosine recombinase family catalyzes reversible reactions.

[0078] The identifier can be constructed according to a multiplication method using template-designated ligation (TDL) as illustrated in FIG. 10a. Template-designated ligation can form the identifier by utilizing a single-stranded nucleic acid sequence called a "template" or "staple" to facilitate the aligned ligation of components. The template is simultaneously hybridized to components from adjacent layers and remains adjacent to each other (3' end to 5' end) while the ligase ligates them. In the example of FIG. 10a, three layers or sets of single-stranded components are combined. A first layer of components sharing a common sequence a at the 3' end that is complementary to sequence a* (e.g., layer X or layer 1), a second layer of components sharing common sequences b and c at the 5' and 3' ends that are complementary to sequences b* and c*, respectively (e.g., layer Y or layer 2), a third layer of components sharing a common sequence d at the 5' end that may be complementary to sequence d* (e.g., layer Z or layer 3), and a set of two templates or “staples” having a first staple comprising sequence a*b* (5' to 3') and a second staple comprising sequence c*d* ('5' to 3'). In this example, one or more components of each layer may be selected and mixed in reaction with the staples, which may facilitate the formation of an identifier by ligating one component from each layer in an order defined by complementary annealing. For TDL, refer to Chemical Methods Section B. DNA size selection (e.g., gel extraction, see Chemical Methods Section E) or polymerase chain reaction (PCR) using primers on the outermost layer side (see Chemical Methods Section D) can be implemented to separate the identifier product from other byproducts that may be formed in the reaction. Sequential nucleic acid capture using two probes, one for each of the two outermost layers, can also be implemented to separate the identifier product from other byproducts that may be formed in the reaction (see Chemical Methods Section F).

[0079] Figure 10b shows a histogram of the copy number (abundance) of 256 individual nucleic acid sequences, each assembled with a 6-layer TDL. The outer layers (the first and last layers) each contain one component, and each inner layer (the remaining four layers) contains four components. Each outer layer component was 28 bases, including a 10-base hybridization region. Each inner layer component was 30 bases, including a 10-base common hybridization region, a 10-base variable (barcode) region at the 5' end, and a 10-base common hybridization region at the 3' end. The length of each of the three template strands was 20 bases. All 256 individual sequences were assembled in a multiple manner in a single reaction containing all components, the template, T4 polynucleotide kinase (for component phosphorylation), T4 ligase, ATP, and other appropriate reaction reagents. The reaction mixture was incubated at 37°C for 30 minutes, followed by incubation at room temperature for 1 hour. A sequencing adapter was added to the reaction product via PCR, and the product was sequenced using an Illumina MiSeq instrument. The relative copy number of each individual assembled sequence is shown out of a total of 192,910 assembled sequence reads. Another embodiment of this method may use a double-stranded component, wherein the component may initially melt to form a single-stranded version that can be annealed onto a staple. Another embodiment or derivative of this method (i.e., TDL) may be used to construct a combination space of identifiers more complex than can be achieved in the multiplication method.

[0080] The identifier may be configured according to the product system using various other chemical implementations, including the Golden Gate assembly, Gibson assembly, and ligase cyclic reaction assembly.

[0081] FIGS. 11a and 11b schematically illustrate an exemplary method referred to as a "permutation method" for constructing an identifier (e.g., a nucleic acid molecule) from permuted components (e.g., a nucleic acid sequence). FIG. 11a illustrates the architecture of an identifier constructed using the permutation method. An identifier can be constructed by combining single components from each layer in a programmable order. FIG. 11b illustrates an example of a combination space of an identifier that can be constructed using the permutation method. For example, a combination space of size 6 can be generated from three layers, each containing one individual component. The components can be connected in any order. Generally, using M layers, each having N components, a total through the permutation method N M M! A combination space of dogs becomes possible.

[0082] FIG. 11c illustrates an exemplary implementation of a permutation scheme using template-designated ligation (TDL, see Chemical Methods Section B). Components from multiple layers are assembled between fixed left-end and right-end components, also known as edge scaffolds. Since these edge scaffolds are identical for all identifiers in the combination space, they can be added as part of the reaction master mix for the implementation. Templates or staples exist for any possible junction between any two layers or scaffolds such that the order in which components from different layers are incorporated into the reaction identifier depends on the template selected for the reaction. To enable any possible layer permutation for M layers, for all possible junctions (including junctions with scaffolds) M 2 +2MThere may be individual selectable staples. M of these templates (shaded in gray) form junctions between the layer and themselves and may be excluded for the purposes of the permutation assembly described herein. However, including them may enable a larger combination space with identifiers containing repeating components, as illustrated in FIG. 11d-g. DNA size selection (e.g., gel extraction, see Chemical Methods Section E) or polymerase chain reaction (PCR) using primers on the outermost layer side (see Chemical Methods Section D) may be implemented to separate the identifier product from other byproducts that may be formed in the reaction. Sequential nucleic acid capture using two probes, one for each of the two outermost layers, may also be implemented to separate the identifier product from other byproducts that may be formed in the reaction (see Chemical Methods Section F).

[0083] FIGS. 11d-g illustrate exemplary methods for how a permutation scheme can be extended to include specific instances of an identifier having repeating components. FIG. 11d illustrates an example of how the implementation form of FIG. 11c can be used with permutations and repeating components. For example, the identifier may include a total of three components assembled from two separate components. In this example, a component of a layer may appear multiple times in the identifier. Adjacent connections of the same component can be achieved by using a staple that has adjacent complementary hybridization regions for both the 3' end and the 5' end of the same component, e.g., the a*b*(5' to 3') staple in the figure. Typically, for M layers, there are M such staples. Incorporating repeating components into this implementation can generate nucleic acid sequences of length exceeding 2 (i.e., containing 1, 2, 3, 4 or more components) assembled between edge scaffolds, as illustrated in FIG. 11e. FIG. 11e shows that the exemplary implementation method of FIG. 11d can yield a non-target nucleic acid sequence assembled between edge scaffolds in addition to the identifier. Since the appropriate identifier shares the same primer binding site on the edge, it cannot be separated from the non-target nucleic acid sequence by PCR. However, in this example, since each assembled nucleic acid sequence can be designed to have a unique length (e.g., when all components have the same length), DNA size selection (e.g., using gel extraction) can be implemented to separate the target identifier (e.g., the second sequence from the top) from the non-target sequence. For size selection, refer to Chemical Methods Section E. FIG. 11f shows another example in which multiple nucleic acid sequences of the same edge sequence but different lengths can be generated in the same reaction in which the identifier is constructed with repeated components. In this method, a template can be used to assemble components of one layer with components of another layer in an alternating pattern. When using the method illustrated in FIG. 11e, size selection can be used to select the identifier of the designed length. FIG. 11g shows an example in which constructing an identifier with repeated components can generate multiple nucleic acid sequences that have the same edge sequence and have the same length for some nucleic acid sequences (e.g., the third and fourth from the top, and the sixth and seventh from the top). In this example, even if PCR and DNA size selection are implemented, it may be impossible to construct one without constructing the other, so nucleic acid sequences sharing the same length may both be excluded from individual identifiers.

[0084] Figures 12a–12d show a larger number M An arbitrary number among the possible components of k An exemplary method referred to as the "MchooseK" method for constructing an identifier (e.g., a nucleic acid molecule) having assembled components (e.g., a nucleic acid sequence) is schematically illustrated. Figure 12a illustrates the architecture of an identifier constructed using the MchooseK method. Using this method, an identifier is constructed by assembling one component from each layer in any subset of all layers (e.g., selecting a component from k layers out of M possible layers). Figure 12b illustrates an example of a combination space of identifiers that can be constructed using the MchooseK method. In this assembly method, the combination space is for M layers, N components per layer, and k identifier lengths. N K MchooseK Possible identifiers may be included. For example, if there are 5 layers each containing one component, up to 10 individual identifiers each containing 2 components may be assembled.

[0085] The MchooseK method can be implemented using template-designated ligation (see Chemical Methods Section B) as illustrated in Fig. 12c. Similar to the TDL implementation for the permutation method (Fig. 11c), the components of this example are assembled between edge scaffolds that may or may not be included in the reaction master mix. The components can be divided into M layers, for example, M = 4 layers with predefined ranks from 2 to M, where the left edge scaffold may be rank 1 and the right edge scaffold may be rank M+1. Each template contains a nucleic acid sequence for the 3' to 5' linkage of any two components, from lower rank to higher rank. These templates (( M+1) 2 +M+1) / 2There are individual identifiers for any K component from individual layers can be constructed by combining the selected components in a ligation reaction with corresponding K+1 staples used to bring the K components together with the edge scaffolds in rank order. This reaction setup can generate a nucleic acid sequence corresponding to the target identifier between the edge scaffolds. Alternatively, the target identifier can be assembled by combining a reaction mixture containing all templates with the selected components. This alternative method can generate various nucleic acid sequences having the same edge sequence but different lengths (where all components are the same length), as exemplified in FIG. 12D. The target identifier (bottom) can be separated from the byproduct nucleic acid sequence by size. For nucleic acid size selection, refer to Chemical Methods Section E.

[0086] FIGS. 13a and 13b schematically illustrate an exemplary method referred to as a "partitioning method" for constructing an identifier with partitioned components. FIG. 13a shows an example of a combination space of identifiers that can be constructed using the partitioning method. Individual identifiers can be constructed by assembling one component of each layer in a fixed order by optionally placing a partition (a specially classified component) between two components of different layers. For example, a set of components can be composed of four layers, each containing one partition component and one component. Components from each layer can be combined in a fixed order, and a single partition component can be assembled at various locations between the layers. The identifier of this combination space can create eight possible combination spaces of identifiers, including a partition component between components of the first and second layers, a partition between components of the second and third layers, and so on, without including a partition component. In general, a combination space that can be constructed using M layers, each having N components, and p partition components N K (p+1) M-1 There are several possible identifiers. This method can generate identifiers of various lengths.

[0087] FIG. 13b shows an example of an implementation of a partitioning method using template-designated ligation (see Chemical Methods Section B). The template comprises a nucleic acid sequence for ligating one component from each of M layers together in a fixed order. For each partition component, there is an additional template pair that allows the partition component to be ligated between components from any two adjacent layers. For example, the template pair is formed such that one template in a pair (e.g., containing sequence g*b*(5' to 3')) can ligate the 3' end of layer 1 (containing sequence b) to the 5' end of the partition component (containing sequence q), and a second template in the pair (e.g., containing sequence c*h*(5' to 3')) can ligate the 3' end of the partition component (containing sequence h) to the 5' end of layer 2 (containing sequence c). To insert a partition between any two components of adjacent layers, a standard template for ligating the layers together may be excluded from the reaction, and a template pair for ligating the partition at that location may be selected from the reaction. In the present example, targeting the partition component between Layer 1 and Layer 2 may be selected using the template pair c*h*(5' to 3') and g*b*(5' to 3') rather than template c*b*(5' to 3'). The components may be assembled between edge scaffolds that may be included in the reaction mixture (along with the corresponding templates for ligating the first and M-th layers, respectively). Generally, the total approximately M-1+2*p*(M-1)A number of selectable templates can be used in this method for M layers and p partition components. The implementation of such a partitioning scheme can generate various nucleic acid sequences in reactions of different lengths that have the same edge sequence. Target identifiers can be separated from byproduct nucleic acid sequences through DNA size selection. Specifically, there may be exactly one nucleic acid sequence product having exactly M layer components. If the layer components are designed to be sufficiently large relative to the partition components, multiple partitioned identifiers from multiple reactions can be separated in the same size selection step by defining a global size selection region, so that the identifier (and none of the non-target byproducts) can be selected regardless of the specific partitioning of the components within the identifier. For nucleic acid size selection, refer to Chemical Methods Section E.

[0088] FIGS. 14a and FIGS. 14b schematically illustrate an exemplary method referred to as an "unconstrained string" (or USS) method for constructing an identifier composed of any string of components from a plurality of possible components. FIG. 14a shows an example of a combination space of 3-component (or 4-scaffold) length identifiers that can be constructed using an unrestricted string method. The unrestricted string method constructs individual identifiers of length K components using one or more individual components taken from one or more layers, respectively, where each individual component may appear at one of the K component positions of the identifier (repetition allowed). For example, for two layers each containing one component, there are 8 possible 3-component length identifiers. Generally, for M layers each having one component, there are M of length K components. KThere are several possible identifiers. FIG. 14b shows an example of an implementation of an unrestricted string method using template-designated ligation (see Chemical Methods Section B). In this method, K+1 single-stranded and aligned scaffold DNA components (including two edge scaffolds and K-1 inner scaffolds) are present in the reaction mixture. Individual identifiers include a single component connected between all pairs of adjacent scaffolds. For example, the process continues until all K adjacent scaffold junctions, such as a component ligated between scaffolds A and B, or a component ligated between scaffolds C and D, are occupied by the component. In the reaction, selected components from different layers are introduced into the scaffolds along with selected staple pairs and are instructed to be assembled into the appropriate scaffolds. For example, the staple pairs a*L* (5' to 3') and A*b* (5' to 3') specify that a Layer 1 component with a 5' terminal region 'a' and a 3' terminal region 'b' should be ligated between the L and A scaffolds. Generally M layers and K+1 For dog scaffolds, 2* M * K Number of selectable staples are used, length K Any USS identifier can be constructed. Because the staple connecting the component to the scaffold on the 5' end is separated from the staple connecting the same component to the scaffold on the 3' end, nucleic acid byproducts can be formed as target identifiers in the reaction with the same edge scaffold, but K Less than 1 component ( K+1 Less than 1 scaffold) or K Components exceeding one ( K+1 It includes scaffolds exceeding [number]. The target identifier is exactly K 1 component ( K+1Since it can be formed with scaffolds, it can be selected through techniques such as DNA size selection, where all components are designed to be of the same length and all scaffolds are designed to be of the same length. For nucleic acid size selection, refer to Chemical Methods Section E. In a specific embodiment of an unrestricted string method in which there may be one component per layer, said component comprises only a single individual nucleic acid sequence that performs all three roles: (1) an identification barcode, (2) a hybridization region for staple-mediated ligation of the 5' end to the scaffold, and (3) a hybridization region for staple-mediated ligation of the 3' end to the scaffold.

[0089] The internal scaffold illustrated in FIG. 14b may be designed to use the same hybridization sequence for both the staple-mediated 5' ligation of the scaffold to a component and the staple-mediated 3' ligation of the scaffold to another (not necessarily individual) component. Thus, the 1-scaffold, 2-staple stacked hybridization event illustrated in FIG. 14b represents a statistical forward-backward hybridization event occurring between the scaffold and each staple to enable both 5' component ligation and 3' component ligation. In other embodiments of the string method that are not limited, the scaffold may be designed with two connected hybridization regions, namely, an individual 3' hybridization region for staple-mediated 3' ligation and an individual 5' hybridization region for staple-mediated 5' ligation.

[0090] FIGS. 15a and 15b schematically illustrate an exemplary method referred to as a "component deletion method" for constructing an identifier by deleting nucleic acid sequences (or components) from a parent identifier. FIG. 15a shows an example of a combination space of possible identifiers that can be constructed using the component deletion method. In this example, the parent identifier may be composed of multiple components. The parent identifier may include approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 or more components. Individual identifiers are formed by selectively deleting any number of components from N possible components, thereby creating a size 2 N Generate the "entire" combination space of, or remove a fixed number of K components from N possible components to obtain the size N choose K of " It can be configured by generating "NchooseK". In the example with a mother identifier having 3 components, the total combination space can be 8 and the 3choose2 combination space can be 3.

[0091] Figure 15b illustrates an exemplary implementation of a component deletion method using double-strand targeted cleavage and repair (DSTCR). The parent sequence may be a single-stranded DNA substrate containing a component adjacent to a nuclease-specific target site (which may be four bases or less in length), wherein the parent sequence may be cultured with one or more double-strand-specific nucleases corresponding to the target site. The individual component may be targeted for deletion using a complementary single-stranded DNA (or cleavage template) that binds to the parent component DNA (and the adjacent nuclease site), thereby forming a double-stranded sequence stable to the parent that can be cleaved by the nuclease at both ends. Another single-stranded DNA (or repair template) hybridizes to the separated ends of the parent (where the component sequence was located) and is connected either directly or by a replacement sequence for ligation, so that the parent no longer contains the active nuclease target site. We refer to this method as "double-strand targeted cleavage (DSTC)." Size selection can be used to select an identifier with a specific number of deleted components. For nucleic acid size selection, refer to Chemical Methods Section E.

[0092] Alternatively or additionally, the parent identifier may be a double-stranded or single-stranded nucleic acid substrate containing components separated by a spacer sequence so that the two components are not located on the same side of the sequence. The parent identifier may be cultured with a Cas9 nuclease. Individual components may be targeted for deletion using a guide ribonucleic acid (cutting template) that binds to the edges of the components and enables Cas9-mediated cleavage at the lateral sites. A single-stranded nucleic acid (repair template) may be hybridized to the separated ends of the parent identifier (e.g., between the ends where the component sequences were located) to bring them together for ligation. Ligation may be performed directly or by connecting the ends with a replacement sequence so that the ligated sequence of the parent no longer contains the spacer sequence that can be targeted by Cas9. We refer to this method as "sequence-specific targeted cleavage and repair" or "SSTCR".

[0093] The identifier can be constructed by inserting components into the parent identifier using a derivative of DSTCR. The parent identifier may be a single-stranded nucleic acid substrate containing nuclease-specific target sites (which may be four bases or less in length), each embedded within a separate nucleic acid sequence. The parent identifier may be incubated with one or more double-stranded specific nucleases corresponding to the target sites. Individual target sites of the parent identifier may be targeted for component insertion using a complementary single-stranded nucleic acid (cleavage template) that binds to the target site and a separate surrounding nucleic acid sequence of the parent identifier to form a double-stranded site. The double-stranded site may be cleaved by the nuclease. Another single-stranded nucleic acid (repair template) may hybridize to the separated ends of the parent identifier to assemble them for ligation, and the mock-ligated sequence, connected by the component sequence, no longer contains the active nuclease target site. Alternatively, components may be inserted into the parent identifier using a derivative of SSTCR. The parent identifier can be a double-stranded or single-stranded nucleic acid, and the parent identifier can be cultured with Cas9 nuclease. Individual sites on the parent identifier can be targeted for cleavage using guide RNA (cleavage template). A single-stranded nucleic acid (repair template) can be hybridized to the separated ends of the parent identifier and assembled for ligation, linked by component sequences, so that the ligated sequence of the parent identifier no longer contains the active nuclease target site. Size selection can be used to select identifiers with a specific number of component insertions.

[0094] FIG. 16 schematically illustrates a modal identifier having recombinase recognition sites. Various patterns of recognition sites can be recognized by various recombinases. All recognition sites for a specific set of recombinases are arranged so that the nucleic acids between them can be removed when the recombinase is applied. The nucleic acid strands illustrated in FIG. 16 are 2 depending on the subset of recombinases applied. 5 = 32 different sequences can be adopted. In some embodiments, as illustrated in FIG. 16, unique molecules can be produced by using recombinases that cut, move, invert, and transpose segments of DNA to generate other nucleic acid molecules. Generally, using N recombinases results in 2 from the parent N Can be generated for all possible identifiers. In some embodiments, multiple orthogonal pairs of recognition sites from different recombinases may be arranged on a parent identifier in an overlapping manner so that the application of one recombinase influences the type of recombination event that occurs when a downstream recombinase is applied (see Roquet et al., Synthetic recombinase-based state machines in living cells, Science 353 (6297): aad8559 (2016), incorporated herein by reference). Such a system can generate different identifiers for all alignments of N recombinases, N!. The recombinases may be tyrosine family, such as Flp and Cre, or large serine recombinase family, such as PhiC31, BxbI, TP901, or A118. Using recombinases from the large serine recombinase family may be advantageous because they facilitate irreversible recombination and thus generate identifiers more efficiently than other recombinases.

[0095] In some cases, a single nucleic acid sequence can be programmed to become multiple distinct nucleic acid sequences by applying multiple recombinases in a distinct order. Approximately ~e1 M! individual nucleic acid sequences can be generated by applying M recombinases in different subsets and sequences thereof, where the number of recombinases M can be 7 or less for a large serine recombinase family. Where the number of recombinases M can be greater than 7, the number of sequences that can be generated is approximately 3.9M, and, for example, may be referred to Roquet et al., Synthetic recombinase-based state machines in living cells, Science 353 (6297): aad8559 (2016), the entirety of which is incorporated herein by reference. Additional methods for producing different DNA sequences from one common sequence may include targeted nucleic acid editing enzymes such as CRISPR-Cas, TALENS, and zinc finger nucleases. Sequences generated by recombinases, targeted editing enzymes, etc. may be used in conjunction with any prior methods, for example, methods disclosed in any drawings and disclosures of this application.

[0096] If the bitstream of information to be encoded is larger than can be encoded by any single nucleic acid molecule, the information can be split and indexed into nucleic acid sequence barcodes. Furthermore, any subset of size k nucleic acid molecules is selected from a set of N nucleic acid molecules and log2( N choose k It can generate bit information. Barcodes can be assembled into nucleic acid molecules within a subset of size k to encode longer bit streams. For example, M barcodes M *log2( N choose k It can be used to generate bit information. Given the number of available nucleic acid molecules N in the set and the number of available barcodes M, the size is such that the total number of molecules in the pool is minimized to encode information. k = k0 A subset of may be selected. A method for encoding digital information may include steps for splitting a bit stream and encoding individual elements. For example, a bit stream containing 6 bits may be split into three components, each consisting of 2 bits. Each 2-bit component may form an information cassette with a barcode and may be grouped or pooled together to form a hyperpool of information cassettes.

[0097] A barcode can facilitate information indexing when the amount of digital information to be encoded exceeds the amount that can fit into a single pool. Information containing longer bit strings and / or multiple bytes can be encoded by stratifying the approach disclosed in FIG. 3, for example, by including a tag having a unique nucleic acid sequence encoded using a nucleic acid index. An information cassette or identifier library may include a nitrogen-containing base or nucleic acid sequence containing a unique nucleic acid sequence that provides position and bit value information in addition to a barcode or tag in which a given sequence represents a component or components of a corresponding bit stream. An information cassette may include one or more unique nucleic acid sequences as well as a barcode or tag. A barcode or tag on an information cassette may provide a reference to the information cassette and all sequences contained in the information cassette. For example, a tag or barcode on an information cassette may indicate which part of the bit stream or bit component of the bit stream the unique sequence encodes information (e.g., bit value and bit position information).

[0098] Using barcodes allows more bit-unit information to be encoded in a pool than the size of the possible identifier combination space. For example, a 10-bit sequence can be separated into two sets of bytes, each consisting of 5 bits. Each byte can be mapped to a set of 5 possible individual identifiers. Initially, the identifier generated for each byte may be the same, but they may be stored in separate pools, or a reader of the information may not know which byte a specific nucleic acid sequence belongs to. However, each identifier can be barcoded or tagged with a label corresponding to the byte to which the encoded information applies (e.g., Barcode 1 can be attached to a sequence in the nucleic acid pool to provide the first 5 bits, and Barcode 2 can be attached to a sequence within the nucleic acid pool to provide the second 5 bits), and then the identifiers corresponding to the 2 bytes can be combined into a single pool (e.g., a "hyper-pool" or one or more identifier libraries). Each identifier library in one or more combined identifier libraries may contain individual barcodes that identify a given identifier as belonging to that given identifier library. A method for adding a barcode to each identifier in an identifier library may include using PCR, Gibson, ligation, or any other approach that allows a given barcode (e.g., barcode 1) to be attached to a given nucleic acid sample pool (e.g., attaching barcode 1 to nucleic acid sample pool 1 and barcode 2 to nucleic acid sample pool 2). Samples from a hyper-pool can be read by a sequencing method, and sequencing information can be parsed using a barcode or tag. An identifier library with M sets of barcodes and N possible identifiers (combination space) and a method using barcodes can encode a bit stream of length equal to the product of M and N.

[0099] In some embodiments, identifier libraries may be stored in an array of wells. An array of wells may be defined as having n columns and q rows, and each well may contain two or more identifier libraries in a hyper-pool. The information encoded in each well is greater than the information contained in each well. nxq A single large continuous information of a larger size can be constructed. A fraction can be taken from one or more wells of an array of wells, and the encoding can be read using sequencing, hybridization, or PCR.

[0100] A nucleic acid sample pool, a hyper-pool, an identifier library, a group of identifier libraries, or a well containing a nucleic acid sample pool or a hyper-pool may include a unique nucleic acid molecule (e.g., an identifier) ​​corresponding to an information bit and multiple supplementary nucleic acid sequences. The supplementary nucleic acid sequences may not correspond to the encoded data (e.g., not to the bit value). The supplementary nucleic acid samples may mask or encode information stored in the sample pool. The supplementary nucleic acid sequences may be derived from biological sources or synthetically produced. Supplementary nucleic acid sequences derived from biological sources may include randomly fragmented nucleic acid sequences or reasonably fragmented sequences. In particular, if the synthetically encoded information (e.g., the combination space of identifiers) is made to resemble natural genetic information (e.g., a fragmented genome), the biologically derived supplementary nucleic acid may hide or obscure the data-containing nucleic acid in the sample pool by providing natural genetic information along with the synthetically encoded information. In one example, the identifier is derived from a biological source, and the supplementary nucleic acid is derived from a biological source. The sample pool may contain multiple sets of identifiers and supplementary nucleic acid sequences. Each set of identifiers and supplementary nucleic acid sequences may originate from different organisms. In one example, the identifiers originate from one or more organisms, and the supplementary nucleic acid sequences originate from a single, different organism. The supplementary nucleic acid sequences may also originate from one or more organisms, and the identifiers may originate from a single organism different from the organism from which the supplementary nucleic acid originates. Both the identifiers and the supplementary nucleic acid sequences may originate from multiple different organisms. A key may be used to distinguish the identifiers from the supplementary nucleic acid sequences.

[0101] The supplementary nucleic acid sequence may store metadata regarding the recorded information. The metadata may include additional information for determining and / or acknowledging the source of the original information and / or the intended recipient of the original information. The metadata may include additional information regarding the format of the original information, the tools and methods used to encode and record the original information, and the date and time the original information was recorded in the identifier. The metadata may include additional information regarding the format of the original information, the tools and methods used to encode and record the original information, and the date and time the original information was recorded in the nucleic acid sequence. The metadata may include additional information regarding modifications applied to the original information after the information was recorded in the nucleic acid sequence. The metadata may include annotations to the original information or one or more references to external information. Alternatively or additionally, the metadata may be stored in one or more barcodes or tags attached to the identifier.

[0102] Identifiers in the identifier pool may have the same, similar, or different lengths. Supplementary nucleic acid sequences may have a length that is shorter than, substantially the same as, or greater than that of the identifiers. Supplementary nucleic acid sequences may have an average length that is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more bases of the average length of the identifiers. In one example, the supplementary nucleic acid sequence is the same or substantially the same length as the identifier. The concentration of the supplementary nucleic acid sequence may be lower, substantially the same, or higher than that of the identifiers in the identifier library. The concentration of the supplementary nucleic acid is approximately 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1x10⁻¹⁰⁴ 4 %, 1 x10 5%, 1 x10 6 %, 1 x10 7 %, 1 x10 8 It may be lower than or equal to % or less. The concentration of supplementary nucleic acid is approximately 1%, 10%, 20%, 40%, 60%, 80%, 100%, 125%, 150%, 175%, 200%, 1000%, 1 x 10⁻⁶ compared to the identifier concentration. 4 %, 1 x10 5 %, 1 x10 6 %, 1 x10 7 %, 1 x10 8 It can be greater than or equal to %. Higher concentrations can help obfuscate or hide data. In one example, the concentration of supplementary nucleic acid sequences is substantially higher than the concentration of identifiers in the identifier pool (e.g., 1 x 10⁻¹⁰). 8 % higher).

[0103] Method for copying and accessing data stored in nucleic acid sequences

[0104] In another aspect, the present disclosure provides a method for copying information encoded in nucleic acid sequence(s). A method for copying information encoded in nucleic acid sequence(s) may include (a) providing an identifier library and (b) configuring one or more copies of the identifier library. The identifier library may include a subset of a plurality of identifiers from a larger combination space. Each individual identifier of the plurality of identifiers may correspond to an individual symbol of a string of symbols. An identifier may include one or more components. A component may include a nucleic acid sequence.

[0105] In another aspect, the present disclosure provides a method for accessing information encoded in a nucleic acid sequence. The method for accessing information encoded in a nucleic acid sequence may include (a) providing an identifier library, and (b) extracting a portion or a subset of identifiers present in the identifier library from the identifier library. The identifier library may include a subset of a plurality of identifiers from a larger combination space. Each individual identifier of the plurality of identifiers may correspond to an individual symbol of a string of symbols. An identifier may include one or more components. A component may include a nucleic acid sequence.

[0106] Information may be recorded in one or more identifier libraries as described elsewhere in this specification. Identifiers may be constructed using methods described elsewhere in this specification. Stored data may be replicated by creating copies of individual identifiers in an identifier library or in one or more identifier libraries. A portion of the identifiers may be replicated, or the entire library may be replicated. Replication may be performed by amplifying the identifiers in the identifier library. When one or more identifier libraries are combined, a single identifier library or multiple identifier libraries may be replicated. If an identifier library contains a supplementary nucleic acid sequence, the supplementary nucleic acid sequence may or may not be replicated.

[0107] Identifiers in an identifier library may be configured to include one or more common primer binding sites. One or more binding sites may be located at the edges of each identifier or woven across the entire identifier. Primer binding sites allow identifier library-specific primer pairs or universal primer pairs to bind to the identifier and amplify it. All identifiers within the identifier library, or all identifiers in one or more identifier libraries, may be replicated multiple times over multiple PCR cycles. Traditional PCR may be used to replicate identifiers, and identifiers may be replicated exponentially with each PCR cycle. The number of identifier replicas may increase exponentially with each PCR cycle. Linear PCR may be used to replicate identifiers, and identifiers may be replicated linearly with each PCR cycle. The number of identifier replicas may increase linearly with each PCR cycle. Identifiers may be ligated to a circular vector prior to PCR amplification. The circular vector may contain barcodes at each end of the identifier insertion site. PCR primers for identifier amplification may be designed to be primed to the vector such that the barcode edges are included along with the identifier of the amplification product. During amplification, identifiers containing uncorrelated barcodes at each edge may be duplicated due to recombination between identifiers. Uncorrelated barcodes may be detected during identifier reading. Identifiers containing uncorrelated barcodes may be considered false positives and may be ignored during the information decoding process. Refer to Chemical Methods Section D.

[0108] Information can be encoded by assigning each bit of information to a unique nucleic acid molecule. For example, three sample sets (X, Y, and Z) each containing two nucleic acid sequences can be assembled into eight unique nucleic acid molecules to encode 8 bits of data.

[0109] N1 = X1Y1Z1

[0110] N2 = X1Y1Z2

[0111] N3 = X1Y2Z1

[0112] N4 = X1Y2Z2

[0113] N5 = X2Y1Z1

[0114] N6 = X2Y1Z2

[0115] N7 = X2Y2Z1

[0116] N8 = X2Y2Z2

[0117] Then, each bit of the string can be assigned to a corresponding nucleic acid molecule (e.g., N1 can specify the first bit, N2 can specify the second bit, N3 can specify the third bit, etc.). The entire bit string can be assigned to a combination of nucleic acid molecules that are included in a combination or pool of nucleic acid molecules, where the nucleic acid molecule corresponding to the bit value of '1' is included. For example, in UTF-8 coding, the character 'K' can be represented by the 8-bit string code 01001011, which can be encoded as the presence of four nucleic acid molecules (e.g., X1Y1Z2, X2Y1Z1, X2Y2Z1, and X2Y2Z2 in the preceding example).

[0118] Information can be accessed through sequencing or hybridization analysis. For example, primers or probes can be designed to bind to a common region or a barcode region of a nucleic acid sequence. This can enable the amplification of any region of the nucleic acid molecule. The amplified product can be read by analyzing the sequence of the amplified product or through hybridization analysis. In the above example encoding the letter 'K', if the first part of the data is of interest, a primer specific to the barcode region of the X1 nucleic acid sequence and a primer binding to the common region of the Z set can be used to amplify the nucleic acid molecule. This can return the sequence Y1Z2, which can encode 0100. Substrings of the corresponding data can be accessed by further amplifying the nucleic acid molecule using a primer binding to the barcode region of the Y1 nucleic acid sequence and a primer binding to the common sequence of the Z set. This can return the Z2 nucleic acid sequence encoding the substring 01. Alternatively, data can be accessed by checking for the presence of a specific nucleic acid sequence without sequencing. For example, amplification using a primer specific to a Y2 barcode can generate an amplification product for the Y2 barcode but may not generate an amplification product for the Y1 barcode. The presence of a Y2 amplification product can signal a bit value of '1'. Alternatively, the absence of a Y2 amplification product can signal a bit value of '0'.

[0119] PCR-based methods can be used to access and replicate data from identifiers or nucleic acid sample pools. By utilizing common primer binding sites next to identifiers in a pool or hyper-pool, nucleic acids containing information can be easily replicated. Alternatively, other nucleic acid amplification approaches, such as isothermal amplification, can also be used to easily replicate data from sample pools or hyper-pools (e.g., identifier libraries). For nucleic acid amplification, refer to Chemical Methods Section D. If the sample contains a hyper-pool, a specific subset of information (e.g., all nucleic acids associated with a specific barcode) can be accessed and retrieved by using a primer that binds to the specific barcode on one edge of the identifier in the forward direction, and another primer that binds to the common sequence on the opposite edge of the identifier in the reverse direction. Various reading methods can be used to extract information from encoded nucleic acids; for example, microarrays (or any type of fluorescence hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can be further used to read encoded sequences and read digitally encoded data by extension.

[0120] Accessing information stored in nucleic acid molecules (e.g., identifiers) may be performed by selectively removing a portion of non-target identifiers from an identifier library or identifier pool, or, for example, by selectively removing all identifiers from an identifier library from a pool of multiple identifier libraries. As used herein, "access" and "query" may be used interchangeably. Data access may also be performed by selectively capturing target identifiers from an identifier library or identifier pool. Targeted identifiers may correspond to data of interest within larger information. A pool of identifiers may include supplementary nucleic acid molecules. Supplementary nucleic acid molecules may include metadata for encoded information or may be used to encode or mask identifiers corresponding to information. Supplementary nucleic acid molecules may or may not be extracted while accessing target identifiers. FIGS. 17a–17c schematically illustrate an overview of an exemplary method for accessing a portion of information stored in a nucleic acid sequence by accessing multiple specific identifiers from a larger number of identifiers. FIG. 17a illustrates an exemplary method using a polymerase chain reaction, an affinity-tagged probe, and a degradation-targeting probe to access identifiers containing specific components. For PCR-based access, an identifier pool (e.g., an identifier library) may contain identifiers having a common sequence at each end, a variable sequence at each end, or either a common sequence or a variable sequence at each end. The common sequence or the variable sequence may be a primer binding site. One or more primers may bind to the common or variable regions of the identifier edges. Identifiers to which primers are bound may be amplified by PCR. The number of amplified identifiers may be much greater than the number of unamplified identifiers. Amplified identifiers may be identified during reading.Identifiers from an identifier library may include sequences on one or both ends that are distinct from the library, so a single library may be selectively accessed from two or more groups or pools of identifier libraries.

[0121] For affinity-tag based access, in the case of a process that may be referred to as nucleic acid capture, the components constituting the pool identifier may share complementarity with one or more probes. One or more probes may bind to or hybridize with the identifier to be accessed. The probe may include an affinity tag. The affinity tag may be captured on a solid-phase substrate, e.g., a membrane, a well, a column, or a bead. When a bead is used as the solid-phase substrate, the affinity tag may bind to the bead to form a complex comprising the bead, at least one probe, and at least one identifier. The bead may be a magnet and may collect and isolate the identifier to be accessed together with the magnet. The identifier may be removed from the bead under denaturing conditions before reading. Alternatively, or additionally, the bead may collect non-target identifiers and wash them into a separate container to separate them from the rest of the pool to be read. When a column is used, the affinity tag may be bound to the column. The identifier to be accessed may be bound to the column for capture. Column boundary identifiers may be eluted from or denatured from the column prior to reading. Alternatively, non-target identifiers may be optionally targeted to the column, while target identifiers may flow through the column. Identifiers bound to a solid-phase substrate may be removed from the solid-phase substrate by exposing them to conditions such as acid, base, oxidation, reduction, heat, light, metal ion catalysis, substitution, or removal chemistry, or by enzymatic cleavage. In certain embodiments, the identifier to be accessed may be attached to a solid support via a cleavable linkage moiety. For example, the solid-phase substrate may be functionalized to provide a cleavable linker for covalent attachment to the target identifier. The linker moiety may be six or more atoms long. In some embodiments, the cleavable linker may be a TOPS (two oligonucleotides per synthetic) linker, an amino linker, a chemically cleavable linker, or an optically cleavable linker.Accessing targeted identifiers may involve applying one or more probes to the identifier pool simultaneously or applying one or more probes to the identifier pool sequentially. For nucleic acid capture, refer to Chemical Methods Section F.

[0122] In the case of degradation-based access, the components constituting the identifiers in the pool may share complementarity with one or more degradation-targeting probes. The probes may bind to or hybridize with individual components of the identifiers. The probes may be targets for degradation enzymes, such as endonucleases. For example, one or more identifier libraries may be combined. A set of probes may hybridize with one of the identifier libraries. A set of probes may contain RNA, and the RNA may guide a Cas9 enzyme. The Cas9 enzyme may be introduced into one or more identifier libraries. Identifiers hybridized with probes may be degraded by the Cas9 enzyme. Identifiers to be accessed may not be degraded by the degradation enzyme. In another example, the identifier may be single-stranded, and the identifier library may be combined with single-strand-specific endonuclease(s), such as S1 nuclease, which selectively degrades identifiers that are not accessed. Identifiers to be accessed may be hybridized with a complementary set of identifiers to protect them from degradation by single-strand-specific endonuclease(s). The identifier to be accessed can be separated from the degradation products through size selection, such as size-selective chromatography (e.g., agarose gel electrophoresis). Alternatively, or additionally, the undegraded identifier can be selectively amplified (e.g., using PCR) so that the degradation products are not amplified. Since the undegraded identifier hybridizes to each end of the undegraded identifier, it can be amplified using primers that do not hybridize to each end of the degraded or cleaved identifier.

[0123] FIG. 17b illustrates an exemplary method of using a polymerase chain reaction to perform 'OR' or 'AND' operations to access identifiers containing multiple components. For example, if two forward primers combine individual sets of identifiers on the left end, 'OR' amplification for the combination of these sets of identifiers can be achieved by using the two forward primers together in a multiplex PCR reaction having a reverse primer that combines all identifiers on the right end. In another example, if one forward primer combines with a set of identifiers on the left end and one reverse primer combines with a set of identifiers on the right end, 'AND' amplification for the intersection of the two sets of identifiers can be achieved by using the forward primer and the reverse primer together as a primer pair in the PCR reaction.

[0124] FIG. 17c illustrates an exemplary method of using affinity tags to perform an 'OR' or 'AND' operation to access an identifier containing multiple components. For example, if an affinity probe 'P1' captures all identifiers having component 'C1' and another affinity probe 'P2' captures all identifiers having component 'C2', a set of all identifiers having C1 or C2 can be captured by using P1 and P2 simultaneously (corresponding to the 'OR' operation). In another example using the same components and probes, a set of all identifiers having C1 and C2 can be captured by using P1 and P2 sequentially (corresponding to the 'AND' operation).

[0125] A method for reading information stored in a nucleic acid sequence

[0126] In another aspect, the present disclosure provides a method for reading information encoded in a nucleic acid sequence. A method for reading information encoded in a nucleic acid sequence may include: (a) providing an identifier library; (b) identifying an identifier present in the identifier library; (c) generating a string of symbols from an identifier present in the identifier library; and (d) compiling information from the string of symbols. The identifier library may include a subset of a plurality of identifiers from a combination space. Each individual identifier of the subset of identifiers may correspond to an individual symbol of the string of symbols. An identifier may include one or more components. A component may include a nucleic acid sequence.

[0127] Information may be recorded in one or more identifier libraries as described elsewhere in this specification. Identifiers may be configured using methods described elsewhere in this specification. Stored data may be replicated and accessed using any method described elsewhere in this specification.

[0128] An identifier may contain information regarding the location of an encoded symbol, the value of an encoded symbol, or both the location and the value of an encoded symbol. An identifier may contain information related to the location of an encoded symbol, and whether an identifier exists in an identifier library may indicate the value of the symbol. If an identifier exists in the identifier library, it may indicate the value of the first symbol of the binary string (e.g., the first bit value), and if an identifier does not exist in the identifier library, it may indicate the value of the second symbol of the binary string (e.g., the second bit value). In a binary system, basing the bit value on the presence of an identifier in the identifier library can reduce the number of identifiers assembled and reduce write time. For example, the presence of an identifier may indicate a bit value '1' at a mapped location, and the absence of an identifier may indicate a bit value '0' at a mapped location.

[0129] Generating a symbol for information (e.g., a bit value) may include identifying the presence or absence of an identifier to which the symbol (e.g., a bit) can be mapped or encoded. Determining the presence or absence of an identifier may include sequencing the current identifier or using a hybridization array to detect the presence of the identifier. In the example, decoding and reading the encoded sequence may be performed using a sequencing platform. Examples of sequencing platforms, the entirety of which are incorporated herein by reference, are U.S. Patent Application No. 14 / 465,685 filed August 21, 2014 and U.S. Patent Publication No. 2014-0371100 A1 published December 18, 2014, titled "METHOD OF NUCLEIC ACID AMPLIFICATION"; U.S. Patent Application No. 13 / 886,234 filed May 2, 2013 and U.S. Patent Publication No. 2013-0231254 A1 published September 5, 2013, titled "METHOD OF NUCLEIC ACID AMPLIFICATION"; and U.S. Patent Application No. 12 / 400,593 filed March 9, 2009 and U.S. Patent No. US 2009-0253141 published October 8, 2009. It is described in the title of invention A1, "METHODS AND APPARATUSES FOR ANALYZING POLYNUCLEOTIDE SEQUENCES".

[0130] In one example, decoding nucleic acid encoding data can be achieved by base-wise sequencing of nucleic acid strands, e.g., Illumina® sequencing, or by using sequencing techniques indicating the presence or absence of specific nucleic acid sequences, such as fragmentation analysis by capillary electrophoresis. Sequencing may employ the use of reversible terminators. Sequencing may employ the use of natural or non-natural (e.g., engineered) nucleotides or nucleotide analogs. Alternatively or additionally, decoding of nucleic acid sequences may be performed using various analytical techniques, including but not limited to any method of generating optical, electrochemical, or chemical signals. Various sequencing methods, non-limiting examples, may be used, such as polymerase chain reaction (PCR), digital PCR, Sanger sequencing, high-throughput sequencing, synthetic-specific sequencing, single molecule sequencing, ligation-specific sequencing, RNA-Seq (Illumina), next-generation sequencing, digital gene expression (Helicos), clonal single microarray (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, or large-scale parallel sequencing.

[0131] Various reading methods can be used to extract information from encoded nucleic acids. For example, microarrays (or all types of fluorescence hybridization), digital PCR, quantitative PCR (qPCR), and various sequencing platforms can be additionally used to read encoded sequences and further read digitally encoded data.

[0132] The identifier library may additionally include supplementary nucleic acid sequences that provide metadata about the information, encode or mask the information, or provide metadata and mask the information. The supplementary nucleic acid may be identified simultaneously with the identifier identification. Alternatively, the supplementary nucleic acid may be identified before or after the identifier identification. For example, the supplementary nucleic acid is not identified while reading the encoded information. The supplementary nucleic acid sequence may not be distinguishable from the identifier. An identifier index or key may be used to distinguish the identifier from the supplementary nucleic acid molecule.

[0133] The efficiency of data encoding and decoding can be increased by using fewer nucleic acid molecules through recoding the input bit string. For example, if an input string with a high-occurrence '111' substring that can be mapped to three nucleic acid molecules (e.g., identifiers) by the encoding method is received, it can be recorded into a '000' substring that can be mapped to a null set of nucleic acid molecules. An alternative input substring of '000' can also be recoded into '111'. This recording method can reduce the total amount of nucleic acid molecules used to encode data because the number of 'l's in the dataset can be reduced. In this example, the overall size of the dataset may be increased to accommodate a codebook specifying new mapping instructions. Another way to improve encoding and decoding efficiency is to recode input strings to reduce variable length. For example, '111' can be recoded to '00', which can reduce the size of the dataset and reduce the number '1' within the dataset.

[0134] The speed and efficiency of decoding nucleic acid-encoded data can be controlled (e.g., increased) by specifically designing identifiers for detectability. For example, a nucleic acid sequence designed for detectability (e.g., an identifier) ​​may include a nucleic acid sequence containing the majority of nucleotides that are easier to call and detect based on optical, electrochemical, chemical, or physical properties. The engineered nucleic acid sequence may be single-stranded or double-stranded. The engineered nucleic acid sequence may include synthetic or non-natural nucleotides that improve the detectable properties of the nucleic acid sequence. The engineered nucleic acid sequence may include all natural nucleotides, all synthetic or non-natural nucleotides, or a combination of natural, synthetic, and non-natural nucleotides. Synthetic nucleotides may include nucleotide analogs such as peptide nucleic acids, lock nucleic acids, glycol nucleic acids, and threose nucleic acids. Non-natural nucleotides may include dNaM, an artificial nucleoside containing a 3-methoxy-2-naphthal group, and d5SICS, an artificial nucleoside containing a 6-methylisoquinoline-1-thion-2-yl group. The engineered nucleic acid sequence may be designed for a single enhancement property, such as enhanced optical properties, or the engineered nucleic acid sequence may be designed for multiple enhancement properties, such as enhanced optical and electrochemical properties or enhanced optical and chemical properties. Refer to Section H of the Chemical Methods for DNA Design.

[0135] The engineered nucleic acid sequence may include reactive natural, synthetic, and non-natural nucleotides that do not improve the optical, electrochemical, chemical, or physical properties of the nucleic acid sequence. The reactive components of the nucleic acid sequence may enable the addition of chemical moiety that imparts improved properties to the nucleic acid sequence. Each nucleic acid sequence may include a single chemical moiety or multiple chemical moietys. Exemplary chemical moietys may include, but are not limited to, fluorescent moiety, chemiluminescent moiety, acidic or basic moiety, hydrophobic or hydrophilic moiety, and moiety that alters the oxidation state or reactivity of the nucleic acid sequence.

[0136] A sequencing platform may be specifically designed to decode and read information encoded in nucleic acid sequences. The sequencing platform may be dedicated to sequencing single or double-stranded nucleic acid molecules. The sequencing platform may decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of the entire nucleic acid sequence (e.g., components) contained within a nucleic acid molecule (e.g., an identifier). The sequencing platform may include the use of promiscuous reagents, increased read lengths, and the addition of detectable chemical parts to detect specific nucleic acid sequences. Using more promiscuous reagents during sequencing can enable faster base calling, thereby increasing read efficiency and consequently reducing sequencing time. The use of increased read lengths may allow longer sequences of the encoded nucleic acid to be decoded per read. The addition of detectable chemical moiety tags may enable the detection of the presence or absence of the nucleic acid sequence based on the presence or absence of the chemical part. For example, each nucleic acid sequence encoding a bit of information can be tagged with a chemical moiety that generates a unique optical, electrochemical, or chemical signal. The presence or absence of such unique optical, electrochemical, or chemical signal may be indicated by a '0' or '1' bit value. A nucleic acid sequence may contain a single chemical moiety or multiple chemical moietys. Chemical parts may be added to the nucleic acid sequence before the nucleic acid sequence is used to encode data. Alternatively, or additionally, chemical moietys may be added to the nucleic acid sequence after encoding the data, but before decoding the data. Chemical part tags may be added directly to the nucleic acid sequence, or the nucleic acid sequence may contain synthetic or non-natural nucleotide anchors and chemical moiety tags may be added to those anchors.

[0137] Unique codes may be applied to minimize or detect encoding and decoding errors. Encoding and decoding errors may occur due to false negatives (e.g., nucleic acid molecules or identifiers not included in random sampling). An example of an error detection code may be a checksum sequence that calculates the number of identifiers in a set of possible consecutive identifiers included in an identifier library. While reading the identifier library, the checksum may indicate the number of identifiers expected to be retrieved from a set of consecutive identifiers, and identifiers may continue to be sampled for reading until the expected count is met. In some embodiments, a checksum sequence may be included for every consecutive set of R identifiers, where R is of equal size or greater than 1, 2, 5, 10, 50, 100, 200, 500, or 1000, or less than 1000, 500, 200, 100, 50, 10, 5, or 2. The smaller the value of R, the better the error detection performance. In some embodiments, the checksum may be a supplementary nucleic acid sequence. For example, a set containing seven nucleic acid sequences (e.g., components) may be divided into two groups: nucleic acid sequences for forming an identifier by multiplication (components X1-X3 of layer X and Y1-Y3 of layer Y) and nucleic acid sequences for the supplementary checksum (X4-X7 and Y4-Y7). The checksum sequence (X4-X7) may indicate whether zero, one, two, or three sequences of layer X are assembled with each member of layer Y. Alternatively, the checksum sequence (Y4-Y7) may indicate whether zero, one, two, or three sequences of layer Y are assembled with each member of layer X. In this example, the original identifier library with identifiers {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3} can be supplemented to include checksums to become the following pool: {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3, X1Y6, X2Y7, X3Y4, X6Y1, X5Y2, X6Y3}. The checksum sequences can also be used for error correction.For example, if the above dataset contains no X1Y1 but includes X1Y6 and X6Y1, it becomes possible to infer that the X1Y1 nucleic acid molecule is not present in the dataset. A checksum sequence may indicate whether an identifier is missing from a sample of the identifier library or from an accessed portion of the identifier library. If a checksum sequence is missing, it can be amplified and / or isolated through an access method such as PCR or affinity-tagged probe hybridization. In some embodiments, the checksum may not be a supplementary nucleic acid sequence. The checksum may be coded directly into the information to be represented as an identifier.

[0138] For example, in a multiplication method, constructing an identifier as a palindrome using palindrome pairs of components rather than a single component can reduce noise in data encoding and decoding. Then, component pairs from different layers can be assembled together in a palindrome manner (e.g., YXY instead of XY for components X and Y). This palindrome method can be extended to a larger number of layers (e.g., ZYXYZ instead of XYZ) and can detect incorrect cross-reactions between identifiers.

[0139] Adding an excess (e.g., a massive excess) of supplementary nucleic acid sequences to an identifier can prevent sequencing from recovering the encoded identifier. Before decoding the information, the identifier can be enhanced from the supplementary nucleic acid sequences. For example, the identifier can be enhanced by a nucleic acid amplification reaction using primers specific to the identifier's end. Alternatively, or additionally, information can be decoded without enhancing the sample pool through sequencing using specific primers (e.g., synthetic sequencing). In both decoding methods, it can be difficult to enhance or decode the information if a decoding key is unavailable or the identifier configuration is unknown. Alternative approaches, such as using affinity tag-based probes, may also be employed.

[0140] Binary sequence data encoding system

[0141] A system for encoding digital information into nucleic acid (e.g., DNA) may include a system, method, and apparatus for converting files and data (e.g., raw data, compressed zip files, integer data, and other forms of data) into bytes and encoding said bytes into segments or sequences of nucleic acid, generally, DNA, or a combination thereof.

[0142] In one embodiment, the present disclosure provides a system for encoding binary sequence data using nucleic acids. The system for encoding binary sequence data using nucleic acids may include an apparatus and one or more computer processors. The apparatus may be configured to form an identifier library. One or more computer processors may be programmed individually or collectively to (i) convert information into a string of symbols, (ii) map the string of symbols to a plurality of identifiers, and (iii) form an identifier library comprising at least a subset of the plurality of identifiers. An individual identifier of the plurality of identifiers may correspond to an individual symbol among the string of symbols. An individual identifier among the plurality of identifiers may include one or more components. An individual component among the one or more components may include a nucleic acid sequence.

[0143] In another aspect, the present disclosure provides a system for reading binary sequence data using nucleic acids. A system for reading binary sequence data using nucleic acids may include a database and one or more computer processors. The database may store an identifier library that encodes information. One or more computer processors may be programmed individually or collectively to (i) identify identifiers in the identifier library, (ii) generate a plurality of symbols from the identifiers identified in (i), and (iii) compile information from the plurality of identifiers. The identifier library may include a subset of a plurality of identifiers. Each individual identifier of the plurality of identifiers may correspond to an individual symbol in a string of symbols. An identifier may include one or more components. A component may include a nucleic acid sequence.

[0144] A non-limiting embodiment of a method for using a system to encode digital data may include a step of receiving digital information in the form of a byte stream. A step of parsing the byte stream into individual bytes, mapping bit positions within the bytes using a nucleic acid index (or identifier rank), and encoding a sequence corresponding to a bit value 1 or a bit value 0 as an identifier. A step of retrieving digital data may include sequencing a nucleic acid pool containing nucleic acid samples or nucleic acid sequences (e.g., identifiers) mapped to one or more bits, checking whether the identifier is present in the nucleic acid pool by referring to the identifier rank, and decoding position and bit-value information for each sequence into bytes containing a sequence of digital information.

[0145] A system for encoding, writing, copying, accessing, reading, and decoding information encoded and written in nucleic acid molecules may be a single integrated device or multiple devices configured to perform one or more of the aforementioned operations. A system for encoding and writing information in nucleic acid molecules (e.g., identifiers) may include a device and one or more computer processors. One or more computer processors may be programmed to parse the information into symbol strings (e.g., bit strings). The computer processors may generate identifier rankings. The computer processors may classify symbols into two or more categories. One category may contain symbols indicating that a corresponding identifier exists in an identifier library, and another category may contain symbols indicating that a corresponding identifier does not exist in an identifier library. The computer processors may instruct the device to assemble an identifier corresponding to the symbol to be displayed if an identifier exists in the identifier library.

[0146] The device may include multiple regions, sections, or partitions. Reagents and components for assembling identifiers may be stored in one or more regions, sections, or partitions of the device. Layers may be stored in separate regions of a section of the device. Layers may include one or more unique components. Components within one layer may be unique to components within another layer. Regions or sections may include containers, and partitions may include wells. Each layer may be stored in a separate container or partition. Each reagent or nucleic acid sequence may be stored in a separate container or partition. Alternatively, or additionally, reagents may be combined to form a master mix for constructing identifiers. The device may transfer reagents, components, and templates from one section of the device to be combined with another section. The device may provide conditions for completing the assembly reaction. For example, the device may provide heating, stirring, and reaction progress detection functions. The constructed identifier may be directed to undergo one or more subsequent reactions to add a barcode, common sequence, variable sequence, or tag to one or more ends of the identifier. The identifier may then be transferred to a region or partition to create an identifier library. One or more identifier libraries may be stored in each area, section, or individual partition of the device. The device may transfer fluids (e.g., reagents, components, molds) using pressure, vacuum, or suction.

[0147] The identifier library may be stored in a device or moved to a separate database. The database may contain one or more identifier libraries. The database may provide conditions for the long-term storage of the identifier library (e.g., conditions to reduce the degradation of the identifier). The identifier library may be stored in powder, liquid, or solid form. For more stable storage, an aqueous solution of the identifier may be freeze-dried (see Chemical Methods Section G for details on freeze-drying). Alternatively, the identifier may be stored in an oxygen-free environment (e.g., anaerobic storage conditions). The database may provide protection against UV radiation, temperature reduction (e.g., refrigeration or freezing), and degrading chemicals and enzymes. Before being transferred to the database, the identifier library may be freeze-dried or frozen. The identifier library may contain EDTA (ethylenediaminetetraacetic acid) to inactivate nucleases and / or buffers to maintain the stability of nucleic acid molecules.

[0148] The database may be connected to, contain, or be separated from a device that records information to identifiers, copies information, accesses information, or reads information. A portion of the identifier library may be removed from the database before copying, accessing, or reading. The device that copies information from the database may be the same or different from the device that records information. The device that copies information may extract a subsample of the identifier library from the device and combine said subsample with reagents and components to amplify part or all of the identifier library. The device may control the temperature, pressure, and stirring of the amplification reaction. The device may include partitions, and one or more amplification reactions may occur in the partition containing the identifier library. The device may copy two or more identifier pools at a time.

[0149] The copied identifier may be transferred from the copying device to the access device. The access device may be the same device as the copying device. The access device may include separate regions, sections, or partitions. The access device may have one or more columns, bead reservoirs, or magnetic regions for separating identifiers bound to affinity tags (see Chemical Methods Section F regarding nucleic acid capture). Alternatively, or additionally, the access device may have one or more size selection units. The size selection units may include agarose gel electrophoresis or any other method for size selection of nucleic acid molecules (see Chemical Methods Section E for details on nucleic acid size selection). Copying and extraction may be performed in the same region of the device or in different regions of the device (see Chemical Methods Section D for nucleic acid amplification).

[0150] The accessed data may be read from the same device, or the accessed data may be transmitted to another device. The reading device may include a detection unit for detecting and identifying an identifier. The detection unit may be part of a sequencer, a hybridization array, or other unit for identifying the presence or absence of an identifier. The sequencing platform may be specifically designed to decode and read information encoded as a nucleic acid sequence. The sequencing platform may be dedicated to sequencing single or double-stranded nucleic acid molecules. The sequencing platform may decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of the entire nucleic acid sequence (e.g., components) contained within the nucleic acid molecule (e.g., identifier). Alternatively, the sequencing platform may be a system such as Illumina® sequencing or fragmentation analysis by capillary electrophoresis. Alternatively or additionally, decoding of nucleic acid sequences may be performed using various analytical techniques implemented by the device, including, but not limited to, all methods of generating optical, electrochemical, or chemical signals.

[0151] Information storage within nucleic acid molecules can have various applications, including, but not limited to, long-term information storage, sensitive information storage, and medical information storage. For example, an individual's medical information (e.g., medical history and records) can be stored in nucleic acid molecules and delivered to the individual. The information can be stored outside the body (e.g., in wearable devices) or inside the body (e.g., in subcutaneous capsules). When a patient is admitted to a doctor's office or hospital, a sample can be taken from the device or capsule and the information can be decoded using a nucleic acid sequencer. Storing medical records individually as nucleic acid molecules can provide an alternative to computer and cloud-based storage systems. Storing individual medical records as nucleic acid molecules may reduce the frequency or instances of medical records being hacked. The nucleic acid molecules used for capsule-based storage of medical records can be derived from the human genome sequence. The use of the human genome sequence can reduce the immunogenicity of the nucleic acid sequence in the event of capsule failure or leakage.

[0152] computer system

[0153] The present disclosure provides a computer system programmed to implement the method of the present disclosure. FIG. 19 illustrates a computer system (1901) programmed or otherwise configured to encode digital information into a nucleic acid sequence and / or read (e.g., decode) information derived from a nucleic acid sequence. The computer system (1901) may control various aspects of the encoding and decoding procedures of the present disclosure, such as bit value and bit position information for a given bit or byte from an encoded bitstream or bytestream, for example.

[0154] A computer system (1901) includes a central processing unit (CPU, also referred to as a “processor” and “computer processor”) (1905), which may be a single-core or multi-core processor, or a plurality of processors for parallel processing. The computer system (1901) also includes memory or memory location (1910) for communication (e.g., random access memory, read-only memory, flash memory), an electronic storage device (1915) (e.g., a hard disk), a communication interface (1920) for communicating with one or more other systems (e.g., a network adapter), and a peripheral device (1925), e.g., a cache, other memory, data storage, and / or an electronic display adapter. The memory (1910), storage unit (1915), interface (1920), and peripheral device (1925) communicate with the CPU (1905) via a communication bus (solid line), such as a motherboard. The storage unit (1915) may be a data storage unit (or data repository) for storing data. A computer system (1901) can be operably connected to a computer network (“network”) (1930) with the help of a communication interface (1920). The network (1930) may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, the network (1930) is a communication and / or data network. The network (1930) may include one or more computer servers capable of enabling distributed computing, such as cloud computing. In some cases, the network (1930) may implement a peer-to-peer network with the help of the computer system (1901), which may allow a device connected to the computer system (1901) to act as a client or a server.

[0155] The CPU (1905) may execute a series of machine-readable instructions that may be implemented as a program or software. The instructions may be stored in a memory location such as memory (1910). The instructions may be passed to the CPU (1905), and said instructions may subsequently program or configure the CPU (1905) to implement the method of the present disclosure. Examples of operations performed by the CPU (1905) may include fetch, decode, execute, and writeback.

[0156] The CPU (1905) may be part of a circuit, for example, an integrated circuit. One or more other components of the system (1901) may be included in the circuit. In some cases, the circuit is an ASIC (Application-Specific Integrated Circuit).

[0157] The storage unit (1915) can store files, such as drivers, libraries, and stored programs. The storage unit (1915) can store user data, such as user preferences and user programs. In some cases, the computer system (1901) may include one or more additional data storage devices located outside the computer system (1901), such as being located on a remote server that communicates with the computer system (1901) via an intranet or the internet.

[0158] The computer system (1901) may communicate with one or more remote computer systems through a network (1930). For example, the computer system (1901) may communicate with the user's remote computer system or other devices and / or machines available to the user in the process of analyzing data encoded or decoded as a nucleic acid sequence (e.g., a sequencer or other systems for chemically determining the order of nitrogen bases in a nucleic acid sequence). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. The user may access the computer system (1901) through the network (1930).

[0159] The method described herein may be implemented through machine (e.g., computer processor) executable code stored in an electronic storage location of a computer system (1901), such as memory (1910) or an electronic storage device (1915). Machine executable code or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor (1905). In some cases, the code may be retrieved from the storage unit (1915) and stored in memory (1910) for immediate access by the processor (1905). In some situations, the electronic storage unit (1915) may be excluded, and machine executable instructions are stored in memory (1910).

[0160] The code can be pre-compiled and configured for use with a machine equipped with a processor tuned to execute the code, or it can be compiled at runtime. The code can be provided in a programming language that allows one to choose to execute it either as pre-compiled or as compiled.

[0161] Aspects of the systems and methods provided herein, such as the computer system (1901), may be implemented through programming. Various aspects of the technology may generally be considered as “products” or “articles” in the form of related data that are delivered or implemented in machine (or processor) executable code and / or machine-readable media types. Machine-executable code may be stored in memory (e.g., read-only memory, random access memory, flash memory) or electronic storage devices such as hard disks. Media of the “storage” type may include some or all of the type memory of a computer, processor, etc., or various semiconductor memory, tape drives, disk drives, etc., which may provide non-transient storage at any time for software programming. All or part of the software may be delivered from time to time via the Internet or various other communication networks. For example, through such communication, software may be loaded from one computer or processor to another computer or processor, for example, from a management server or host computer to a computer platform of an application server. Accordingly, another type of media that may contain software elements includes optical, electric, and electromagnetic waves, such as those used through physical interfaces between local devices, wired and optical wired networks, and various wireless links. Physical elements that transmit these waves, such as wired and wireless links, optical links, etc., may also be considered as media containing software. As used herein, terms such as “readable media” of a computer or machine, unless limited to non-transient, tangible “storage” media, mean any medium that participates in providing instructions to a processor for execution.

[0162] Accordingly, machine-readable media, such as computer-executable code, may take various forms including, but not limited to, tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device, such as any computer(s) that may be used to implement the database, etc., illustrated in the drawings. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, copper wires including wires that constitute buses within a computer system, and optical fibers. Carrier transmission media may take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communication. Thus, common forms of computer-readable media include floppy disks, flexible disks, hard disks, magnetic tapes, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch card paper, etc. Includes tape, other physical storage media with a hole pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, other memory chips or cartridges, carriers transmitting data or instructions, cables or link waves transmitting such carriers, or other media in which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in delivering one or more sequences of one or more instructions to a processor for execution.

[0163] A computer system (1901) may include or communicate with an electronic display (1935) comprising a user interface (UI) (1940) for providing bits, bytes, and bit streams that are encoded or read by a machine or computer system that encodes or decodes sequence output data, e.g., chromatographs, sequences, and nucleic acids, raw data, files, and compressed or decompressed zip files into DNA stored data. Examples of the UI include, but are not limited to, graphical user interfaces (GUI) and web-based user interfaces.

[0164] The method and system of the present disclosure may be implemented through one or more algorithms. The algorithm may be implemented through software when executed by a central processing unit (1905). For example, the algorithm may be used with a DNA index and raw data or zip file compressed or decompressed data to determine a customized method for coding digital information from raw data or zip file compressed data before encoding the digital information.

[0165] Chemical Methods Section

[0166] A. Overlapping Extension PCR (OEPCR) Assembly

[0167] In OEPCR, components are assembled in a reaction involving a polymerase and dNTPs (deoxynucleotide triphosphates containing dATP, dTTP, dCTP, dGTP, or variants or analogs thereof). The components may be single-stranded or double-stranded nucleic acids. Components to be assembled adjacent to each other may have homology between their complementary 3' ends, complementary 5' ends, or between the 5' end of one component and the 3' end of an adjacent component. These end regions, referred to as "hybridization regions," are intended to facilitate the formation of hybridized splicings between components during OEPCR, where the 3' end of one input component (or its complement) hybridizes to the 3' end of an intended adjacent component (or its complement). Subsequently, a double-stranded product assembled by polymerase extension may be formed. This product can be assembled into more components through subsequent hybridization and extension. Figure 7 illustrates an exemplary schematic diagram of OEPCR for assembling three nucleic acids.

[0168] In some embodiments, OEPCR may involve a cycle between three temperatures: melting temperature, annealing temperature, and extension temperature. The melting temperature is intended not only to convert double-stranded nucleic acids into single-stranded nucleic acids but also to eliminate the formation of secondary structures or hybridizations within or between components. Generally, the melting temperature is high, above 95 degrees Celsius. In some embodiments, the melting temperature may be at least 96, 97, 98, 99, 100, 101, 102, 103, 104, or 105 degrees Celsius. In other embodiments, the melting temperature may be up to 95, 94, 93, 92, 91, or 90 degrees Celsius. While higher melting temperatures may enhance the dissociation of nucleic acids and their secondary structures, they may also lead to side effects such as degradation of nucleic acids or polymerases. The melting temperature can be applied to the reaction for at least 1, 2, 3, 4, 5 seconds or longer, for example, 30 seconds, 1 minute, 2 minutes, or 3 minutes.

[0169] The annealing temperature is intended to promote hybridization formation between the 3' ends of the intended adjacent components (or their complements). In some embodiments, the annealing temperature may correspond to the calculated melting temperature of the intended hybridized nucleic acid formation. In other embodiments, the annealing temperature is 10 times the melting temperature. It may be within the above range. In some embodiments, the annealing temperature may be 25, 30, 50, 55, 60, 65, or 70 degrees Celsius or higher. The melting temperature may vary depending on the order of the intended hybridization regions between the components. Longer hybridization regions have higher melting temperatures, and hybridization regions with higher guanine or cytosine nucleotide content may have higher melting temperatures. Thus, it may be possible to design components for an OEPCR reaction intended to be optimally assembled at a specific annealing temperature. The annealing temperature may be applied to the reaction for at least 1 second, 5 seconds, 10 seconds, 15 seconds, 20 seconds, 25 seconds, or 30 seconds or longer.

[0170] The extension temperature is intended to initiate and promote the extension of the hybridized 3' end nucleic acid chain catalyzed by one or more polymerases. In some embodiments, the extension temperature may be set to a temperature at which the polymerase functions optimally in terms of nucleic acid binding strength, extension rate, extension stability, or fidelity. In some embodiments, the extension temperature may be at least 30, 40, 50, 60, or 70 degrees Celsius. The annealing temperature may be applied to the reaction for at least 1 second, 5 seconds, 10 seconds, 15 seconds, 20 seconds, 25 seconds, 30 seconds, 40 seconds, 50 seconds, or 60 seconds. The recommended extension time may be about 15 to 45 seconds per kilobase of the expected extension.

[0171] In some embodiments of OEPCR, the annealing temperature and the extension temperature may be the same. Thus, a two-step temperature cycle may be used instead of a three-step temperature cycle. Examples of combined annealing and extension temperatures include 60, 65, or 72 degrees Celsius.

[0172] In some embodiments, OEPCR may be performed in a single temperature cycle. Such embodiments may involve the intended assembly of only two components. In other embodiments, OEPCR may be performed in multiple temperature cycles. All specific nucleic acids of OEPCR may be assembled to at most one other nucleic acid in a single cycle. This is because assembly (or extension or elongation) occurs only at the 3' end of the nucleic acid, and each nucleic acid has only one 3' end. Therefore, multiple temperature cycles may be required to assemble multiple components. For example, three temperature cycles may be required to assemble four components. Five temperature cycles may be required to assemble six components. Nine temperature cycles may be required to assemble ten components. In some embodiments, using more temperature cycles than the minimum required may increase assembly efficiency. For example, using four temperature cycles to assemble two components may produce more product than using only one temperature cycle. This is because the hybridization and expansion of components are statistical events that occur in a portion of the total number of components in each cycle. Therefore, the total proportion of assembled components can increase as the number of cycles increases.

[0173] In addition to temperature cycling considerations, the design of nucleic acid sequences in OEPCR can affect the assembly efficiency of each other. Nucleic acids with long hybridization regions can hybridize more efficiently at a given annealing temperature compared to nucleic acids with short hybridization regions. This is because longer hybridization products contain a greater number of stable base pairs and are therefore potentially more stable overall than shorter hybridization products. Hybridization regions can have a base length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.

[0174] Hybridization regions with high guanine or cytosine content can hybridize more efficiently at a given temperature than hybridization regions with low guanine or cytosine content. This is because guanine forms a more stable base pair with cytosine than adenine forms with thymine. Hybridization regions can have a guanine or cytosine content (also called GC content) between 0% and 100%.

[0175] In addition to hybridization region length and GC content, there are other aspects of nucleic acid sequence design that can affect the efficiency of OEPCR. For example, the formation of undesirable secondary structures within components can interfere with the ability to form hybridization products with intended adjacent components. These secondary structures may include hairpin loops. The types of possible secondary structures for nucleic acids and their stability (e.g., melting temperature) can be predicted based on the sequence. Design space search algorithms can be used to determine nucleic acid sequences that meet appropriate length and GC content criteria for efficient OEPCR while avoiding sequences with potentially repressive secondary structures. Design space search algorithms may include genetic algorithms, heuristic search algorithms, metaheuristic search strategies such as contraindication search, branch-and-bound search algorithms, dynamic programming-based algorithms, restricted combinatorial optimization algorithms, gradient descent-based algorithms, random search algorithms, or combinations thereof.

[0176] Similarly, the formation of homodimers (nucleic acid molecules that hybridize with other nucleic acid molecules of the same sequence) and unwanted heterodimers (nucleic acid sequences that hybridize with other nucleic acid sequences, excluding the intended assembly partner) can interfere with OEPCR. Similar to secondary structures within nucleic acids, the formation of homodimers and heterodimers can be predicted and accounted for during nucleic acid design using computational methods and design space search algorithms.

[0177] Longer nucleic acid sequences or higher GC content can increase the formation of unwanted secondary structures, homomers, and heteromers through OEPCR. Therefore, in some embodiments, the use of shorter nucleic acid sequences or lower GC content may result in higher assembly efficiency. These design principles may contradict design strategies that use long hybridization regions or high GC content for more efficient assembly. Accordingly, in some embodiments, OEPCR may be optimized by using a long hybridization region with high GC content and a short non-hybridization region with low GC content. The total length of the nucleic acid may be at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases or more. In some embodiments, there may be an optimal length and an optimal GC content for the hybridization region of the nucleic acid at which assembly efficiency is optimized.

[0178] In OEPCR reactions, a larger number of individual nucleic acids can hinder expected assembly efficiency. This is because a larger number of distinct nucleic acid sequences can generate a higher probability of undesirable molecular interactions, particularly in heterodimeric forms. Therefore, in some embodiments of OEPCR that assemble multiple components, nucleic acid sequence constraints may be stricter for efficient assembly.

[0179] Primers for amplifying the expected final assembled product may be included in the OEPCR reaction. The OEPCR reaction may then be performed with more temperature cycles to improve the yield of the assembled product by not only generating more assemblies between the components but also exponentially amplifying the entire assembled product using the conventional PCR method (see Chemical Methods Section D).

[0180] Additives may be included in the OEPCR reaction to improve assembly efficiency. For example, betaine, dimethyl sulfoxide (DMSO), nonionic detergent, formamide, magnesium, bovine serum albumin (BSA), or a combination thereof may be added. The additive content (weight per volume) may be at least 0%, 1%, 5%, 10%, 20% or more.

[0181] Various polymerases can be used for OEPCR. Polymerases can occur naturally or be synthesized. An exemplary polymerase is Φ29 polymerase or its derivatives. In some cases, a transcriptionase or ligase (i.e., an enzyme that catalyzes bond formation) is used in conjunction with or as an alternative to the polymerase to construct a new nucleic acid sequence. Examples of polymerases include DNA polymerase, RNA polymerase, heat-stable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, platinum Taq polymerase. There are Tbr polymerase, Phusion polymerase, KAPA polymerase, Q5 polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3' to 5' exonuclease activity, and their modifications, modified products, and derivatives. Different polymerases can function stably and optimally at different temperatures. In addition, different polymerases have different characteristics. For example, some polymerases, such as Phusion polymerase, can exhibit 3' to 5' exonuclease activity, which can contribute to higher fidelity during nucleic acid elongation. Some polymerases can replace the major sequence during elongation, while others can degrade it or stop elongation.Some polymerases, such as Taq, include an adenine base at the 3' end of a nucleic acid sequence. This process is called A-tailing, and adding an adenine base can inhibit OEPCR because it interferes with the designed 3' complementarity between the intended adjacent components.

[0182] OEPCR is also called polymerase chain assembly (or PCA).

[0183] B. Ligation Assembly

[0184] In ligation assembly, separate nucleic acids are assembled in a reaction involving one or more ligase enzymes and additional cofactors. Cofactors include adenosine triphosphate (ATP), dithiothreitol (DTT), or magnesium ions (Mg 2+ ...may be included. During ligation, the 3'-end of one nucleic acid strand is covalently linked to the 5'-end of another nucleic acid strand to form an assembled nucleic acid. The components of the ligation reaction can be blunt-terminated double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), or partially hybridized single-stranded DNA. Strategies for assembling the ends of nucleic acids can be used to improve the efficiency of the ligase reaction by increasing the frequency of viable substrates for the ligase enzyme. While blunt-terminated dsDNA molecules tend to form a hydrophobic stack on which the ligase enzyme can act, a more successful strategy for assembling nucleic acids may be to use nucleic acid components with 5' or 3' single-strand overhangs that are complementary to the overhang of the component to be assembled. In the latter case, a more stable nucleic acid double-strand can be formed due to base-base hybridization.

[0185] When an overhang strand is present at one end of a double-stranded nucleic acid, the other strand on the same end may be referred to as a "cavity." The cavity and the protrusion together form a "sticky end," also known as a "cohesive end." The sticky end may be a 3' overhang and a 5' cavity, or a 5' overhang and a 3' cavity. The sticky end between two intended adjacent components can be designed to be complementary, with the overhangs of the two sticky ends hybridizing so that each overhang ends directly adjacent to the beginning of the cavity of the other component. This forms a "nick" (double-stranded DNA break) that can be "sealed" (covalently bonded via a phosphodiester bond) by the action of a ligase. An example schematic of sticky end ligation for assembling three nucleic acids is shown in Fig. 8. The nick of one strand, the other strand, or both may be sealed. Thermodynamically, since the top and bottom strands of a molecule forming the sticky end can move between a bonded and a dissociated state, the sticky end may be a transient formation. However, when a gap along one strand of the sticky end double strand between the two components is sealed, the corresponding covalent bond remains intact even if the member of the opposite strand is separated. The connected strand can then serve as a template to form a gap that can be bonded to and sealed once again by the intended adjacent member of the opposite strand.

[0186] Sticky ends can be produced by cleaving dsDNA with one or more endonucleases. Endonucleases (also called restriction enzymes) can leave sticky ends by targeting specific sites (also called restriction sites) at one or both ends of a dsDNA molecule to produce staggered cleavages (sometimes called digestion). For restriction digestion, refer to Chemical Methods Section C. Digestion can leave palindromic overhangs (overhangs containing sequences that are their own anti-complement). In that case, two components digested by the same endonucleases can form complementary sticky ends that can be assembled with a ligase. If the endonucleases and ligases are compatible, digestion and ligation can occur together in the same reaction. The reaction can take place at a uniform temperature, such as 4, 10, 16, 25, or 37 degrees Celsius. Alternatively, the reaction can be cycled between various temperatures, such as between 16 and 37 degrees Celsius. By cycling between various temperatures, digestion and ligation can proceed at their respective optimal temperatures during different parts of the cycle.

[0187] It may be beneficial to perform digestion and ligation as separate reactions. For example, this occurs when the desired ligase and the desired endonuclease function optimally under different conditions. Alternatively, for example, if the ligated product forms a new restriction site for the endonuclease. In such cases, it may be better to perform ligation separately after restriction digestion, and it may be advantageous to remove the restriction enzyme before ligation. The nucleic acid can be separated from the enzyme via phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. While multiple endonucleases may be used in the same reaction, care must be taken to ensure they function under similar reaction conditions without interfering with each other. Using two endonucleases can create orthogonal (non-complementary) sticky ends at both ends of the dsDNA component.

[0188] Endonuclease digestion will leave a sticky end along with a phosphorylated 5' end. Ligase can function only at the phosphorylated 5' end and cannot function at the unphosphorylated 5' end. Therefore, an intermediate 5' phosphorylation step between digestion and ligation may not be necessary. Digested dsDNA components with palindromic overhangs on the sticky end may self-ligate. To prevent self-ligation, it may be beneficial to dephosphorylate the dsDNA components before ligation.

[0189] Multiple endonucleases can target different restriction sites but leave compatible overhangs (overhangs that are inversely complementary to each other). The ligation products of the sticky ends generated by two of these endonucleases can produce an assembled product that does not contain a restriction site for either endonuclease at the ligation site. These endonucleases form the basis of assembly methods, such as biobrick assembly, which can programmatically assemble multiple components using only two endonucleases by performing repetitive digestion-ligation cycles. Figure 20 illustrates an example of a digestion-ligation cycle using endonucleases BamHI and BglII, which have compatible overhangs.

[0190] In some embodiments, the endonuclease used to generate sticky ends may be an IIS-type restriction enzyme. Since these enzymes cleave a fixed number of bases in a specific direction at the restriction site, the sequence of the overhang they generate can be customized. The overhang sequence does not need to be palindromic. The same type of IIS restriction enzyme may be used to generate multiple different sticky ends in the same or multiple reactions. Furthermore, one or multiple types of IIS restriction enzymes may be used to generate components with compatible overhangs in the same or multiple reactions. The ligation site between two sticky ends generated by a type IIS restriction enzyme may be designed not to form a new restriction site. Additionally, the type IIS restriction enzyme site may be located on dsDNA so that the restriction enzyme cleaves its own restriction site when generating components with sticky ends. Thus, the linkage product between multiple components generated by the IIS restriction enzyme type may not contain a restriction site.

[0191] Type IIS restriction enzymes can be mixed with ligases in the reaction to perform component digestion and ligation together. The reaction temperature can be cycled between two or more values ​​to promote optimal digestion and ligation. For example, digestion can be performed optimally at 37 degrees Celsius, and ligation can be performed optimally at 16 degrees Celsius. More generally, the reaction can be cycled between temperature values ​​of at least 0, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or 65 degrees Celsius or higher. The combined digestion and ligation reaction can be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more components. Examples of assembly reactions that utilize type IIS restriction enzymes to generate sticky ends include Golden Gate Assembly (also known as Golden Gate Cloning) or Modular Cloning (also known as MoClo).

[0192] In some embodiments of ligation, an exonuclease may be used to create a component having a sticky end. A 3' exonuclease may be used to chew back the 3' end of dsDNA to create a 5' overhang. Similarly, a 5' exonuclease may be used to chew back the 5' end of dsDNA to create a 3' overhang. Different exonucleases may have different characteristics. For example, depending on whether they act on ssDNA, whether they act on phosphorylated or non-phosphorylated 5' ends, whether they can start from a nick, or whether they can start activity in the 5' cavity, 3' cavity, 5' overhang, or 3' overhang, exonucleases may have different nuclease activity directions (5' to 3' or 3' to 5'). Various types of exonucleases include lambda exonucleases and RecJ f, exonuclease III, exonuclease I, exonuclease T, exonuclease V, exonuclease VIII, exonuclease VII, nuclease BAL_31, T5 exonuclease, and T7 exonuclease are included.

[0193] Exonucleases can be used in reactions with ligases to assemble multiple components. The reaction can occur at a fixed temperature or in a cycle between multiple temperatures, each of which is ideal for either the ligase or the exonuclease. Polymerases can be involved in assembly reactions with ligases and 5'-to-3' exonucleases. The components of these reactions can be designed so that components intended to be assembled adjacent to each other share homologous sequences at their edges. For example, component X to be assembled with component Y may have a 3' edge sequence of the 5'-z-3' form, and component Y may have a 5' arbitral sequence of the 5'-z-3' form, where z is any nucleic acid sequence. We refer to homologous edge sequences of the form 'Gibson overlap'. 5' exonuclease chews back the 5' ends of dsDNA components with Gibson overlaps, generating compatible 3' overhangs that hybridize with each other. The hybridized 3' ends are then extended by the action of the polymerase to the ends of the template components or to the point where the extended 3' overhang of one component meets the 5' cavity of an adjacent component, forming a nick that can be sealed by a ligase. This assembly reaction, in which a polymerase, a ligase, and an exonuclease are used together, is often referred to as "Gibson assembly." Gibson assembly can be performed by using T5 exonuclease, Phusion polymerase, and Taq ligase, and incubating the reaction mixture at 50 degrees Celsius. In this case, using Taq, a thermophilic ligase, allows the reaction to proceed at 50 degrees Celsius, a temperature suitable for all three types of enzymes in the reaction.

[0194] The term "Gibson assembly" generally refers to any assembly reaction involving polymerases, ligases, and exonucleases. Gibson assembly may be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more components. Gibson assembly may occur as a single-step, isothermal reaction, or a multi-step reaction involving one or more temperature incubations. For example, Gibson assembly may occur at temperatures of at least 30, 40, 50, 60, or 70 degrees. The incubation time for Gibson assembly may be at least 1, 5, 10, 20, 40, or 80 minutes.

[0195] The Gibson assembly reaction can occur optimally when the Gibson overlap between the intended adjacent components is of a specific length and possesses sequence features, such as sequences that avoid undesirable hybridization events like hairpins, homodimers, or unwanted heterodimers. Generally, at least 20 bases of Gibson overlap are recommended. However, the length of the Gibson overlap can be at least 1, 2, 3, 5, 10, 20, 30, 40, 50, 60, 100, or more bases. The GC content of the Gibson overlap can be between 0% and 100%.

[0196] Gibson assembly is generally described by 5' exonucleases, but the reaction can also occur with 3' exonucleases. When 3' exonucleases chew back the 3' ends of dsDNA components, polymerases extend the 3' ends to interfere with glycolysis. This dynamic process can continue until the 5' overhangs (generated by exonucleases) of the two components (which share a Gibson overlap) hybridize and polymerases extend the 3' end of one component far enough to meet the 5' end of the adjacent component, thus leaving a nick that can be sealed by ligase.

[0197] In some embodiments of the ligation, a component having a sticky end can be produced synthetically rather than enzymatically by mixing together two single-stranded nucleic acids or oligos that do not share complete complementarity. For example, two oligos, oligo X and oligo Y, can be designed to fully hybridize along a continuous complementary base string that forms a substring of a larger base string comprising one or both oligos. This complementary base string is referred to as the "index region." When the index region occupies the entire oligo X and only the 5' end of oligo Y, the oligos together form a component having a sticky end with a blunt end on one side and a 3' overhang of oligo Y on the other (Fig. 21a). When the index region occupies the entire oligo X and only the 3' end of oligo Y, the oligos together form a component having a sticky end with a blunt end on one side and a 5' overhang of oligo Y on the other (Fig. 21b). When the index region occupies the entire oligo X and does not occupy either end of the oligo Y (meaning the index region is embedded in the middle of the oligo Y), the oligos together form a component having a sticky end with a 3' overhang from the oligo Y on one side and a 5' overhang from the oligo Y on the other side (Fig. 21c). When the index region occupies only the 5' end of the oligo X and the 5' end of the oligo Y, the oligos together form a component having a sticky end with a 3' overhang from the oligo Y on one side and a 3' overhang from the oligo X on the other side (Fig. 21d). When the index region occupies only the 3' end of the oligo X and the 3' end of the oligo Y, the oligos together form a component having a sticky end with a 5' overhang from the oligo Y on one side and a 5' overhang from the oligo X on the other side (Fig. 21e). In the above examples, the sequence of the overhang is defined by the oligo sequence outside the index region.These overhang sequences can be referred to as hybridization regions because they are regions where components are hybridized for ligation.

[0198] In adhesive end ligation, the index and hybridization regions of the oligo can be designed to facilitate the proper assembly of components. Components with long overhangs can hybridize more efficiently with each other at a given annealing temperature compared to components with short overhangs. The overhangs can have a length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30 or more bases.

[0199] A component having an overhang containing a high guanine or cytosine content can hybridize more efficiently to a complementary component at a given temperature than a component having an overhang containing a low guanine or cytosine content. This is because guanine forms a more stable base pair with cytosine than adenine forms with thymine. The guanine or cytosine content of the overhang (also called GC content) can be between 0% and 100%.

[0200] Similar to overhang sequences, the GC content and length of the oligo index region can also affect ligation efficiency. This is because stably binding the top and bottom strands of each component allows the adhesive end components to be assembled more efficiently. Therefore, the index region can be designed with higher GC content, longer sequences, and other features that promote higher melting temperatures. However, regarding both the index region and the overhang sequence, there are more aspects of oligo design that can affect the efficiency of the ligation assembly. For example, if unwanted secondary structures form within a component, the ability to form an assembled product with the intended adjacent components may be hindered. This can be caused by secondary structures in the index region, the overhang sequence, or both. These secondary structures may include hairpin loops. The types of possible secondary structures for the oligo and their stability (e.g., melting temperature) can be predicted based on the sequence. Using a design space search algorithm, it is possible to determine oligo sequences that meet appropriate length and GC content criteria for effective component formation while avoiding sequences with potentially repressive secondary structures. The design space search algorithm may include genetic algorithms, heuristic search algorithms, metaheuristic search strategies such as contraindication search, branch-and-bound search algorithms, dynamic programming-based algorithms, restricted combinatorial optimization algorithms, gradient descent-based algorithms, random search algorithms, or combinations thereof.

[0201] Similarly, the formation of homodimers (oligos that hybridize with oligos of the same sequence) and unwanted heterodimers (oligos that hybridize with oligos other than the intended assembly partner) can interfere with ligation. Similar to secondary structures within a component, the formation of homodimers and heterodimers can be predicted and accounted for during component design using computational methods and design space search algorithms.

[0202] As the oligo sequence length increases or the GC content increases, the formation of unwanted secondary structures, homomers, and heteromers may increase within the ligation reaction. Therefore, in some embodiments, using shorter oligos or lower GC content may result in higher assembly efficiency. This design principle may be contrary to design strategies that use long oligos or high GC content for more efficient assembly. Thus, there may be optimal lengths and optimal GC content for the oligos constituting each component to optimize the efficiency of the ligation assembly. The total length of the oligo used for ligation may be at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases or more. The total GC content of the oligo used for ligation may be between 0% and 100%.

[0203] In addition to adhesive end ligation, ligation between single-stranded nucleic acids may occur using staple (or template or bridge) strands. This method can be referred to as Staple-Strand Ligation (SSL), Template-Designated Ligation (TDL), or Bridge-Strand Ligation. An exemplary schematic diagram of a TDL for assembling three nucleic acids is shown in Fig. 10a. In a TDL, two single-stranded nucleic acids hybridize adjacently to the template to form a nick that can be sealed by a ligase. The same nucleic acid design considerations for adhesive end ligation apply to TDL as well. Stronger hybridization between the template and the intended complementary nucleic acid sequence can lead to increased ligation efficiency. Therefore, sequence features that improve the hybridization stability (or melting temperature) of both sides of the template can enhance ligation efficiency. These features may include longer sequence lengths and higher GC content. The nucleic acid length of the TDL containing the template may be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases or more. The GC content of the nucleic acid containing the template may be between 0% and 100%.

[0204] In TDL, as with sticky end ligation, care can be taken to design component and template sequences that avoid unwanted secondary structures by using nucleic acid structure prediction software that includes sequence space search algorithms. Since the components of TDL can be single-stranded rather than double-stranded, there may be a higher likelihood of unwanted secondary structures (compared to sticky end ligation) occurring due to exposed bases.

[0205] TDL can also be performed using blunt-terminated dsDNA components. In this reaction, for the staple strand to properly ligate two single-stranded nucleic acids, the staple may first need to replace or partially replace the entire single-stranded complement. To facilitate the TDL reaction with the dsDNA component, the dsDNA can be melted by incubating at a high temperature initially. The reaction is then cooled so that the staple strand can be annealed to the appropriate nucleic acid complement. This process can be carried out much more efficiently by using a template at a relatively high concentration compared to the dsDNA component, thus allowing the template to compete with the appropriate full-length ssDNA complement for binding. Once the two ssDNA strands are assembled by the template and ligase, the assembled nucleic acid can serve as a template for the opposite full-length ssDNA complement. Therefore, the ligation between blunt-terminated dsDNA and TDL can be improved through multiple cycles of melting (incubation at higher temperatures) and annealing (incubation at lower temperatures). This process is called the ligase cycle (LCR). The appropriate melting and annealing temperatures depend on the nucleic acid sequence. The melting and annealing temperatures may be at least 4, 10, 20, 20, 30, 40, 50, 60, 70, 80, 90, or 100 degrees Celsius. The number of temperature cycles may be at least 1, 5, 10, 15, 20, 15, 30, or more.

[0206] All ligations can be performed in a fixed-temperature or multi-temperature reaction. The ligation temperature may be at least 0, 4, 10, 20, 20, 30, 40, 50, or 60 degrees Celsius. The optimal temperature for ligase activity may vary depending on the ligase type. Additionally, the rate at which components adjoin or hybridize in the reaction may vary depending on the corresponding nucleic acid sequence. Higher incubation temperatures lead to faster diffusion rates and a higher frequency of transient adjoining or hybridization of components. However, increasing temperature may break hydrogen bonds between base pairs, potentially reducing the stability of adjoining or hybridized component dimers. The optimal temperature for ligation may depend on the number of nucleic acids to be assembled, the sequences of the corresponding nucleic acids, the ligase type, and other factors such as reaction additives. For example, two sticky end components with a 4-base complementary overhang can be assembled faster at 4°C using T4 ligase than at 25°C using T4 ligase. However, two sticky end components with a 25-base complementary overhang can be assembled faster at 2°C using T4 ligase than at 4°C using T4 ligase, and can be faster than ligation using a 4-base overhang at any temperature. In some embodiments of ligation, it may be beneficial to heat the components for annealing and slowly cool them before adding the ligase.

[0207] Ligation can be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleic acids. The ligation incubation time can be up to 30 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 1 hour, or longer. The longer the incubation time, the better the ligation efficiency may be.

[0208] Lagging may require nucleic acids with 5' phosphorylated ends. Nucleic acid components without 5' phosphorylated ends can be phosphorylated by reaction with polynucleotide kinases such as T4 polynucleotide kinase (or T4 PNK). Other cofactors, such as ATP, magnesium ions, or DTT, may be present in the reaction. The polynucleotide kinase reaction may occur at 37 degrees Celsius for 30 minutes. The polynucleotide kinase reaction temperature may be at least 4, 10, 20, 20, 30, 40, 50, or 60 degrees Celsius. The polynucleotide kinase reaction incubation time may be up to 1 minute, 5 minutes, 10 minutes, 20 minutes, 30 minutes, or 60 minutes or more. Alternatively, nucleic acid components may be synthetically designed and manufactured using modified 5' phosphorylation (as opposed to enzymatically). Only nucleic acids assembled at the 5' end may require phosphorylation. For example, the template of TDL may not be phosphorylated because it is not intended to be assembled.

[0209] To improve ligation efficiency, additives may be included in the ligation reaction. For example, dimethyl sulfoxide (DMSO), polyethylene glycol (PEG), 1,2-propanediol (1,2-Prd), glycerol, Tween-20, or a combination thereof may be added. PEG6000 may be a particularly effective ligation enhancer. PEG6000 can increase ligation efficiency by acting as a crowding agent. For example, PEG6000 can form aggregated nodules that occupy space in the ligase reaction solution and bring the ligase and components closer together. The additive content (weight per volume) may be at least 0%, 1%, 5%, 10%, or 20%.

[0210] Various ligases can be used for ligation. Ligases can occur naturally or be synthesized. Examples of ligases include T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, Taq DNA ligase, 9o N TM These include DNA ligase, E. coli DNA ligase, and SplintR DNA ligase. Different ligases can function stably and optimally at various temperatures. For example, Taq DNA ligase is heat-resistant, but T4 DNA ligase is not. Furthermore, different ligases have different characteristics. For example, T4 DNA ligase can ligate blunt-terminated dsDNA, but T7 DNA ligase may not.

[0211] Sequencing adapters can be attached to a nucleic acid library using ligation. For example, ligation can be performed using a common adhesive end or staple at the end of each member of the nucleic acid library. Sequencing adapters can be ligated asymmetrically if the adhesive end or staple at one end of the nucleic acid differs from that at the other end. For example, a forward sequencing adapter can be ligated to one end of a nucleic acid library member, and a reverse sequencing adapter can be ligated to the other end of a nucleic acid library member. Alternatively, adapters can be attached to a blunt-end double-strand nucleic acid library using blunt-end ligation. Fork adapters can be used to asymmetrically attach adapters to a nucleic acid library where each end has the same blunt end or adhesive end (e.g., an A-tail).

[0212] Ligation is thermal inactivation (e.g., 65 It can be inhibited by incubation for more than 20 minutes), addition of a denaturant, or addition of a chelating agent such as EDTA.

[0213] C. Restricted digestion

[0214] Restriction digestion is a reaction in which a restriction endonuclease (or restriction enzyme) recognizes a homologous restriction site of a nucleic acid and subsequently cleaves (or digests) the nucleic acid containing said restriction site. Type I, Type II, Type III, or Type IV restriction enzymes may be used for restriction digestion. Type II restriction enzymes may be the most efficient restriction enzymes for nucleic acid degradation. Type II restriction enzymes can recognize palindromic restriction sites and cleave the nucleic acid within the recognition site. Examples of such restriction enzymes (and their restriction sites) include AatII (GACGTC), AfeI (AGCGCT), ApaI (GGGCCC), DpnI (GATC), EcoRI (GAATTC), and NgeI (GCTAGC). Some restriction enzymes, such as DpnI and AfeI, cleave the central restriction site, leaving a dsDNA product with blunt ends. Other restriction enzymes, such as EcoRI and AatII, cleave the restriction site off-center, leaving sticky ends (or staggered ends) in the dsDNA product. Some restriction enzymes can target discontinuous restriction sites. For example, the restriction enzyme AlwNI recognizes the restriction site CAGNNNCTG, where N can be A, T, C, or G. The length of the restriction site can be at least 2, 4, 6, 8, 10, or more bases.

[0215] Some Type II restriction enzymes cleave nucleic acids outside the restriction site. Enzymes may be subclassified as Type IIS or Type IIG restriction enzymes. These enzymes can recognize non-palindromic restriction sites. Examples of such restriction enzymes include BbsI, which recognizes GAAAC and produces staggered cleavages 2 (same-stranded) and 6 (opposite-stranded) bases further downstream. Another example includes BsaI, which recognizes GGTCTC and produces staggered cleavages 1 (same-stranded) and 5 (opposite-stranded) bases further downstream. These restriction enzymes can be used in Golden Gate Assembly or Modular Cloning (MoClo). Some restriction enzymes, such as BcgI (Type IIG restriction enzyme), can produce staggered cleavages at both ends of the recognition site. Restriction enzymes can cleave nucleic acids by separating at least 1, 5, 10, 15, 20, or more bases at the recognition site. Since the above restriction enzyme can generate a staggered cleavage outside the recognition site, the sequence of the resulting nucleic acid overhang can be arbitrarily designed. This is in contrast to restriction enzymes that generate a staggered cleavage within the recognition site, where the sequence of the resulting nucleic acid overhang binds to the sequence of the restriction site. The nucleic acid overhang generated by restriction digestion may be at least 1, 2, 3, 4, 5, 6, 7, or 8 bases long. The 5' end generated when the restriction enzyme cleaves the nucleic acid contains a phosphate group.

[0216] One or more nucleic acid sequences may be included in the restriction digestion reaction. Similarly, one or more restriction enzymes may be used together in the restriction digestion reaction. The restriction digest may contain additives and cofactors including potassium ions, magnesium ions, sodium ions, BSA, S-adenosyl-L-methionine (SAM), or combinations thereof. The restriction digestion reaction may be incubated at 37°C for 1 hour. The restriction digestion reaction may be incubated at temperatures of 0, 10, 20, 30, 40, 50, or 60°C or higher. The optimal digestion temperature may vary depending on the enzyme. The restriction digestion reaction may be incubated for up to 1 minute, 10 minutes, 30 minutes, 60 minutes, 90 minutes, or 120 minutes or more. Prolonged incubation time may increase digestion.

[0217] D. Nucleic Acid Amplification

[0218] Nucleic acid amplification can be performed via polymerase chain reaction or PCR. In PCR, a starting nucleic acid pool (called a template pool or template) may be combined with a polymerase, primers (short nucleic acid probes), nucleotide triphosphates (e.g., dATP, dTTP, dCTP, dGTP, and their analogs or variants), and additional cofactors and additives, such as betaine, DMSO, and magnesium ions. The template may be a single-stranded or double-stranded nucleic acid. The primer may be a short nucleic acid sequence synthetically constructed to complement and hybridize to the target sequence in the template pool. The primer binds to each identifier nucleic acid sequence containing the target sequence in the template pool, thereby selecting only the identifier nucleic acid sequence containing the target sequence. Generally, there are two primers in a PCR reaction: one to complement the primer binding site on the top strand of the target template, and the other to complement the primer binding site on the bottom strand of the target template downstream of the first binding site. The 5' to 3' directions in which these primers bind to the target must face each other to successfully replicate the nucleic acid sequence between them and amplify it exponentially. "PCR" can typically refer specifically to the above form of reaction, but it may also be used more generally to refer to any nucleic acid amplification reaction.

[0219] In some embodiments, PCR may involve a cycle between three temperatures: melting temperature, annealing temperature, and extension temperature. The melting temperature is intended to convert double-stranded nucleic acids into single-stranded nucleic acids and to eliminate the formation of hybridization products and secondary structures. Generally, the melting temperature is high, above 95 degrees Celsius. In some embodiments, the melting temperature may be at least 96, 97, 98, 99, 100, 101, 102, 103, 104, or 105 degrees Celsius. In other embodiments, the melting temperature may be up to 95, 94, 93, 92, 91, or 90 degrees Celsius. Higher melting temperatures enhance the dissociation of nucleic acids and their secondary structures, but may also cause side effects such as degradation of nucleic acids or polymerase. The melting temperature can be applied to the reaction for at least 1, 2, 3, 4, 5 seconds or longer, for example, 30 seconds, 1 minute, 2 minutes, or 3 minutes. For PCR using complex or long templates, a longer initial melting temperature step may be recommended.

[0220] The annealing temperature is intended to promote hybridization formation between the primer and the target template. In some embodiments, the annealing temperature may correspond to the calculated melting temperature of the primer. In other embodiments, the annealing temperature may be within 10 degrees Celsius of the melting temperature. In some embodiments, the annealing temperature may be 25, 30, 50, 55, 60, 65, or 70 degrees Celsius or higher. The melting temperature may vary depending on the sequence of the primers. Longer primers may have higher melting points, and primers with a higher guanine or cytosine nucleotide content may have higher melting points. Thus, it may be possible to design primers intended to assemble optimally at a specific annealing temperature. The annealing temperature may be applied to the reaction for at least 1 second, 5 seconds, 10 seconds, 15 seconds, 20 seconds, 25 seconds, or 30 seconds or longer. To ensure annealing, the primer concentration may be high or saturated. The primer concentration may be 500 nanomolar (nM). The primer concentration may be up to 1 nM, 10 nM, 100 nM, 1000 nM, or higher.

[0221] The extension temperature is intended to initiate and promote the extension of the 3' end nucleic acid chain of a primer catalyzed by one or more polymerases. In some embodiments, the extension temperature may be set to a temperature at which the polymerase functions optimally in terms of nucleic acid binding strength, extension rate, extension stability, or fidelity. In some embodiments, the extension temperature may be at least 30, 40, 50, 60, or 70 degrees Celsius. The annealing temperature may be applied to the reaction for at least 1 second, 5 seconds, 10 seconds, 15 seconds, 20 seconds, 25 seconds, 30 seconds, 40 seconds, 50 seconds, or 60 seconds. The recommended extension time may be about 15 to 45 seconds per kilobase of the expected extension.

[0222] In some embodiments of PCR, the annealing temperature and the extension temperature may be the same. Therefore, a 2-step temperature cycle may be used instead of a 3-step temperature cycle. Examples of combined annealing and extension temperatures include 60, 65, or 72 degrees Celsius.

[0223] In some embodiments, PCR may be performed in a single temperature cycle. These embodiments may involve converting a targeted single-stranded template nucleic acid into a double-stranded nucleic acid. In other embodiments, PCR may be performed in multiple temperature cycles. If PCR is efficient, the number of target nucleic acid molecules is expected to double with each cycle, resulting in an exponential increase in the number of target nucleic acid templates in the original template pool. The efficiency of PCR may vary. Therefore, the actual proportion of target nucleic acids replicated in each round may be more or less than 100%. Undesirable artifacts, such as mutants and recombinant nucleic acids, may be introduced with each PCR cycle. To reduce such potential damage, high-fidelity and highly processable polymerases may be used. Additionally, a limited number of PCR cycles may be used. PCR may include up to 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, or more cycles.

[0224] In some embodiments, multiple individual target nucleic acid sequences may be amplified together in a single PCR. If each target sequence has a common primer binding site, all nucleic acid sequences may be amplified using the same set of primers. Alternatively, the PCR may include multiple primers intended to target each individual nucleic acid. Such PCR may be referred to as multiplex PCR. The PCR may include up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more individual primers. In PCR using multiple individual nucleic acid targets, each PCR cycle may alter the relative distribution of the target nucleic acids. For example, a uniform distribution may be distorted or become non-uniform. To reduce such potential damage, an optimal polymerase (e.g., one with high fidelity and sequence robustness) and optimal PCR conditions may be used. Factors such as annealing, extended temperature, and time may be optimized. Additionally, a limited number of PCR cycles may be used.

[0225] In some embodiments of PCR, a target sequence can be mutated using a primer having a base mismatch for the target primer binding site within the template. In some embodiments of PCR, a sequence can be attached to a target nucleic acid using a primer having an additional sequence (known as an overhang) at the 5' end. For example, a nucleic acid library for sequencing can be prepared and / or amplified using a primer containing a sequencing adapter at the 5' end. A primer targeting the sequencing adapter can be used to amplify the nucleic acid library to a sufficient enrichment for a specific sequencing technique.

[0226] In some embodiments, linear-PCR (or asymmetric-PCR) is used, where the primers target only one strand of the template (not both). In linear PCR, the nucleic acid replicated in each cycle is not complementary to the primer, so the primer does not bind to it. Therefore, since the primer replicates only the original target template in each cycle, linear (the opposite of exponential) amplification occurs. Amplification in linear PCR is not as fast as conventional (exponential) PCR, but the maximum yield can be higher. Theoretically, the primer concentration in linear PCR may not be a limiting factor that increases the number of cycles and yield, as in conventional PCR. Post-linear exponential PCR (or late-PCR) is a modified version of linear PCR in which particularly high yields may be possible.

[0227] In some embodiments of nucleic acid amplification, the melting, annealing, and extension processes may occur at a single temperature. Such PCR may be referred to as isothermal PCR. Isothermal PCR can utilize temperature-independent methods to separate or replace fully complemented nucleic acid strands for primer binding. Strategies include loop-mediated isothermal amplification, strand substitution amplification, helicase-dependent amplification, and nicking enzyme amplification reactions. Isothermal nucleic acid amplification may occur at temperatures of up to 20, 30, 40, 50, 60, or 70 degrees or higher.

[0228] In some embodiments, PCR may additionally include a fluorescent probe or dye to quantify the amount of nucleic acid in a sample. For example, the dye may be inserted into the double-stranded nucleic acid. An example of such a dye is SYBR Green. The fluorescent probe may also be a nucleic acid sequence attached to a fluorescent unit. The fluorescent unit may be emitted upon hybridization of the probe to the target nucleic acid and subsequent modification from the extension polymerase unit. Examples of such probes include the Taqman probe. These probes may be used in conjunction with PCR and optical measurement tools (for excitation and detection) to quantify the nucleic acid concentration in a sample. This process may be referred to as quantitative PCR (qPCR) or real-time PCR (rtPCR).

[0229] In some embodiments, PCR may be performed on a single-molecule template (a process that can be called single-molecule PCR) rather than a pool of multiple template molecules. For example, emulsion-PCR (ePCR) may be used to encapsulate a single nucleic acid molecule within a droplet in an oil emulsion. The droplet may also contain PCR reagents, and the droplet may be maintained in a temperature-controlled environment capable of the temperature cycling required for PCR. In this way, multiple self-encapsulating PCR reactions can occur simultaneously with high throughput. The stability of the oil emulsion may be enhanced by using a surfactant. The movement of the droplet may be controlled by pressure through microfluidic channels. Microfluidic devices may be used for droplet generation, droplet splitting, droplet merging, material introduction droplet injection, and droplet culture. The droplet size of the oil emulsion may be at least 1 picoliter (pL), 10 pL, 100 pL, 1 nanoliter (nL), 10 nL, 100 nL, or greater.

[0230] In some embodiments, single-molecule PCR may be performed on a solid-phase substrate. Examples include the Illumina solid-phase amplification method or a variation thereof. A template pool may be exposed to a solid-phase substrate, whereby the solid-phase substrate can immobilize the templates at a specific spatial resolution. Bridge amplification can then occur within the spatial proximity of each template, thereby amplifying single molecules on the substrate in a high-throughput manner.

[0231] High-throughput single-molecule PCR can be useful for amplifying different nucleic acid pools that may interfere with each other. For example, if multiple different nucleic acids share a common sequence region, recombination between the nucleic acids may occur along this common region during the PCR reaction, generating new recombinant nucleic acids. Single-molecule PCR prevents these potential amplification errors by compartmentalizing different nucleic acid sequences so they cannot interact. Single-molecule PCR can be particularly useful for preparing nucleic acids for sequencing analysis. Single-molecule PCR mats are also useful for the absolute quantification of multiple targets within a template pool. For example, digital PCR (or dPCR) estimates the number of starting nucleic acid molecules in a sample using the frequency of distinct single-molecule PCR amplification signals.

[0232] In some embodiments of PCR, groups of nucleic acids may be amplified non-discriminately using primers for primer binding sites common to all nucleic acids. For example, primers for primer binding sites are located on all sides of the pool of nucleic acids. Synthetic nucleic acid libraries may be generated or assembled using these common sites for general amplification. However, in some embodiments, PCR may be used to selectively amplify a subset of targeted nucleic acids from the pool, for example, using primers with primer binding sites that appear only in the targeted nucleic acid subset. Synthetic nucleic acid libraries may be generated or assembled such that nucleic acids belonging to a sub-library of potential interest all share a common primer binding site at their respective edges (common within the sub-library but distinct from other sub-libraries) for the selective amplification of sub-libraries from a more comprehensive library. In some embodiments, PCR may be combined with a nucleic acid assembly reaction (e.g., ligation or OEPCR) to selectively amplify fully assembled or potentially fully assembled nucleic acids from partially assembled or misassembled (or unintended or undesirable) byproducts. For example, assembly may involve assembling the nucleic acid with the primer binding sites of each edge sequence such that only the entire assembled nucleic acid product contains the two primer binding sites required for amplification. In the above example, the partially assembled product must not be amplified because it may contain neither of the edge sequences containing primer binding sites, or only one. Similarly, a mis-assembled (or unintended or undesirable) product may contain none, only one, or both of the edge sequences separated by incorrect orientation or incorrect base amount. Therefore, the mis-assembled product must not be amplified or amplified to produce a product of incorrect length.In the latter case, the amplified misassembled product of the wrong length can be separated from the amplified fully assembled product of the correct length through nucleic acid size selection methods such as gel extraction after DNA electrophoresis on an agarose gel (see Chemical Methods Section E).

[0233] PCR may include additives to increase nucleic acid amplification efficiency. For example, betaine, dimethyl sulfoxide (DMSO), nonionic detergents, formamide, magnesium, bovine serum albumin (BSA), or combinations thereof may be added. The additive content (weight per volume) may be at least 0%, 1%, 5%, 10%, 20% or more.

[0234] Various polymerases can be used in PCR. Polymerases can occur naturally or be synthesized. An exemplary polymerase is Φ29 polymerase or its derivatives. In some cases, a transcriptionase or ligase (i.e., an enzyme that catalyzes bond formation) is used in conjunction with or as an alternative to the polymerase to construct a new nucleic acid sequence. Examples of polymerases include DNA polymerase, RNA polymerase, heat-stable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, platinum Taq polymerase. There are Tbr polymerase, Phusion polymerase, KAPA polymerase, Q5 polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3' to 5' exonuclease activity, and their modifications, modified products, and derivatives. Different polymerases can function stably and optimally at different temperatures. In addition, different polymerases have different characteristics. For example, some polymerases, such as Phusion polymerase, can exhibit 3' to 5' exonuclease activity, which can contribute to higher fidelity during nucleic acid elongation. Some polymerases can replace the major sequence during elongation, while others can degrade it or stop elongation.Some polymerases, such as Taq, include an adenine base at the 3' end of the nucleic acid sequence. Additionally, some polymerases can have higher fidelity and progression than others and may be more suitable for PCR applications, such as sequence analysis preparation, where it is important that the amplified nucleic acid yield has minimal mutations and that the distribution of individual nucleic acids remains uniform throughout the amplification.

[0235] E. Select Size

[0236] Nucleic acids of a specific size can be selected from a sample using size selection techniques. In some embodiments, size selection may be performed using gel electrophoresis or chromatography. A liquid sample of nucleic acid may be loaded at one end of a stationary phase or gel (or matrix). A voltage difference may be applied across the gel such that the negative terminal of the gel is the terminal where the nucleic acid sample is loaded and the positive terminal of the gel is the opposite terminal. Since nucleic acids have a negatively charged phosphate backbone, they can move through the gel toward the positive end. The size of the nucleic acid can determine the relative migration speed through the gel. Therefore, nucleic acids of various sizes are degraded in the gel as they migrate. The voltage difference may be 100 V or 120 V. The voltage difference may be up to 50 V, 100 V, 150 V, 200 V, 250 V, or more. The greater the voltage difference, the higher the nucleic acid migration speed and size resolution may be. However, if the voltage difference becomes too large, the nucleic acid or the gel may be damaged. To isolate larger nucleic acids, a larger voltage difference may be recommended. Typical migration times range from 15 to 60 minutes. Migration times can be up to 10, 30, 60, 90, or over 120 minutes. Similar to increasing voltage, longer migration times may improve nucleic acid resolution but may increase nucleic acid damage. Longer migration times may be recommended to isolate larger nucleic acids. For example, a voltage difference of 120 V and a migration time of 30 minutes may be sufficient to isolate a 200-base nucleic acid from a 250-base nucleic acid.

[0237] The characteristics of the gel or matrix can influence the size selection process. Gels generally contain polymeric materials, such as agarose or polyacrylamide, dispersed in a conductive buffer such as TAE (Tris-acetate-EDTA) or TBE (Tris-borate-EDTA). The content (weight per volume) of the material within the gel (e.g., agarose or acrylamide) can be up to 0.5%, 1%, 2%, 3%, 5%, 10%, 15%, 20%, 25%, or higher. Higher content may result in slower migration rates. Higher content may be desirable for isolating smaller nucleic acids. Agarose gels may be better for resolving double-stranded DNA (dsDNA). Polyacrylamide gels may be more suitable for analyzing single-stranded DNA (ssDNA). The desirable gel composition may depend on the nucleic acid type and size, compatibility of additives (e.g., dyes, stains, denaturation solutions, or loading buffers), as well as the expected downstream application (e.g., ligation, PCR, or sequencing after gel extraction). Agarose gels may be simpler to extract than polyacrylamide gels. Although TAE is not as good a conductor as TBE, it may be better for gel extraction because borate (enzyme inhibitor) residues during the extraction process can inhibit downstream enzymatic reactions.

[0238] The gel may additionally contain a denaturing solution such as SDS (sodium dodecyl sulfate) or urea. For example, SDS can be used to denature proteins or to separate nucleic acids from potentially bound proteins. Urea can be used to denature the secondary structure of DNA. For example, urea can convert dsDNA into ssDNA, or urea can convert folded ssDNA (e.g., hairpin) into unfolded ssDNA. A urea-polyacrylamide gel (additionally containing TBE) may be used to accurately separate ssDNA.

[0239] Samples can be incorporated into gels of various types. In some embodiments, the gel may include wells into which samples can be manually loaded. A single gel may have multiple wells for running multiple nucleic acid samples. In other embodiments, the gel may be attached to a microfluidic channel that automatically loads nucleic acid sample(s). Each gel may be located downstream of multiple microfluidic channels, or the gel itself may occupy a separate microfluidic channel. The size of the gel may affect the sensitivity of nucleic acid detection (or visualization). For example, a thin gel or gel inside a microfluidic channel (e.g., a bioanalyzer or tape station) may improve nucleic acid detection sensitivity. The nucleic acid detection step may be critical for selecting and extracting nucleic acid fragments of the correct size.

[0240] A ladder can be loaded onto a gel for nucleic acid size reference. The ladder may contain markers of various sizes that can be compared with a nucleic acid sample. Different ladders may have different size ranges and resolutions. For example, a 50-base ladder may have markers at 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, and 600 bases. The ladder may be useful for detecting and selecting nucleic acids within the 50 to 600 base size range. The ladder may also be used as a standard for estimating the concentration of nucleic acids of various sizes within a sample.

[0241] To facilitate the gel electrophoresis (or chromatography) process, the nucleic acid sample and ladder may be mixed with a loading buffer. The loading buffer may contain dyes and markers to help track nucleic acid migration. The loading buffer may additionally contain a reagent (e.g., glycerol) that is denser than the running buffer (e.g., TAE or TBE) to cause the nucleic acid sample to sink to the bottom of the sample loading well (where it can be immersed in the running buffer). The loading buffer may additionally contain a denaturant such as SDS or urea. The loading buffer may additionally contain reagents to improve the stability of the nucleic acid. For example, the loading buffer may contain EDTA to protect the nucleic acid from nucleases.

[0242] In some embodiments, the gel may contain a dye that binds to nucleic acids and can be used to optically detect nucleic acids of various sizes. The dye may be specific to dsDNA, ssDNA, or both. Different dyes may be compatible with different gel materials. Some dyes may require stimulation by a light source (or electromagnetic waves) to be visualized. The light source may be UV (ultraviolet) or blue light. In some embodiments, the dye may be added to the gel before electrophoresis. In other embodiments, the dye may be added to the gel after electrophoresis. Examples of dyes include EtBr (Ethidium Bromide), SYBR Safe, SYBR Gold, silver dye, or methylene blue. For example, a reliable method for visualizing dsDNA of a specific size may be to use an agarose TAE gel with SYBR Safe or EtBr dyes. For example, a reliable method for visualizing ssDNA of a specific size may be to use a urea-polyacrylamide TBE gel containing methylene blue or silver staining.

[0243] In some embodiments, the movement of nucleic acids through the gel may be induced by methods other than electrophoresis. For example, nucleic acids may be moved through the gel using gravity, centrifugation, vacuum, or pressure to be separated by size.

[0244] Nucleic acids of a specific size can be extracted from the gel using a blade or razor to cut out the gel bands containing the nucleic acids. By using appropriate optical detection techniques and a DNA ladder, cleavage can be accurately detected at specific bands, and nucleic acids that might belong to different, undesirable size bands can be successfully excluded. The gel bands can be incubated with a buffer to lyse, thereby releasing the nucleic acids into the buffer solution. The rate of lysis can be accelerated by heat or physical stirring. Alternatively, the gel bands can be incubated in the buffer for a sufficiently long time to allow DNA to diffuse into the buffer without requiring gel lysis. The buffer can then be separated from the remaining solid gel, for example, by aspiration or centrifugation. The nucleic acids can then be purified from the solution using standard purification or buffer exchange techniques, such as phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. The nucleic acids may also be concentrated at this stage.

[0245] As an alternative to gel ablation, nucleic acids of a specific size can be separated from the gel by flowing down. The moving nucleic acids may be embedded in the gel or pass through branches (or wells) at the ends of the gel. The movement process can be timed or optically monitored so that a sample is collected from the basin when a group of nucleic acids of a specific size enters the basin. Collection can be performed, for example, via aspiration. The nucleic acids from the collected solution can then be purified using standard purification or buffer exchange techniques, such as phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. The nucleic acids may also be concentrated at this stage.

[0246] Other methods for nucleic acid size selection may include mass spectrometry or membrane-based filtration. In some embodiments of membrane-based filtration, nucleic acids pass through a membrane (e.g., a silica membrane) that can preferentially bind to dsDNA, ssDNA, or both. The membrane may be designed to preferentially capture nucleic acids of at least a specific size. For example, the membrane may be designed to filter out nucleic acids consisting of 20, 30, 40, 50, 70, or fewer than 90 bases. The membrane-based size selection technique may not be as strict as gel electrophoresis or chromatography.

[0247] F. Nucleic acid capture

[0248] Nucleic acids tagged with affinity tags can be used as sequence-specific probes for nucleic acid capture. The probe can be designed to complement a target sequence within a nucleic acid pool. Subsequently, the probe can be incubated with the nucleic acid pool and hybridized to the target. The incubation temperature may be lower than the probe's melting point to promote hybridization. The incubation temperature may be 5, 10, 15, 20, or 25 degrees Celsius or lower than the probe's melting point. The hybridized target can be captured on a solid-phase substrate that specifically binds to the affinity tag. The solid-phase substrate may be a membrane, well, column, or bead. Multiple washes can remove all nucleic acids that have not hybridized to the target. Washing may occur at a temperature lower than the probe's melting point to promote stable fixation of the target sequence during washing. The washing temperature may be 5, 10, 15, 20, or 25 degrees Celsius or lower than the probe's melting point. In the final elution step, nucleic acid targets can be recovered from the solid-phase substrate as well as from the affinity-tagged probe. The elution step may occur at a temperature higher than the melting temperature of the probe to facilitate the release of nucleic acid targets into the elution buffer. The elution temperature may be 5, 10, 15, 20, or 25 degrees Celsius higher than the melting temperature of the probe.

[0249] In certain embodiments, an oligonucleotide bound to a solid-phase substrate may be removed from the solid-phase substrate by exposing it to conditions such as, for example, acid, base, oxidation, reduction, heat, light, metal ion catalysis, substitution, or removal chemistry, or by enzymatic cleavage. In certain embodiments, the oligonucleotide may be attached to a solid support via a cleavable linkage moiety. For example, the solid support may be functionalized to provide a cleavable linker for covalent attachment to the targeted oligonucleotide. In some embodiments, the linker moiety may have a length of six or more atoms. In some embodiments, the cleavable linker may be a TOPS (two oligonucleotides per synthesis) linker, an amino linker, or an optically cleavable linker.

[0250] In some embodiments, biotin can be used as an affinity tag immobilized on a solid-phase substrate by streptavidin. Biotinylated oligonucleotides can be designed and manufactured for use as nucleic acid capture probes. Oligonucleotides can be biotinylated at the 5' or 3' ends. They can also be biotinylated within thymine residues. Increased biotin in the oligo can lead to stronger capture of the streptavidin substrate. Biotin at the 3' end of the oligo can block the expansion of the oligo during PCR. The biotin tag can be a variant of standard biotin. For example, biotin variants can be biotin-TEG (triethylene glycol), bibiotin, PC biotin, destiobiotin-TEG, and biotin azide. Bibiotin can increase biotin-streptavidin affinity. Biotin-TEG attaches a biotin group to a nucleic acid separated by a TEG linker. This can prevent biotin from interfering with the function of the nucleic acid probe, for example, hybridization to a target. A nucleic acid biotin linker can also be attached to the probe. The nucleic acid linker may contain a nucleic acid sequence that is not intended to hybridize to a target.

[0251] Biotinylated nucleic acid probes can be designed considering how well they hybridize to a target. Nucleic acid probes designed with a higher melting temperature may hybridize more strongly to the target. Longer nucleic acid probes, as well as probes with a higher GC content, may hybridize more strongly due to the increased melting temperature. The length of the nucleic acid probe can be at least 5, 10, 15, 20, 30, 40, 50, or 100 bases or more. The nucleic acid probe can have a GC content between 0 and 100%. Care must be taken to ensure that the melting temperature of the probe does not exceed the temperature tolerance of the streptavidin substrate. Nucleic acid probes can be designed to prevent inhibitory secondary structures such as hairpins, homomers, and heteromers containing off-target nucleic acids. There may be a trade-off between the probe melting temperature and off-target binding. There may be an optimal probe length and GC content that results in a high melting temperature and low off-target binding. A synthetic nucleic acid library can be designed so that the nucleic acid contains an efficient probe binding site.

[0252] The solid-phase streptavidin substrate may be magnetic beads. Magnetic beads may be immobilized using magnetic strips or plates. The magnetic strips or plates may be in contact with the container to fix the magnetic beads to the container. Conversely, the magnetic strips or plates may be removed from the container to release the magnetic beads into the solution from the container walls. Various bead characteristics may affect the application. The size of the beads may vary. For example, the beads may have a diameter between 1 and 3 micrometers (µm). The diameter of the beads may be up to 1, 2, 3, 4, 5, 10, 15, 20, or greater micrometers. The surface of the beads may be hydrophobic or hydrophilic. The beads may be coated with a blocking protein, for example, BSA. Before use, the beads may be washed or pretreated with additives such as a blocking solution to prevent non-specific binding of nucleic acids.

[0253] Biotinylated probes can be bound to magnetic streptavidin beads before incubation with a nucleic acid sample pool. This process can be referred to as direct capture. Alternatively, biotinylated probes can be incubated with a nucleic acid sample pool before adding magnetic streptavidin beads. This process can be referred to as indirect capture. The indirect capture method can improve the target yield. Short nucleic acid probes may require a shorter time to bind to magnetic beads.

[0254] Optimal incubation of nucleic acid samples and nucleic acid probes may occur at temperatures 1 to 10 degrees Celsius or lower than the probe's melting point. Incubation temperatures may be above 5, 10, 20, 30, 40, 50, 60, 70, and 80 degrees Celsius. The recommended incubation time is 1 hour. Incubation times may be up to 1, 5, 10, 20, 30, 60, 90, 120 minutes, or longer. Longer incubation times may improve capture efficiency. To allow biotin-streptavidin binding, an additional 10-minute incubation may be performed after adding streptavidin beads. This additional time may be up to 1, 5, 10, 20, 30, 60, 90, 120 minutes, or longer. Incubation may take place in a buffer solution containing additives such as sodium ions.

[0255] If the nucleic acid pool consists of single-stranded nucleic acids (as opposed to double-stranded ones), the hybridization of the probe to the target may be enhanced. To prepare an ssDNA pool from a dsDNA pool, linear PCR may need to be performed using a single primer that typically binds to the edges of all nucleic acid sequences in the pool. If the nucleic acid pool is synthetically generated or assembled, this common primer binding site may be included in the synthetic design. The product of the linear PCR will be ssDNA. More starting ssDNA templates for nucleic acid capture can be generated through more cycles of linear PCR. Refer to Section D of the Chemical Methods of PCR.

[0256] After the nucleic acid probe hybridizes to the target and binds to magnetic streptavidin beads, the beads can be fixed by a magnet and multiple washes may occur. Three to five washes may be sufficient to remove non-target nucleic acids, but more or fewer washes may be used. Each incremental wash may further reduce non-target nucleic acids but may also reduce the yield of target nucleic acids. Low incubation temperatures may be used during the washing step to promote proper hybridization of the target nucleic acid to the probe. Temperatures of 60, 50, 40, 30, 20, 10, or 5 degrees or lower may be used. The wash buffer may contain Tris buffer solution containing sodium ions.

[0257] Optimal elution of the hybridized target from the magnetic bead-binding probe may occur at a temperature equal to or higher than the probe's melting temperature. Higher temperatures facilitate the separation of the target and the probe. The elution temperature can be up to 30, 40, 50, 60, 70, 80, or 90 degrees Celsius or higher. The elution incubation time can be up to 1, 2, 5, 10, 30, or 60 minutes or longer. A typical incubation time is about 5 minutes, but longer incubation times can improve yield. The elution buffer can be water or a Tris buffer solution containing additives such as EDTA.

[0258] Nucleic acid capture of a target sequence containing at least one of a set of individual sites can be performed as a single reaction using multiple distinct probes for each of these sites. Nucleic acid capture of a target sequence containing all members of a series of individual sites can be performed as a series of capture reactions, that is, as a single reaction for each individual site using probes for specific sites. Although the target yield may be low after a series of capture reactions, the captured target can subsequently be amplified via PCR. If the nucleic acid library is synthetically designed, the target can be designed using common primer binding sites for PCR.

[0259] Synthetic nucleic acid libraries can be generated or assembled using common probe binding sites for general nucleic acid capture. These common sites can be used to selectively capture fully assembled or potentially fully assembled nucleic acids in an assembly reaction, thereby filtering out partially assembled or misassembled (or unintended or undesirable) byproducts. For example, assembly may involve assembling nucleic acids with probe binding sites on each edge sequence such that only fully assembled nucleic acid products contain the two essential probe binding sites required to pass through a series of two capture reactions using each probe. In the above example, a partially assembled product may contain none of the probe sites or only one, and therefore should not be ultimately captured. Similarly, a misassembled (or unintended or undesirable) product may contain none of the edge sequences or only one. Thus, the misassembled product may not be ultimately captured. To increase strictness, common probe binding sites may be included on each component of the assembly. In a series of subsequent nucleic acid capture reactions using probes for each component, only the fully assembled product (including each component) can be separated from the byproducts of the assembly reaction. Subsequent PCR can enhance target strengthening, and subsequent size selection can improve target strictness.

[0260] In some embodiments, nucleic acid capture may be used to selectively capture a targeted nucleic acid subset from a pool. For example, this is possible by using a probe having a binding site that appears only in said targeted nucleic acid subset. A synthetic nucleic acid library may be generated or assembled such that nucleic acids belonging to a sub-library of potential interest share a common primer binding site (common within the sub-library but distinct from other sub-libraries) for the selective capture of a sub-library from a more comprehensive library.

[0261] G. Freeze-drying

[0262] Freeze-drying is a dehydration process. Both nucleic acids and enzymes can be freeze-dried. Freeze-dried materials may have a longer shelf life. Additives, such as chemical stabilizers, can be used to maintain functional products (e.g., active enzymes) through the freeze-drying process. Disaccharides such as sucrose and trehalose can be used as chemical stabilizers.

[0263] H. DNA Design

[0264] The sequences of nucleic acids (e.g., components) for constructing a synthetic library (e.g., an identifier library) can be designed to avoid complexity in synthesis, sequencing, and assembly. Furthermore, they can be designed to reduce the cost of constructing the synthetic library and improve the shelf life of the synthetic library.

[0265] Nucleic acids can be designed to avoid long strings of homopolymers (or repeating base sequences) that may be difficult to synthesize. Nucleic acids can be designed to avoid homopolymers with lengths of 2, 3, 4, 5, 6, 7 or more. Furthermore, nucleic acids can be designed to prevent the formation of secondary structures, such as hairpin loops, that can interfere with the synthesis process. For example, predictive software can be used to generate nucleic acid sequences that do not form stable secondary structures. Nucleic acids for building synthetic libraries can be designed to be short. Longer nucleic acids are more difficult and costly to synthesize. The longer the nucleic acid, the higher the probability of mutations occurring during synthesis. Nucleic acids (e.g., components) can have up to 5, 10, 15, 20, 25, 30, 40, 50, 60 or more bases.

[0266] Nucleic acids serving as components in an assembly reaction can be designed to facilitate the assembly reaction. For detailed considerations regarding nucleic acid sequences for OEPCR and ligation-based assembly reactions, refer to Chemical Methods Sections A and B. Efficient assembly reactions generally involve hybridization between adjacent components. Sequences can be designed to facilitate these intra-target hybridization events while avoiding potential out-of-target hybridization. Target hybridization can be enhanced using nucleic acid base modifications, such as lock nucleic acids (LNAs). These modified nucleic acids can be used, for example, as staples in staple strand ligation or as sticky ends in sticky strand ligation. Other modified bases that can be used to construct synthetic nucleic acid libraries (or identifier libraries) include 2,6-diaminopurine, 5-bromo dU, deoxyuridine, inverted dT, inverted dideoxy-T, dideoxy-C, 5-methyl dC, deoxynosine, Super T, Super G, or 5-nitroindole. The nucleic acid may contain one or more of the same or different modified bases. Some of the modified bases are natural base analogs with higher melting temperatures (e.g., 5-methyl dC and 2,6-diaminopurine) and may be useful for promoting specific hybridization events in assembly reactions. Some of the modified bases are universal bases capable of binding to all natural bases (e.g., 5-nitroindole) and may be useful for promoting hybridization with nucleic acids that may have variable sequences within the desired binding site. In addition to their beneficial role in assembly reactions, these modified bases may be useful for primers (e.g., for PCR) and probes (e.g., for nucleic acid capture) as they can promote the specific binding of primers and probes to target nucleic acids within the nucleic acid pool. For additional nucleic acid design considerations regarding nucleic acid amplification (or PCR) and nucleic acid capture, refer to Chemical Methods Sections D and F.

[0267] Nucleic acids can be designed to facilitate sequencing. For example, nucleic acids can be designed to prevent common sequencing problems such as secondary structures, homopolymer extensions, repetitive sequences, and sequences with excessively high or low GC content. Errors can occur in specific sequencers or sequencing methods. Nucleic acid sequences (or components) constituting a synthetic library (e.g., an identifier library) can be designed to have specific Hamming distances from each other. In this way, even if base resolution errors occur at a high rate during sequencing, the range of sequences containing errors can still be mapped back to the most likely nucleic acid (or component). Nucleic acid sequences can be designed with Hamming distances of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base mutations. A minimum required distance between designed nucleic acids can also be defined using alternative distance measures from the Hamming distance.

[0268] Some sequencing methods and instruments may require input nucleic acids containing specific sequences, such as adapter sequences or primer binding sites. These sequences may be referred to as "method-specific sequences." A general preparation workflow for the sequencing instruments and methods may include assembling method-specific sequences into a nucleic acid library. However, if it is known in advance that a synthetic nucleic acid library (e.g., an identifier library) will be sequenced using a specific instrument or method, these method-specific sequences may be designed as nucleic acids (e.g., components) containing the library (e.g., the identifier library). For example, a sequencing adapter may be assembled to a member of the synthetic nucleic acid library at the same reaction step as when the member of the synthetic nucleic acid library is assembled from individual nucleic acid components.

[0269] Nucleic acids can be designed to prevent sequences that can promote DNA damage. For example, sequences containing site-specific nuclease sites can be avoided. As another example, UVB (ultraviolet-B) light can inhibit sequencing and PCR by causing adjacent thymines to form pyrimidine dimers. Therefore, if you intend to store a synthetic nucleic acid library in an environment exposed to UVB, it may be advantageous to design the nucleic acid sequence to avoid adjacent thymines (i.e., TT).

[0270] All information contained in the Chemical Methods section is intended to support and enable the techniques, methods, protocols, systems, and processes described herein.

[0271] Method of assembling identifiers from components with azide-alkyne variations

[0272] Identifiers can be generated by ligating two or more nucleic acid components together using chemical and / or biological ligation methods. In some embodiments, chemical ligation methods, such as "click chemistry," may have advantages over biological methods, such as enzymatic ligation.

[0273] Cycloaddition, or CuAAC (Copper-Catalyzed Azide-Alkyne Cycloaddition), is a variation of the Huisgen 1,3-dipolar cycloaddition reaction. In this reaction, an alkyne reacts with an azide group to form a triazole phosphodiester mimic. Current methods utilize Cu(I) ions to enhance the specificity, rate, and yield of this reaction. The reaction can be fast due to some alkynes, which report a reaction completion time of approximately one minute. Reaction times can be 30, 60, 90, 120, 150, or over 180 seconds. The reaction is also robust and may exhibit tolerance over a wide pH range.

[0274] Chemical ligation using click chemistry can occur between two single-stranded nucleic acid components with the help of a template (or staple or splint) oligonucleotide. Alternatively, chemical ligation may occur between double-stranded nucleic acid components where there is a common complementary overhang (or sticky end). Chemical ligation using click chemistry can be used to construct identifiers according to the aforementioned multiplication method (Fig. 6), permutation method (Fig. 11), MchooseK method (Fig. 12), partition method (Fig. 13), or unrestricted string method (Fig. 14).

[0275] To ligate components using click chemistry, one component must have one or more alkyne groups and the other component must have one or more azide groups. Any modification may be located at the 5' or 3' end of one nucleic acid component, provided that the complementary modification is located on the adjacent component so that the 3' end of one component ligates the 5' end of the other component.

[0276] Various different types of alkyne-azide bonds can be used in click chemistry. Alkyne-azide linkages compatible with molecular biology methods such as PCR may be particularly suitable for identifier generation. If a specific identifier pool contains one or more alkyne-azide bonds, the identifiers can be replicated in their natural form (including phosphodiester bonds between bases) using PCR.

[0277] How to assemble identifiers from multipart components

[0278] The components constituting an identifier can be divided into two or more parts with different functions. For example, each component may consist of two parts: a long part for hybridization to a nucleic acid probe for data access, and a short part for sequencing readout. Since the two parts can be separated and intended to be assembled into the identifiers at each edge, the final identifier product has two functionally different regions. One region on one side is for chemical access, and the other region is for sequencing.

[0279] FIG. 22 provides an exemplary schematic diagram of this concept for the ligation assembly of the adhesive ends of an identifier, where the components of each layer are assembled together according to a multiplication method. The first layer nucleates the identifier assembly process with combined two-part components, and subsequent layers consist of unconnected two-part components assembled into identifiers at both edges. The symbols on the adhesive ends indicate their respective order. The adhesive ends with different symbols are orthogonal. An asterisk next to a symbol indicates a reverse complement. For example, 'a' and 'a*' are reverse complements to each other, so they hybridize during ligation to form a product.

[0280] How to build an identifier using the default editor

[0281] A base editor may be used to construct a new identifier by programmatically mutating a base located at a specific locus within the parent identifier. In one embodiment, the base editor may be a dCas9 protein fused to cytidine deaminase that converts cystosine (C) to uracil (U). The parent identifier may be designed with several orthogonal target loci for the binding of guide RNA (gRNA). The target loci may contain one or more cytosines within the active range of the dCas9-deaminase bound to the locus. The active range may be 1, 2, 3, 4, 5, or 6 or more bases within the locus. Subsequent incubation of the parent identifier with the dCas9-deaminase and a subset of gRNA for the specific locus may result in one or more mutations from cystosine to uracil at each target locus. Additionally, since DNA polymerase recognizes uracil as thymine, performing PCR on a mutated identifier may result in a complementary mutation (from guanine to adenine). A parent identifier with N orthogonal target loci can be programmatically converted into 2N individual daughter identifier sequences by applying dCas9-deaminase and various subsets of N gRNAs (each targeting an individual parental locus). Thus, the combination space of possible identifiers constructed in this scheme can store N bits of information for N gRNA inputs.

[0282] In some embodiments, any given target locus of the parent sequence may contain cytosines targeted on both the upper and lower strands to promote increased mutation efficiency. Additionally, for efficient gRNA targeting to occur, each locus must be adjacent to a PAM site. However, the PAM sequence may vary depending on the use of various engineered Cas9 variants.

[0283] The dCas9-deaminase fusion may include a linker sequence between the two fused proteins. The optimal linker length for efficient targeting mutations may be 16 amino acids. The linker length may be at least 0, 1, 5, 10, 15, 20, or 25 amino acids. One of several cytidine deaminases may be used. Examples of cytidine deaminases include APOBEC1, AID, CDA1, or APOBEC3G. Active Cas9 nicase may be used instead of dCas9, but DNA repair enzymes may also need to be included in the identifier construction reaction.

[0284] In another embodiment of constructing an identifier using a base editor, adenine deaminase fused to dCas9 (opposite to or in addition to cytidine deaminase fused to dCas9) can be used to mutate adenine to inosine at a defined locus of a parent identifier accessible by gRNA. Inosine is interpreted as guanine by DNA polymerase. Therefore, PCR of the base-edited locus can result in a thymine mutation complementary to cytosine on the opposite strand.

[0285] Method to delete information stored in DNA

[0286] The ability to reliably erase (or delete) stored data using nucleic acids can be beneficial for security, privacy, and regulatory reasons. Data erasure may involve breaking covalent bonds within nucleic acids, irreversibly modifying nucleic acids to interfere with sequencing capabilities, encapsulating or adsorbing them in an irreversible manner, or adding more nucleic acids or other materials to render the original collection of nucleic acids unreadable or unreadable. These methods may be performed in a selective or non-selective manner. The selection process may be separate from the deletion process. For example, a library of identifiers may be started, and a subset of identifiers to be deleted may be pulled down using sequence-specific probes. As another example, the purification of selected identifiers based on size or mass-to-charge ratio may be performed in conjunction with other selective or non-selective deletion methods.

[0287] Selective methods for deleting nucleic acids from a library include the use of sequence-specific probes to pull down a subset of nucleic acids for deletion, the use of CRISPR-based methods to cleave selected nucleic acids containing one or more target sequences, and purification techniques for selecting nucleic acids based on size or mass-to-charge ratio.

[0288] Non-selective methods for removing information-encoding nucleic acids from a library include sonication, autoclaving, treatment with bleach, bases, acids, ethidium bromide, or other DNA modifying agents, irradiation (e.g., using ultraviolet light), combustion, and non-specific nuclease degradation (in vitro or in vivo), e.g., using DNase I. Other methods may be used to obfuscate, hide, or physically protect the access or sequencing of nucleic acids. Methods may include encapsulation, dilution, the addition of random nucleic acids to obfuscate the original nucleic acid, and the addition of other agents to prevent downstream sequencing of the nucleic acid. In one embodiment, data stored in the nucleic acid may be obfuscated due to amplification by an error-prone polymerase, e.g., a polymerase lacking a correction function.

[0289] For data stored in nucleic acids with a defined value period, it may be advantageous to use a method to automatically delete the data at a specific point in time. For example, data may be scheduled to be deleted after a required regulatory period. As another example, data may be scheduled to be deleted if it is in transit and does not reach its destination in time. In one embodiment, the planned deletion of nucleic acids may involve the use of a degrading agent that acts at a defined rate or immediately at a specific point in time. In another embodiment, the planned deletion of nucleic acids may involve the use of nucleic acid capsules or protective cases that degrade over time. In another embodiment, nucleic acids may be stored at various temperatures or environments to promote degradation at various rates. For example, high temperatures or high humidity are used to increase the degradation rate. In another embodiment, nucleic acids may be converted into less stable forms for faster degradation. For example, DNA may be converted into less stable RNA.

[0290] Confirmation of nucleic acid deletions can be achieved through sequencing, PCR, or quantitative PCR.

[0291] Identifier Design and Ranking Methods for Efficient Random Access

[0292] The system and method described herein allow for efficient random access retrieval of any bit distribution from encoded and stored information. When data is stored with component-specific primers used at edge layers (or end sequences) to amplify a targeted subset of identifiers in a library, parts of the encoded information can be efficiently retrieved. Efficient access may involve reducing the number of PCR steps required to retrieve a selected portion of information from the stored data. For example, using the method described herein, identifiers in a stored dataset can be accessed in fewer than L / 2 sequential PCR steps, where L is the number of layers containing the identifiers. The identifier architecture and identifier ranking system influence the random access properties of the identifier pool. The rank of an identifier corresponds to the position of the bits represented by the identifier. Identifier ranks can be determined lexicographically from the order of each possible component that may appear in each layer, which can be strategically defined. For example, layers at the edges of an identifier may be assigned a higher priority than layers in the middle of the identifier, so random access (e.g., using PCR primers that bind the edge layers of the identifier) ​​will return identifiers with consecutive ranks corresponding to a sequence of encoded bits or related stretches. The higher the "priority," the lower the depth of access; for example, high-priority elements are easier to access than low-priority elements.

[0293] The identifier architecture and identifier ranking system allow random access to a specific subset of identifiers in an identifier pool. In some implementations, each identifier nucleic acid sequence in the identifier pool corresponds to a symbol value and a symbol position within a symbol string. Additionally, the presence or absence of an identifier nucleic acid sequence in the pool may indicate the symbol value of each corresponding symbol position within the symbol string.

[0294] In a specific implementation, symbols having adjacent symbol locations encode similar digital information. As used herein, similar digital information may include data of the same structure (i.e., image data or binary code strings). Similar digital information may also refer to the data contained therein. For example, all image data locations encoded with a specific intensity of red may be grouped together at adjacent symbol locations. Alternatively, symbols having consecutive symbol locations may not encode similar digital information. For example, consecutive symbol locations may correspond to various features of the data (i.e., image data), such as x-coordinates, y-coordinates, intensity values, or intensity value ranges. FIG. 23 shows an example of an identifier generated by a multiplication method of three layers A, B, and C, where each layer has two components 1 and 2. The components of each of the three layers A, B, and C are assembled in their respective order. The rank of each identifier may be determined by assigning a specific order to each layer, then assigning a specific order to each component within each layer, and sorting the identifiers alphabetically. FIG. 23a shows the resulting ranking obtained by defining the lexicographical alignment of layers in the same way that layers are aligned in physical identifiers. When querying this pool of identifiers with a PCR reaction using primers that combine the pool of identifiers (e.g., component A1 and component C1), the accessed identifiers have discontinuous rankings, which can make it impossible to randomly access a continuous string of bits in a single PCR reaction. In the specific implementation described herein, the edges of the identifiers (e.g., component A1 and component C1) are referred to as "terminal sequences" or "terminal molecules." However, since bits within a continuous stretch often encode relevant information, it is ideal to randomly access a continuous stretch of bits (represented by a continuously ranked identifier).Each bit within a continuous bit stretch is accessed using a probe and hybridized to the target terminal sequence of each identifier nucleic acid sequence among multiple identifier nucleic acid sequences, thereby selecting the identifier nucleic acid sequence corresponding to each symbol having a continuous symbol position. FIG. 23b shows how the lexicographical order of layers A, B, and C can be modified to enable querying of adjacent bit stretches in a single PCR reaction using primers that bind the edges (or terminal sequences) of the identifier. The strategy is not to use a lexicographical layer order identical to the physical order of the layers. Instead, the strategy is to assign a higher priority lexicographical order to the layers at the edges (or terminal sequences) of the identifier and a lower priority to the layers in the middle of the identifier.

[0295] The component distribution of the partitioning method underlying the combination space can affect the number of symbols accessible in a PCR reaction. Figure 24 shows an example of an identifier generated by a multiplication method of three layers A, B, and C, where the components are not uniformly distributed across the layers. Specifically, two layers have two components 1 and 2, and one layer has three components 1, 2, and 3. According to the aforementioned identifier ranking principle, the physical order is A, B, C, but the lexicographical order of the layers is A, C, B. For example, since layers at the edges of an identifier may be assigned a higher priority than layers in the middle of the identifier, random access (e.g., using PCR primers that combine the edge layers of the identifier) ​​will return an identifier with a consecutive rank corresponding to a sequence of encoded bits or related stretches. Specifically, the first and second terminal sequences of a specific identifier nucleic acid sequence are shared between multiple identifier nucleic acid sequences corresponding to adjacent bit stretches. FIG. 24a shows that when more components are placed in the middle layer(s) of the identifier, a larger pool of identifiers accessed by a PCR query (using primers that combine edge components (or terminal sequences), respectively) can be generated. Accordingly, more bits can be accessed at once. FIG. 24b shows that when more components are placed in the edge layer(s) of the identifier, the pool of identifiers accessed by an equivalent PCR query can be smaller. Accordingly, bits can be accessed at a higher resolution.

[0296] The number of multiplicative layers for constructing identifiers can also affect the number of symbols accessible per PCR query. FIG. 25 shows an example of identifiers generated by a multiplicative of five layers (A, B, C, D, E), where each layer has two components (1 and 2). In addition to the aforementioned identifier ranking principle, the lexicographical order of the layers assigns the highest priority to the outermost layers (A and E), the next highest priority to the second-to-outermost layers (B and D), and the lowest priority to the middle layer (layer C). As used herein, priority represents the depth (or level) of data access, with higher priority corresponding to shallow depth and lower priority corresponding to deep depth. For example, access to a book in a collection of books (i.e., layers A and E) is considered the highest priority, access to a chapter within a book (i.e., layers B and D) is considered the next highest priority, and access to a paragraph within a chapter of a book (i.e., layer C) is considered the lowest priority. If there are more layers, the lexicographical sorting of the layers continues in this manner, allowing fewer PCR queries to be used to search for consecutive or related bit stretches. All identifiers associated with the components of the outermost layers (A1 and E1) can be queried in a single PCR reaction. Then, higher resolution (i.e., lower priority or deeper) queries can be performed through an additional PCR reaction using primers that combine the components of the second and outermost layers (B1 and D1). If the identifier architecture has more layers, sequential PCR reactions continue in this manner to obtain higher resolution queries. However, instead of using two sequential PCR reactions to query all identifiers associated with the four components of A1, B1, D1, and E1, this can be done.(Especially when designed so that the components have sufficiently short sequences) PCR primers can be designed to bind A1-B1 and E1-D1 together, but do not bind any components on their own, so the resulting PCR query is the same identifier as when B1 and D1 are sequentially PCR queried following A1 and E1.

[0297] Method of encoding information using DNA and multiple repositories

[0298] Information can be encoded as a DNA identifier using a "multi-bin scheme." In one implementation of this scheme, there are b bins, each maintaining a disjoint set of identifiers. Each bin is a unique identifier that can be referred to as a label or a bin label. It is labeled as a bit symbol. l The bitstream of a bit is It is divided into "words," and each word is of length It has bits. Any word w can be an empty label.

[0299] Specifically, the multi-bin method may be a "multi-bin positional encoding method." In this multi-bin method, a unique identifier is constructed to indicate the position of each word w in the bitstream and is placed in a unique bin with the label w. In a multi-bin implementation of this method, l To encode bit information An identifier is generated, and each bit is encoded into exactly one identifier existing in exactly one bin. We refer to this as "multi-bin positional encoding."

[0300] The previously described multi-bin location encoding method can be illustrated through the following example. Consider 35 bins labeled with unique symbols of the English alphabet, including punctuation marks. The encoding of a paragraph of English text is performed in the following manner: For each symbol x, all occurrences of x are identified within the paragraph. Integer addresses are obtained by numbering each character of the text in ascending order. All identifiers corresponding to the address of a specific symbol x are generated and collected into a single storage labeled x. Therefore, every location in the text where x occurs is represented by an identifier of the storage labeled x.

[0301] FIG. 26 illustrates an example of a multi-bin location encoding scheme, wherein the symbol location of each type of symbol stream is written to a bin reserved for that type of symbol. The figure is labeled "1" " shows examples of phrases. In this example, 9 types of symbols "A", "B", "C", "D", "E", "F", "G", "H", and " Assume a 9-character alphabet consisting of "(representing a space). Each symbol in this alphabet is assigned a unique bin corresponding to that symbol and named after it. For example, the empty bin "D" is labeled 7. For example, the label for the bin "F" appears as label 6. Phrases to be encoded are identified by the symbols of the alphabet and mapped in a one-to-one correspondence with the identifier library as indicated by label 3. Whenever a symbol appears, the corresponding identifier is added to the repository reserved for that symbol. For example, the bin A contains the phrase to be encoded (" A BE A CH C A F Because the "A" symbol appears 3 times in ", (add emphasis), 3 identifiers (Label 4) are included. Furthermore, the three identifiers in the empty "A" indicate the locations where the corresponding symbol appears. The phrase mapped to the characters "B" and "G" (" Storage "D" and "G" are empty because they do not appear in "").

[0302] In another implementation of the multi-bin method, l The bitstream of bits is implicitly encoded in an identifier distribution for b bins labeled 1, 2, 쪋, b. In this scheme, the length is l A mapping is designed between the set of all bitstreams, which are bits, and the set of all d identifier distributions into b bins. The distribution of d identifiers for b bins is a vector of integer labels. (b 1 , b 2 , ..., b d ) 0 ≤ b i < b and: each non-negative integer b i is the label of the unique bin assigned to the i-th identifier. Since each assigned bin label can be freely selected from b possible labels, b d There are several possible distributions.

[0303] Figure 27 illustrates an example of a multi-bin scheme based on the use of identifier distributions for information encoding. Figure 27 shows an example of an identifier library consisting of two identifiers (labeled as 1) and a bin collection consisting of three named bins (0, 1, 2). Each row of the bin (each row consists of three named bins 0, 1, and 2) shows an example of the distribution of two identifiers divided into three bins. The table (labeled as 6) is fixed but shows an arbitrary bitstream mapped to each distribution. For example, the fourth row (labeled as 5), consisting of three bins, shows a distribution where two identifiers are placed in the bin named 1 and bins 0 and 2 are empty. This distribution is arbitrarily mapped to bitstream 0011. Similarly, the second row of three bins shows a distribution where two identifiers are placed in the bin named 0 and 1 and the third bin is empty. This distribution is mapped to bitstream 0001 (labeled as 3). The following line shows a distribution where the bin named 1 is empty. This corresponds to bitstream 0010. Given such a bitstream, the corresponding distribution is constructed and preserved. In this way, any bitstream can be encoded using this multi-bin identifier distribution scheme with a sufficient number of bins and identifiers.

[0304] In another embodiment of the multi-bin scheme, identifiers may exist in more than one bin. In this scheme, a bitstream of l bits is implicitly encoded in the identifier distribution for bins labeled 1, 2, ..., b. In this scheme, each bin contains a subset of identifiers. Thus, in this scheme, a mapping is designed between the set of all bitstreams of length l bits and the set of all b-subsets among the set of all identifier subsets. A b-subset refers to a set containing b elements. For example, if there are a total of d identifiers in the combination space, the set of all identifier subsets is 2 dIt includes sets and is denoted by D. This method uses a mapping between any bitstream of length l and any subset of D containing b sets, and It is possible to encode a bitstream of a length not greater than that. In another embodiment, each bin contains an individual subset, in which case the method is of length It can encode bitstreams that are not larger than

[0305] Figure 28 illustrates an example of a multi-bin scheme based on the use of identifier distributions for encoding information, where the identifier may appear in more than one bin. We call this scheme "Identifier Distributions with Reuse." Figure 28 shows an example involving an identifier library of two identifiers (labeled 8 and 9) and three bins (bins 0, 1, and 2). The two identifiers and three bins consist of six bits (b0, b1, b2, b3, b4, b5, where each b xIt is used to code (where x corresponds to a single bit within the bitstream and x represents each bit position in the bitstream). At the top of the diagram, a subset of possible identifiers corresponding to bits b0b1 (labeled 4), b2b3, and b4b5 are shown, respectively. A subset of identifiers can be contained in any bin. Thus, each of the three bins can contain four options: no identifier, a single identifier (labeled 8), a different identifier (labeled 9), or both identifiers (8 and 9). Since this example contains three bins, each subset is displayed three times in each row (label 2). Each of the three bins can contain exactly one subset, but any subset triple is allowed. This is indicated by lines connecting the subsets (label 3); that is, each path from left to right corresponds to a collection of subsets to be contained in the three stores. Each identifier distribution is mapped to a specific bitstream as shown in the table (labeled 7). In one embodiment, the bitstream can be inferred by naming the subsets for each bin 00, 01, 10, and 11. Thus, for example, the distribution labeled 5 chooses to include a subset of bin identifiers in each of the three bins, and since the name of this subset is 00, it corresponds to the bitstream 000000. Similarly, the distribution labeled 6 will correspond to the bitstream 010110 because it chooses to include subset 01 in bin 0, subset 01 in bin 1, and subset 10 in bin 2. The figure shows a few more examples of the 64 possible distributions (indicated by dashed items in the figure).

[0306] Multi-bin encoding methods can be applied to the secure storage of data because decoding data encoded in this way may require access to and decoding of all bins. For example, to map a multi-bin encoded identifier library back to a source bitstream, it may be necessary to obtain the set of identifiers present in each bin, as the multi-bin method maps the bitstream to an individual distribution of identifiers within the multi-bins, which generally does not allow decoding any meaningful substring of the source bitstream from an appropriate subset of bins.

[0307] In another embodiment, the source bitstream may be encoded using a multi-bin scheme that uses a plurality of orthogonal identifier libraries. The resulting multi-bin libraries may be combined in a manner that enables decoding from any subset of bins of some minimum cardinality. For example, the source bitstream may be encoded using five orthogonal libraries and three bins each. The resulting 15 bins may be combined in a manner that enables decoding of the bitstream from any subset of three bins. In practice, the bins may be physical locations, such as tubes, wells, or spots on a substrate.

[0308] In some embodiments, a bin may be a physical location, e.g., a tube, a well, or a spot on a substrate. In other embodiments, a bin may be a more abstract association shared by all identifiers in a collection, such as a specific barcode sequence.

[0309] Information encoding method using DNA and integer partitioning

[0310] We use the term "integer partitioning" method to refer to the encoding strategy that stores information when partitioning random sequences of DNA. Figure 29 illustrates an example of an integer partitioning method summarized in five steps. DNA is represented as a string consisting of gray or black bars and symbols. Each DNA depicted represents a distinct species. A "species" is defined as one or more DNA molecule(s) of the same sequence. When "species" is used in a plural sense, it can be assumed that all species included in the plural have individual sequences, but this can sometimes be explicitly indicated by using "individual species" instead of "species."

[0311] In Step 1 of the Method Example, a pool of a very large number of species, each referred to as a "count," is started. The counts can be designed so that a common sequence is at the edges (black and light gray bars) and individual sequences are in the middle (N⪋N). ​​These starting count pools can be prepared in a rapid and inexpensive manner using a degenerate oligonucleotide synthesis strategy. In Step 2, the counts are divided into bins (rectangles in Step 2). It does not matter which count is divided into which bin; what matters is the number of counts divided into each bin. Thus, splitting can occur by randomly sampling a single count from the starting pool and then assigning it to a specific bin (e.g., one of the five bins in Step 2). A single count can be sampled from the pool in a small droplet. A bin is a reaction vessel. For example, a bin can be a microfluidic channel on a substrate or a chamber within a site. Counts can be assigned to a chamber via a microfluidic device or assigned to a site on the substrate via printing. Each bin contains an individual DNA species referred to as a barcode. The barcode can be designed to have a common sequence on the edges (light gray bars and dark gray bars) and a central individual sequence (B0, B1, B2, B3, B4, ....) identifying each bin. In step 3, the common edge sequence of the barcode is assembled into the common edge sequence of the count. For example, the common edge sequence of the barcode can be configured to be assembled via adhesive end ligation or Gibson assembly. In step 4, the assembled DNA molecules from each bin are combined into a final pool for storage, as directed to step 5. The species in the final pool contains all information about how the counts are partitioned into each bin. This information can be recovered through sequencing. In the given example, the sequencing data may mean that 9 counts are partitioned into 5 bins, so the first bin (B0) has 2 counts, the second bin (B1) has 3 counts, the third bin (B2) has 1 count, the fourth bin (B3) has 1 count, and the fifth bin (B4) has 2 counts. This is equivalent to mathematically rewriting the integer "9" as an ordered sum "2+3+1+1+2" known as a "configuration." If the parameters of this method are fixed to always have a total of 9 counts and 5 bins, there are 13choose4 possible compositions, so the specific composition recorded in this example contains log2(13choose4) bits of information. At any point in this process, multiple copies of each species may exist or be generated without interfering with the information being stored (e.g., using PCR). This allows the final pool to be amplified to prevent degradation and facilitate sequencing.

[0312] Generally, when an integer partition system has n partition counts and k fixed parameter values ​​for bins, the method log 2 [(n+k-1)choose(k-1)] It can be implemented to store bit information. Mathematically, it is said that information measures the number of "weak configurations" of the system. However, this applies only if the barcode sequence of each storage is known. If the barcode sequence of each bin is unknown (for example, if the barcode itself is a random sequence), the method is It can be implemented to store, where Pj(n) is the number obtained by exactly dividing n into j parts.

[0313] Method for designing a data pipeline to encode information in DNA

[0314] The input bitstream to be written to DNA is processed by a computer encoding-decoding pipeline abbreviated as "codec". Fig. 30 illustrates a high-level block diagram of an exemplary encoding portion of the codec. Upon receiving a source bitstream and a request to write it to DNA, the codec divides the source bitstream into one or more blocks of a size not larger than a fixed length known as the block size. The codec determines the appropriate block size based on the source bitstream (i.e., symbol string), processing requirements, and the intended application of the bitstream content (i.e., digital information). For example, a 100 Gbit bitstream can be divided into 100 blocks, each with a length of 1 Gbit, or 1,000 blocks, each with a length of 100 Mbit, or divided in other ways.

[0315] The codec can calculate the hash of each block using one or more hashing algorithms. It can add the hash and other metadata (e.g., block length, block address) to the block.

[0316] The codec can apply one or more error detection and correction algorithms to each block and calculate one or more error protection bytes. The codec can then combine the original block with the error protection information to obtain an error protection block. For example, the codec can apply convolution coding to the bits of the block, apply Reed-Solomon or erase coding to the byte chunks of the block, and add Reed-Solomon or erase error protection bytes to each chunk of the block. The codec can add error protection metadata to each block.

[0317] When calculating error protection information, the codec may select a specific logarithmic field size to perform the error protection calculation. The field size may represent the source word length, which can be any number of bits such as 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 64, or 128 bits. A source word is a continuous string of bits (fixed length) that constitutes the source bitstream. The codec may select a specific field size and word length based on computational complexity and error protection considerations. For example, an 8-bit word length may be computationally efficient, while a 16-bit word length may provide better error protection. The codec may use a search algorithm to identify the optimal set of parameter values ​​based on one or more objective functions. For example, a codec may use the number of independent response compartments within the writer hardware system, a specific configuration of parameter values, or the number of unique identifiers required to encode a bitstream under some other function or some combination of functions as a cost function.

[0318] The codec may apply an additional encoding step to the error protection block to improve write or read performance. The codec may map each word in the error protection block to a new codeword. The codec may use a search algorithm to generate a set of codewords with a specific set of attributes. For example, the codec may generate codewords of variable length or with an equal fixed number of "1" bit values, codewords with specified Hamming distances from each other, or some combination of these features. When determining the best codeword length, weights, Hamming distances, or other features of the codeword, the codec may use a set of parameters including source word length, writer hardware speed, and the total number of available components. The codec may include another layer of error detection or correction information along with these codewords. For example, the codec may generate a codeword of length n with exactly k "1" bit values, where two bits known as the high bit or the low bit serve as parity bits; the high bit is set when the parity bit is 1, and the low bit is set otherwise. One or more pairs of these error protection bits can protect various parts of a codeword.

[0319] The codec can select a specific set of codewords to ensure optimized chemical conditions during encoding or decoding. For example, the codec can generate fixed-weight codewords so that an equal number of identifiers are assembled in each reaction compartment of the recorder system and at nearly the same concentration within each compartment and across the entire compartment. The codec can select codeword lengths and partitioning schemes so that each reaction compartment assembles an equal number of identifiers and encodes an integer number of codewords.

[0320] A codec can choose to encode some or all bits of a source bitstream using multiple sets of identifiers. Identifiers can originate from or belong to the same identifier library. Identifiers can encode the source bitstream or a combination of bits from the source bitstream. By using multiple sets of identifiers to encode combinations of bits, the codec can reduce the sample size required to reliably decode all bits.

[0321] A codec can generate one or more output blocks for each source block. Output blocks can describe sets of identifiers to be assembled into other types of data structures, including lists or trees. A codec can generate one or more command files that instruct a device to assemble the identifiers assigned to it. For example, a codec can generate command files that control a liquid handling robot or an inkjet printer using ink containing components. A codec can communicate with a device and optimize block files based on the device's information. For example, a device can report an assembly error rate, and a codec can generate a new block file with higher error protection performance. A codec can transmit block files or commands to a file or over a network. A codec can execute computation processes on one or more computers.

[0322] How to specify instructions to the information creator

[0323] We refer to any system that builds an identifier library as a "writer." For example, some embodiments of a writer may use a print-based method to arrange components for constructing identifiers together. A print-based method may involve the use of one or more printheads, each capable of printing one or more nucleic acid molecules onto a substrate.

[0324] The identifier library to be assembled is specified and transmitted to the writer via a set of specification files. The block data file specifies the set of identifiers that the writer will generate. The block data file can be compressed using a data compression algorithm. The identifiers constituting the block may be specified in the form of serialized data structures such as trees, lists, and bitmaps, but are not limited thereto.

[0325] For example, an identifier library to be generated using a multiplication method may be specified by a component library partitioning method (the way components are divided into layers of the identifier architecture) and a block metadata file containing a list of possible component names to be used in each layer. The block data file may contain the identifier to be generated, which is composed of a serialized tree data structure where each path from the root to the leaf of the tree represents an identifier and each node along the path specifies a component name to be used in the layer of that identifier. The block data file may be composed of the serialization of this tree by traversing the tree in the order of starting from the root, visiting the left child node of each node, visiting the node itself, and then visiting the right child node.

[0326] FIG. 31 illustrates an example of a data structure and serialization for representing an identifier library. An identifier library encoding some bitstreams appears (Label 11). Each path from the tree root to the leaf represents a single identifier, and the components of the identifier are specified by the names of the nodes encountered along the path. Label 6 shows a serialized representation of the data structure, consisting mainly of component names and distinguishing symbols. The serialized format begins with the specifications of the constructor-specific partition scheme (Label 5). In this case, the constructor configuration is used as four layers containing 3, 2, 3, and 5 components in each layer. The remaining items of the serialization sketch the path of the data structure as indicated by 1. The segment indicated by 4 in the serialization sketches a path starting from the root of the tree and descending to Node 0 of the first layer, Node 0 of the second layer, Node 0 of the third layer, and Leaf 0 of the last layer. Since the partition scheme has four layers, the algorithm infers that a complete identifier can be output at this stage. More generally, this serialization segment (labeled as 7) specifies all alternative components of the final layer. A separator symbol (a period in this example) is included in the serialization to indicate that all alternatives to be included in the identifier library of a specific layer have been listed. This then triggers the algorithm to move up the layers as indicated by the path in the tree (labeled as 3). The next segment of component identifiers in the serialization (labeled as 16) describes the next set of identifiers. In this way, the entire identifier library can be represented as a flat serial file in a compact manner.

[0327] Calculation method using identifiers

[0328] It may be possible to perform calculations on data encoded in an identifier library using chemical operations. Doing so may be advantageous, as these operations can be performed in parallel on a subset or the entire archive. Additionally, since the calculations can be performed in-vitro without decoding the data, confidentiality can be guaranteed while allowing for calculations. In some implementations, calculations involving Boolean logic operations such as AND, OR, NOT, and NAND are performed on an encoded bitstream using an identifier representing each bit position, where the presence of the identifier encodes a bit value of '1' and the absence of the identifier encodes a bit value of '0'.

[0329] In some embodiments, all identifiers consist of single-stranded nucleic acid molecules (or initially consist of double-stranded nucleic acid molecules that are separated into single-stranded forms). For any single-stranded identifier x, the identifier is represented by the inverse complement of x by x*. For any set of single-stranded identifiers S, the set of inverse complements for each identifier in S is denoted as S*. The set of all possible single-stranded identifiers in the library is denoted as U, and the set of inverse complements is denoted as U*. We refer to these sets as the universe and the universe*. U s and U s * represents a second pair of sets of universes and universes*, and each identifier of these sets is reinforced by an additional nucleic acid sequence known as a search region that can be targeted or selected by chemical methods.

[0330] Computation on a given identifier library can be implemented by a series of chemical operations involving hybridization and cleavage. An abstraction of these operations is described below. Each operation takes the identifier pool as input, performs the operation, and returns the identifier pool as output.

[0331] As an example, as shown in the table below, the first library (L1) and the second library (L2) may each contain 8 bits. The results of bitwise "OR" operations between the two libraries and bitwise "AND" operations between the two libraries are also shown. Details of these operations (and additional operations) performed by chemical steps are described in more detail below.

[0332]

[0333] Table 1

[0334] Each bit of each library is encoded as an identifier containing a symbol location. If there is no identifier for a symbol location, it is represented as 0, and if there is an identifier for a symbol location, it is represented as 1. In this example, the library identifier is double-stranded.

[0335] To perform an OR operation on two libraries, L1 and L2, the two library pools are combined. The identifiers of the two libraries may remain in a double-stranded state for the OR operation. Since the OR operation indicates whether there is a 1 in L1 or L2, the combination of the two pools is a fully determined OR operation output (as shown in the OR column above). For the same symbol location, there are up to twice the number of identifier copies (compared to the original library), which still indicates that there is a 1 at that symbol location (i.e., symbol location b5). In some implementations, double-stranded identifiers can be denatured to generate two single strands (i.e., one sense or "positive" strand and one antisense or "negative" strand for each double-stranded identifier). We refer to the resulting two complementary single strands as the "positive" and "negative" strands. In some implementations, a subsection of the library may be selected, an OR operation may be performed, and the result of the OR operation may replace the existing bit values ​​of one or both of the existing libraries.

[0336] To perform an AND operation on the two libraries L1 and L2, the double-stranded identifier is first denatured to generate two single strands (i.e., one sense strand and one antisense strand for each double-stranded identifier). Once again, we refer to the two resulting complementary single strands as "positive" and "negative" strands. The positive and negative strands are separated into separate pools. In practice, this can be achieved using probes tagged with affinity for the positive or negative strands (see Chemical Methods for Nucleic Acid Capture, Section F). For this purpose, the identifier can be designed to include a common probe target. Then, the positive strand (e.g., the sense strand) of the double-stranded identifier from th...

Claims

Claim 1 A method for recording information as a nucleic acid sequence, the method comprising: obtaining a first fixed-point number; obtaining a library of component nucleic acid sequences defining a combination space of identifier nucleic acid sequences each comprising an aligned subset of component nucleic acid sequences; identifying a first subset of identifier nucleic acid sequences in the combination space as a first codeword having a codeword size corresponding to the number of identifier nucleic acid sequences of the first subset; and forming a first set of one or more identifier nucleic acid molecules having individual identifier nucleic acid sequences of the first subset, wherein the ratio of the number of individual identifier nucleic acid sequences expressed in the first set to the codeword size approximates the first fixed-point number. Claim 2 A method according to claim 1, wherein a library of component nucleic acid sequences comprises a plurality of layers, each layer comprises a subset of component nucleic acid sequences, and each identifier nucleic acid sequence comprises one component nucleic acid sequence from each layer. Claim 3 A method according to claim 1 or 2, wherein the first fixed-point number has a value x, the codeword size is w, k identifier nucleic acid molecules are formed in the first set, the ratio is k / w, and is equal to x. Claim 4 In paragraph 3, the method in which k / w is within ±20% of x. Claim 5 A method according to claim 1 or 2, wherein the codeword size is at least 8. Claim 6 In paragraph 5, the method wherein the codeword size is at least 256. Claim 7 In claim 6, the method wherein the codeword size is at least 512. Claim 8 In claim 7, the method wherein the codeword size is at least 1024. Claim 9 The method of claim 1 or 2 further comprises the steps of: obtaining a second fixed-point number; identifying a second subset of identifier nucleic acid sequences in a combination space as a second codeword having the codeword size of a first codeword and corresponding to the number of identifier nucleic acid sequences in the second subset; and forming a second set of one or more identifier nucleic acid molecules having individual identifier nucleic acid sequences of the second subset— wherein the ratio of the number of individual identifier nucleic acid sequences of the second set to the codeword size approximates the second fixed-point number. Claim 10 A method according to claim 9, further comprising the step of pooling a first set and a second set to obtain a sum pool, and adding the first fixed-point number and the second fixed-point number by diluting the pooled set to obtain a scaled sum pool. Claim 11 A method according to claim 9, further comprising the step of pooling a first set and a second set to obtain a factor pool, and multiplying the first fixed-point number and the second fixed-point number by applying a chemical AND operation to the first set and the second set of identifier nucleic acid molecules to obtain a product pool. Claim 12 A method according to claim 11, wherein the chemical AND operation comprises the steps of converting an identifier nucleic acid molecule into a single-strand identifier nucleic acid molecule, hybridizing a complementary identifier nucleic acid molecule, and selecting a fully hybridized double-strand nucleic acid molecule to obtain a product pool. Claim 13 A method according to claim 12, wherein the selection comprises using at least one of an enzyme that selectively degrades single-stranded nucleic acid molecules or an enzyme that selectively degrades double-stranded nucleic acid molecules having sequence mismatch. Claim 14 The method of claim 9 further comprises the steps of pooling a first set and a second set to obtain a sum pool, and applying a chemical OR operation to a first set and a second set of identifier nucleic acid molecules to obtain a product pool. Claim 15 A method according to claim 14, comprising the step of mixing the first set and the second set. Claim 16 In claim 9, the method further comprises the steps of pooling a first set and a second set to obtain a factor pool, and applying a chemical NIMPLY operation to a first set and a second set of identifier nucleic acid molecules to obtain a product pool. Claim 17 In claim 16, the chemical NIMPLY operation comprises the steps of converting an identifier nucleic acid molecule into a single-strand identifier nucleic acid molecule—the second set of single-strand identifier nucleic acid molecules comprising an affinity tag—providing a second set of molar excess single-strand identifier nucleic acid molecules, hybridizing a complementary identifier nucleic acid molecule, and obtaining a multiplication pool using a specific capture mechanism for the affinity tag by selecting a fully hybridized double-strand nucleic acid molecule. Claim 18 The method of claim 9 further comprises the steps of pooling a first set and a second set to obtain a factor pool, and applying a chemical NOT operation to the first set and a second set of identifier nucleic acid molecules to obtain a product pool. Claim 19 In claim 18, the chemical NOT operation comprises the steps of converting an identifier nucleic acid molecule into a single-strand identifier nucleic acid molecule—the second set of single-strand identifier nucleic acid molecules comprising an affinity tag—providing a molar excess of a first set of single-strand identifier nucleic acid molecules, hybridizing a complementary identifier nucleic acid molecule, and obtaining a multiplication pool using a specific capture mechanism for the affinity tag by selecting a fully hybridized double-strand nucleic acid molecule. Claim 20 The method of claim 9 further comprises the steps of pooling a first set and a second set to obtain a factor pool, and applying a chemical XOR operation to a first set and a second set of identifier nucleic acid molecules to obtain a product pool. Claim 21 In claim 20, the chemical XOR operation comprises the step of performing two NIMPLY operations followed by an OR operation.

Citation Information

Patent Citations

  • Methods of storing information using nucleic acids

    KR1020150037824A

  • DNA-based data storage

    KR1020200071720A

  • Chemical methods for storing nucleic acid-based data

    KR1020200132921A

  • Compositions and methods for storing nucleic acid-based data

    KR1020210029147A

  • Methods for storing and reading digital data on a set of DNA strands

    US20170187390A1