Noise reduction for data storage in DNA
By using a print-based system to assemble nucleic acid components into identifiers, the method addresses the inefficiencies and high costs of base-by-base synthesis, improving the fidelity and reducing noise in nucleic acid data storage and retrieval.
Patent Information
- Application Number
- JP2025522516
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2025-11-05
AI Technical Summary
Current methods for encoding digital information into nucleic acid sequences are expensive and error-prone due to base-by-base nucleic acid synthesis, leading to inefficiencies in data storage and retrieval.
A method involving the use of a print-based system to dispense nucleic acid components onto a substrate, allowing them to self-assemble into identifiers, which are then combined into a pool representing a symbol string, with techniques to mitigate PCR-generated recombination and chimeric identifier formation, thereby improving fidelity and reducing noise.
This approach reduces the cost and enhances the accuracy of encoding and decoding digital information in nucleic acid molecules, facilitating efficient and reliable data storage and retrieval.
Smart Images

Figure 2025536329000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 417,809, entitled "METHODS FOR NOISE REDUCTION," filed October 20, 2022, the entire contents of which are incorporated herein by reference. [Background technology]
[0002] background
[0002] Nucleic acid digital data storage is a robust approach to encoding and storing information over long periods of time, storing data at higher densities than magnetic tape or hard drive storage systems. Furthermore, digital data stored in nucleic acid molecules stored in cold and dry conditions can be retrieved for over 60,000 years.
[0003]
[0003] To access the digital data stored in nucleic acid molecules, the nucleic acid molecules can be sequenced. Thus, nucleic acid digital data storage can be an ideal method for storing data that is not accessed frequently but may contain large amounts of information to be stored or archived for long periods of time.
[0004]
[0004] Current methods rely on encoding digital information (e.g., binary code) into a base-by-base nucleic acid sequence so that the base-to-base relationships in the sequence are directly converted into digital information (e.g., binary code). Sequencing digital data stored in a base-by-base sequence that can be read into a bitstream or byte of digitally encoded information can be error-prone and expensive to encode because the cost of de novo base-by-base nucleic acid synthesis can be expensive. Opportunities for new methods of implementing nucleic acid digital data storage may provide approaches to encoding and retrieving data that are cheaper and easier to implement commercially. Summary of the Invention [Means for solving the problem]
[0005] overview The disclosed systems, assemblies, and methods generally relate to creating DNA molecules that store digital information. For example, component nucleic acid molecules (e.g., components) are selected and individually dispensed onto a substrate material, such as webbing. The components are printed or dispensed at the same locations (e.g., coordinates) on the substrate so that they are co-located. The components are configured to self-assemble or otherwise sort themselves in a predetermined order to form identifier nucleic acid molecules (e.g., identifiers). Each identifier corresponds to a specific symbol (e.g., a bit or series of bits) or the position (e.g., rank or address) of that symbol in a string (e.g., a bit stream). To assemble the components, the system can print or dispense a reaction mixture at the same locations, allowing the components to align themselves and form the identifier. Alternatively or additionally, the system can provide the conditions necessary to physically link the components, such as a specific temperature to align the components. Once formed, multiple identifiers can be combined into a pool of identifiers, where the pool represents at least a portion of the entire string.
[0006]
[0006] Described herein are techniques for writing information to nucleic acid molecules (e.g., DNA) and storing and / or reading information encoded in nucleic acid molecules. Using the disclosed methods and systems, computer data or information can be encoded into multiple identifiers, each of which can represent one or more bits of the original information. The techniques include methods for mitigating PCR-generated recombination, for example, by targeted removal of incompletely assembled products and / or blocking extension from non-primer DNA molecules. These techniques aim to improve the fidelity of the identifier library during the amplification process by reducing chimeric identifier formation. Ultimately, reducing chimeric identifiers in the root library can improve the decodability of the identifier library.
[0007] In one aspect, the present disclosure provides a method for writing information to a nucleic acid molecule with reduced noise. The method includes determining a symbol string representing the information and generating a plurality of oligonucleotides comprising a plurality of identifiers. Each identifier in the plurality of identifiers corresponds to a symbol in the symbol string. Generating the plurality of oligonucleotides includes (a) assembling a plurality of building blocks, each individual building block of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding a first volume comprising a template-independent polymerase and an amount of dideoxynucleotides (ddNTPs) to a reaction volume containing the plurality of building blocks; (c) incubating the reaction volume to add ddNTPs to the 3' ends of at least some of the building blocks; (d) adding reagents to the reaction volume to chemically link two or more building blocks of the plurality of building blocks, thereby generating a plurality of identifiers and a plurality of fragments; and (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification, wherein PCR amplification of any oligonucleotides containing ddNTPs is inhibited.
[0008] In one aspect, the present disclosure provides a method for writing information to a nucleic acid molecule. The method includes determining a symbol string representing the information and generating a plurality of oligonucleotides comprising a plurality of identifiers. Each identifier in the plurality of identifiers corresponds to a symbol in the symbol string. Generating the plurality of oligonucleotides includes (a) assembling a plurality of building blocks, each individual component of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding a first volume comprising a polymerase and an amount of acyclonucleotides to a reaction volume containing the plurality of components; (c) incubating the reaction volume to add acyclonucleotides to the 3' ends of at least some of the components; (d) adding reagents to the reaction volume to chemically link two or more components of the plurality of components, thereby generating a plurality of identifiers and a plurality of fragments; and (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification, wherein PCR amplification of any oligonucleotides comprising acyclonucleotides is inhibited.
[0009] In one aspect, the present disclosure provides a method for writing information to a nucleic acid molecule. The method includes determining a symbol string representing the information and generating a plurality of oligonucleotides comprising a plurality of identifiers. Each identifier in the plurality of identifiers corresponds to a symbol in the symbol string. Generating the plurality of oligonucleotides includes (a) assembling a plurality of building blocks, each individual building block of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding reagents to a reaction volume to chemically link two or more building blocks of the plurality of components, thereby generating a plurality of identifiers and a plurality of fragments; (c) adding a first volume comprising a polymerase and a plurality of 3'-DNA flaps to the reaction volume containing the plurality of building blocks; (d) incubating the reaction volume to add 3'-DNA flaps to the 3' ends of at least some of the building blocks and fragments; and (e) adding PCR primers to the reaction volume and subsequently conducting PCR amplification, wherein PCR amplification of any oligonucleotides comprising the 3'-DNA flaps is inhibited.
[0010] In an aspect, the present disclosure provides a method for writing information to a nucleic acid molecule, the method including: (a) determining a symbol string representing the information; (b) constructing a plurality of components, each individual component of the plurality of components being a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (c) chemically linking two or more components of the plurality of components, thereby generating a plurality of identifiers, each identifier of the plurality of identifiers including two or more components, each identifier having a first end and a second end, each component located at the first end of the identifier being a first edge component, and each component located at the second end of the identifier being a second edge component; and (d) chemically modifying each end of the first edge component, the second edge component, or both, such that the first edge component and / or the second edge component is protected from exonuclease activity, wherein each identifier of the plurality of identifiers corresponds to an individual symbol in the symbol string.
[0011] Incorporation by Reference
[0011] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over such conflicting material.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative implementations, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Brief explanation of the drawings]
[0013] [Figure 1]
[0013] An exemplary system for storing digital information in DNA by using inkjet printing to rapidly and high-throughputly assemble DNA identifiers from components is presented. The system and its different embodiments are hereinafter referred to as the "Printer-Finisher System" or PFS. [Figure 2]
[0014] An example printer subsystem is shown in more detail below: The printheads are designed to overprint different components at the same coordinates on the web. [Figure 3A]
[0015] 1 shows an example of a printhead for a printer. [Figure 3B] 1 shows an example of a printhead for a printer. [Figure 3C] 1 shows an example of a printhead for a printer. [Figure 3D] 1 shows an example of a printhead for a printer. [Figure 4]
[0016] 1 illustrates potential placement of printheads within a printer. [Figure 5]
[0017] 1 shows an example of a spot imager configuration in a printer subsystem. [Figure 6]
[0018] An example of a finisher subsystem is shown in more detail below: In addition to dispensing the reaction mixture onto each coordinate of the substrate, the finisher may also include a portion that dispenses a reaction inhibitor onto each coordinate of the substrate before solidification. [Figure 7]
[0019] 1 shows an example of a loop of rollers for passing the web through a finisher during the incubation stage. [Figure 8]
[0020] 1 shows the effect of reaction mixture glycerol composition and finisher humidity on predicted equilibrium volume during incubation. [Figure 9]
[0021] 1 shows an exemplary pooling system that aggregates all responses from the web into one container. [Figure 10]
[0022] 1 illustrates a schematic diagram of one embodiment of a data transfer pipeline through a PFS. [Figure 11]
[0023] 1 illustrates one embodiment of a PFS that includes four modules: a chassis module, a print engine module, an incubator module, and a pooling module. [Figure 12]
[0024] 1 shows an embodiment of a PFS that pools reaction droplets into an emulsion. [Figure 13]
[0025] 1 shows an embodiment of a PFS in which reaction droplets are printed onto a webbing and then coated with oil (or another immiscible liquid). [Figure 14]
[0026] 1 shows an embodiment of a PFS in which the reaction droplets contain beads that bind to printed DNA components. [Figure 15]
[0027] An example is shown of how DNA components bound on beads can be processed into identifiers using emulsions. [Figure 16]
[0028] FIG. 2 illustrates the operating principle of an implementation of the noise reduction techniques described herein. [Figure 17]
[0029] Image of an agarose gel after electrophoresis of a 360 bp block of amplification product using two primers. Lane 3: Unmodified primers. Lane 6: Same primer set as lane 3, but with 3' ddNTP. [Figure 18A]
[0030] 1 is a graph representing Agilent® Tapestation® (automated gel electrophoresis) results showing sample intensity versus fragment length for the products of a 9-layer ligation. [Figure 18B]
[0030] Figure 1 is a graph representing Agilent® Tapestation® (automated gel electrophoresis) results showing sample intensity versus fragment length for the products of a 9-layer ligation treated with dNTPs only. [Figure 18C]
[0030] Figure 1 is a graph representing Agilent® Tapestation® (automated gel electrophoresis) results showing sample intensity versus fragment length for the products of a nine-layer ligation treated with ddNTPs and dNTPs. [Figure 19]
[0031] FIG. 1 shows a comparison of 5′ and 3′ overhangs. [Figure 20A]
[0032] 1 is a flow diagram illustrating an exemplary post-processing workflow having a noise reduction step for performing ablation. [Figure 20B] 4 is a flow diagram illustrating an exemplary post-processing workflow having a noise reduction step for a puller implementation. [Figure 21]
[0033] FIG. 1 shows the principle of overhang-specific oligonucleotides containing a 3′ flap as a mechanism to prevent amplification from unligated products. [Figure 22]
[0034] FIG. 1 illustrates the principle of an exemplary nuclease-based noise reduction process. [Figure 23]
[0035] FIG. 1 illustrates the principle of an exemplary hairpin loop-based FLI protection and noise reduction process. [Figure 24]
[0036] FIG. 1 shows an example of a protelomerase recognition site and the resulting edge protection. [Figure 25]
[0037] FIG. 1 illustrates the principle of an exemplary protelomerase-based FLI protection and noise reduction process. [Figure 26]
[0038] Figure 1 shows the RP and SP configurations resulting from phosphorothioate linkages. Image from NEB website. [Figure 27]
[0039] FIG. 1 shows reverse dT binding. [Figure 28]
[0040] FIG. 1 shows sugar modifications that can protect FLI from nucleases. DETAILED DESCRIPTION OF THE INVENTION
[0014] Detailed Description definition
[0041] As used herein, the term "component" generally refers to a nucleic acid sequence. A component may be a separate nucleic acid sequence. A component may be ligated or assembled with one or more other components to generate other nucleic acid sequences or molecules.
[0015]
[0042] As used herein, the term "layer" generally refers to a group or pool of components. Each layer may include a distinct set of components, such that the components of one layer differ from the components of another layer. Components from one or more layers may be assembled to generate one or more identifiers.
[0016]
[0043] As used herein, the term "identifier" generally refers to a nucleic acid molecule or sequence that represents the position and value of a bit string within a larger bit string. More generally, an identifier may refer to any object that represents or corresponds to a symbol within a symbol string. In some implementations, an identifier may include one or more ligated components.
[0017]
[0044] As used herein, the term "identifier library" generally refers to a collection of identifiers that correspond to symbols in a string representing digital information. In some implementations, the absence of a given identifier in an identifier library may indicate a symbol value at a particular location. One or more identifier libraries may be combined in pools, groups, or sets of identifiers. Each identifier library may include a unique barcode that identifies the identifier library.
[0018]
[0045] The term "nucleic acid" as used herein generally refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or variants thereof. A nucleic acid can include one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. A nucleotide can include A, C, G, T, or U, or variants thereof. A nucleotide can include any subunit that can be incorporated into a growing nucleic acid chain. Such a subunit can be specific to A, C, G, T, or U, or one of the more complementary A, C, G, T, or U, or any other subunit that can be complementary to a purine (i.e., A or G or variants thereof) or a pyrimidine (i.e., C, T, or U or variants thereof). In some examples, a nucleic acid can be single-stranded or double-stranded, and in some cases, a nucleic acid is circular.
[0019]
[0046] As used herein, the term "nucleic acid molecule" or "nucleic acid sequence" generally refers to a polymeric form of nucleotides or polynucleotides that can be of various lengths, either deoxyribonucleotides (DNA) or ribonucleotides (RNA) or their analogs. The term "nucleic acid sequence" can refer to the alphabetical representation of a polynucleotide; alternatively, the term can apply to the physical polynucleotide itself. This alphabetical representation can be entered into a database in a computer having a central processing unit and used to map the nucleic acid sequence or nucleic acid molecule to symbols or bits that encode digital information. A nucleic acid sequence or oligonucleotide can contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0020]
[0047] As used herein, "oligonucleotide" generally refers to a single-stranded nucleic acid sequence, typically composed of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), and, if the polynucleotide is RNA, guanine (G), and thymine (T) or uracil (U).
[0021]
[0048] Examples of modified nucleotides include diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylketone, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyl Examples of such uracil include, but are not limited to, uracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mansylqueosin, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid (v), wibutoxysin, pseudouracil, queosin, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, and 2,6-diaminopurine. Nucleic acid molecules can also be modified at the base moiety (e.g., typically at one or more atoms available to form a hydrogen bond with a complementary nucleotide and / or typically at one or more atoms not available to form a hydrogen bond with a complementary nucleotide), sugar moiety, or phosphate backbone. Nucleic acid molecules can also contain amine-modifying groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexylamide-dCTP (aha-dCTP), to allow covalent attachment of amine-reactive moieties, such as N-hydroxysuccinimide ester (NHS).
[0022]
[0049] As used herein, the term "primer" generally refers to a strand of nucleic acid that serves as the starting point for nucleic acid synthesis, such as in the polymerase chain reaction (PCR). In one example, during replication of a DNA sample, an enzyme that catalyzes replication initiates replication at the 3' end of a primer added to the DNA sample and copies the opposite strand.
[0023]
[0050] As used herein, the term "polymerase" or "polymerase enzyme" generally refers to any enzyme capable of catalyzing a polymerase reaction. Examples of polymerases include, but are not limited to, nucleic acid polymerases. Polymerases can be naturally occurring or synthetic. An example of a polymerase is Φ29 polymerase or a derivative thereof. In some cases, transcriptases or ligases are used in conjunction with or as alternatives to polymerases to construct new nucleic acid sequences (i.e., enzymes that catalyze bond formation). Examples of polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 Examples of such enzymes include DNA polymerases, Φ29 (Phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tca polymerase, Tih polymerase, Tfi polymerase, platinum Taq polymerase, Tbr polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, and Klenow fragments having 3' to 5' exonuclease activity, as well as variants, modified products, and derivatives thereof.
[0024]
[0051] Digital information, such as computer data in the form of binary code, may include sequences or strings of symbols. Binary code, for example, can encode or represent text or computer processor instructions using a binary number system having two binary symbols called bits, typically 0 and 1. Digital information can be represented in the form of non-binary code, which may include a series of non-binary symbols. Each encoding symbol can be reassigned to a unique bit sequence (or "byte"), which can be arranged into a byte sequence or byte stream. The bit value of a given bit can be one of two symbols (e.g., 0 or 1). A byte, which may include a sequence of N bits, can have a total of 2N unique byte values. For example, a byte containing 8 bits can generate a total of 28 or 256 possible unique byte values, and each of the 256 bytes may correspond to one of the 256 possible distinct symbols, characters, or instructions that can be encoded in the byte. Raw data (e.g., text files and computer instructions) can be represented as a byte sequence or byte stream. Zip files or compressed data files containing raw data can also be stored as a byte stream; these files can be stored as a byte stream in a compressed format and then decompressed to the raw data before being read by a computer.
[0025] overview
[0052] Previous methods for encoding digital information into nucleic acids using inkjet printer systems rely on base-by-base synthesis of nucleic acids, which can be both expensive and time-consuming. For example, inkjet printer-based techniques have previously been used for oligonucleotide synthesis on microreactor chips. However, these techniques utilize base-by-base synthesis, which requires the use of a four-step (deprotection, coupling, capping, and oxidation) solid-phase phosphoramidite cycle reaction to add a single oligonucleotide during each synthesis round. The novel methods described herein can encode digital information using a combinatorial arrangement of building blocks, where each building block (e.g., a nucleic acid sequence) is dispensed (e.g., printed) onto a substrate, and the reaction mixture and / or conditions are provided such that each of the building blocks is physically linked in a single reaction.
[0026]
[0053] The information may be stored in a nucleic acid sequence. In some aspects of the present disclosure, provided herein are methods for encoding digital information into an identifier constructed from one or more components. Each component may include a nucleic acid sequence. A print-based system known as a print finisher system (or PFS) can be used to juxtapose and assemble the components to construct the identifier. A PFS may include two subsystems: a printer and a finisher. A PFS may include a single system, a printer, that dispenses both the components and the reaction mixture onto a substrate. In some implementations, the two subsystems may be attached and dependent on each other for their respective functions. In other implementations, the two subsystems may be separate and capable of functioning independently.
[0027] Methods for encoding and writing information into nucleic acid sequences
[0054] In one aspect, the present disclosure provides a method for encoding information into a nucleic acid sequence. The method for encoding information into a nucleic acid sequence may include (a) converting the information into a symbol string, (b) mapping the symbol string to a plurality of identifiers, and (c) constructing an identifier library including at least a subset of the plurality of identifiers. Each identifier of the plurality of identifiers may include one or more components. Each component of the one or more components may include a nucleic acid sequence. Each symbol at each position in the symbol string may correspond to a distinct identifier. An individual identifier may correspond to an individual symbol at each position in the symbol string. Furthermore, one symbol at each position in the symbol string may correspond to the absence of an identifier. For example, in a binary string (e.g., a bit) of "0" and "1," each occurrence of "0" may correspond to the absence of an identifier.
[0028]
[0055] In another aspect, the present disclosure provides a method of nucleic acid-based computer data storage. The method of nucleic acid-based computer data storage may include (a) receiving computer data, (b) synthesizing nucleic acid molecules including a nucleic acid sequence that encodes the computer data, and (c) storing the nucleic acid molecules having the nucleic acid sequence. The computer data may be encoded in at least a subset of the synthesized nucleic acid molecules, rather than in the sequence of each of the nucleic acid molecules.
[0029]
[0056] In another aspect, the present disclosure provides a method for writing and storing information in a nucleic acid sequence. The method may include (a) receiving or encoding a virtual identifier library representing the information, (b) physically constructing the identifier library, and (c) storing one or more physical copies of the identifier library in one or more separate locations. Each identifier in the identifier library may include one or more components. Each of the one or more components may include a nucleic acid sequence.
[0030]
[0057] In another aspect, the present disclosure provides a method for nucleic acid-based computer data storage. The method for nucleic acid-based computer data storage may include (a) receiving computer data, (b) synthesizing a nucleic acid molecule comprising at least one nucleic acid sequence that encodes the computer data, and (c) storing the nucleic acid molecule comprising the at least one nucleic acid sequence. The synthesis of the nucleic acid molecule may be in the absence of base-by-base nucleic acid synthesis.
[0031]
[0058] In another aspect, the present disclosure provides a method for writing and storing information in a nucleic acid sequence. The method for writing and storing information in a nucleic acid sequence may include (a) receiving or encoding a virtual identifier library representing the information, (b) physically constructing the identifier library, and (c) storing one or more physical copies of the identifier library in one or more separate locations. Each identifier in the identifier library may include one or more components. Each of the one or more components may include a nucleic acid sequence.
[0032] Methods for reading information stored in nucleic acid sequences
[0059] In another aspect, the present disclosure provides a method for reading information encoded in a nucleic acid sequence. The method for reading information encoded in a nucleic acid sequence may include (a) providing an identifier library, (b) identifying identifiers present in the identifier library, (c) generating a symbol string from the identifiers present in the identifier library, and (d) compiling information from the symbol string. The identifier library may include a subset of a plurality of identifiers from the combinatorial space. Each individual identifier in the subset of identifiers may correspond to an individual symbol in the symbol string. The identifier may include one or more components. The components may include a nucleic acid sequence.
[0033]
[0060] Information may be written to one or more identifier libraries as described elsewhere herein. Identifiers may be constructed using any of the methods described elsewhere herein. Stored data may be copied and accessed using any of the methods described elsewhere herein.
[0034]
[0061] The identifier may include information about the location of the coding symbol, the value of the coding symbol, or both the location and value of the coding symbol. The identifier may include information about the location of the coding symbol, and the presence or absence of the identifier in the identifier library may indicate the value of the symbol. The presence of the identifier in the identifier library may indicate a first symbol value (e.g., a first bit value) in the binary string, and the absence of the identifier in the identifier library may indicate a second symbol value (e.g., a second bit value) in the binary string. In a binary system, bit values based on the presence or absence of the identifier in the identifier library can reduce the number of identifiers to be assembled and therefore reduce write time. In one example, the presence of the identifier may indicate a bit value of "1" at the mapped position, and the absence of the identifier may indicate a bit value of "0" at the mapped position.
[0035]
[0062] Generating symbols (e.g., bit values) for the information may include identifying the presence or absence of an identifier to which the symbols (e.g., bits) can be mapped or encoded. Determining the presence or absence of the identifier may include sequencing the current identifier or detecting the presence of the identifier using a hybridization array. In one example, decoding and reading the encoded sequence may be performed using a sequencing platform. Examples of sequencing platforms are described in U.S. Patent Application Publication No. 14 / 465,685, filed August 21, 2014; U.S. Patent Application Publication No. 13 / 886,234, filed May 2, 2013; and U.S. Patent Application Publication No. 12 / 400,593, filed March 9, 2009, each of which is incorporated herein by reference in its entirety.
[0036]
[0063] In one example, decoding of nucleic acid-encoded data can be achieved by utilizing sequencing techniques that indicate the presence or absence of specific nucleic acid sequences, such as Illumina® sequencing, base-by-base sequencing of nucleic acid strands, or fragmentation analysis by capillary electrophoresis. Sequencing can utilize the use of reversible terminators. Sequencing can utilize the use of natural or unnatural (e.g., engineered) nucleotides or nucleotide analogs. Alternatively or additionally, decoding of nucleic acid sequences can be performed using various analytical techniques, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal. Various sequencing approaches can be used, including, but not limited to, polymerase chain reaction (PCR), digital PCR, Sanger sequencing, high-throughput sequencing, sequencing-by-synthesis, single-molecule sequencing, sequencing-by-ligation, RNA-Seq (Illumina), next-generation sequencing, digital gene expression (Helicos), clonal single-strand microarrays (Solexa), shotgun sequencing, maximum Gilbert sequencing, or massively parallel sequencing.
[0037]
[0064] Various readout methods can be used to extract information from the encoded nucleic acids. In one example, microarrays (or any type of fluorescent hybridization), digital PCR, quantitative PCR (qPCR) and various sequencing platforms can also be used to read out the encoded sequence and thus the digitally encoded data.
[0038]
[0065] The identifier library may further include supplemental nucleic acid sequences that provide metadata about the information, encrypt or mask the information, or provide metadata and mask the information at the same time. The supplemental nucleic acids may be identified simultaneously with the identification of the identifier. Alternatively, the supplemental nucleic acids may be identified before or after identifying the identifier. In one example, the supplemental nucleic acids are not identified during reading of the encoded information. The supplemental nucleic acid sequences may be indistinguishable from the identifier. An identifier index or key can be used to distinguish the supplemental nucleic acid molecules from the identifier.
[0039]
[0066] Recoding input bit sequences to allow for the use of fewer nucleic acid molecules can improve the efficiency of encoding and decoding data. For example, if an encoding method receives an input sequence with high occurrences of the "111" subsequence, which can map to three nucleic acid molecules (e.g., an identifier), it can be recoded to a "000" subsequence, which can map to a null set of nucleic acid molecules. An alternative input subsequence for "000" can also be recoded to "111." This recoding method can reduce the number of "1"s in the dataset and therefore the total number of nucleic acid molecules used to encode the data. In this example, the total size of the dataset can be increased to accommodate a codebook that specifies new mapping instructions. An alternative method for improving encoding and decoding efficiency can be to recode input sequences to reduce variable length. For example, "111" can be recoded to "00," reducing the size of the dataset and the number of "1"s in the dataset.
[0040]
[0067] The speed and efficiency of decoding nucleic acid-encoded data can be controlled (e.g., increased) by specifically designing identifiers to facilitate detection. For example, a nucleic acid sequence (e.g., identifier) designed for ease of detection can include a nucleic acid sequence containing a majority of nucleotides that are easier to call and detect based on optical, electrochemical, chemical, or physical properties. Engineered nucleic acid sequences can be either single-stranded or double-stranded. Engineered nucleic acid sequences can include synthetic or non-natural nucleotides that improve the detectable properties of the nucleic acid sequence. Engineered nucleic acid sequences can include all natural nucleotides, all synthetic or non-natural nucleotides, or combinations of natural, synthetic, and non-natural nucleotides. Synthetic nucleotides can include nucleotide analogs such as peptide nucleic acids, locked nucleic acids, glycol nucleic acids, and threose nucleic acids. Non-natural nucleotides can include dNaM, an artificial nucleoside containing a 3-methoxy-2-naphthoyl group, and d5SICS, an artificial nucleoside containing a 6-methylisoquinoline-1-thion-2-yl group. The engineered nucleic acid sequences can be designed for a single enhanced property, such as an enhanced optical property, or the engineered nucleic acid sequences can be designed to have multiple enhanced properties, such as enhanced optical and electrochemical properties or enhanced optical and chemical properties.
[0041]
[0068] Engineered nucleic acid sequences may contain reactive natural, synthetic, and non-natural nucleotides that do not improve the optical, electrochemical, chemical, or physical properties of the nucleic acid sequence. The reactive components of the nucleic acid sequence may allow for the addition of chemical moieties that impart improved properties to the nucleic acid sequence. Each nucleic acid sequence may contain a single chemical moiety, or may contain multiple chemical moieties. Exemplary chemical moieties may include, but are not limited to, fluorescent moieties, chemiluminescent moieties, acidic or basic moieties, hydrophobic or hydrophilic moieties, and moieties that alter the oxidation state or reactivity of the nucleic acid sequence.
[0042]
[0069] Sequencing platforms can be specifically designed for decoding and reading information encoded in nucleic acid sequences. Sequencing platforms can be dedicated to sequencing single-stranded or double-stranded nucleic acid molecules. Sequencing platforms can decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of entire nucleic acid sequences (e.g., components) incorporated within a nucleic acid molecule (e.g., identifiers). Sequencing platforms can include the use of promiscuous reagents, increased read length, and the addition of detectable chemical moieties to detect specific nucleic acid sequences. The use of more promiscuous reagents during sequencing can increase read efficiency by enabling faster base calling, which can shorten sequencing time. The use of increased read lengths can allow longer sequences of encoded nucleic acids to be decoded read-by-read. The addition of detectable chemical moiety tags can allow the presence or absence of a nucleic acid sequence to be detected by the presence or absence of the chemical moiety. For example, each nucleic acid sequence encoding one bit of information can be tagged with a chemical moiety that generates a unique optical, electrochemical, or chemical signal. The presence or absence of that unique optical, electrochemical, or chemical signal can indicate a "0" or "1" bit value. The nucleic acid sequence can contain a single chemical moiety or multiple chemical moieties. The chemical moiety can be added to the nucleic acid sequence before using the nucleic acid sequence to encode data. Alternatively, or in addition, the chemical moiety can be added to the nucleic acid sequence after encoding the data but before decoding the data. The chemical moiety tag can be added directly to the nucleic acid sequence, or the nucleic acid sequence can contain a synthetic or non-natural nucleotide anchor and the chemical moiety tag can be added to the anchor.
[0043]
[0070] A unique code can be applied to minimize or detect encoding and decoding errors. Encoding and decoding errors can arise from false negatives (e.g., nucleic acid molecules or identifiers not included in random sampling). One example of an error detection code can be a checksum sequence that counts the number of identifiers in a contiguous set of possible identifiers included in an identifier library. While reading the identifier library, the checksum can indicate the number of identifiers from that contiguous set of identifiers expected to be obtained, and identifiers can continue to be sampled for reading until the expected number is met. In some implementations, a checksum sequence can be included in every contiguous set of R identifiers, where R is equal in size or greater than 1, 2, 5, 10, 50, 100, 200, 500, or 1000, or less than 1000, 500, 200, 100, 50, 10, 5, or 2. The smaller the value of R, the better the error detection. In some implementations, the checksum can be a complementary nucleic acid sequence. For example, a set containing seven nucleic acid sequences (e.g., components) can be divided into two groups: nucleic acid sequences for constructing identifiers in a product scheme (components X1-X3 of layer X and components Y1-Y3 of layer Y) and nucleic acid sequences for supplemental checksums (X4-X7 and Y4-Y7). Checksum sequences X4-X7 can indicate whether 0, 1, 2, or 3 sequences from layer X are assembled with each member of layer Y. Alternatively, checksum sequences Y4-Y7 can indicate whether 0, 1, 2, or 3 sequences from layer Y are assembled with each member of layer X. In this example, the original identifier library with identifiers {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3} can be supplemented to include a checksum resulting in the following pool: {X1Y1, X1Y3, X2Y1, X2Y2, X2Y3, X1Y6, X2Y7, X3Y4, X6Y1, X5Y2, X6Y3}. The checksum sequence can also be used for error correction. For example, the absence of X1Y1 from the above dataset and the presence of X1Y6 and X6Y1 can make it possible to infer that the X1Y1 nucleic acid molecule is missing from the dataset.The checksum sequence can indicate whether the identifier is missing from the sampling of the identifier library or the accessed portion of the identifier library. In the case of a missing checksum sequence, an accessing method such as PCR or affinity-tagged probe hybridization can amplify and / or isolate it. In some implementations, the checksums may not be complementary nucleic acid sequences. The checksums may be directly encoded into the information so that they are represented by the identifier.
[0044]
[0071] Noise in data encoding and decoding can be reduced by constructing identifiers palindromically, for example by using palindromic pairs of components rather than single components in the product scheme. Pairs of components from different layers can then be assembled together palindromically (e.g., YXY instead of XY for components X and Y). This palindrome can be extended to a larger number of layers (e.g., ZYXYZ instead of XYZ), allowing for the detection of erroneous cross-reactivity between identifiers.
[0045]
[0072] Adding excess (e.g., vast excess) complementary nucleic acid sequences to the identifier may prevent sequencing from recovering the encoded identifier. Before decoding the information, the identifier can be enriched from the complementary nucleic acid sequences. For example, the identifier can be enriched by a nucleic acid amplification reaction using primers specific to the identifier end. Alternatively, or in addition, the information can be decoded without enriching the sample pool by sequencing using specific primers (e.g., sequencing-by-synthesis). With both decoding methods, it can be difficult to enrich or decode the information without having the decoding key or knowing something about the composition of the identifier. Alternative access methods can also be used, such as using affinity tag-based probes.
[0046] System for encoding binary array data
[0073] Systems for encoding digital information into nucleic acids (e.g., DNA) can include systems, methods, and devices for converting files and data (e.g., raw data, compressed zip files, integer data, and other forms of data) into bytes and encoding the bytes into segments or sequences of nucleic acids, typically DNA, or combinations thereof.
[0047]
[0074] In one aspect, the present disclosure provides a system for encoding binary sequence data using nucleic acids. The system for encoding binary sequence data using nucleic acids may include a device and one or more computer processors. The device may be configured to construct an identifier library. The one or more computer processors may be individually or collectively programmed to (i) convert information into a symbol string, (ii) map the symbol string to a plurality of identifiers, and (iii) construct an identifier library including at least a subset of the plurality of identifiers. Each identifier in the plurality of identifiers may correspond to an individual symbol in the symbol string. Each identifier in the plurality of identifiers may include one or more components. Each component in the one or more components may include a nucleic acid sequence.
[0048]
[0075] In another aspect, the present disclosure provides a system for reading binary sequence data using nucleic acids. The system for reading binary sequence data using nucleic acids may include a database and one or more computer processors. The database may store an identifier library encoding information. The one or more computer processors may be individually or collectively programmed to (i) identify identifiers in the identifier library, (ii) generate a plurality of symbols from the identifiers identified in (i), and (iii) compile information from the plurality of symbols. The identifier library may include a subset of the plurality of identifiers. Individual identifiers in the plurality of identifiers may correspond to individual symbols in the symbol string. The identifier may include one or more components. The components may include nucleic acid sequences.
[0049]
[0076] A non-limiting implementation of a method of using the system to encode digital data can include receiving digital information in the form of a byte stream, parsing the byte stream into individual bytes, mapping bit positions within the bytes using a nucleic acid index (or identifier rank), and encoding sequences corresponding to either a bit value of 1 or a bit value of 0 into identifiers. Searching the digital data can include sequencing a nucleic acid sample or pool containing sequences of nucleic acids (e.g., identifiers) that map to one or more bits, referencing the identifier rank to determine if the identifier is present in the nucleic acid pool, and decoding the position and bit value information for each sequence into bytes comprising the sequence of digital information.
[0050]
[0077] A system for encoding, writing, copying, accessing, reading, and decoding information encoded and written to nucleic acid molecules may be a single integrated unit or may be multiple units configured to perform one or more of the aforementioned operations. A system for encoding and writing information to nucleic acid molecules (e.g., identifiers) may include a device and one or more computer processors. The one or more computer processors may be programmed to parse the information into a symbol string (e.g., a string of bits). The computer processor may generate an identifier rank. The computer processor may classify the symbols into two or more categories. One category may include symbols represented by the presence of a corresponding identifier in an identifier library, and the other category may include symbols represented by the absence of a corresponding identifier in the identifier library. The computer processor may instruct the device to assemble an identifier corresponding to the symbol represented for the presence of the identifier in the identifier library.
[0051]
[0078] The device may include multiple regions, sections, or partitions. Reagents and components for assembling the identifier may be stored in one or more regions, sections, or partitions of the device. Layers may be stored in separate regions of a section of the device. A layer may contain one or more unique components. Components in one layer may be unique from components in another layer. A region or compartment may include a container, and a partition may include a well. Each layer may be stored in a separate container or partition. Each reagent or nucleic acid sequence may be stored in a separate container or partition. Alternatively or additionally, reagents may be combined to form a master mix for identifier construction. The device may transfer reagents, components, and templates from one section of the device and combine them in another section. The device may provide conditions for completing the assembly reaction. For example, the device may provide heating, agitation, and detection of reaction progress. The constructed identifier may be directed to one or more subsequent reactions to add barcodes, consensus sequences, variable sequences, or tags to one or more ends of the identifier. The identifier may then be directed to a region or partition to generate an identifier library. One or more identifier libraries can be stored in each region, section, or individual partition of the device. The device can transfer fluids (e.g., reagents, components, templates) using pressure, vacuum, or suction.
[0052]
[0079] The identifier library can be stored on the device or transferred to a separate database. The database can contain one or more identifier libraries. The database can provide conditions for long-term storage of the identifier library (e.g., conditions to reduce identifier degradation). The identifier library can be stored in powder, liquid, or solid form. Aqueous solutions of identifiers can be lyophilized for more stable storage. Alternatively, the identifiers can be stored in the absence of oxygen (e.g., anaerobic storage conditions). The database can provide UV protection, low temperature (e.g., refrigeration or freezing), and protection from degrading chemicals and enzymes. The identifier library can be lyophilized or frozen before being transferred to the database. The identifier library can contain ethylenediaminetetraacetic acid (EDTA) to inactivate nucleases and / or buffers to maintain the stability of the nucleic acid molecules.
[0053]
[0080] The database can be coupled to, contain, or separate from a device that writes information to the identifiers, copies the information, accesses the information, or reads the information. Portions of the identifier library can be deleted from the database before copying, accessing, or reading. The device that copies information from the database can be the same as or different from the device that writes the information. The device that copies the information can extract an aliquot of the identifier library from the device and combine the aliquot with reagents and components to amplify part or all of the identifier library. The device can control the temperature, pressure, and agitation of the amplification reaction. The device can include partitions, and one or more amplification reactions can occur in the partition containing the identifier library. The device can copy more than one pool of identifiers at a time.
[0054]
[0081] The copied identifiers can be transferred from the copy device to an access device. The access device can be the same device as the copy device. The access device can include separate regions, sections, or partitions. The access device can have one or more columns, bead reservoirs, or magnetic regions for separating identifiers bound to affinity tags. Alternatively or in addition, the access device can have one or more size selection units. The size selection unit can include agarose gel electrophoresis or any other method for size-selecting nucleic acid molecules. The copying and extraction can be performed in the same region of the device or in different regions of the device.
[0055]
[0082] The accessed data can be read out on the same device, or the accessed data can be transferred to another device. The reading device can include a detection unit for detecting and identifying the identifier. The detection unit can be part of a sequencer, hybridization array, or other unit for identifying the presence or absence of an identifier. The sequencing platform can be specifically designed for decoding and reading information encoded in a nucleic acid sequence. The sequencing platform can be dedicated to sequencing single-stranded or double-stranded nucleic acid molecules. The sequencing platform can decode nucleic acid-encoded data by reading individual bases (e.g., base-by-base sequencing) or by detecting the presence or absence of an entire nucleic acid sequence (e.g., component) incorporated within the nucleic acid molecule (e.g., identifier). Alternatively, the sequencing platform can be a system such as Illumina® sequencing or capillary electrophoretic fragmentation analysis. Alternatively or additionally, decoding the nucleic acid sequence can be performed using various analytical techniques implemented by the device, including, but not limited to, any method that generates an optical, electrochemical, or chemical signal.
[0056]
[0083] Information storage in nucleic acid molecules can have a variety of applications, including, but not limited to, long-term information storage, confidential information storage, and medical information storage. In one example, a person's medical information (e.g., medical history and records) can be stored in a nucleic acid molecule and carried by the person. The information can be stored externally (e.g., in a wearable device) or internally (e.g., in a subcutaneous capsule). When the patient is transported to a clinic or hospital, a sample can be taken from the device or capsule, and the information can be decoded using a nucleic acid sequencer. Personal storage of medical records in nucleic acid molecules can provide an alternative to computer- and cloud-based storage systems. Personal storage of medical records in nucleic acid molecules can reduce the incidence or spread of hacked medical records. Nucleic acid molecules used in capsule-based storage of medical records can be derived from human genome sequences. The use of human genome sequences can reduce the immunogenicity of nucleic acid sequences in the event of capsule failure and leakage.
[0057] Chemical methods for nucleic acid-based data storage
[0084] In one aspect, the disclosure provides a method of writing information to a nucleic acid sequence, the method comprising: (a) generating a symbol string representing the information; (b) constructing a plurality of components, each individual component of the plurality of components comprising a nucleic acid sequence; (c) generating at least one sticky end of each individual component of the plurality of components; (d) chemically linking two or more components of the plurality of components via at least one sticky end of each individual component of the two or more components, thereby generating a plurality of identifiers, each identifier of the plurality of identifiers comprising two or more components, each individual identifier of the plurality of identifiers corresponding to an individual symbol in the symbol string; and (e) selectively capturing or amplifying an identifier library comprising at least a subset of the plurality of identifiers.
[0058]
[0085] In some embodiments, each symbol in a symbol string is one of one or more possible symbol values. In some embodiments, each symbol in a symbol string is one of two possible symbol values. In some embodiments, one symbol value at each position in a symbol string may be represented by the absence of a distinct identifier in an identifier library. In some embodiments, the two possible symbol values are bit values of 0 and 1, and individual symbols in a symbol string with a bit value of 0 may be represented by the absence of a distinct identifier in the identifier library and individual symbols in a symbol string with a bit value of 1 may be represented by the presence of a distinct identifier in the identifier library, or vice versa.
[0059]
[0086] In some embodiments, (d) comprises chemically linking two or more components from two or more layers, each of the two or more layers comprising a distinct set of components. In some embodiments, an individual identifier from the identifier library comprises one component from each of the two or more layers. In some embodiments, the two or more components are assembled in a fixed order. In some embodiments, the two or more components are assembled in any order. In some embodiments, the two or more components are assembled with one or more partition components disposed between two components from different layers of the two or more layers. In some embodiments, an individual identifier comprises one component from each layer of a subset of the two or more layers. In some embodiments, an individual identifier comprises at least one component from each of the two or more layers.
[0060]
[0087] In some embodiments, (c) comprises using an endonuclease to generate at least one sticky end of an individual component of the plurality of components. In some embodiments, at least one sticky end is at a 5' end of an individual component. In some embodiments, at least one sticky end is at a 3' end of an individual component. In some embodiments, (c) comprises generating two sticky ends of an individual component. In some embodiments, at least one sticky end is at least 1 nucleotide in length. In some embodiments, at least one sticky end is 6 nucleotides in length.
[0061]
[0088] In some embodiments, the plurality of nucleic acid sequences stores information metadata or conceals information. In some embodiments, two or more identifier libraries are combined, and each identifier library of the two or more identifier libraries is tagged with a distinct barcode. In some embodiments, each individual identifier in an identifier library comprises a distinct barcode, or a subset identifier of an identifier library comprises a distinct barcode. In some embodiments, the plurality of identifiers or the plurality of components comprising the identifiers are selected to facilitate read, write, access, copy, and delete operations.
[0062]
[0089] In some embodiments, chemically linking comprises ligating two or more components of the plurality of components together using a reagent comprising a ligase. In some embodiments, the ligase is T4 ligase, T7 ligase, T3 ligase, or E. coli ligase. In some embodiments, the reagent further comprises an additive. In some embodiments, the additive increases the efficiency of the ligase. In some embodiments, the additive comprises polyethylene glycol (PEG). In some embodiments, the PEG is PEG400, PEG6000, PEG8000, or any combination thereof. In some embodiments, the final concentration of PEG molecules is at least about 1% weight / volume (w / v). In some embodiments, the ligation reaction time is at least 1 minute. In some embodiments, the ligation is at 30° C. or higher. In some embodiments, the ligation reaction efficiency is at least about 20%. In some embodiments, the method further comprises inactivating the ligase using a buffer containing EDTA or guanidine thiocyanate. In some embodiments, the final concentration of the ligase is at least about 5 CEU / μL.In some embodiments, the reagent further comprises a glycerol molecule.
[0063]
[0090] In some embodiments, the chemical linking in (d) comprises using overlap extension polymerase chain reaction (PCR). In some embodiments, the individual components are deoxyribonucleic acid (DNA) or ribonucleic acid. In some embodiments, the individual components are rehydrated. In some embodiments, the individual components are rehydrated from dehydrated components. In some embodiments, the method further comprises dehydrating the identifier library by dehydrating each individual identifier of at least a subset of the plurality of identifiers. In some embodiments, each individual identifier of at least a subset of the plurality of identifiers is dehydrated. In some embodiments, the method further comprises rehydrating each individual identifier of at least a subset of the plurality of identifiers. In some embodiments, the method further comprises adding a preservation additive to the identifier library to prevent degradation of the identifiers.
[0064]
[0091] In some embodiments, the plurality of identifiers are copied by PCR. In some embodiments, the PCR has at least 10 cycles. In some embodiments, the plurality of identifiers are amplified by PCR to a concentration of 10 nanograms / microliter. In some embodiments, the PCR is emulsion PCR. In some embodiments, the plurality of identifiers are copied by linear amplification. In some embodiments, after PCR, linear amplification is used to create more copies of the plurality of identifiers. In some embodiments, a subset of the plurality of identifiers is accessed in one or more PCR reactions. In some embodiments, a subset of the plurality of identifiers is accessed using one or more affinity-tagged probes.
[0065]
[0092] In some embodiments, the identifiers of a subset of the plurality of identifiers have a set of common components. In some embodiments, the identifiers are purified by gel electrophoresis. In some embodiments, the identifiers are purified by affinity-tagged probes. In some embodiments, the identifiers are amplified using PCR. In some embodiments, the identifiers are designed to avoid thymine-thymine or cytosine-cytosine dinucleotides.
[0066] Chemical methods for assembling building blocks
[0093] The reactions and methods provided herein can be used in the systems described herein to assemble an identifier from one or more components. For example, different reaction mixtures for different chemical methods provided herein can be used in the finisher of the system to assemble different components.
[0067] A. Overlap extension PCR (OEPCR) assembly
[0094] In OEPCR, components can be assembled in a reaction containing a polymerase and dNTPs (deoxynucleotide triphosphates, including dATP, dTTP, dCTP, dGTP, or variants or analogs thereof). Components can be single-stranded or double-stranded nucleic acids. Components assembled adjacent to each other can have complementary 3' ends, complementary 5' ends, or homology between the 5' end of one component and the 3' end of the adjacent component. These terminal regions, called "hybridization regions," are intended to facilitate the formation of hybridized junctions between components during OEPCR, such that the 3' end of one input component (or its complement) hybridizes to the 3' end of its intended adjacent component (or its complement). An assembled double-stranded product is then formed by polymerase extension. This product can then be assembled into more components by subsequent hybridization and extension.
[0068]
[0095] In some implementations, OEPCR can involve cycling between three temperatures: a melting temperature, an annealing temperature, and an extension temperature. The melting temperature is intended to convert double-stranded nucleic acids to single-stranded nucleic acids and eliminate the formation of secondary structures or hybridization within or between components. Typically, the melting temperature is high, e.g., greater than 95°C. In some implementations, the melting temperature can be at least 96, 97, 98, 99, 100, 101, 102, 103, 104°C, or at least 105°C. In other implementations, the melting temperature can be up to 95, 94, 93, 92, 91°C, or up to 90°C. Higher melting temperatures improve the dissociation of nucleic acids and their secondary structures, but may also cause side effects such as degradation of nucleic acids or polymerase. The melting temperature can be applied to the reaction for at least 1, 2, 3, 4 seconds, or at least 5 seconds or more, e.g., 30 seconds, 1 minute, 2 minutes, or 3 minutes.
[0069]
[0096] The annealing temperature is intended to promote the formation of hybridization between complementary 3' ends of intended adjacent components (or their complements). In some implementations, the annealing temperature can correspond to the calculated melting temperature of the intended hybridized nucleic acid formation. In other implementations, the annealing temperature can be within 10°C or more of the melting temperature. In some implementations, the annealing temperature can be at least 25, 30, 50, 55, 60, 65°C, or at least 70°C. The melting temperature can depend on the sequence of the intended hybridization region between the components. Longer hybridization regions have higher melting temperatures, and hybridization regions with a higher percent content of guanine or cytosine nucleotides can have higher melting temperatures. Thus, it may be possible to design components for an OEPCR reaction that are intended to optimally assemble at a specific annealing temperature. The annealing temperature can be applied to the reaction for at least 1, 5, 10, 15, 20, 25 seconds, or at least 30 seconds or more.
[0070]
[0097] The extension temperature is intended to initiate and promote nucleic acid chain extension of the hybridized 3' end catalyzed by one or more polymerase enzymes. In some implementations, the extension temperature can be set at a temperature at which the polymerase functions optimally in terms of nucleic acid binding strength, extension rate, extension stability, or fidelity. In some implementations, the extension temperature can be at least 30, 40, 50, 60°C, or at least 70°C or higher. The annealing temperature can be applied to the reaction for at least 1, 5, 10, 15, 20, 25, 30, 40, 50 seconds, or at least 60 seconds or longer. The recommended extension time can be approximately 15-45 seconds per kilobase of expected extension.
[0071]
[0098] In some implementations of OEPCR, the annealing temperature and the extension temperature can be the same. Therefore, two-stage temperature cycles can be used instead of three-stage temperature cycles. Examples of combinations of annealing and extension temperatures include 60, 65, or 72°C.
[0072]
[0099] In some implementations, OEPCR can be performed in one temperature cycle. Such implementations may involve the intended assembly of only two components. In other implementations, OEPCR can be performed in multiple temperature cycles. Any given nucleic acid in an OEPCR can assemble with at most one other nucleic acid in one cycle. This is because assembly (or extension or elongation) can only occur at the 3' end of the nucleic acid, and each nucleic acid has only one 3' end. Therefore, assembly of multiple components may require multiple temperature cycles. For example, assembling four components may involve three temperature cycles. Assembling six components may involve five temperature cycles. Assembling ten components may involve nine temperature cycles. In some implementations, assembly efficiency can be increased by using more temperature cycles than the minimum required. For example, using four temperature cycles to assemble two components can yield more product than using only one temperature cycle. This is because hybridization and extension of building blocks is a statistical event that occurs on a fraction of the total number of building blocks in each cycle, and therefore the total percentage of assembled building blocks can increase with increasing cycles.
[0073]
[0100] In addition to temperature cycling considerations, the design of nucleic acid sequences in OEPCR can affect the efficiency of their mutual assembly. Nucleic acids with long hybridization regions may hybridize more efficiently at a given annealing temperature than nucleic acids with short hybridization regions. This is because longer hybridized products contain a greater number of stable base pairs and therefore may result in a more stable overall hybridized product than shorter hybridized products. The hybridization region may have a length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or at least 10 or more bases.
[0074]
[0101] Hybridization regions with high guanine or cytosine content may hybridize more efficiently at a given temperature than hybridization regions with low guanine or cytosine content. This is because guanine forms more stable base pairs with cytosine than adenine does with thymine. Hybridization regions can have a guanine or cytosine content (also known as GC content) between 0% and 100%. For example, the hybridization region may have a guanine or cytosine content of 0% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35% to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 60% to 65%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 90% to 95%, or 95% to 100%.
[0075]
[0102] In addition to the length and GC content of the hybridization region, there are many additional aspects of nucleic acid sequence design that can affect the efficiency of OEPCR. For example, the formation of undesired secondary structures within a component can interfere with its ability to form a hybridization product with its intended neighboring component. These secondary structures can include hairpin loops. The types of possible secondary structures of nucleic acids and their stability (e.g., quantitation temperature) can be predicted based on the sequence. Design space search algorithms can be used to determine nucleic acid sequences that meet appropriate length and GC content criteria for efficient OEPCR while avoiding sequences with potentially inhibitory secondary structures. Design space search algorithms can include genetic algorithms, heuristic search algorithms, metaheuristic search strategies such as tabu search, branch-and-bound search algorithms, dynamic programming-based algorithms, constrained combinatorial optimization algorithms, gradient descent-based algorithms, randomized search algorithms, or combinations thereof.
[0076]
[0103] Similarly, the formation of homodimers (nucleic acid molecules that hybridize with nucleic acid molecules of the same sequence) and unwanted heterodimers (nucleic acid sequences that hybridize with other nucleic acid sequences other than their intended assembly partners) can interfere with OEPCR. Similar to secondary structures within nucleic acids, the formation of homodimers and heterodimers can be predicted and accounted for during nucleic acid design using computational methods and design space exploration algorithms.
[0077]
[0104] Longer nucleic acid sequences or higher GC content may result in increased formation of undesired secondary structures, homodimers, and heterodimers by OEPCR. Therefore, in some implementations, using shorter nucleic acid sequences or lower GC content may result in higher assembly efficiency. These design principles may preclude design strategies that use long hybridization regions or high GC content for more efficient assembly. Therefore, in some implementations, OEPCR may be optimized by using long hybridization regions with high GC content but short non-hybridization regions with low GC content. The total length of the nucleic acid may be at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or at least 100 bases or more. In some implementations, there may be an optimal length and optimal GC content for the hybridization region of the nucleic acid that optimizes assembly efficiency.
[0078]
[0105] More different nucleic acids in OEPCR reaction may hinder expected assembly efficiency.This is because more different nucleic acid sequences may increase the probability of undesired molecular interactions, especially in the form of heterodimers.Therefore, in some implementations of OEPCR that assemble a large number of components, the constraints on nucleic acid sequence may be more stringent for efficient assembly.
[0079]
[0106] Primers for amplifying the expected final assembly product can be included in the OEPCR reaction, which can then be run for more temperature cycles to improve the yield of assembled product, not only by generating more assemblies between the components, but also by exponentially amplifying the fully assembled product in the manner of conventional PCR.
[0080]
[0107] Additives can be included in the OEPCR reaction to improve assembly efficiency. For example, betaine, dimethyl sulfoxide (DMSO), non-ionic surfactants, formamide, magnesium, bovine serum albumin (BSA), or combinations thereof can be added. The additive content (weight per volume) can be at least 0%, 1%, 5%, 10%, or at least 20% or more.
[0081]
[0108] A variety of polymerases can be used in OEPCR. The polymerase can be naturally occurring or synthetic. An example of a polymerase is Φ29 polymerase or its derivatives. In some cases, a transcriptase or ligase is used in conjunction with or as a substitute for a polymerase to construct new nucleic acid sequences (i.e., an enzyme that catalyzes the formation of bonds). Examples of polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase Examples of polymerases include Tca polymerase, Tih polymerase, Tfi polymerase, Platinum-Taq polymerase, Tbr polymerase, Phusion polymerase, KAPA polymerase, Q5 polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, and Klenow fragment polymerase with 3' to 5' exonuclease activity, as well as variants, modifications, and derivatives thereof. Different polymerases may be stable and function optimally at different temperatures. Furthermore, different polymerases have different properties. For example, some polymerases, such as Phusion polymerase, may exhibit 3' to 5' exonuclease activity, which may contribute to higher fidelity during nucleic acid elongation. Some polymerases can replace leading sequences during elongation, while others may degrade them or terminate elongation. Some polymerases, such as Taq, incorporate an adenine base at the 3' end of a nucleic acid sequence.This process, called A-tailing, can be inhibitory to OEPCR because the addition of an adenine base can disrupt the designed 3' complementarity between intended adjacent components. OEPCR can also be called polymerase cycling assembly (or PCA).
[0082] B. Ligation Assembly
[0109] In ligation assembly, separate nucleic acids are assembled in a reaction containing one or more ligase enzymes and additional cofactors. Cofactors may include adenosine triphosphate (ATP), dithiothreitol (DTT), or magnesium ions (Mg2+). During ligation, the 3' end of one nucleic acid strand is covalently linked to the 5' end of another nucleic acid strand, thus forming an assembled nucleic acid. The components in the ligation reaction can be blunt-ended double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), or partially hybridized single-stranded DNA. Strategies for joining the ends of nucleic acids together can be used to increase the frequency of viable substrates for the ligase enzyme and thus improve the efficiency of the ligase reaction. Blunt-ended dsDNA molecules tend to form hydrophobic stacks that the ligase enzyme can act on, but a more successful strategy for joining nucleic acids may be to use nucleic acid components with 5' or 3' single-stranded overhangs that are complementary to the overhangs of the components to be assembled. In the latter case, base-base hybridization may result in the formation of a more stable nucleic acid duplex.
[0083]
[0110] When a double-stranded nucleic acid has an overhanging strand at one end, the other strand at that end may be referred to as a "cavity." Together, the cavity and overhang form a "sticky end," also known as a "sticky end." A sticky end can be either a 3' overhang and a 5' cavity or a 5' overhang and a 3' cavity. Sticky ends between two intended adjacent components can be designed so that the overhangs of both sticky ends are complementary, hybridizing directly adjacent to the beginning of a cavity on the other component. This creates a "nick" (a break in double-stranded DNA) that can be "sealed" (covalently linked via a phosphodiester bond) by the action of a ligase. Nicks can be sealed in either or both strands. Thermodynamically, the top and bottom strands of the molecules forming the sticky end can move between associated and dissociated states; therefore, sticky ends can be transient. However, once a nick along one strand of a sticky-ended duplex between two components is sealed, that covalent bond remains even when members of the opposite strand dissociate. The linked strands can then become templates to which intended adjacent members of the opposite strand can bind, forming a nick that can be resealed.
[0084]
[0111] Sticky ends can be created by digesting dsDNA with one or more endonucleases. Endonucleases (which may be called restriction enzymes) can target specific sites (which may be called restriction sites) at one or both ends of a dsDNA molecule and create staggered cuts (sometimes called digestion), thereby leaving sticky ends. Digestion can leave palindromic overhangs (overhangs with a sequence that is the reverse complement of itself). If so, two components digested with the same endonuclease can form complementary sticky ends along which they can be assembled with a ligase. If the endonuclease and ligase are compatible, digestion and ligation can be performed together in the same reaction. The reaction can occur at a uniform temperature, such as 4, 10, 16, 25, or 37°C. Alternatively, the reaction can be cycled between multiple temperatures, such as 16°C to 37°C. Cycling between multiple temperatures can allow digestion and ligation to proceed at their respective optimal temperatures during different portions of the cycle.
[0085]
[0112] It may be beneficial to perform digestion and ligation in separate reactions, for example, when the desired ligase and desired endonuclease function optimally under different conditions. Or, for example, when the ligated product forms a new restriction site for the endonuclease. In these cases, it may be better to perform restriction digestion followed by ligation separately, and perhaps it may be even more beneficial to remove the restriction enzyme before ligation. The nucleic acid may be separated from the enzyme by phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. Multiple endonucleases may be used in the same reaction, but care should be taken to ensure that the endonucleases do not interfere with each other and function under similar reaction conditions. Two endonucleases can be used to create orthogonal (non-complementary) sticky ends at both ends of a dsDNA building block.
[0086]
[0113] Endonuclease digestion leaves sticky ends with phosphorylated 5' ends. Ligase can only function on phosphorylated 5' ends, not on non-phosphorylated 5' ends. Therefore, an intermediate 5' phosphorylation step between digestion and ligation may not be necessary. Digested dsDNA building blocks with palindromic overhangs at their sticky ends can ligate to themselves. To prevent self-ligation, it may be beneficial to dephosphorylate the dsDNA building blocks before ligation.
[0087]
[0114] Multiple endonucleases may target different restriction sites but leave compatible overhangs (overhangs that are the reverse complement of each other). Ligation products of sticky ends created using two such endonucleases can result in an assembly product that does not contain a restriction site for either endonuclease at the ligation site. Such endonucleases form the basis of assembly methods such as BioBrick assembly, which can programmably assemble multiple components using only two endonucleases by performing repeated digestion-ligation cycles. Figure 20 shows an example of a digestion-ligation cycle using the endonucleases BamHI and BglII with compatible overhangs.
[0088]
[0115] In some implementations, the endonuclease used to create sticky ends can be a type IIS restriction enzyme. Because these enzymes cleave a fixed number of bases in a specific direction from their restriction site, the sequences of the overhangs they generate can be customized. The overhang sequences do not need to be palindromic. The same type of IIS restriction enzyme can be used to create multiple different sticky ends in the same reaction or multiple reactions. Furthermore, one or more type IIS restriction enzymes can be used to create building blocks with compatible overhangs in the same reaction or multiple reactions. The ligation site between two sticky ends generated by a type IIS restriction enzyme can be designed so that it does not form a new restriction site. In addition, the type IIS restriction enzyme site can be positioned on dsDNA so that the restriction enzyme cleaves its own restriction site when it generates a building block with a sticky end. Thus, the ligation product between multiple building blocks generated by a type IIS restriction enzyme does not need to contain a restriction site.
[0089]
[0116] A type IIS restriction enzyme can be mixed with a ligase in a reaction to perform digestion and ligation of components together. To promote optimal digestion and ligation, the temperature of the reaction can be cycled between two or more values. For example, digestion can be optimally performed at 37°C, and ligation can be optimally performed at 16°C. More commonly, the reaction can be cycled between temperature values of at least 0, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60°C, or at least 65°C or higher. Digestion and ligation reactions can be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more components. Examples of assembly reactions that utilize type IIS restriction enzymes to generate sticky ends include Golden Gate assembly (also known as Golden Gate cloning) or modular cloning (also known as MoClo).
[0090]
[0117] In some implementations of ligation, exonucleases can be used to create building blocks with sticky ends. 3' exonucleases can be used to push back the 3' end from dsDNA, thereby creating a 5' overhang. Similarly, 5' exonucleases can be used to feed back the 5' end from dsDNA, thus creating a 3' overhang. Different exonucleases can have different properties. For example, exonucleases can differ in the direction of their nuclease activity (5' to 3' or 3' to 5'), whether they act on ssDNA, whether they act on phosphorylated or non-phosphorylated 5' ends, whether they can initiate on a nick, or whether they can initiate their activity on a 5' cavity, a 3' cavity, a 5' overhang, or a 3' overhang. Various types of exonucleases include lambda exonuclease, RecJf, exonuclease III, exonuclease I, exonuclease T, exonuclease V, exonuclease VIII, exonuclease VII, nuclease BAL_31, T5 exonuclease, and T7 exonuclease.
[0091]
[0118] An exonuclease can be used in a reaction with a ligase to assemble multiple building blocks. The reaction can occur at a fixed temperature or cycle between temperatures ideal for the ligase or exonuclease, respectively. A polymerase can be included in the assembly reaction with a ligase and a 5' to 3' exonuclease. Building blocks in such a reaction can be designed so that building blocks intended to assemble adjacent to each other share homologous sequences at their edges. For example, building block X to be assembled with building block Y can have a 3' edge sequence of the form 5'-z-3', and building block Y can have a 5' edge sequence of the form 5'-z-3', where z is any nucleic acid sequence. We refer to such a form of homologous edge sequence as a "Gibson overlap." When a 5' exonuclease chews back the 5' end of a dsDNA building block with a Gibson overlap, it forms compatible 3' overhangs that hybridize with each other. The hybridized 3' end can then be extended by the action of a polymerase to the end of the template component or to the point where the extended 3' overhang of one component meets the 5' cavity of the adjacent component, thereby forming a nick that can be sealed by a ligase. Such assembly reactions, in which a polymerase, ligase, and exonuclease are used together, are often referred to as "Gibson assembly." Gibson assembly can be performed using T5 exonuclease, Phusion polymerase, and Taq ligase, and by incubating the reaction at 50°C. In the example above, the use of the thermophilic ligase Taq allows the reaction to proceed at 50°C, a temperature suitable for all three enzymes in the reaction.
[0092]
[0119] The term "Gibson assembly" generally refers to any assembly reaction involving a polymerase, a ligase, and an exonuclease. Gibson assembly can be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, or at least 10 or more components. Gibson assembly can occur as a single-step isothermal reaction or a multi-step reaction with one or more temperature incubations. For example, Gibson assembly can occur at temperatures of at least 30, 40, 50, 60 degrees, or at least 70 degrees or higher. The incubation time for Gibson assembly can be at least 1, 5, 10, 20, 40 minutes, or at least 80 minutes.
[0093]
[0120] Gibson assembly reactions occur optimally when the Gibson overlap between intended adjacent components is of a specific length and has sequence features, such as sequences that avoid undesired hybridization events, such as hairpins, homodimers, or unwanted heterodimers. Generally, a Gibson overlap of at least 20 bases is recommended. However, the Gibson overlap can be at least 1, 2, 3, 5, 10, 20, 30, 40, 50, 60, or at least 100 bases long. The GC content of the Gibson overlap can be anywhere from 0% to 100%. For example, the GC content of the Gibson duplication can be 0% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35% to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 60% to 65%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 90% to 95%, or 95% to 100%.
[0094]
[0121] Although Gibson assembly is generally described with a 5' exonuclease, the reaction can also occur with a 3' exonuclease. As the 3' exonuclease chews back the 3' end of a dsDNA building block, the polymerase counters the action by extending the 3' end. This dynamic process can continue until the 5' overhangs (generated by the exonuclease) of two building blocks (that share a Gibson overlap) hybridize and the polymerase extends the 3' end of one building block far enough to fill the 5' end of its neighboring building block, thereby leaving a nick that can be sealed by ligase.
[0095]
[0122] In some implementations of ligation, sticky-ended building blocks can be synthetically created by mixing together two single-stranded nucleic acids or oligos that do not share perfect complementarity, rather than enzymatically.
[0096]
[0123] The index and hybridization regions of oligos in sticky end ligation can be designed to facilitate proper assembly of the components. Components with long overhangs can hybridize to each other more efficiently at a given annealing temperature than components with short overhangs. The overhangs can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or at least 30 bases in length.
[0097]
[0124] Components with overhangs containing a high guanine or cytosine content may hybridize more efficiently to their complementary components at a given temperature than components with overhangs containing a low guanine or cytosine content. This is because guanine forms more stable base pairs with cytosine than adenine does with thymine. Overhangs can have a guanine or cytosine content (also known as GC content) between 0% and 100%.
[0098]
[0125] Similar to the overhang sequence, the GC content of an oligo and the length of its index region can also affect ligation efficiency. This is because sticky-end components can assemble more efficiently if the top and bottom strands of each component are stably linked. Therefore, index regions can be designed with higher GC content, longer sequences, and other features that promote higher melting temperatures. However, for both index regions and overhang sequences, there are many additional aspects of oligo design that can affect the efficiency of ligation assembly. For example, the formation of undesired secondary structures within a component can hinder its ability to form an assembled product with its intended neighboring components. This can occur due to secondary structures in either the index region, the overhang sequence, or both. These secondary structures can include hairpin loops. The types of possible secondary structures and the stability (e.g., quantitation temperature) of an oligo can be predicted based on the sequence. Design space exploration algorithms can be used to determine oligo sequences that meet appropriate length and GC content criteria for the formation of effective components while avoiding sequences with potentially inhibitory secondary structures. The design space exploration algorithm may include a genetic algorithm, a heuristic search algorithm, a metaheuristic search strategy such as tabu search, a branch and bound search algorithm, a dynamic programming based algorithm, a constrained combinatorial optimization algorithm, a gradient descent based algorithm, a randomized search algorithm, or a combination thereof.
[0099]
[0126] Similarly, the formation of homodimers (oligos that hybridize with oligos of the same sequence) and unwanted heterodimers (oligos that hybridize with other oligos than the intended assembly partner) can interfere with ligation. Similar to secondary structures within the building blocks, homodimer and heterodimer formation can be predicted and accounted for during oligo design using computational methods and design space exploration algorithms.
[0100]
[0127] Longer oligo sequences or higher GC content can lead to increased formation of undesired secondary structures, homodimers, and heterodimers within the ligation reaction. Therefore, in some implementations, using shorter oligos or lower GC content can result in higher assembly efficiency. These design principles may preclude design strategies that use longer oligos or higher GC content for more efficient assembly. Therefore, there may be an optimal length and GC content for each component oligo to optimize ligation assembly efficiency. The total length of the oligos used for ligation can be at least 10, 20, 30, 40, 50, 60, 70, 80, 90 bases, or at least 100 bases or more. The total GC content of the oligos used for ligation can be anywhere from 0% to 100%. For example, the total GC content of the oligos used in ligation can be 0% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35% to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 60% to 65%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 90% to 95%, or 95% to 100%.
[0101]
[0128] In addition to sticky-end ligation, ligation can also occur between single-stranded nucleic acids using staple (or template or bridge) strands. This method may be called staple-strand ligation (SSL), template-directed ligation (TDL), or bridge-strand ligation. In TDL, two single-stranded nucleic acids hybridize adjacently on a template, thus forming a nick that can be sealed by a ligase. The same nucleic acid design considerations for sticky-end ligation also apply to TDL. Stronger hybridization between the template and its intended complementary nucleic acid sequence can result in increased ligation efficiency. Therefore, sequence features that improve the hybridization stability (or melting temperature) on both sides of the template can improve ligation efficiency. These features may include longer sequence length and higher GC content. The length of the nucleic acid in the template-containing TDL can be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or at least 100 bases or more. The GC content of the template-containing nucleic acid can be any of 0% to 100%. For example, the GC content of the template-containing nucleic acid can be 0% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35% to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 60% to 65%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 90% to 95%, or 95% to 100%.
[0102]
[0129] In TDL, as with sticky end ligation, care can be taken to design building block and template sequences that avoid undesired secondary structures by using nucleic acid structure prediction software with sequence space search algorithms. Because the building blocks in TDL may be single-stranded rather than double-stranded, the incidence of undesired secondary structures may be higher (compared to sticky end ligation) due to exposed bases.
[0103]
[0130] TDL can also be performed using blunt-ended dsDNA building blocks. In such reactions, the staples may first need to displace or partially displace their complete single-stranded complements in order for the staple strands to properly crosslink two single-stranded nucleic acids. To facilitate the TDL reaction with dsDNA building blocks, the dsDNA can first be melted by incubation at high temperature. The reaction can then be cooled to allow the staple strands to anneal to their appropriate nucleic acid complements. This process can be made more efficient by using a relatively high concentration of template compared to the dsDNA building blocks, thus allowing the template to outcompete the appropriate full-length ssDNA complement for ligation. Once two ssDNA strands are assembled by their template and ligase, the assembled nucleic acid can serve as a template for the opposite full-length ssDNA complement. Therefore, ligation of blunt-ended dsDNA with TDL can be improved by multiple rounds of melting (incubation at high temperature) and annealing (incubation at low temperature). This process may be referred to as ligase silencing reaction (LCR). Appropriate melting and annealing temperatures depend on the nucleic acid sequence. Melting and annealing temperatures can be at least 4, 10, 20, 20, 30, 40, 50, 60, 70, 80, 90, or 100° C. The number of temperature cycles can be at least 1, 5, 10, 15, 20, 15, 30, or more.
[0104]
[0131] All ligations can be performed in fixed-temperature or multi-temperature reactions. Ligation temperatures can be at least 0, 4, 10, 20, 30, 40, 50, or 60°C or higher. The optimal temperature for ligase activity can vary depending on the type of ligase. Furthermore, the rate at which components mate or hybridize in a reaction can vary depending on their nucleic acid sequences. Higher incubation temperatures can promote faster diffusion and therefore increase the frequency at which components transiently mate or hybridize. However, increasing the temperature can also disrupt base pairing and therefore reduce the stability of the adjacent or hybridized component duplexes. The optimal temperature for ligation can depend on other factors, such as the number of nucleic acids to be constructed, the sequence of those nucleic acids, the type of ligase, and reaction additives. For example, two sticky-end components with four-base complementary overhangs can assemble faster at 4°C using T4 ligase than at 25°C using T4 ligase. However, two sticky-end components with 25-base complementary overhangs can assemble faster at 25°C with T4 ligase than at 4°C with T4 ligase, and likely faster than ligation with 4-base overhangs at any temperature. In some implementations of ligation, it may be beneficial to heat and slowly cool the components to anneal before adding ligase.
[0105]
[0132] Ligation can be used to assemble at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleic acids. Ligation incubation times can be up to 30 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes, 20 minutes, 30 minutes, 1 hour, or longer. Longer incubation times can improve ligation efficiency.
[0106]
[0133] Ligation may require nucleic acids with 5' phosphorylated ends. Nucleic acid building blocks without 5' phosphorylated ends can be phosphorylated by reaction with a polynucleotide kinase such as T4 polynucleotide kinase (or T4 PNK). Other cofactors, such as ATP, magnesium ions, or DTT, may be present in the reaction. The polynucleotide kinase reaction may occur at 37°C for 30 minutes. The polynucleotide kinase reaction temperature may be at least 4, 10, 20, 20, 30, 40, 50, or 60°C. The incubation time for the polynucleotide kinase reaction may be up to 1, 5, 10, 20, 30, 60, or more minutes. Alternatively, nucleic acid building blocks may be synthetically designed and manufactured (rather than enzymatically) using modified 5' phosphorylation. Only nucleic acids being assembled at the 5' end may require phosphorylation. For example, templates in TDLs may not be phosphorylated because they are not intended to be assembled.
[0107]
[0134] To improve ligation efficiency, additives can be included in the ligation reaction. For example, dimethyl sulfoxide (DMSO), polyethylene glycol (PEG), 1,2-propanediol (1,2-Pr), glycerol, Tween-20, or a combination thereof can be added. PEG 6000 can be a particularly effective ligation enhancer. PEG 6000 can increase ligation efficiency by acting as a crowding agent. For example, PEG 6000 can form clumped nodules that occupy space in the ligase reaction solution and bring the ligase and components into closer proximity. The additive content (weight per volume) can be at least 0%, 1%, 5%, 10%, 20%, or more.
[0108]
[0135] Various ligases can be used for ligation. Ligases can be naturally occurring or synthetic. Examples of ligases include T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, Taq DNA ligase, 9oNTM DNA ligase, E. coli DNA ligase, and Splint® DNA ligase. Different ligases can be stable and function optimally at different temperatures. For example, Taq DNA ligase is thermostable, while T4 DNA ligase is not. Furthermore, different ligases have different properties. For example, T4 DNA ligase can ligate blunt-ended dsDNA, while T7 DNA ligase cannot.
[0109]
[0136] Ligation can be used to attach sequencing adaptors to a library of nucleic acids. For example, ligation can be performed using a common sticky end or staple at the end of each member of the nucleic acid library. Sequencing adaptors can be asymmetrically ligated when the sticky end or staple at one end of the nucleic acid is different from the sticky end or staple at the other end. For example, a forward sequencing adaptor can be ligated to one end of a member of the nucleic acid library, and a reverse sequencing adaptor can be ligated to the other end of the member of the nucleic acid library. Alternatively, blunt-end ligation can be used to add adaptors to a library of blunt-ended double-stranded nucleic acids. Forked adaptors can be used to asymmetrically add adaptors to a nucleic acid library with either blunt or sticky ends where each end (such as an A-tail) is equivalent.
[0110]
[0137] Ligation can be inhibited by heat inactivation (eg, incubation at 65° C. for at least 20 minutes), addition of denaturing agents, or addition of chelating agents such as EDTA.
[0111] C. Restriction Digest
[0138] A restriction digest is a reaction in which restriction endonucleases (or restriction enzymes) recognize their cognate restriction sites on a nucleic acid and subsequently cleave (or digest) the nucleic acid containing the restriction site. Type I, II, III, or IV restriction enzymes can be used in a restriction digest. Type II restriction enzymes may be the most efficient restriction enzymes for nucleic acid digestion. Type II restriction enzymes can recognize palindromic restriction sites and cleave the nucleic acid within the recognition site. Examples of such restriction enzymes (and their restriction sites) include AatII (GACGTC), AfeI (AGCGCT), ApaI (GGGCCC), DpnI (GATC), EcoRI (GAATTC), NgeI (GCTAGC), etc. Some restriction enzymes, such as DpnI and AfeI, can cleave their restriction sites centrally, leaving blunt-ended dsDNA products. Other restriction enzymes, such as EcoRI and AatII, cleave their restriction sites off-center, thereby leaving dsDNA products with sticky (or staggered) ends. Some restriction enzymes can target discontinuous restriction sites. For example, the restriction enzyme AlwNI recognizes the restriction site CAGNNNCTG, where N can be A, T, C, or G. Restriction sites can be at least 2, 4, 6, 8, 10 or more bases in length.
[0112]
[0139] Some type II restriction enzymes cleave nucleic acids outside their restriction sites. Enzymes can be subclassified as either type IIS or type IIG restriction enzymes. These enzymes can recognize non-palindromic restriction sites. An example of such an enzyme is BbsI, which recognizes GAAAC and produces staggered cuts of 2 (same strand) and 6 (opposite strand) bases further downstream. Another example is BsaI, which recognizes GGTCTC and produces staggered cuts of 1 (same strand) and 5 (opposite strand) bases further downstream. These restriction enzymes can be used in Golden Gate assembly or modular cloning (MoClo). Some restriction enzymes, such as BcgI (type IIG restriction enzyme), can produce staggered cuts on both ends of their recognition site. Restriction enzymes can cleave nucleic acids at least 1, 5, 10, 15, 20, or more bases from their recognition site. Because these restriction enzymes can produce staggered cuts outside their recognition sites, the sequences of the resulting nucleic acid overhangs can be arbitrarily designed. This is in contrast to restriction enzymes, which produce staggered cuts within their recognition sites, where the sequence of the resulting nucleic acid overhang is coupled to the sequence of the restriction site. The nucleic acid overhangs created by restriction digests can be at least 1, 2, 3, 4, 5, 6, 7, 8, or more bases in length. When a restriction enzyme cleaves a nucleic acid, the resulting 5' end contains a phosphate.
[0113]
[0140] One or more nucleic acid sequences can be included in a restriction digestion reaction. Similarly, one or more restriction enzymes can be used together in a restriction digestion reaction. The restriction digest can contain additives and cofactors, including potassium ions, magnesium ions, sodium ions, BSA, S-adenosyl-L-methionine (SAM), or combinations thereof. The restriction digestion reaction can be incubated at 37°C for 1 hour. The restriction digestion reaction can be incubated at a temperature of at least 0, 10, 20, 30, 40, 50, or 60°C. The optimal digestion temperature can depend on the enzyme. The restriction digestion reaction can be incubated for up to 1, 10, 30, 60, 90, 120 minutes, or more. Longer incubation times can result in increased digestion.
[0114] D. Nucleic Acid Amplification
[0141] Nucleic acid amplification can be performed using polymerase chain reaction (PCR). In PCR, a starting pool of nucleic acids (called the template pool or template) can be combined with a polymerase, primers (short nucleic acid probes), nucleotide triphosphates (e.g., dATP, dTTP, dCTP, dGTP, and their analogs or variants), and additional cofactors and additives such as betaine, DMSO, and magnesium ions. Templates can be single-stranded or double-stranded nucleic acids. Primers can be short nucleic acid sequences synthetically constructed to complement and hybridize with target sequences in the template pool. Typically, there are two primers in a PCR reaction: one complementary to the primer binding site on the top strand of the target template and the other complementary to the primer binding site on the bottom strand of the target template downstream of the first binding site. The 5' to 3' orientation of these primers binding to their targets must be opposite each other to successfully replicate and exponentially amplify the nucleic acid sequence between them. "PCR" may typically refer specifically to this form of reaction, but may also be used more generally to refer to any nucleic acid amplification reaction.
[0115]
[0142] In some implementations, PCR can involve cycling between three temperatures: a melting temperature, an annealing temperature, and an extension temperature. The melting temperature is intended to convert double-stranded nucleic acids into single-stranded nucleic acids and eliminate the formation of hybridization products and secondary structures. Typically, the melting temperature is high, e.g., greater than 95°C. In some implementations, the melting temperature can be at least 96, 97, 98, 99, 100, 101, 102, 103, 104, or 105°C. In other implementations, the melting temperature can be up to 95, 94, 93, 92, 91, or 90°C. A higher melting temperature improves the dissociation of nucleic acids and their secondary structures, but may also cause side effects such as degradation of nucleic acids or polymerase. The melting temperature can be applied to the reaction for at least 1, 2, 3, 4, 5, or more seconds, e.g., 30 seconds, 1 minute, 2 minutes, or 3 minutes. For PCR with complex or long templates, a longer initial melting temperature step may be recommended.
[0116]
[0143] The annealing temperature is intended to promote hybridization between the primers and their target templates. In some implementations, the annealing temperature can match the calculated melting temperature of the primers. In other implementations, the annealing temperature can be within 10°C or more of the melting temperature. In some implementations, the annealing temperature can be at least 25, 30, 50, 55, 60, 65, or 70°C. The melting temperature can depend on the sequence of the primers. Longer primers can have higher melting temperatures, and primers with a higher percentage of guanine or cytosine nucleotides can have higher melting temperatures. Therefore, it may be possible to design primers that are optimally assembled at a specific annealing temperature. The annealing temperature can be applied to the reaction for at least 1, 5, 10, 15, 20, 25, or 30 seconds or more. To help ensure annealing, the primer concentration can be high or saturating. The primer concentration can be 500 nanomolar (nM). Primer concentrations can be up to 1 nM, 10 nM, 100 nM, 1000 nM or more.
[0117]
[0144] The extension temperature is intended to initiate and promote nucleic acid chain extension of the 3' end of the primer, catalyzed by one or more polymerase enzymes. In some implementations, the extension temperature can be set to a temperature at which the polymerase functions optimally in terms of nucleic acid binding strength, extension rate, extension stability, or fidelity. In some implementations, the extension temperature can be at least 30, 40, 50, 60, or 70°C or higher. The annealing temperature can be applied to the reaction for at least 1, 5, 10, 15, 20, 25, 30, 40, 50, or 60 seconds or more. The recommended extension time can be approximately 15 to 45 seconds per kilobase of expected extension.
[0118]
[0145] In some implementations of PCR, the annealing temperature and the extension temperature can be the same. Therefore, two-stage temperature cycles can be used instead of three-stage temperature cycles. Examples of combinations of annealing and extension temperatures include 60, 65, or 72°C.
[0119]
[0146] In some implementations, PCR can be performed in one temperature cycle. Such implementations can include converting target single-stranded template nucleic acids into double-stranded nucleic acids. In other implementations, PCR can be performed in multiple temperature cycles. If PCR is efficient, the number of target nucleic acid molecules is expected to double with each cycle, thereby exponentially increasing the number of target nucleic acid templates from the original template pool. PCR efficiency can vary. Therefore, the actual percentage of target nucleic acids replicated each time can be greater or less than 100%. Each PCR cycle can introduce undesirable artifacts, such as mutant and recombinant nucleic acids. To reduce this potential harm, polymerases with high fidelity and high processivity can be used. Furthermore, a limited number of PCR cycles can be used. PCR can include up to 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, or more cycles.
[0120]
[0147] In some implementations, multiple different target nucleic acid sequences can be amplified together in a single PCR. If each target sequence has a common primer binding site, all nucleic acid sequences can be amplified with the same primer set. Alternatively, the PCR can include multiple primers intended to target each different nucleic acid. Such PCR can be referred to as multiplex PCR. The PCR can include up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more different primers. In PCR using multiple different nucleic acid targets, each PCR cycle can change the relative distribution of the target nucleic acids. For example, a uniform distribution can become skewed or unevenly distributed. To reduce this potential detriment, an optimal polymerase (e.g., with high fidelity and sequence robustness) and optimal PCR conditions can be used. Factors such as annealing and extension temperatures and times can be optimized. Furthermore, a limited number of PCR cycles can be used.
[0121]
[0148] In some implementations of PCR, a primer with a base mismatch at its target primer binding site in the template can be used to mutate a target sequence. In some implementations of PCR, a primer with extra sequence at its 5' end (known as an overhang) can be used to add sequence to its target nucleic acid. For example, primers containing a sequencing adapter at their 5' end can be used to prepare and / or amplify a nucleic acid library for sequencing. Primers targeting the sequencing adapter can be used to amplify a nucleic acid library to enrich it sufficiently for a particular sequencing technology.
[0122]
[0149] In some implementations, linear PCR (or asymmetric PCR) is used, in which primers target only one strand of the template (not both strands). In linear PCR, the replicated nucleic acid from each cycle is not complementary to the primer, so the primer does not bind to it. Therefore, the primer only replicates the original target template in each cycle, resulting in linear (as opposed to exponential) amplification. Amplification from linear PCR may not be as fast as conventional (exponential) PCR, but the maximum yield may be greater. Theoretically, primer concentration in linear PCR may not be a limiting factor with increasing cycles and increasing yield, as in conventional PCR. Linear exponential PCR (or LATE-PCR) is a modified version of linear PCR that may be capable of particularly high yields.
[0123]
[0150] In some implementations of nucleic acid amplification, the melting, annealing, and extension processes can occur at a single temperature. Such PCR can be called isothermal PCR. Isothermal PCR can utilize a temperature-independent method to dissociate or displace perfectly complementary nucleic acid strands from each other in favor of primer binding. Strategies include loop-mediated isothermal amplification, strand displacement amplification, helicase-dependent amplification, and nicking enzyme amplification reactions. Isothermal nucleic acid amplification can occur at temperatures up to 20, 30, 40, 50, 60, or 70°C or higher.
[0124]
[0151] In some implementations, PCR may further include a fluorescent probe or dye for quantifying the amount of nucleic acid in a sample. For example, the dye may be intercalated into double-stranded nucleic acid. An example of such a dye is SYBR Green. The fluorescent probe may also be a nucleic acid sequence attached to a fluorescent unit. The fluorescent unit may be released upon hybridization of the probe to the target nucleic acid and subsequent modification from the elongating polymerase unit. An example of such a probe is a Taqman probe. Such a probe may be used in conjunction with PCR and optical measurement tools (for excitation and detection) to quantify the concentration of nucleic acid in a sample. This process may be referred to as quantitative PCR (qPCR) or real-time PCR (rtPCR).
[0125]
[0152] In some implementations, PCR can be performed on a single molecular template (a process that can be called single-molecule PCR) rather than a pool of multiple template molecules. For example, emulsion-PCR (ePCR) can be used to encapsulate a single nucleic acid molecule within an aqueous droplet in an oil emulsion. The droplets can also contain PCR reagents, and the droplets can be maintained in a temperature-controlled environment that allows for the temperature cycling required for PCR. In this way, multiple self-contained PCR reactions can occur simultaneously with high throughput. The stability of the oil emulsion can be improved with surfactants. The movement of the droplets can be controlled by pressure through microfluidic channels. Microfluidic devices can be used to generate droplets, split droplets, merge droplets, inject materials into droplets, and incubate droplets. The size of the aqueous droplets in the oil emulsion can be at least 1 picoliter (pL), 10 pL, 100 pL, 1 nanoliter (nL), 10 nL, 100 nL, or more.
[0126]
[0153] In some implementations, single molecule PCR can be performed on a solid-phase substrate. For example, Illumina solid-phase amplification or its variants can be used. The template pool can be exposed to a solid-phase substrate, which can immobilize the templates with a certain spatial resolution. Bridge amplification can then occur within the spatial vicinity of each template, thereby amplifying single molecules on the substrate in a high-throughput manner.
[0127]
[0154] High-throughput single-molecule PCR can be useful for amplifying pools of distinct nucleic acids that may interfere with each other. For example, if multiple different nucleic acids share a common sequence region, recombination between the nucleic acids along this common region can occur during the PCR reaction, resulting in new recombinant nucleic acids. Single-molecule PCR prevents this potential amplification error because it compartmentalizes different nucleic acid sequences so they do not interact with each other. Single-molecule PCR can be particularly useful for preparing nucleic acids for sequencing. Single-molecule PCR methods are also useful for absolute quantitation of several targets within a template pool. For example, digital PCR (or dPCR) uses the frequency of distinct single-molecule PCR amplification signals to estimate the number of starting nucleic acid molecules in a sample.
[0128]
[0155] In some implementations of PCR, a group of nucleic acids can be indiscriminately amplified using primers for primer binding sites common to all nucleic acids. For example, primers for primer binding sites flanking all nucleic acids in a pool. Synthetic nucleic acid libraries can be created or constructed using these common sites for general amplification. However, in some implementations, PCR can be used to selectively amplify targeted subsets of nucleic acids from a pool. For example, by using primers with primer binding sites that appear only on the targeted subset of nucleic acids. Synthetic nucleic acid libraries can be created or constructed for selective amplification of sub-libraries from a more general library, such that nucleic acids belonging to a potential sub-library of interest all share common primer binding sites on their edges (common within the sub-library but distinct from other sub-libraries). In some implementations, PCR can be combined with nucleic acid assembly reactions (such as ligation or OEPCR) to selectively amplify fully assembled or potentially fully assembled nucleic acids from partially assembled or misassembled (or unintended or undesired) by-products. For example, assembly can involve assembling nucleic acids using primer binding sites on each edge sequence so that only fully assembled nucleic acid products contain the two primer binding sites necessary for amplification. In the example above, partially assembled products can contain none or only one edge sequence with a primer binding site and therefore should not be amplified. Similarly, incorrectly assembled (or unintended or undesired) products can contain none or only one edge sequence, or both edge sequences but may be in the wrong orientation or separated by the wrong amount of bases. Thus, the incorrectly assembled products should not be amplified to produce products of the wrong length.In the latter case, the amplified misassembled products of the wrong length can be separated from the amplified fully assembled products of the correct length by nucleic acid size selection methods, such as DNA electrophoresis in an agarose gel, followed by gel extraction.
[0129]
[0156] To improve the efficiency of nucleic acid amplification, additives can be included in PCR. For example, betaine, dimethyl sulfoxide (DMSO), non-ionic surfactants, formamide, magnesium, bovine serum albumin (BSA), or combinations thereof can be added. The content of additives (weight per volume) can be at least 0%, 1%, 5%, 10%, 20%, or more.
[0130]
[0157] A variety of polymerases can be used in PCR. The polymerase can be naturally occurring or synthetic. An example of a polymerase is Φ29 polymerase or its derivatives. In some cases, a transcriptase or ligase (i.e., an enzyme that catalyzes the formation of bonds) is used in conjunction with or as a substitute for a polymerase to construct new nucleic acid sequences. Examples of polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, and bacteriophage T4. DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, Ex-Taq polymerase, LA-Taw polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase Examples of polymerases include Tca polymerase, Tih polymerase, Tfi polymerase, Platinum-Taq polymerase, Tbr polymerase, Phusion polymerase, KAPA polymerase, Q5 polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, KOD polymerase, Bst polymerase, Sac polymerase, and Klenow fragment polymerase with 3' to 5' exonuclease activity, as well as variants, modifications, and derivatives thereof. Different polymerases may be stable and function optimally at different temperatures. Furthermore, different polymerases have different properties. For example, some polymerases, such as Phusion polymerase, may exhibit 3' to 5' exonuclease activity, which may contribute to higher fidelity during nucleic acid elongation. Some polymerases can replace leading sequences during elongation, while others may degrade them or terminate elongation. Some polymerases, such as Taq, incorporate an adenine base at the 3' end of a nucleic acid sequence.Additionally, some polymerases may have higher fidelity and processivity than others and may be better suited for PCR applications such as sequencing preparations where it is important that the amplified nucleic acid yield has minimal variation and that the distribution of different nucleic acids remains uniform throughout the amplification.
[0131] E. Size Selection
[0158] Nucleic acids of specific sizes can be selected from a sample using size-selection techniques. In some implementations, size selection can be performed using gel electrophoresis or chromatography. A liquid sample of nucleic acids can be loaded at one end of a stationary phase or gel (or matrix). A voltage difference can be placed across the gel so that the negative end of the gel is the end where the nucleic acid sample is loaded and the positive end of the gel is the opposite end. Because nucleic acids have a negatively charged phosphate backbone, they migrate across the gel to the positive end. The size of the nucleic acids determines the relative speed of migration through the gel. Thus, nucleic acids of different sizes resolve on the gel as they migrate. The voltage difference can be 100 V or 120 V. The voltage difference can be up to 50 V, 100 V, 150 V, 200 V, 250 V, or more. Larger voltage differences can increase the speed of nucleic acid migration and size resolution. However, larger voltage differences can also damage the nucleic acids or the gel. To separate nucleic acids of larger sizes, a larger voltage difference may be recommended. Typical migration times may be 15 to 60 minutes. Migration times may be up to 10, 30, 60, 90, 120 minutes, or longer. Longer migration times, like higher voltages, may result in better nucleic acid resolution, but may also result in increased nucleic acid damage. To separate nucleic acids of larger sizes, a longer migration time may be recommended. For example, a voltage difference of 120V and a migration time of 30 minutes may be sufficient to separate nucleic acids of 200 bases from nucleic acids of 250 bases.
[0132]
[0159] The properties of the gel or matrix can affect the size selection process. Gels typically contain polymeric materials such as agarose or polyacrylamide dispersed in a conductive buffer such as TAE (Tris-acetate-EDTA) or TBE (Tris-borate-EDTA). The content (weight per volume) of material (e.g., agarose or acrylamide) in the gel can be up to 0.5%, 1%, 2%, 3%, 5%, 10%, 15%, 20%, 25%, or more. Higher contents can slow migration rates. Higher contents may be preferable for resolving smaller nucleic acids. Agarose gels may be better at separating double-stranded DNA (dsDNA). Polyacrylamide gels may be better at separating single-stranded DNA (ssDNA). The preferred gel composition may depend on the type and size of nucleic acids, the compatibility of additives (e.g., dyes, stains, denaturing solutions, or loading buffers), and the anticipated downstream application (e.g., gel extraction followed by ligation, PCR, or sequencing). Agarose gels may be easier to gel extract than polyacrylamide gels. TAE, although not as good a conductor as TBE, may also be better for gel extraction because carryover of borate (an enzyme inhibitor) in the extraction process can inhibit downstream enzymatic reactions.
[0133]
[0160] The gel may further contain a denaturing solution such as SDS (sodium dodecyl sulfate) or urea. SDS can be used, for example, to denature proteins or separate nucleic acids from potentially bound proteins. Urea can be used to denature secondary structures in DNA. For example, urea can convert dsDNA to ssDNA, or urea can convert folded ssDNA (e.g., hairpins) to unfolded ssDNA. Urea-polyacrylamide gels (further containing TBE) can be used to precisely resolve ssDNA.
[0134]
[0161] Samples can be incorporated into gels of different formats. In some implementations, gels can include wells into which samples can be manually loaded. A single gel can have multiple wells for running multiple nucleic acid samples. In other implementations, gels can be attached to microfluidic channels into which nucleic acid samples are automatically loaded. Each gel can be downstream of several microfluidic channels, or each gel can occupy its own separate microfluidic channel. The dimensions of the gel can affect the sensitivity of nucleic acid detection (or visualization). For example, thin gels or gels within microfluidic channels (e.g., in bioanalyzers or tapestries) can improve the sensitivity of nucleic acid detection. The nucleic acid detection step can be important for selecting and extracting nucleic acid fragments of the correct size.
[0135]
[0162] A ladder can be loaded onto a gel for nucleic acid size reference. The ladder can include markers of different sizes to which nucleic acid samples can be compared. Different ladders can have different size ranges and resolutions. For example, a 50-base ladder can have markers of 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, and 600 bases. The ladder can be useful for detecting and selecting nucleic acids within the size range of 50 and 600 bases. The ladder can also be used as a standard for estimating the concentration of nucleic acids of different sizes in a sample.
[0136]
[0163] The nucleic acid sample and ladder may be mixed with a loading buffer to facilitate the gel electrophoresis (or chromatography) process. The loading buffer may contain dyes and markers to facilitate tracking the migration of nucleic acids. The loading buffer may further contain a reagent (such as glycerol) that is denser than the running buffer (e.g., TAE or TBE) to ensure that the nucleic acid sample sinks to the bottom of the sample loading well (so that it can be immersed in the running buffer). The loading buffer may further contain a denaturant such as SDS or urea. The loading buffer may further contain a reagent to improve the stability of the nucleic acid. For example, the loading buffer may contain EDTA to protect the nucleic acid from nucleases.
[0137]
[0164] In some implementations, the gel may contain a stain that binds to nucleic acids and can be used to optically detect nucleic acids of different sizes. The stain may be specific for dsDNA, ssDNA, or both. Different stains may be compatible with different gel materials. Some stains may require excitation from a light source (or electromagnetic waves) to visualize. The light source may be UV (ultraviolet) or blue light. In some implementations, the stain may be added to the gel before electrophoresis. In other implementations, the stain may be added to the gel after electrophoresis. Examples of stains include ethidium bromide (EtBr), SYBR Safe, SYBR Gold, silver stain, or methylene blue. A reliable method for visualizing dsDNA of specific sizes may be, for example, using an agarose-TAE gel containing SYBR Safe or EtBr stain. A reliable method for visualizing ssDNA of specific sizes may be, for example, using a urea-polyacrylamide TBE gel with methylene blue or silver stain.
[0138]
[0165] In some implementations, the movement of nucleic acids through the gel can be driven by methods other than electrophoresis, for example, gravity, centrifugation, vacuum, or pressure can be used to drive the nucleic acids through the gel and resolve them according to their size.
[0139]
[0166] Nucleic acids of a specific size can be extracted from the gel using a blade or razor, and the gel band containing the nucleic acid can be excised. Appropriate optical detection techniques and DNA ladders can be used to ensure that excision is precise at the specific band and effectively excludes nucleic acids that may belong to different, undesired size bands. The gel band can be incubated with a buffer to dissolve it, thus releasing the nucleic acid into the buffer. Heat or physical agitation can facilitate dissolution. Alternatively, the gel band can be incubated in the buffer long enough to allow diffusion of the DNA into the buffer without the need for gel dissolution. The buffer can then be separated from the remaining solid-phase gel, for example, by aspiration or centrifugation. The nucleic acids can then be purified from the solution using standard purification or buffer exchange techniques, such as phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. The nucleic acids can also be concentrated during this process.
[0140]
[0167] As an alternative to gel excision, nucleic acids of a specific size can be separated from the gel by allowing them to flow out. The migrating nucleic acids can pass through basins (or wells) that are either embedded in the gel or at the end of the gel. The migration process can be timed or optically monitored so that samples are collected from the basin once a specific size population of nucleic acids enters the basin. Collection can be performed, for example, by aspiration. The nucleic acids can then be purified from the collected solution using standard purification or buffer exchange techniques, such as phenol-chloroform extraction, ethanol precipitation, magnetic bead capture and / or silica membrane adsorption, washing, and elution. Nucleic acids can also be concentrated during this process.
[0141]
[0168] Other methods for nucleic acid size selection may include mass spectrometry or membrane-based filtration. In some implementations of membrane-based filtration, nucleic acids are passed through a membrane (e.g., a silica membrane) that can preferentially bind either dsDNA, ssDNA, or both. The membrane can be designed to preferentially capture nucleic acids of at least a specific size. For example, the membrane can be designed to filter out nucleic acids less than 20, 30, 40, 50, 70, 90, or more bases. The membrane-based size selection techniques may not be as rigorous as gel electrophoresis or chromatography.
[0142] F. Nucleic acid capture
[0169] Affinity-tagged nucleic acids can be used as sequence-specific probes for nucleic acid capture. The probe can be designed to complement a target sequence in a pool of nucleic acids. The probe can then be incubated with the nucleic acid pool and hybridized to its target. The incubation temperature can be lower than the melting temperature of the probe to facilitate hybridization. The incubation temperature can be 5, 10, 15, 20, 25°C or lower than the melting temperature of the probe. The hybridized target can be captured on a solid substrate that specifically binds to the affinity tag. The solid substrate can be a membrane, well, column, or bead. Multiple washes can remove all non-hybridized nucleic acids from the target. Washes can be performed at temperatures lower than the melting temperature of the probe to facilitate stable immobilization of the target sequence during washing. The wash temperature can be 5, 10, 15, 20, 25°C or lower than the melting temperature of the probe. A final elution step can recover the nucleic acid target from the solid phase-substrate and the affinity-tagged probe. The elution step can be performed at a temperature above the melting temperature of the probe to facilitate the release of the nucleic acid target into the elution buffer. The elution temperature can be up to 5, 10, 15, 20, 25°C or higher than the melting temperature of the probe.
[0143]
[0170] In certain implementations, oligonucleotides bound to a solid substrate can be removed from the solid substrate by exposure to conditions such as acid, base, oxidation, reduction, heat, light, metal ion catalysis, substitution or elimination chemistry, or by enzymatic cleavage. In certain embodiments, oligonucleotides can be attached to a solid support via a cleavable linking moiety. For example, a solid support can be functionalized to provide a cleavable linker for covalent attachment to a targeting oligonucleotide. In some embodiments, the linker moiety can be six or more atoms in length. In some embodiments, the cleavable linker can be a TOPS (two oligonucleotides per synthesis) linker, an amino linker, or a photocleavable linker.
[0144]
[0171] In some implementations, biotin can be used as an affinity tag immobilized on a solid-phase substrate by streptavidin. Biotinylated oligonucleotides for use as nucleic acid capture probes can be designed and manufactured. Oligonucleotides can be biotinylated at the 5' or 3' end. They can also be internally biotinylated on thymine residues. Increasing the amount of biotin on an oligo can result in stronger capture on a streptavidin substrate. Biotin at the 3' end of an oligo can prevent the oligo from being extended during PCR. The biotin tag can be a variant of standard biotin. For example, biotin variants can be biotin-TEG (triethylene glycol), double biotin, PC-biotin, desthiobiotin-TEG, and biotin azide. Double biotin can increase biotin-streptavidin affinity. Biotin-TEG adds a biotin group to the nucleic acid separated by a TEG linker. This can prevent biotin from interfering with the function of the nucleic acid probe, such as its hybridization to a target. A nucleic acid biotin linker can also be added to the probe. The nucleic acid linker can contain a nucleic acid sequence that is not intended to hybridize to a target.
[0145]
[0172] Biotinylated nucleic acid probes can be designed based on how well they hybridize to their targets. Nucleic acid probes with higher design melting temperatures may hybridize more strongly to their targets. Longer nucleic acid probes and probes with higher GC content may hybridize more strongly due to increased melting temperatures. Nucleic acid probes can be at least 5, 10, 15, 20, 30, 40, 50, or 100 bases in length. Nucleic acid probes can have a GC content of 0 to 100%. Care can be taken to ensure that the probe's melting temperature does not exceed the temperature tolerance range of the streptavidin substrate. Nucleic acid probes can be designed to avoid inhibitory secondary structures such as hairpins, homodimers, and heterodimers with off-target nucleic acids. There can be a trade-off between probe melting temperature and off-target binding. There may be an optimal probe length and GC content that results in a high melting temperature and low off-target binding. Synthetic nucleic acid libraries can be designed so that their nucleic acids contain efficient probe binding sites.
[0146]
[0173] The solid-phase streptavidin substrate can be magnetic beads. The magnetic beads can be immobilized using a magnetic strip or plate. The magnetic strip or plate can be contacted with a container to immobilize the magnetic beads to the container. Conversely, the magnetic strip or plate can be removed from the container to release the magnetic beads from the container wall into solution. Different bead characteristics can affect their application. The beads can have a variety of sizes. For example, the beads can be anywhere from 1 to 3 micrometers (μm) in diameter. The beads can have diameters up to 1, 2, 3, 4, 5, 10, 15, 20 micrometers, or more. The bead surface can be hydrophobic or hydrophilic. The beads can be coated with a blocking protein, such as BSA. Prior to use, the beads can be washed or pretreated with an additive, such as a blocking solution, to prevent nonspecific binding to nucleic acids.
[0147]
[0174] The biotinylated probes can be bound to magnetic streptavidin beads before incubation with the nucleic acid sample pool. This process can be called direct capture. Alternatively, the biotinylated probes can be incubated with the nucleic acid sample pool before adding the magnetic streptavidin beads. This process can be called indirect capture. The indirect capture method can improve target yield. Shorter nucleic acid probes can require less time to bind to the magnetic beads.
[0148]
[0175] Optimal incubation of the nucleic acid probe with the nucleic acid sample may be performed at a temperature 1°C to 10°C or more below the melting temperature of the probe. The incubation temperature may be up to 5, 10, 20, 30, 40, 50, 60, 70, 80°C, or more. The recommended incubation time may be 1 hour. The incubation time may be up to 1, 5, 10, 20, 30, 60, 90, 120 minutes, or more. Longer incubation times may result in better capture efficiency. An additional 10 minutes of incubation can be performed after adding streptavidin beads to allow biotin-streptavidin coupling. This additional time may be up to 1, 5, 10, 20, 30, 60, 90, 120 minutes, or more. The incubation may be performed in a buffer solution containing additives such as sodium ions.
[0149]
[0176] If the nucleic acid pool is single-stranded (rather than double-stranded), hybridization of the probe to its target can be improved. Preparing a ssDNA pool from a dsDNA pool can involve performing linear PCR using one primer that generally binds to the edge of all nucleic acid sequences in the pool. If the nucleic acid pool is synthetically created or constructed, this common primer binding site can be included in the synthetic design. The product of linear PCR will be ssDNA. More cycles of linear PCR can be used to generate more starting ssDNA templates for nucleic acid capture.
[0150]
[0177] After the nucleic acid probes are hybridized to their targets and bound to magnetic streptavidin beads, the beads can be immobilized with a magnet and washed several times. Three to five washes may be sufficient to remove non-target nucleic acids, although more or fewer washes may be used. Each incremental wash may further reduce non-target nucleic acids but may also reduce the yield of target nucleic acid. To facilitate proper hybridization of the target nucleic acid to the probes during the wash steps, a low incubation temperature can be used. Temperatures as low as 60, 50, 40, 30, 20, 10, or 5°C or lower can be used. The wash buffer may contain a Tris buffer solution containing sodium ions.
[0151]
[0178] Optimal elution of hybridized targets from magnetic bead-bound probes can occur at temperatures equal to or higher than the melting temperature of the probe. Higher temperatures promote dissociation of targets from the probe. Elution temperatures can be up to 30, 40, 50, 60, 70, 80, or 90°C or higher. Elution incubation times can be up to 1, 2, 5, 10, 30, 60 minutes, or more. A typical incubation time can be about 5 minutes, although longer incubation times can improve yield. The elution buffer can be water or a Tris buffer solution containing additives such as EDTA.
[0152]
[0179] The nucleic acid capture of a target sequence that comprises at least one or more of a set of different sites can be carried out in a single reaction with a plurality of different probes for each of these sites.The nucleic acid capture of a target sequence that comprises all members of a set of different sites can be carried out in a series of capture reactions, with one reaction carried out for each distinct site using the probe for that specific site.Although the target yield after a series of capture reactions may be low, the captured target can then be amplified by PCR.When a nucleic acid library is synthetically designed, the target can be designed using a common primer binding site for PCR.
[0153]
[0180] Synthetic nucleic acid libraries can be created or constructed using common probe binding sites for general nucleic acid capture. These common sites can be used to selectively capture fully assembled or potentially fully assembled nucleic acids from assembly reactions, thereby eliminating partially assembled or misassembled (or unintended or undesired) by-products. For example, assembly can involve assembling nucleic acids with probe binding sites on each edge sequence such that only fully assembled nucleic acid products contain the two required probe binding sites necessary to pass through a series of two capture reactions using each probe. In this example, partially assembled products may contain none or only one of the probe sites and therefore should not ultimately be captured. Similarly, misassembled (or unintended or undesired) products may contain none or only one of the edge sequences. Thus, the misassembled products may not ultimately be captured. To increase stringency, common probe binding sites can be included in each component of the assembly. Subsequent series of nucleic acid capture reactions using probes for each component can isolate only the fully assembled product (containing each component) from any by-products of the assembly reaction. Subsequent PCR can improve target enrichment, and subsequent size selection can improve target stringency.
[0154]
[0181] In some implementations, nucleic acid capture can be used to selectively capture a targeted subset of nucleic acids from a pool, for example, by using probes with binding sites that appear only on the targeted subset of nucleic acids. For selective capture of a sub-library from a more general library, synthetic nucleic acid libraries can be created or constructed such that the nucleic acids belonging to a potential sub-library of interest all share a common probe binding site (common within the sub-library but distinct from other sub-libraries).
[0155] G. Freeze-drying
[0182] Lyophilization is a dehydration process. Both nucleic acids and enzymes can be freeze-dried. The freeze-dried material may have a longer shelf life. Additives such as chemical stabilizers can be used to maintain the functionality of the product (e.g., active enzymes) throughout the freeze-drying process. Disaccharides, such as sucrose and trehalose, can be used as chemical stabilizers.
[0156] H.DNA design
[0183] The sequences of nucleic acids (e.g., building blocks) for constructing synthetic libraries (e.g., identifier libraries) can be designed to avoid the complexities of synthesis, sequencing, and assembly. Furthermore, they can be designed to reduce the cost of constructing synthetic libraries and improve the longevity over which synthetic libraries can be stored.
[0157]
[0184] Nucleic acids can be designed to avoid long stretches of homopolymers (or repetitive base sequences), which can be difficult to synthesize. Nucleic acids can be designed to avoid homopolymers of 2, 3, 4, 5, 6, 7, or more bases in length. Furthermore, nucleic acids can be designed to avoid the formation of secondary structures, such as hairpin loops, which can inhibit their synthesis process. For example, prediction software can be used to generate nucleic acid sequences that do not form stable secondary structures. Nucleic acids for constructing synthetic libraries can be designed to be short. Longer nucleic acids can be more difficult and expensive to synthesize. Longer nucleic acids can also have a higher chance of mutation during synthesis. Nucleic acids (e.g., building blocks) can be up to 5, 10, 15, 20, 25, 30, 40, 50, 60, or more bases.
[0158]
[0185] Nucleic acids that serve as components in an assembly reaction can be designed to facilitate the assembly reaction. Efficient assembly reactions typically involve hybridization between adjacent components. Sequences can be designed to facilitate these on-target hybridization events while avoiding potential off-target hybridization. Nucleic acid base modifications, such as locked nucleic acids (LNAs), can be used to enhance on-target hybridization. These modified nucleic acids can be used, for example, as staples in staple strand ligation or sticky ends in sticky strand ligation. Other modified bases that can be used to construct a synthetic nucleic acid library (or identifier library) include 2,6-diaminopurine, 5-bromo-dU, deoxyuridine, reverse dT, reverse dideoxy-T, dideoxy-C, 5-methyl-dC, deoxynosine, super-T, super-G, or 5-nitroindole. Nucleic acids can contain one or more of the same or different modified bases. Some of the modified bases are natural base analogs (e.g., 5-methyl dC and 2,6-diaminopurine) that have higher melting temperatures and therefore may be useful for promoting specific hybridization events in assembly reactions. Some of the modified bases are universal bases (e.g., 5-nitroindole) that can bind to all natural bases and therefore may be useful for promoting hybridization with nucleic acids that may have variable sequences within the desired binding site. In addition to their beneficial role in assembly reactions, these modified bases may be useful in primers (e.g., for PCR) and probes (e.g., for nucleic acid capture) because they can promote specific binding of primers and probes to their target nucleic acids within a nucleic acid pool.
[0159]
[0186] Nucleic acids can be designed to facilitate sequencing. For example, nucleic acids can be designed to avoid typical sequencing complications, such as secondary structures, homopolymer stretches, repetitive sequences, and sequences with too high or too low GC content. Certain sequencers or sequencing methods may be error-prone. Nucleic acid sequences (or components) constituting a synthetic library (e.g., an identifier library) can be designed at a constant Hamming distance from each other. In this way, even if base resolution errors occur rapidly in sequencing, stretches of error-containing sequences can still be mapped back to their most likely nucleic acids (or components). Nucleic acid sequences can be designed at a Hamming distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more base mutations. Alternative distance metrics to Hamming distance can also be used to define the minimum required distance between designed nucleic acids.
[0160]
[0187] Some sequencing methods and devices may require input nucleic acids that contain specific sequences, such as adapter sequences or primer binding sites. These sequences may be referred to as "method-specific sequences." A typical preparatory workflow for the sequencing devices and methods may include assembling method-specific sequences into a nucleic acid library. However, if it is known in advance that a synthetic nucleic acid library (e.g., an identifier library) will be sequenced by a specific instrument or method, these method-specific sequences may be designed into the nucleic acids (e.g., components) that comprise the library (e.g., an identifier library). For example, sequencing adapters can be assembled onto members of a synthetic nucleic acid library in the same reaction step as when the members of the synthetic nucleic acid library themselves are assembled from individual nucleic acid components.
[0161]
[0188] Nucleic acids can be designed to avoid sequences that may promote DNA damage. For example, sequences containing sites for site-specific nucleases can be avoided. As another example, UVB (ultraviolet-B) light can cause adjacent thymines to form pyrimidine dimers, which can then inhibit sequencing and PCR. Therefore, if a synthetic nucleic acid library is intended to be stored in an environment exposed to UVB, it can be beneficial to design its nucleic acid sequence to avoid adjacent thymines (i.e., TT).
[0162] System for building an identifier library
[0189] As previously mentioned, a print-based system known as a print finisher system (or PFS) can be used to juxtapose and assemble the components to construct the identifier.
[0163]
[0190] Provided herein is a system for assembling an identifier from one or more components for storing information, the system including: (a) a printer for dispensing one or more components onto a substrate, each of the one or more components comprising a nucleic acid sequence; and (b) a finisher for assembling the one or more components onto the substrate, the finisher providing a reaction mixture and / or conditions necessary to physically link the one or more nucleic acid sequences.
[0164]
[0191] In some implementations, the printer further includes a plurality of printheads, each printhead of the plurality of printheads including one or more components. In some implementations, the printer includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more printheads. In some implementations, each of the plurality of printheads includes a different component. In some implementations, each printhead includes at least one nozzle. In some implementations, each printhead includes a row of nozzles. In some embodiments, each printhead includes at least 1, 2, 3, 4, or more rows of nozzles. In some implementations, a printhead can be considered a set of nozzles that each deliver the same ink. In some embodiments, the rows of nozzles deliver the same ink. In some implementations, a particular subset of nozzles in a row of nozzles delivers a different ink than other nozzles in the row of nozzles. In some implementations, the array of nozzles includes at least 20, 40, 60, 80, 100, 150, 200, 250, 300, 350, 400, or more nozzles. In some embodiments, some or all of the nozzles in the array of nozzles may be discontinuous. In some implementations, the print head dispenses droplets including the component onto the substrate. In some implementations, the print head dispenses droplets including the reaction mixture onto the substrate. In some implementations, the droplets have a volume of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 picoliters. In some implementations, the droplets have a volume of at least 10, 20, 30, 40, 50, 60, 70, or 80 picoliters. In some implementations, the printer further includes a printer base. In some implementations, the printer further includes a register, a spot imager, and / or a spot dryer. In some implementations, the one or more components are in solution, in some implementations, the one or more components are dry components, in some implementations, the reaction mixture includes a ligase.A ligase can be used to ligate different components comprising nucleic acid sequences. In some implementations, the conditions are temperature conditions. In some implementations, the substrate passes through the printer and / or the finisher in a linear motion. In some implementations, the linear motion is controlled by a reel-to-reel system. In some implementations, the spot imager is a camera. In some implementations, the one or more components further comprise a dye. In some implementations, the reaction mixture comprises a dye. The dye can be any nucleic acid dye. The dye can be a visible dye.
[0165]
[0192] In some implementations, the substrate further comprises a polymer material. In some implementations, the print head is a MEMS (microelectromechanical systems) thin-film piezoelectric inkjet head or a MEMS thermal inkjet head. In some implementations, the one or more components include an additive. In some implementations, the additive provides compatibility between the one or more components and the print head. In some implementations, the additive is a solute, a wetting agent, or a surfactant. In some implementations, the spot imager uses a line-scan inspection principle. In some implementations, the finisher further comprises a finisher base.
[0166]
[0193] In some implementations, the finisher further includes a spot humidifier, a spot imager, and / or a pooling subsystem. In some implementations, the finisher further includes a print head. In some implementations, the finisher print head dispenses volumes of at least 1 pL, 5 pL, 10 pL, 50 pL, 100 pL, or 200 pL. In some implementations, the finisher includes a fixed internal temperature optimal for reaction incubation. In some implementations, the finisher includes a loop of rollers.
[0167] Printer-Based Systems
[0194] A PFS may involve the use of one or more printheads, each capable of printing one or more nucleic acid molecules onto a substrate. Given an identifier library to be generated, the task of assembling all identifiers that encode a given bitstream may be divided into subtasks, each subtask involving generating a portion of the identifier library. These portions may be referred to as "sectors" of the identifier library. The size of a sector may be selected so that any errors in the PFS's generation of a sector can be detected or corrected by the PFS. Errors may be caused by several factors, including, but not limited to, a malfunctioning printhead, unintended mixing of components during or after printing, variations in the amount of reagent or nucleic acid dispensed by the printhead, misalignment between the printhead and the target coordinates (or spots) on the substrate, or drying or wetting due to high or low humidity. Some of these factors may lead to errors in which one or more identifiers that should be generated are not generated. This type of error may be referred to as a missing identifier error.
[0168]
[0195] Depending on the cause, some missing identifier errors may be detected by the PFS. For example, the PFS can automatically inspect all or a portion of the print sectors using one or more cameras. The PFS can capture one or more images of each print sector continuously or at programmable intervals and computationally process those images to determine whether each designated reaction has been printed on the substrate. In another embodiment, the PFS can monitor one or more nozzles on one or more print heads continuously or at programmable intervals and capture images or videos of the nozzles as they print reactions on the substrate. The PFS can subject the captured video or images to image processing to determine whether all intended reagent and nucleic acid droplets have been delivered to the reactions. The monitoring camera can use visible light or light in other frequency bands. In another embodiment, the PFS can cyclically print one or more test patterns from all nozzles on all print heads within a test area of the substrate. The PFS can visually capture or analyze the results of the test pattern printing using a spot imager, camera, or some other device with an output suitable for analysis. In another embodiment, the PFS can print a test pattern and analyze it using one or more chemical verification methods, such as gel electrophoresis.
[0169]
[0196] If, after visual analysis, the PFS concludes that some or all of the components necessary to assemble all of the specified identifiers were not printed in the reaction, the PFS can report this conclusion in an error log. The control software controlling the PFS can continuously analyze this log during or after printing and choose to reprint sectors containing such missing identifier errors. From the log, the control software can identify faulty printheads or nozzles and print the remaining sectors using spare printheads or nozzles. In one embodiment, the control software can also exclude sectors without identifier errors from downstream processing steps; such incomplete sectors are not included in the final identifier library.
[0170]
[0197] The identifier library to be assembled is specified and sent to the PFS via a set of specification files. The identifier library to be generated may be specified in a collection of smaller units called blocks. The specification files include a write specification file containing the scheme used to assemble the identifier library from DNA components, a list of scheme-specific parameters, and a list of block specification file names. The block specification may include a block metadata file and a block data file. The block metadata file describes information about the block, such as its length, hash, and other constructor-defined parameters. The block data file specifies the set of identifiers to be generated by the PFS. The block data file may be compressed using a data compression algorithm. The identifiers comprising the blocks may be specified in the form of a serialized data structure, such as, but not limited to, a tree, a trie, a list, or a bitmap.
[0171]
[0198] For example, an identifier library generated using a product scheme may be specified in a block metadata file containing the component library partition scheme and a list of possible component names used at each layer. The block data file may organize the generated identifiers as a serialized trie data structure, where each path from the root to a leaf of the trie represents an identifier and each node along the path specifies the component name used at that layer for that identifier. The block data file may contain a serialization of this trie by starting from the root and traversing it in order, visiting each node's left child node, then the node itself, and then its right child node.
[0172]
[0199] The PFS can monitor an input queue of incoming specification files. When it detects a new specification, it can read the write specification and program itself with the necessary components to supply the appropriate printhead or nozzles. The PFS can read block metadata and data files and process them to generate print instructions for the printhead. The PFS can send these instructions for each block to the printhead and get status information for each sector from the printhead. Sectors that were not printed correctly or completely can be reported in a log and can be automatically reprinted.
[0173] Illustrative PFS
[0200] 1 illustrates a system for storing digital information in DNA by assembling DNA identifiers from components at high speed and throughput using inkjet printing, such as thermal inkjet printing, bubble inkjet printing, and piezoelectric inkjet printing. The system, hereinafter referred to as a "printer-finisher system" or PFS, and its different implementations may include two subsystems: a printer 120 and a finisher 130. In some implementations, the two subsystems 120, 130 may be attached and dependent on each other for their respective functions. In other implementations, the two subsystems 120, 130 may be separate and capable of functioning independently.
[0174]
[0201] The printer 120 includes an array of printheads 122, each containing a DNA component in solution (or, in some implementations, dried DNA components). We may refer to each aqueous solution of a distinct DNA component as an "ink" or a "color." The printheads 122 can programmably (on-demand) dispense pL-scale droplets onto coordinates on a substrate (or web or webbing). The coordinates can be 1 micrometer (μm) diameter / spacing, 10 μm diameter / spacing, 50 μm diameter / spacing, 100 μm diameter / spacing, 150 μm diameter / spacing, 200 μm diameter / spacing, or greater. The input to the printer system 120 includes aqueous components / substrate. The output from the printer system 120 includes dried multilayer spots on the substrate. The environment of the printer 120 may be dry (evaporated).
[0175]
[0202] The finisher 130 includes an instrument component (e.g., a print head) for dispensing a reaction mixture (e.g., a ligase mix) to assemble the components into identifiers. Input to the finisher system 130. The finisher 130 can dispense the reaction mixture to each coordinate on the substrate (or web or webbing). The finisher 130 can then incubate the reaction before consolidating the assembled identifiers from the substrate into a single pool 132, thus allowing assembly. In some implementations, the reaction mixture can be dispensed as part of the printer rather than the finisher. In other implementations, the reaction mixture can be dispensed to each coordinate before the DNA components. In some embodiments, a visible dye can be incorporated into the reaction mixture.
[0176]
[0203] The substrate (or web) 136 can automatically move through the printer and finisher in a linear (one-dimensional) motion. Linear motion at a constant velocity can be achieved with a reel-to-reel system (roller-to-roller) 134. In some implementations, linear motion at a constant velocity can be achieved with recirculating or continuous webbing. In some embodiments, linear motion at a constant velocity can be achieved using webbing following a Snell path. See, for example, FIG. 7. In some implementations, linear motion at a constant velocity can be achieved using webbing that follows a spiral path. In some implementations, linear motion at a constant velocity can be achieved using webbing that follows a 180° twist path. For example, the webbing rotates 180° at each roller with the system, and the webbing passes all of the rollers in a right-up, upward motion. In other implementations, the substrate can be fixed, and the print head can move over the substrate in two dimensions (e.g., in a raster pattern).
[0177]
[0204] FIG. 2 shows the printer subsystem 120 in more detail. The printer base 121 includes a printer base with a web drive that hosts a print engine 122, a spot imager 126, and a spot dryer 128. The print engine prints and overprints to support an addressing scheme. The print engine 122 may include a print head. The print head is designed to overprint, or juxtapose, or overlay different components at the same coordinates on the web 136. A single nozzle, a single print head, multiple nozzles, multiple print heads, or any combination thereof, can overprint components at the same coordinates. In addition to the print head, the printer may optionally include a register 124, a spot imager 126, and a spot dryer 128.
[0178]
[0205] Alignment includes spot alignment (in the case of a multi-pass system). The register 124 is intended to maintain alignment between the substrate coordinates and the print head. This can be achieved by labeling the substrate with special markings that allow the register to track the substrate's movement in real time. In other implementations, alignment can be achieved by dead reckoning the substrate position from encoders on the rollers. Control of alignment along the web can be done by adjusting the timing of dispensing operations from the print head. Alignment across the web can require either the substrate or the print head to move using actuators.
[0179]
[0206] The spot imager 126 provides verification of component addition. The spot imager 126 may be a camera intended to verify proper dispensing of a component or reaction mixture. To facilitate the function of the spot imager 126, a visible dye may be incorporated into the component ink or reaction mixture.
[0180]
[0207] The spot dryer 128 is intended to dry the printed droplets so that they can dry between print heads or as they exit the printer (e.g., if the substrate is intended to be rolled as it exits the printer). Drying droplets between print heads can be useful to prevent overflow of liquid at specific coordinates during the overprint process. Each print head can dispense droplets of at least 1 pL, 5 pL, 10 pL, 20 pL, 30 pL, 40 pL, 50 pL, or more. In some implementations, at least 1, 5, 10, 20, 50, 100, or more print heads can dispense to the same coordinate.
[0181]
[0208] The printer subsystem may optionally include a substrate and coating module 129. The substrate and coating module 129 includes the web material plus coating / patterning. The substrate may include a material or may be coated with a material such as a low-bond plastic, such as polyethylene terephthalate (PET) or polypropylene.
[0182]
[0209] 3A-3D show an example of a printhead 300 in a printer (e.g., printer 120 of FIG. 1). The printhead can contain one, two, three, four, or more inks (separate component solutions). In this particular example, consider a printhead 300 that can contain up to four inks, one ink provided for each row of nozzles. Furthermore, the printhead can contain multiple nozzles per ink, e.g., 300 nozzles. In certain cases, the set of web coordinates addressable by some or all nozzles may be discontinuous because the nozzles may not be properly aligned so that each ink can overprint on the same coordinates on the substrate in a linear fashion past the printhead. Or, the nozzles for different inks may not be properly spaced to print at the desired pitch. To address these issues, the printhead can be mounted at an angle (relative to web motion) to enable overprinting of the component inks at the desired pitch. As shown in FIGS. 3B-3D, a rotation of approximately 9 degrees is sufficient to enable overprinting of four inks at a 167 μm pitch. Specifically, FIG. 3C shows four rows of printer head nozzles 302, 304, 306, and 308. Each of the rows 302, 304, 306, and 308 can dispense a different component. A substrate 312 (extending diagonally upward and to the right from the line indicated by arrow 312) moves linearly beneath the print head 300. Due to the 8.7-degree rotation of the print head, coordinate 314 on the substrate 312 passes directly beneath the nozzles in rows 302, 304, 306, and 308 along line 307, so that each nozzle can deposit a component on coordinate 314. As shown in FIG. 3D, multiple print heads 300, 310, and 320 can be arranged in parallel to enable simultaneous printing on multiple substrates. In one example, the print heads can be actuated into proper alignment for overprinting. The print heads can be MEMS (microelectromechanical systems) thin-film piezoelectric inkjet heads or MEMS thermal inkjet heads. Additives may be added to the component inks to promote compatibility with the printhead.For example, solutes such as Tris can be added to increase conductivity. As an example, humectants or surfactants (e.g., glycerol) can be added to improve jetting quality and printhead nozzle life.
[0183]
[0210] Figure 4 shows a potential arrangement of printheads within a printer. Assume the substrate is passed longitudinally so that printheads on different tracks (T1-T4) print on independent coordinates, but printheads along the same track can print (overprint) on the same coordinates on the substrate. The substrate can pass through the printer multiple times, each time using a new printhead (or the same printhead filled with new ink) to receive more DNA components per coordinate. However, if a sufficient number of printheads are positioned along each track, only a single pass may be required to incorporate enough components for the desired number of identifiers to be constructed. For example, if identifiers each contain eight components (8 10 If a substrate is constructed from a 10-layer product scheme (enough to enable multiple identifiers and store over a gigabit of data), with each printhead capable of printing four components, mounting 20 printheads along the track may be sufficient to allow the entire component set to be collocated in a single pass over the substrate. Multiple tracks allow for more efficient use of the substrate (web), allowing it to be shorter and allowing identifiers to be built at a higher throughput. If there is a larger width (laterally) in the substrate than the tracks, the substrate (or printhead chassis) can be shifted laterally after each pass to allow printing onto empty substrates along the width of the substrate rather than along its length. In another embodiment, separate printer-based systems can print on different portions of the same substrate.
[0184]
[0211] FIG. 5 shows an example of a spot imager configuration in a printer subsystem. The spot imager can use a line-scan inspection principle. For example, the spot imager can include a computer system 520, a display 510, a line-scan camera 530, a rotating drum 540, and an encoder 550. The computer system 520 communicates with the line-scan camera 530. For example, the computer system 520 can send control signals to the line-scan camera 520, and the line-scan camera 530 can send image data back to the computer system 520. The computer system 520 and the line-scan camera system 530 can communicate via a wireless or wired connection. The image data collected via the line-scan camera 530 is displayed on the display 510. As shown in FIG. 5, the line-scan camera 530 can capture an image of the drum 540, which can then be displayed on the display 510.
[0185]
[0212] FIG. 6 shows the finisher subsystem 130 in more detail. The finisher subsystem 130 includes a finisher base 140 with a web drive, an incubation buffer and dispense host, a spot humidifier 144, a spot imager 146, and a pooling subsystem 148. In addition to dispensing the reaction mixture onto each coordinate of the substrate, the finisher may also include a portion 142 that dispenses a reaction inhibitor onto each coordinate of the substrate 136 before solidification. These dispense components may be printheads. They may be on-demand printheads, but continuous printing may also suffice, as each coordinate along the web may be expected to receive a dispense. The dispense volume must be sufficient to cover the area of each coordinate where DNA components have previously been dispensed. The dispensed volume can be at least 1 pL, 5 pL, 10 pL, 20 pL, 30 pL, 40 pL, 50 pL, 60 pL, 70 pL, 80 pL, 90 pL, 100 pL, 150 pL, 200 pL, or more. The printhead can be a MEMS (microelectromechanical systems) thin-film piezoelectric inkjet head or a MEMS thermal inkjet head. Additives can be added to the dispensed liquid (e.g., master mix or inhibitor mix) to promote compatibility with the printhead. For example, solutes such as Tris can be added to increase conductivity. As another example, humectants or surfactants can be added to improve jetting quality and printhead nozzle life. Additionally, humectants such as glycerol or polyethylene glycol (PEG) can be added to control evaporation both at the nozzle-air interface and after the droplets are dispensed. These humectants can further benefit the reaction mixture by increasing the yield of the reaction product.
[0186]
[0213] Similar to the printer subsystem, the finisher may also include a register and spot imager 146 for aligning the web with the printhead and verifying proper dispensing, respectively. To facilitate the function of the spot imager, a visible dye may be incorporated into the dispensed fluid.
[0187]
[0214] The finisher may further include several roller loops (a configuration of rollers intended to loop the webbing) 134 after the reaction mixture is dispensed, so that the reaction on the web (substrate) 136 can be incubated for a longer period before reaction intensification. The finisher may include a fixed internal temperature optimal for reaction incubation, such as 4, 12, 25, 37°C, or higher. To slow and control evaporation of the dispensed reaction mixture during the incubation phase, the finisher may include a fixed high humidity level. The humidity level of the finisher subsystem 130 can be controlled by a spot humidifier 144, which controls the maintenance of a wet spot throughout the incubation period (e.g., while the substrate passes over the rollers 134).
[0188]
[0215] Finally, the finisher may include a pooling system 148 for combining all of the identifier assembly reactions into one vessel after incubation. Reaction inhibition may occur before or during this step.
[0189]
[0216] FIG. 7 shows an example of a loop of rollers 710, 720 for passing the web through the finisher during the incubation phase. Looping the web allows for longer incubation times in a more confined space. For example, if the web is moving through the system at 180 mm / s, a 5-minute incubation time requires an incubated web length of approximately 60 m, but several loops may allow this length to be incubated in a more confined space rather than a linear tunnel of approximately 60 m. Shorter incubation times may allow for shorter incubated web lengths. For example, a 45-second incubation time may allow for an incubated web length of approximately 9 m, and a 10-second incubation time may allow for an incubated web length of approximately 2 m. With these shorter incubated web lengths, fewer web loops may be required to confine the incubation to a smaller space.
[0190]
[0217] Due to the geometry of the roller loops, the webbing 740 can pass right-side up through certain rollers 720 and upside down through other rollers 710.
[0191]
[0218] The bottom of the figure shows a cross section of roller 710 along the path of travel of the web. The rollers can be designed to include valleys (or grooves, pockets, or any other depressions) 730 between contact points in the substrate 740 so that the reaction (e.g., coordinates where components are dispensed) can pass through the valleys without interference. Alternatively, the web can be rotated 180 degrees between rollers so that the web always passes over the rollers in a right-side-up configuration (i.e., a 180° twist path). Alternatively, the webbing can travel a spiral path through the incubator, such that the circular path of the webbing around the set of rollers ensures that the side of the webbing containing the reaction does not come into contact with the rollers. As an analogy, consider winding a ribbon around a cylinder or applying grip tape to a tennis racket.
[0192]
[0219] In some implementations, the webbing is a recirculating or continuous webbing. In some implementations, the webbing is a reel-to-reel system (roller-to-roller). In some implementations, the webbing follows a Snell path. See, for example, FIG. 7. In some implementations, the webbing follows a spiral path. In some implementations, the webbing follows a 180° twist path. For example, the webbing rotates 180° at each roller with the system, and the webbing passes through all rollers in a right-side-up configuration.
[0193]
[0220] Figure 8 shows the effect of reaction mixture glycerol composition and finisher humidity on the expected equilibrium volume during incubation. Particles represent water molecules transitioning between the liquid and gas phases. Droplets 820 represent dispensed reactions on web 810. The outer shaded area represents water, the central shaded area represents glycerol, and the inner shaded area represents solutes (e.g., DNA, enzyme / ligase, salt / magnesium, Tris). High humidity and high glycerol conditions result in an equilibrium reaction composition most similar to the original composition. However, changes in reaction composition at equilibrium can be beneficial. For example, increasing the relative amount of DNA components can result in higher identifier production yields. Similarly, increasing glycerol content can result in a crowding effect that promotes identifier generation. While reaction efficiency can be adversely affected by increasing the concentration of certain solutes (such as salts), the initial solutes present in the reaction mixture can be intentionally low and designed to exist at optimal concentrations after the reaction droplets evaporate to their equilibrium composition and volume.
[0194]
[0221] FIG. 9 shows a pooling system that consolidates all reactions from a web into one container. A series of rollers 902 navigates the web 910 through a spray wash 914 and collection reservoir 942 designed to capture reactions and their identifier products from the web 910. To prevent excessive volume buildup in this process, the collection fluid can be continuously or repeatedly flowed through a membrane designed to capture nucleic acids. For example, the membrane can be a silica membrane, and the collection fluid can be a DNA binding buffer 912 to promote binding of nucleic acids to the membrane. The collection fluid can further contain additives to inhibit reactions so that they do not proceed in the solidification volume. For example, if the reaction is a ligation reaction, the collection fluid can contain EDTA (e.g., 25 mM), which chelates magnesium ions from the ligase and thus inhibits the reaction. In one embodiment, the binding buffer can be recirculated through one or more binding columns to minimize the volume of the binding buffer. The web 910 may be wetted with a liquid to remove DNA from the web 910, which may be combined with immersing the web 910 in a liquid in a collection vessel. Agitation and / or heating of the web 910 or the liquid (e.g., mechanical, fluid, or ultrasonic) may be used to facilitate release of the DNA from the web 910. The scraper 918 may be a physical scraper, a liquid jet, or a gas (e.g., air) jet, also to facilitate removal of DNA from the web 910. One or more sprays may be used to facilitate release of the DNA from the web 910.
[0195]
[0222] After the DNA is captured on the membrane, it can be removed from the system (machine) for elution and further evaluation. Further evaluation can include running the DNA on a gel and selecting a band size that corresponds to the expected identifier length (thereby purifying the identifier from other potential off-target products). In this example, the target identifier is 300 bp in length. The DNA output can optionally be passed through a gel or other filtration 940 to result in DNA data 930, which can be lyophilized.
[0196]
[0223] Another embodiment of this system involves adding and inhibiting a reaction mixture before or during pooling, and instead allowing the reaction to occur in a pooling step. In this embodiment, components are annealed but not assembled during the incubation process, and then combined together in a pool containing the reaction mixture and the appropriate environmental conditions (e.g., temperature, pH, salt) for assembling the components into identifiers. This embodiment allows for shorter incubation times on the web 910 and less stringent hardware requirements in the finisher, since once the annealed components are pooled, the remainder of the reaction can proceed outside the system (machine). In this embodiment, special care can be taken to ensure that components are strongly annealed to each other before and during pooling to prevent undesired cross-assembly between components of different identifiers in the pooled reactions. This can include using components with long sticky ends (and hybridization regions) for strong annealing and using lower temperatures in the pooling step to maintain annealed products and limit diffusion of unannealed products.
[0197]
[0224] FIG. 10 shows a schematic diagram of one embodiment of a data transfer pipeline through a PFS. FIG. 10 begins with a source stream 1002 containing 1 Tb of data. The source stream 1002 is transferred to a codec 1004 and provided to a job module 1006. The job module 1006 creates a job file, block records, and block data for each source stream and / or codec file. This information is provided to a block monitor 1008. The job module 1006 is monitored by a job monitor 1016, which communicates with the block monitor 1008. The block monitor 1008 monitors for new blocks, validates them, and adds them to the pipeline for printing. Block data 1010 from the job module 1006 is separated and sent to a block reader 1012, which processes the ink and printhead configuration needed to print the block data. The block data is then converted into a printable frame 1014 containing the block data 1010 and a "chirp" configured to test the accuracy of the data transfer. The frame 1014 is then sent to a document printer module 1018, which communicates with a printer 1034. For example, the document printer module 1018 sends the frame 1014 to the printer 1034 for printing, and the printer 1034 sends feedback to the document printer 1018. Any faults 1020 are communicated to a finish controller 1022, which writes them to a text file or other storage method 1024. In addition to electronically communicating with the document printer 1018, the printer 1034 also receives physical web sectors 1036. The web sectors 1036 are located by a marker at one corner. Each web sector has a unique ID code. The printer 1032 deposits the components 1032 onto the web. The web then continues to a finisher 1026, which communicates with the finish controller 1022. The finish controller 1022 sends information about the frame or partial frame for finishing to the finisher 1026, and the finisher 1026 sends feedback back to the finish controller 1022.Feedback from both the printer and finisher systems 1034, 1026 facilitates recording of frame to sector assignments, coordination of web registration with print and quality control, and recording of failed frames. After leaving the finisher 1026, the web is printed to completion 1028, resulting in a substrate with DNA spots 1030, which can then be sent to a polling system or any other suitable storage method.
[0198]
[0225] FIG. 11 illustrates one embodiment of a PFS that includes four modules: a chassis module, a print engine module, an incubator module, and a pooling module. The function of the chassis module may be to provide a base system that drives, stabilizes, and controls the movement of the webbing through all modules of the system. The function of the print engine module may be to print DNA components and other materials and reagents into reaction droplets on the webbing. The function of the incubator module may be to provide time and environmental control to improve product (e.g., assembled DNA or identifier) yield in the reaction droplets. The function of the puller module may be to remove the reaction droplets from the webbing and consolidate them into one container.
[0199]
[0226] In some embodiments, the reaction droplets can assemble the DNA identifiers by enzymatic ligation. In some embodiments, the reaction droplets can assemble the DNA identifiers by click chemistry.
[0200]
[0227] In some embodiments, the incubator module may include 100, 50, 25, 10, 5, 1, or 1 meter or less of webbing. In some embodiments, the PFS may not have an incubator module.
[0201]
[0228] In some embodiments, the print engine or incubator can include an intermittent printhead or dispensing sub-module to replenish the volume in the reactant droplets as they evaporate on the webbing.
[0202]
[0229] In some embodiments, the webbing passing through the PFS can unwind from a roll before the print engine and rewind onto the roll after the puller, hi some embodiments, the webbing can form a continuous loop that returns to the print engine after the puller.
[0203]
[0230] FIG. 12 shows an embodiment of a PFS that pools reaction droplets into an emulsion 1260. The emulsion 1260 can contain oil or any liquid immiscible with the reaction droplets, allowing the reaction droplets 1250 to maintain their contents even after pooling. The webbing 1220 of the PFS can be coated with oil (e.g., via rollers 1230 and 1240) before passing under the print head 1210. The reaction droplets 1250 can contain surfactants and other additives to control their size and shape in the emulsion. Surfactants and additives can also promote stability within the emulsion and prevent coalescence between different reaction droplets. The pooled emulsified reaction droplets can be passed through a microfluidic device. The pooled emulsified reaction droplets can be incubated. Furthermore, the pooled emulsified reaction droplets can be coalesced and isolated from the emulsion.
[0204]
[0231] FIG. 13 illustrates an embodiment of a PFS in which reaction droplets 1350 are coated with oil (or another immiscible liquid) 1370 after being printed onto a webbing 1320. The oil coating can be performed by an oil dispensing submodule 1380, which prints, dispenses, or sprays oil onto the reaction droplets 1350 as the webbing 1320 passes under the printhead cluster 1310 via rollers 1330 and 1340. The oil can reduce or prevent evaporation of the reaction droplets on the webbing 1320. The reaction droplets can contain surfactants and other additives. The oil-coated reaction droplets 1370 can be pooled into an emulsion 1390. The pooled emulsified reaction droplets can be passed through a microfluidic device. The pooled emulsified reaction droplets can be incubated. The pooled emulsified reaction droplets can then be coalesced and isolated from the emulsion.
[0205]
[0232] Figure 14 shows an embodiment of a PFS in which reaction droplets 1450 contain beads that bind to printed DNA building blocks. The beads may be coated with silica, carboxyl groups, or amine or imidazole moieties that bind to DNA. Alternatively or additionally, the beads may be coated with streptavidin, which binds to the DNA building blocks via a biotin linkage. The biotin may be linked to the DNA building blocks with a light- or UV-cleavable linker.
[0206]
[0233] The webbing 1420 may be ubiquitously coated with beads before passing under the print head 1410, or may be patterned with beads (e.g., via rollers 1430 and 1440). Alternatively or in addition, beads may be deposited or printed into each of the reaction droplets 1450. The reaction droplets may contain additives that promote DNA binding to the beads. The beads may be 1, 2, 3, 5, 10, 20, 50, 100 or more per reaction droplet.
[0207]
[0234] The reaction droplets 1450 may be pooled in a solution 1460 that prevents further association of DNA with the beads. The solution 1460 may contain a blocking agent, such as BSA. The DNA-bound beads in the pooled solution may be separated from the solution and dried 1470. Separation may be performed by centrifugation. In another embodiment, the beads may be magnetic and separated with a magnet.
[0208]
[0235] The pooled DNA-binding beads (dried 1470 or solution 1460) can be further encapsulated into emulsified reaction droplets. In one embodiment, the DNA-binding beads are each encapsulated in a reaction droplet using microfluidics. In another embodiment, the DNA-binding beads are each encapsulated in a reaction droplet by mixing the reaction solution with oil (or another immiscible liquid) so that the droplets form spontaneously. The ratio of spontaneously formed reaction droplets to DNA-binding beads can be adjusted so that a reaction droplet cannot contain more than one DNA-binding bead. Reaction droplets can contain surfactants or other additives to control their size or prevent coalescence with other reaction droplets.
[0209]
[0236] The reaction droplets may contain reagents that dissociate the DNA on the beads. The reaction droplets may contain reagents that ligate the DNA components together to form the identifier. The reaction droplets may contain the enzyme ligase and ligation cofactors such as ATP, DTT, or salt.
[0210]
[0237] If the DNA is attached to the beads via a photocleavable or UV-cleavable linkage, the DNA can be released from the beads by exposing the emulsion to electromagnetic radiation of the appropriate wavelength (e.g., light or UV).
[0211]
[0238] 15 shows an example of how DNA components bound to beads can be processed into identifiers using emulsion. In step 1510, DNA-bound beads are provided. The DNA-bound beads are then emulsified in 1520 so that the DNA-bound beads encapsulated in reaction mixture droplets are immersed in oil. The DNA is then dissociated to obtain mixture 1530. The dissociated DNA mixture is incubated to obtain assembled DNA in 1540.
[0212]
[0239] While exemplary implementations have been shown and described herein, it will be apparent to those skilled in the art that such implementations are provided by way of example only. Numerous variations, modifications, and substitutions will occur to those skilled in the art. It is understood that various alternatives to the implementations described herein may be used.
[0213] Example of changes to reduce PFS size
[0240] As previously described in Figure 11, a PFS may include the following four modules: a chassis, a print engine, an incubator, and a puller. For a PFS that encodes 1 Tb of information into DNA, the approximate size of each module may be as listed in the table below.
[0214] [Table 1]
[0215]
[0241] To reduce the size of the PFS, individual modules can be reduced in size or modules can be removed. Examples of modifications to reduce size can include: (1) Increase the printhead capacity of the print engine. Either custom printheads or additional printheads can be used to triple (or increase by a larger factor) the number of nozzle rows. This can triple the number of reactions printed as well as the print width on the webbing. (2) Use of recirculating webbing. For example, PFS can use 21 kilometers of polypropylene webbing to print enough reactions to encode 1 Tb of information. Recirculating webbing can be used as an alternative to roll-to-roll webbing to eliminate the use of webbing reels (or rolls). Recovery tests show that DNA can be easily removed from the web in the puller. (3) Shortening of ligation reaction time, which may facilitate the use of smaller incubators or no incubators at all. To shorten ligation reaction time without sacrificing yield, chemistry can be optimized to achieve higher ligation rates. (4) Ligation is performed at room temperature and ambient conditions, eliminating the need for an incubator module. (5) An oil emulsion can be used to maintain reaction droplet volume or allow ligation to begin or continue after pooling, eliminating the need for an incubator module.
[0216]
[0242] While exemplary embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, modifications, and substitutions will occur to those skilled in the art. It is understood that various alternatives to the embodiments described herein may be used.
[0217] Uses of combinatorial DNA assembly methods and systems
[0243] The methods and systems described herein for combinatorially assembling building blocks into large, defined sets of identifiers are described in the context of information technology (e.g., data storage, computing, and cryptography), however, these systems and methods can be used more generally in any application of high-throughput combinatorial DNA assembly.
[0218]
[0244] In one embodiment, we can create a library of combinatorial DNAs encoding amino acid chains. These amino acid chains can represent either peptides or proteins. DNA fragments for assembly can include codon sequences. The junctions along which the fragments assemble can be functionally or structurally inactive codons common to all members of the combinatorial library. Alternatively, the junctions along which the fragments assemble can be introns that are ultimately removed from messenger RNA and subsequently translated into processed peptide chains. Specific fragments can be not codons but rather barcode sequences that (in combination with other assembled barcodes) uniquely tag each combined string of codons. The assembled products (barcodes + strings of codons) can be pooled together and encapsulated into droplets for in vitro expression assays or pooled together and transformed into cells for in vivo expression assays. The assay can have a fluorescent output so that droplets / cells can be binned by fluorescence intensity and then their DNA barcodes can be sequenced to correlate each codon string with a specific output.
[0219]
[0245] In another embodiment, we can create a library of combinatorial DNA encoding RNA. For example, the constructed DNA can represent a combination of microRNAs or CRISPR gRNAs. Pooled in vitro or in vivo RNA expression assays can be performed as described above using either droplets or cells and barcodes to track which droplets or cells contain which RNA sequences. However, some pooled assays can be performed outside of droplets or cells, where the output itself is RNA sequencing data. Examples of such pooled assays include RNA aptamer screening and testing (e.g., SELEX).
[0220]
[0246] In another embodiment, we can create a library of combinatorial DNA encoding genes in metabolic pathways. Each DNA fragment can contain a gene expression construct. The junctions along which the fragments are assembled can represent inactive DNA sequences between genes. Pooled in vitro or in vivo gene pathway expression assays can be performed as described above using either droplets or cells and barcodes to track which droplets or cells contain which gene pathways.
[0221]
[0247] In another embodiment, we can create a combinatorial DNA library with different combinations of gene regulatory elements. Examples of gene regulatory elements include 5' untranslated regions (UTRs), ribosome binding sites (RBSs), introns, exons, promoters, terminators, and transcription factor (TF) binding sites. Pooled in vitro or in vivo gene expression assays can be performed as described above using either droplets or cells and barcodes to track which droplets or cells contain which gene regulatory constructs.
[0222]
[0248] In another embodiment, a library of combinatorial DNA aptamers can be generated and assays can be performed to test the ability of the DNA aptamers to bind to a ligand.
[0223] Exemplary Implementations of Print Finisher Technology
[0249] Provided herein are systems and assemblies for storing digital information by assembling identifier nucleic acid molecules from at least a first component nucleic acid molecule and a second component nucleic acid molecule. The system may include: (a) a first print head configured to dispense first droplets of a first solution containing the first component nucleic acid molecule onto coordinates on a substrate; (b) a second print head configured to dispense second droplets of a second solution containing the second component nucleic acid molecule onto coordinates on the substrate such that the first and second component nucleic acid molecules are juxtaposed on the substrate; and (c) a finisher that dispenses a reaction mixture onto the coordinates on the substrate to physically link the first and second component nucleic acid molecules, provides the conditions necessary for physically linking the first and second component nucleic acid molecules, or both. Generally, the first and second print heads can be part of a system including any number of rows of print heads and corresponding nozzles that print or dispense various components.
[0224]
[0250] In some implementations, the identifier nucleic acid molecule represents the position and value of a symbol in a symbol string. For example, each symbol in the string can have a corresponding identifier representing the corresponding symbol position. In particular, an identifier can be created when the symbol's corresponding value is 1, and an identifier representing a symbol with a value of 0 may not be created. If all identifiers for the symbols in a symbol string are created, the identifier molecules of the symbol string can be combined in a pool such that the presence of a particular identifier in the pool represents a value of 1 for the corresponding symbol position, and the absence of a particular identifier in the pool represents a value of 0 for the corresponding symbol position. An alternative approach can be taken, where an identifier can be created for a corresponding symbol value of 0, and an identifier representing a symbol with a value of 1 may not be created. In some implementations, the finisher includes a third print head configured to dispense the reaction mixture onto coordinates on the substrate. The finisher can further include an incubator, a pooling system, or both. The incubator can provide the specific temperature conditions or set of conditions necessary for the reaction to proceed to assemble the components and form the identifier nucleic acid molecule.
[0225]
[0251] In some implementations, the finisher dispenses the reaction mixture onto the coordinates before the first printhead dispenses the first droplet onto the coordinates, before the second printhead dispenses the second droplet onto the coordinates, or both. Generally, the finisher can dispense the reaction mixture onto the coordinates before any droplets are dispensed, after the first droplet is dispensed, but at any time before the last droplet is dispensed or after all droplets are dispensed.
[0226]
[0252] In some implementations, the system includes at least one roller that moves the substrate past the first print head, the second print head, and the finisher. In some implementations, the roller provides linear motion of the substrate. Generally, the roller can provide two-dimensional or three-dimensional movement of the substrate, which can be through one or multiple passes through each of the first and second print heads and the finisher. In some implementations, the roller is part of a reel-to-reel system that achieves linear motion of the substrate at a constant velocity.
[0227]
[0253] In some implementations, the substrate forms a continuous loop of material, and at least one roller is part of a set of rollers that passes coordinates on the substrate through the first printhead, the second printhead, and the finisher multiple times. In general, it may be desirable to configure the system so that at least one roller does not contact any of the coordinates on the substrate to prevent possible abrasion or contamination of the material being dispensed onto the substrate. In particular, the substrate has a first surface onto which the first droplets, the second droplets, and the reaction mixture are dispensed, and a second surface opposite the first surface, and at least one roller contacts the second surface but does not contact the first surface. Alternatively, even if at least one of the rollers contacts the first surface, the roller may be grooved so that it does not contact any of the coordinates onto which the material is dispensed.
[0228]
[0254] In some implementations, the system includes a second roller including at least one valley, wherein the second roller contacts the first surface such that the at least one valley is aligned with the coordinate. In some implementations, the system includes a second roller, wherein the substrate is rotated 180 degrees between the at least one roller and the second roller or in a helical path such that the second roller contacts the second surface and does not contact the first surface.
[0229]
[0255] In some implementations, the coordinates have a diameter or spacing from other coordinates on the substrate of between 1 micrometer and 200 micrometers. In some implementations, the first and second droplets each have a volume of between 1 pL and 50 pL.
[0230]
[0256] In some implementations, the system includes a register that tracks the movement of the substrate in real time to maintain alignment between the coordinates of the substrate and the first and second print heads. In some implementations, the first and second solutions incorporate a dye, and the system includes a spot imager that includes a camera that verifies proper dispensing of the first and / or second droplets.
[0231]
[0257] In some implementations, the system includes a spot dryer that dries the first and second droplets on the substrate. In some implementations, the first print head includes a first plurality of nozzles that dispense droplets of the first solution at different coordinates on the substrate. In some implementations, the first print head includes a second plurality of nozzles that dispense droplets of the third solution at different coordinates on the substrate.
[0232]
[0258] In some implementations, the system includes a substrate. In some implementations, the substrate includes a low-bond plastic. In some implementations, the substrate includes polyethylene terephthalate (PET) or polypropylene.
[0233]
[0259] In some implementations, the first and second print heads are mounted in the system at an angle relative to the motion of the substrate, which angle enables overprinting on coordinates. In some implementations, the first print head is a MEMS thin-film piezoelectric inkjet head or a MEMS thermal inkjet head. In some implementations, the first and second print heads are positioned along the same track to dispense droplets on coordinates, and the system includes an additional print head positioned along at least one additional track to dispense droplets on other coordinates within the corresponding track.
[0234]
[0260] In some implementations, the finisher has a fixed internal temperature optimal for reaction incubation. In some implementations, the finisher has a fixed humidity level to control evaporation of the reaction mixture during incubation. In some implementations, the finisher includes a heater that heats the substrate before incubation to prevent condensation. In some implementations, the finisher includes a pooling system that consolidates multiple reactions from different coordinates on the substrate into a container. In some implementations, the finisher dispenses a reaction inhibitor onto the coordinates of the substrate before solidification.
[0235]
[0261] In some implementations, the container contains a pooling solution and a reaction inhibitor, hi some implementations, the reaction inhibitor is ethylenediaminetetraacetic acid (EDTA).
[0236]
[0262] In some implementations, the system includes a membrane that captures nucleic acids from fluid collected from different coordinates on the substrate. In some implementations, the system includes a scraper that removes nucleic acids from the substrate. In some implementations, multiple reactions from different coordinates are pooled together into an emulsion that allows the multiple reactions to maintain their contents after pooling.
[0237]
[0263] In some implementations, the substrate is coated with an immiscible liquid or oil. In some implementations, the system includes an oil dispenser that dispenses the oil on coordinates. In some implementations, the substrate is coated or patterned with beads that bind to the first and second component nucleic acid molecules. In some implementations, the system includes a bead dispenser that dispenses the beads on coordinates.
[0238]
[0264] In some implementations, the reaction mixture includes a ligase. In some implementations, the first solution, the second solution, and the reaction mixture include an additive. In some implementations, the additive is configured to enable compatibility of the first solution with a first print head, compatibility of the second solution with a second print head, or compatibility of the reaction mixture with a finisher. In some implementations, the additive mitigates evaporation of the first solution, the second solution, or the reaction mixture. In some implementations, the additive includes at least one of a wetting agent, a surfactant, and a biocide.
[0239]
[0265] In some implementations, the system includes a computer processor configured to execute instructions for operating the system, which may include (1) a set of instructions for moving the substrate past the print heads, such as by controlling a set of rollers, and (2) another set of instructions for specifying the time for each print head or corresponding nozzle to dispense fluid.
[0240]
[0266] In one aspect, the disclosure provides a system for assembling nucleic acid molecules, the system including: (a) a first print head configured to dispense first droplets of a first solution comprising a first component nucleic acid molecule onto coordinates on a substrate; (b) a second print head configured to dispense second droplets of a second solution comprising a second component nucleic acid molecule onto coordinates on the substrate such that the first and second component nucleic acid molecules are juxtaposed on the substrate; and (c) a finisher that dispenses a reaction mixture onto the coordinates on the substrate to physically link the first and second component nucleic acid molecules, provides the conditions necessary to physically link the first and second component nucleic acid molecules, or both.
[0241]
[0267] In some implementations, the finisher includes a third print head configured to dispense the reaction mixture onto coordinates on the substrate. The finisher may further include an incubator, a pooling system, or both. Generally, the finisher can dispense the reaction mixture at any time. Specifically, the reaction mixture can be dispensed onto coordinates before the first print head dispenses a first droplet onto the coordinates, before the second print head dispenses a second droplet onto the coordinates, or both.
[0242]
[0268] In some implementations, the assembled nucleic acid molecules include DNA encoding genes, peptides, or RNA. The assembled nucleic acid molecules can include DNA aptamer libraries.
[0243] Noise Reduction
[0269] While various implementations of noise reduction techniques have been shown and described herein, such implementations are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the implementation of the techniques described herein ("this specification") may be employed.
[0244]
[0270] Described herein are technologies, including systems, devices, and methods, for writing, storing, reading, and computing digital information using nucleic acid molecules (e.g., DNA or RNA) with low or no noise, including, for example, devices and methods for writing and / or reading digital information from nucleic acid sequences, such as using next-generation sequencing (NGS) methods, or using nanochannels and sensors to detect one or more components of a translocated nucleic acid strand.
[0245]
[0271] Using the disclosed methods and systems, computer data or information can be encoded into multiple "identifiers," each of which can represent, for example, one or more bits of the original information, as described above. These identifiers can include two or more building blocks, called "building blocks," each of which has a nucleic acid sequence. A set of prefabricated building blocks can be divided into "layers." For example, each individual identifier can be constructed by ligating one building block from each of multiple layers in a fixed order. The order can be predetermined or random. Identifiers can be assembled using chemical techniques, such as ligation techniques, as described above. Identifiers are generated using the techniques disclosed herein to minimize or eliminate the presence of unligated or incompletely ligated oligonucleotides, called "fragments," in the final product. These unintentionally cleaved identifiers / oligonucleotides, which are a result of the ligation process, can later be amplified (e.g., using PCR) and cause noise in the system. The presence of unligated products and incompletely ligated products (unligated components and / or fragments) can result in spurious products generated during PCR by recombination (a process in which two oligonucleotides are ligated at overlapping regions, generating new, unintended chimeric products during PCR). The presence of unligated or incompletely ligated products can contribute to noise by acting as primers, generating unintended truncated identifiers. These products can act as competing templates for PCR, reducing the signal of the full-length identifier (FLI). Furthermore, the presence of these unligated products can interfere with qPCR quantification.
[0246]
[0272] Existing techniques include out-of-solution processes to remove these products, such as DNA purification methods (silica and / or anion exchange columns, solid-phase reversible immobilization (SPRI) beads and / or gel extraction) and / or nuclease treatment. These existing techniques, which focus on removing unligated and incompletely ligated products, are inefficient and may result in the loss of full-length identifiers while still affecting PCR, leaving enough unligated and incompletely ligated products behind. In systems attempting to process 10 or more nucleotides, purification and sequence accuracy issues can become even more pronounced.
[0247]
[0273] The techniques described herein address this issue. These techniques include methods that involve blocking PCR extension as a mechanism for improving decodeability and reducing noise. This technique includes methods that involve blocking PCR extension to improve qPCR quantification. This technique includes the use of ddNTPs to prevent recombination. These techniques offer flexibility and can be used in conjunction with any existing purification method without affecting signal. PCR after ddNTP treatment is compatible with most polymerases.
[0248]
[0274] In one aspect, the present disclosure provides techniques for writing information (e.g., digital information) into nucleic acid molecules in a manner that reduces or minimizes potential noise resulting from incomplete ligation by modifying the nucleic acid with ddNTPs, polymerases, and / or other enzymes that block the generation or amplification of chimeric fragments. An exemplary method includes: (a) generating a symbol string to represent the digital or other information; (b) assembling a plurality of components described herein, wherein each individual component of the plurality of components comprises a nucleic acid sequence; and (c) chemically linking two or more components of the plurality of components, e.g., via at least one sticky end of each of the two or more components, thereby generating a plurality of identifiers, wherein each identifier of the plurality of identifiers comprises two or more components. Each identifier of the plurality of identifiers corresponds to an individual symbol in the symbol string. The method includes selectively capturing or amplifying an identifier library comprising at least a subset of the plurality of identifiers.
[0249]
[0275] Information, e.g., digital information, stored in nucleic acids (e.g., identifiers) can be accessed by sequencing or hybridization assays. For example, primers or probes can be designed to bind to common or barcoded regions of nucleic acid sequences. This technology can provide amplification of any region of a nucleic acid molecule. The amplification product can then be read by sequencing the amplification product or by hybridization assays.
[0250]
[0276] PCR-based methods can be used to access and copy data from identifiers or nucleic acid sample pools. Primers are relatively short, single-stranded nucleic acids that can be used to initiate DNA synthesis. Common primer binding sites adjacent to identifiers within a pool or hyperpool can be used to easily copy information-containing nucleic acids. Alternatively, other nucleic acid amplification approaches, such as isothermal amplification, can be used to easily copy data from a sample pool or hyperpool (e.g., an identifier library).
[0251]
[0277] Current methods for removing incompletely assembled fragments (e.g., gel extraction) can result in significant loss of full-length product, but still retain large amounts of unligated product that can act as primers during PCR, thereby generating noise. Therefore, preventing PCR extension from these small fragments can be an important mechanism for reducing noise and improving identifier decodability.
[0252]
[0278] This specification includes techniques for modifying nucleic acid molecules to (1) prevent or reduce unwanted PCR amplification of unligated products (building blocks or fragments) and (2) reduce noise by removing unwanted products from the reaction volume prior to PCR amplification.
[0253] Preventing or reducing unwanted PCR amplification
[0279] The present disclosure includes a technique for modifying nucleic acid molecules (e.g., building blocks or fragments) to incorporate dideoxynucleotides (ddNTPs) at the 3' end of all nucleic acid molecules. The 3' end of a nucleic acid molecule is the region where PCR primers attach to the target molecule. ddNTPs are inhibitors of DNA polymerase and are used in Sanger sequencing. Sanger sequencing is a method of DNA sequencing based on the random incorporation of fluorescently labeled ddNTPs into DNA. These ddNTPs are labeled with each of the four deoxynucleotides (ddATP, ddCTP, ddGTP, ddTTP) and are therefore used to identify terminal nucleotides in a DNA sequence. Sequential PCR followed by gel electrophoresis yields labeled DNA fragments of various lengths that provide sequence information for the sample DNA. In contrast, the technique described herein does not use fluorescently labeled ddNTPs to determine the nucleotide sequence, but instead utilizes ddNTPs to selectively control and inhibit PCR extension from modified fragment molecules, thereby minimizing the formation of spurious products generated by recombination (e.g., nucleic acid molecules composed of two or more distinct sequence segments joined together during PCR). The edge components (nucleic acid molecules that make up or include the ends of the fully assembled identifier (FLI)) remain unmodified by ddNTPs at their ends, allowing PCT amplification of the full-length identifier.
[0254]
[0280] The principle of this technology is shown in Figure 16. Reducing or eliminating spurious products can reduce noise after PCR, thus improving the decodability of the information encoded in the identifier. Furthermore, halting PCR extension from these nucleic acids or fragments can reduce or prevent the amplification of partial products resulting from unligated products acting as primers. These partial products can act as competing templates during PCR (reducing primer availability for full-length products), potentially reducing the signal of the full-length identifier. Therefore, blocking PCR extension from unligated or incompletely ligated products can also improve PCR efficiency and improve the yield of full-length identifiers.
[0255]
[0281] Figure 17 shows an attempt to amplify an exemplary 360 base pair (bp) product using two sets of primers containing the same sequence. In lane 3, the primers contain no modifications, while in lane 6, both primers contain ddNTPs at their 3' ends. No signal is seen in lane 6, demonstrating that the presence of ddNTPs at the 3' ends of primers prevents amplification from those primers. The techniques described herein include methods for adding ddNTPs to the 3' ends of molecules that unintentionally act as primers, such as incompletely assembled oligonucleotides that are not intended to be formed through the process.
[0256]
[0282] Described herein are three exemplary methods for generating these ddNTP ends: (1) terminal transferase-based incorporation of ddNTPs, (2) end-filling using DNA polymerase, and (3) blunt-terminating and ddATP tailing.
[0257] 1. Terminal transferase-based incorporation of ddNTPs
[0283] In some implementations, the techniques described herein involve the use of a template-independent polymerase for the incorporation of ddNTPs. Template-independent polymerases, such as members of the X family of polymerase enzymes, do not require a template molecule to synthesize DNA (e.g., by synthesizing the reverse complement of a DNA strand). One suitable polymerase is terminal transferase (TdT), a template-independent polymerase that adds nucleotides to the 3' end of a DNA molecule. TdT can act on protruding, recessed, and / or blunt-ended DNA molecules and lacks exonuclease activity. Thus, the addition of ddNTPs via TdT (and / or another template-independent polymerase) can serve as a mechanism for adding ddNTPs to the ends of molecules, thereby preventing their ability to act as primers. Therefore, this technique differs from other techniques that use TdT, such as the controlled de novo synthesis of oligonucleotides for data encoding using TdT-mediated iterative DNA assembly from deoxynucleotide triphosphates (dNTPs).
[0258]
[0284] To facilitate detection of successful incorporation of ddNTPs, a combination of ddNTPs and dNTPs can be used. In an exemplary process, a reaction volume includes multiple components in one or more layers. The components include multiple edge components. The components have 3' overhangs. A volume containing a template-independent polymerase (e.g., TdT) and a volume containing one or more ddNTPs (e.g., ddATP, ddCTP, ddGTP, or ddTTP) are added. In some implementations, a volume containing one or more dNTPs (e.g., dATP, dCTP, dGTP, or dTTP) is added. TdT reacts with the components, adding ddNTPs and / or dNTPs, for example, to the 3' end of the components. In some implementations, the ends of edge components configured to terminate FLI are modified (e.g., using a hairpin loop at the end) to protect the FLI from the noise reduction process described below, for example. To ligate the components, T4 ligase, ATP, and / or other suitable ligation reagents are added to the volume. The ligation reaction yields a volume containing a set of oligonucleotides containing building blocks, fragments, and FLIs. In some implementations, after the ligation reaction, the termini are deprotected, and hairpins at the ends of the edge building blocks are removed, for example, using DNA-dependent protein kinase catalytic subunit (DNA-PKcs) and Artemis (structure-specific endonuclease). Subsequently, PCR primers are hybridized to the oligonucleotides. The oligonucleotides are then subjected to PCR amplification. Any chimeric building blocks or fragments with terminal ddNTPs are not amplified due to inhibition of DNA polymerase by ddNTPs, and modified building blocks or fragments cannot act as PCR primers. This process increases the signal (FLI) to noise (building blocks, fragments) ratio of the DNA encoding / decoding process. The chemical process can be performed, for example, in a batch process or a microfluidic device. After PCR amplification, the oligonucleotides can be sequenced to read the information encoded therein.
[0259]
[0285] The results of automated electrophoresis (e.g., Agilent® Tapestation®) are shown in Figures 18A-18C. Figure 18A shows the results of a nine-layer (nucleic acid building block) ligation. Figure 18B shows the results of a nine-layer ligation of building blocks treated with dNTPs only. Figure 18C shows the results of a nine-layer ligation of building blocks treated with ddNTPs and dNTPs. Figures 18A-18C show that ddNTPs can be successfully incorporated by TdT because ddNTPs prevent the extension of long 3' ends resulting from TdT incorporation of dNTPs, as evidenced by the peak intensities of over 60, representing a large number of molecules greater than approximately 1500 bp (dashed box in Figure 18B). Without ddNTPs, the majority of molecules in the 50-500 bp range (dashed box in Figure 18A) would "shift" to the 1500+ bp range (Figure 18B). Successful incorporation of ddNTPs is shown in Figure 18C, where ddNTPs prevent uncontrolled ligation, resulting in large multilayered oligonucleotides (as evidenced by the lack of a peak shift from the 50 bp-500 bp range to the 1500+ bp range, as shown in Figure 18B).
[0260]
[0286] The use of a combination of ddNTPs and dNTPs was also tested with a 16-layer, 129,572-identifier library generated by an exemplary DNA writer as described hereinabove. Table 2 presents data demonstrating an increase in the formation of desired products (1-sym observation), a decrease in the formation of undesired (e.g., chimeric) products (0-sym observation), and an approximately 25% improvement in SNR after ddNTP treatment with TdT in DNA writer-generated libraries.
[0261] [Table 2]
[0262] 2. End-fill incorporation of ddNTP
[0287] In some implementations, the techniques described herein involve end-fill-in incorporation of ddNTPs. Exemplary end-fill-in methods include the use of a DNA polymerase (e.g., T4 polymerase, Terminator DNA polymerase, TdT, Klenow fragment, T7 polymerase, Sulfolobus DNA polymerase IV, DNA polymerase I (E. coli)), a reaction buffer, and ddNTPs.
[0263]
[0288] In some implementations, nucleic acid molecules (e.g., identifier components) are designed so that they have 3' overhangs. In some implementations, nucleic acid molecules (e.g., identifier components) are designed so that they include 5' overhangs (Figure 19). In some implementations of technologies that use polymerases, 5' overhangs may be required because the polymerase synthesizes DNA in a 5'-3' fashion. Because the sequence organization is unaffected, any adjustments to component synthesis (e.g., converting a 3' overhang to a 5' overhang) do not affect any applicable encoding / decoding methods.
[0264]
[0289] In an exemplary process, a reaction volume includes multiple components in one or more layers, and the components include multiple edge components. The components have 5' overhangs. A volume containing a polymerase (e.g., T4 polymerase, terminator DNA polymerase, TdT) and a volume containing one or more ddNTPs (e.g., ddATP, ddCTP, ddGTP, or ddTTP) are added. In some implementations, a volume containing one or more dNTPs (e.g., dATP, dCTP, dGTP, or dTTP) is added. The ddNTPs / dNTPs are added to the 3' ends of the components by hybridization to the 5' overhangs. In some implementations, the ends of edge components configured to terminate FLI are modified to protect the FLI from the noise reduction process described below (e.g., using a hairpin loop at the end, for example). Exemplary reaction conditions for an end-fill-in reaction include incubation at 12°C for 15 minutes. After incubation, the product is purified (e.g., using gel extraction, silica columns, alcohol precipitation, or magnetic SPRI beads) to ensure that ddNTPs are not carried over to future reactions. To ligate the building blocks, T4 ligase, ATP, and / or other appropriate ligation reagents are added to the volume. The ligation reaction results in a volume containing a set of oligonucleotides containing building blocks, fragments, and FLIs. In some implementations, after the ligation reaction, the termini are deprotected, and hairpins at the ends of edge building blocks are removed, for example, using DNA-dependent protein kinase catalytic subunit (DNA-PKcs) and Artemis (structure-specific endonuclease). Subsequently, PCR primers are hybridized to the oligonucleotides. The oligonucleotides are then subjected to PCR amplification. Any chimeric building blocks or fragments with terminal ddNTPs will not be amplified due to inhibition of DNA polymerase by ddNTPs, and modified building blocks or fragments will not be able to act as PCR primers. This process increases the signal (FLI) to noise (building blocks, fragments) ratio of the DNA encoding / decoding process. The chemical process can be carried out in a batch process or in a microfluidic device.After PCR amplification, the oligonucleotides can be sequenced and the information encoded therein can be read.
[0265]
[0290] End-filling of the reaction can be incorporated into downstream nucleic acid processing workflows, for example, without the need for additional purification steps. Figure 20A illustrates the integration of a noise reduction step in an "excision run" workflow, e.g., a workflow involving excision of a gel band containing the nucleic acid of interest as described above. Figure 20B illustrates the integration of a noise reduction step in a "pooler run" workflow, e.g., a workflow using the pooling system described herein. Furthermore, because the end-fill-in reaction is performed before qPCR, end-fill-in incorporation of ddNTPs can provide a more specific signal for quantification of full-length identifiers. In some implementations, ddNTP incorporation can be performed after gel extraction, for example, depending on the efficiency of the downstream processing method used.
[0266] 3. Smoothing and ddATP Tailing
[0291] In some implementations, the techniques described herein involve blunting and ddATP tailing. This method of ddNTP incorporation involves two steps. The first step involves removing the overhang (blunting). This removal can be performed using a nuclease, such as mung bean nuclease, nuclease P1, exonuclease I, exonuclease III, micrococcal nuclease, S1 nuclease, or a polymerase containing exonuclease activity. In some implementations, blunting can be achieved by filling in the overhang, for example, as described above. The second step involves adding ddATP to the 3' end of the currently blunt molecule. This process can be performed by a polymerase such as Taq polymerase, Klenow fragment, etc.
[0267]
[0292] In an exemplary process, a reaction volume includes multiple components in one or more layers. The components include multiple edge components. The components have 3' or 5' overhangs. In a first step, a volume containing a nuclease (e.g., mung bean nuclease, nuclease P1, or a polymerase containing exonuclease activity) is added to remove the overhangs, e.g., by degrading or hydrolyzing the single-stranded extensions. In a second step, a volume containing a polymerase (e.g., Taq polymerase, Klenow fragment) and a volume containing one or more ddATPs are added. In some implementations, the ends of the edge components configured to terminate the FLI are configured to protect the FLI from the noise reduction process described below (e.g., using a hairpin loop at the end). To ligate the components, T4 ligase, ATP, and / or other suitable ligation reagents are added to the volume. The ligation reaction results in a volume containing a set of oligonucleotides comprising the components, fragments, and FLI. In some implementations, after the ligation reaction, the termini are deprotected, and the hairpins at the ends of the edge components are removed, for example, using DNA-dependent protein kinase catalytic subunit (DNA-PKcs) and Artemis (a structure-specific endonuclease). Subsequently, PCR primers are hybridized to the oligonucleotides. The oligonucleotides are then subjected to PCR amplification. Any chimeric components or fragments with terminal ddATP are not amplified due to inhibition of DNA polymerase by ddATP, and modified components or fragments cannot act as PCR primers. This process increases the signal (FLI) to noise (components, fragments) ratio of the DNA encoding / decoding process. The chemical process can be performed in a batch process or on a microfluidic device. After PCR amplification, the oligonucleotides can be sequenced to read the information encoded therein.
[0268]
[0293] Each of the techniques described herein can be used alone or in combination with one or more of the other techniques described herein. In some implementations, modified nucleotides, such as acyclonucleotides (or acyclic nucleotides), can be used that can similarly block PCR extension from unligated products. In acyclonucleotides, the (single) phosphate nucleotide has a linkage arrangement between the sugar group and the phosphate group that is not cyclic, resulting in the loss of the 3'-OH group required for chain elongation. In principle, any modification that results in the loss or replacement of the 3'-OH group of the nucleotide (e.g., a C3 spacer, e.g., a C3 propyl spacer or replacement with a phosphate group that can be incorporated internally or at either end of the oligonucleotide during chemical synthesis) can be used for this purpose.
[0269]
[0294] In an exemplary process, a reaction volume includes multiple components in one or more layers. The components include multiple edge components. The components have 5' overhangs. A volume containing a polymerase (e.g., T4 polymerase, terminator DNA polymerase, TdT) and a volume containing one or more acyclonucleotides (e.g., acyclic A, C, T, or G) are added. Hybridization to the 5' overhangs adds the acyclonucleotides to the 3' ends of the components. In some implementations, the ends of edge components configured to terminate FLI are modified (e.g., using a hairpin loop at the end) to protect the FLI from the noise reduction process described below, for example. After incubation, the product is purified (e.g., using gel extraction, silica columns, alcohol precipitation, or magnetic SPRI beads) to ensure that ddNTPs are not carried over to future reactions. To ligate the components, T4 ligase, ATP, and / or other appropriate ligation reagents are added to the volume. The ligation reaction yields a volume containing a set of oligonucleotides containing building blocks, fragments, and FLIs. In some implementations, after the ligation reaction, the termini are deprotected, and hairpins at the ends of the edge building blocks are removed, for example, using DNA-dependent protein kinase catalytic subunit (DNA-PKcs) and Artemis (a structure-specific endonuclease). Subsequently, PCR primers are hybridized to the oligonucleotides. The oligonucleotides are then subjected to PCR amplification. Any chimeric building blocks or fragments with terminal acyclonucleotides are not amplified due to inhibition of DNA polymerase by the acyclonucleotides, and modified building blocks or fragments cannot act as PCR primers. This process increases the signal (FLI) to noise (building blocks, fragments) ratio of the DNA encoding / decoding process. The chemical process can be performed in a batch process or on a microfluidic device. After PCR amplification, the oligonucleotides can be sequenced to read the information encoded therein.
[0270]
[0295] Other mechanisms of blockage involve altering the structure of the DNA strand to block the activity of DNA polymerase. One example of this is creating an oligonucleotide containing a 3' flap that can block polymerase translocation and prevent amplification from unligated products (Figure 21).
[0271]
[0296] In an exemplary process, the reaction volume includes multiple components in one or more layers. The components include multiple edge components. The components have 3' or 5' overhangs. In some implementations, the ends of the edge components configured to terminate the FLI are modified (e.g., using a hairpin loop at the end) to protect the FLI from, for example, the noise reduction process described below.
[0272]
[0297] In some implementations, the building blocks are ligated to form the identifier. To ligate the building blocks, T4 ligase, ATP, and / or other suitable ligation reagents are added to the volume. The ligation reaction results in a volume containing a set of oligonucleotides comprising the building block, the fragment, and the FLI. After ligation, a volume containing a polymerase (e.g., T4 polymerase, terminator DNA polymerase, TdT) and a volume containing one or more single-stranded DNA molecules comprising a non-hybridizing region (a "flap") are added. In some implementations, the flap is added to the 3' end of the building block by hybridization to the 5' overhang. In some implementations, after the ligation reaction, the ends are deprotected, and hairpins at the ends of the edge building blocks are removed, for example, using DNA-dependent protein kinase catalytic subunit (DNA-PKcs) and Artemis (a structure-specific endonuclease). Subsequently, PCR primers are hybridized to the oligonucleotides. The oligonucleotides are then subjected to PCR amplification. Any chimeric components or fragments with terminal flaps will not be amplified due to DNA polymerase inhibition by the flaps, and modified components or fragments cannot act as PCR primers. This process increases the signal (FLI) to noise (components, fragments) ratio of the DNA encoding / decoding process. The chemical process can be performed in batch processing or on a microfluidic device. After PCR amplification, the oligonucleotides can be sequenced to read the information encoded therein.
[0273]
[0298] In addition to its use to improve decipherability, the technology described herein provides the ability to "encrypt" DNA-encoded information / data by adding shorter fragments of DNA that promote recombination. This method can increase noise and prevent decoding under standard processes. Therefore, in order to successfully decipher such data, a noise reduction method, such as the ddNTP treatment described herein, is required. Similarly, large targeting oligonucleotides containing terminal ddNTPs can be added to samples to competitively block the amplification of either the entire data set or the region downstream of a specific sequence, providing a mechanism for reducing decipherability for data security purposes.
[0274] Nuclease resistance
[0299] This specification describes a technique for reducing noise when decoding information stored in nucleic acids. As described above, noise can be caused by the presence of incomplete products of the nucleic acid synthesis process for encoding information into nucleic acid molecules (e.g., the assembly of a nucleic acid identifier from two or more components). Removal of shorter fragments in existing workflows involves a subjective manual process, namely gel extraction. This approach removes a large amount of shorter fragments from the sample (but leaves a significant number remaining) at the expense of a significant amount of full-length product. Furthermore, there is a high variability in the yield recovery using this method. Therefore, alternatives to this approach, as described herein, may be beneficial.
[0275]
[0300] In some implementations, nucleases can be used. Nucleases are highly active and powerful enzymes that can degrade DNA by cleaving the phosphodiester bonds between nucleic acid nucleotides. Nucleic acid molecules, such as identifiers or their components, can be modified to provide resistance to nucleases. This specification describes modifications that can be added to edge components that can selectively protect full-length identifiers (FLIs) from nuclease activity while removing unprotected molecules of DNA, such as incompletely assembled products (Figure 22). This process can increase the signal-to-noise ratio (SNR) and the likelihood of successful decoding.
[0276]
[0301] Described herein are techniques for preventing FLIs from being destroyed by nucleases. This prevention can be achieved by protecting the edge components of the FLI from nuclease activity. In some implementations, for example, as described above, the edge components are modified using one or more of the techniques described below prior to assembly of the identifier.
[0277]
[0302] In some implementations, the technology described herein includes edge elements with hairpin loops. As described above, designing edge elements with hairpin loops results in the FLI being uniquely capped by hairpins, as shown, for example, in Figure 23. These hairpins provide resistance to one or more nucleases, but expose unprotected molecules that are subject to cleavage and degradation by the nuclease, leaving the FLI.
[0278]
[0303] In some implementations, the techniques described herein include protelomerase-based FLI protection. In some implementations, protelomerase can be used to protect edge components from nuclease activity. Protelomerase is an enzyme that cleaves at specific sites, resulting in the generation of covalently closed ends, as shown in FIG. 24. This method is based on the same principle as the use of hairpin-like modifications to provide nuclease resistance. Protelomerase treatment results in edge components with covalently closed ends that provide protection from nucleases for FLI, as shown in FIG. 25. In some implementations, protelomerase recognition sequences can be added to edge components after ligation. In some implementations, protelomerase recognition sequences can be added to edge components before ligation.
[0279]
[0304] In some implementations, the technology described herein includes FLI protection based on phosphorothioate bonds. In some implementations, phosphorothioate (pt) bonds can be used to protect edge components from nuclease activity. A phosphorothioate (pt) bond is a phosphodiester linkage in which a non-bridging oxygen is replaced with sulfur. Substituting the oxygen with sulfur does not change the reactivity of the bond, but because the phosphorus is now attached to a distinct group, a chiral center with two possible configurations, "RP" and "SP" (Figure 26). In certain configurations, phosphorothioate (pt) bonds confer resistance to certain nucleases. For example, exonuclease III has been shown to cleave SP bonds but not RP bonds. In some implementations, at least 3 pt bond modifications are incorporated at each end of the FLI. In some implementations, at least 5 pt bond modifications are incorporated at each end of the FLI. In some implementations, at least 2, 3, 4, 5, 10, or more pt bond modifications are incorporated at each end of the FLI. Incorporating at least these numbers of pt linkages increases the likelihood that a preferred configuration will exist within the target molecule. As with the previous modification, edge components containing multiple phosphorothioate linkages can protect FLI from nuclease treatment and can be used to remove incompletely assembled products.
[0280]
[0305] In some implementations, the techniques described herein include an inverted dT modification. In some implementations, the inverted dT can be used to protect edge components from nuclease activity. The inverted dT is a modification at the 3' end of an oligonucleotide (Figure 27). This modification creates a 3'-3' bond that prevents DNA polymerase from further extending the DNA sequence. Furthermore, this modification confers resistance to nucleases. Designing and implementing identifiers (e.g., edge components) that include an inverted dT at the end of the identifier can be a mechanism to enrich for FLI after nuclease treatment.
[0281]
[0306] In some implementations, the techniques described herein involve the incorporation of modified sugar residues. In some implementations, modified sugar residues can be used to protect edge components from nuclease activity (Figure 28). Similar to the techniques described above, the presence of modified sugar residues on edge components can provide nuclease resistance to FLI, although incompletely assembled products remain degradable by nucleases. Examples of sugar modifications include ribose, 2'-methoxy, and 2'-methoxyethyl modifications.
[0282]
[0307] In some implementations, the techniques described herein involve circularization of the identifier. In some implementations, once the identifier is fully ligated, circularization of the identifier molecule can be used to protect the edge components from nuclease activity. This technique involves designing and implementing overhangs on the edge components that allow the FLI to circularize itself and / or generating unique overhangs that allow the FLI to ligate to a DNA backbone. Nucleases that degrade linear DNA can be used to remove incompletely assembled fragments from the sample. This approach can be performed during ligation of the components or can be modified to provide circularization during post-processing. The latter approach involves designing and implementing restriction enzyme sites on the edge components. Ligation blunting (e.g., using nucleases) can be used to remove overhangs from incompletely assembled fragments (to prevent indiscriminate ligation). Subsequent restriction enzyme digestion can generate overhangs that allow the FLI to circularize (after the ligation reaction), thereby protecting the edge components from nuclease activity.
[0283]
[0308] In some implementations, the techniques described herein involve helicase-based targeted removal of incompletely assembled fragments. In some implementations, helicases can be used to remove incompletely assembled fragments. Helicases are proteins that move directionally along the nucleic acid phosphodiester backbone, separating two hybridized nucleic acid strands. This technique involves targeting overhangs inherently present in incompletely assembled products by utilizing overhang-targeting helicases. These helicases can specifically unwind DNA containing overhangs that provide access to nucleases that act on single-stranded DNA. Overhangs that lack FLI (e.g., due to blunt ends or circularization) lack overhangs and are therefore protected.
[0284] Exemplary Implementation
[0309] Item 1. A method for writing information into a nucleic acid molecule with reduced noise, comprising: determining a symbol string representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, wherein individual identifiers in the plurality of identifiers correspond to individual symbols in the symbol string; generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, wherein each individual building block of the plurality of oligonucleotides is a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding a first volume comprising a template-independent polymerase and an amount of dideoxynucleotides (ddNTPs) to a reaction volume comprising the plurality of building blocks; (c) incubating the reaction volume to add ddNTPs to the 3' ends of at least some of the building blocks; (d) adding reagents to the reaction volume to chemically link two or more building blocks of the plurality of building blocks, thereby generating a plurality of identifiers and a plurality of fragments; and (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification, wherein PCR amplification of any oligonucleotides that include ddNTPs is inhibited.
[0285]
[0310] Item 2. The method of item 1, wherein the plurality of components includes a plurality of edge components, each edge component having an end, and each edge component is configured such that the end of the edge component is the end of the identifier.
[0286]
[0311] Item 3. The method according to any one of Items 1 to 2, wherein the first volume comprises deoxynucleotides (dNTPs).
[0287]
[0312] Item 4. The method according to any one of Items 1 to 3, wherein the polymerase comprises terminal transferase (TdT) that catalyzes the addition of nucleotides to the 3' ends of the multiple building blocks.
[0288]
[0313] Item 5. The method according to any one of Items 1 to 4, wherein the 3' ends of the multiple components comprise 3' overhangs.
[0289]
[0314] Item 6. The method according to any one of Items 1 to 5, wherein the ddNTP is one or more of ddATP, ddGTP, ddTTP, and ddCTP.
[0290]
[0315] Item 7. The method according to any one of Items 1 to 6, wherein the polymerase comprises T4 polymerase or terminator DNA polymerase.
[0291]
[0316] Item 8. The method according to any one of Items 1 to 7, comprising forming 5'-end overhangs on a plurality of components.
[0292]
[0317] Item 9. The method according to any one of Items 1 to 8, comprising removing overhangs from the 3' and / or 5' ends of the multiple building blocks prior to addition of the ddNTP molecules.
[0293]
[0318] Item 10. The method of Item 9, wherein the overhangs are removed using a nuclease.
[0294]
[0319] Item 11. The method of Item 10, wherein the nuclease comprises one or more of mung bean nuclease, nuclease P1, exonuclease I, exonuclease III, micrococcal nuclease, S1 nuclease, or a polymerase containing exonuclease activity.
[0295]
[0320] Item 12. The method according to any one of Items 1 to 11, wherein adding the ddNTP molecule comprises adding ddATP to the 3' end of the building block using a polymerase.
[0296]
[0321] Item 13. The method according to Item 12, wherein the polymerase comprises taq polymerase or the Klenow fragment.
[0297]
[0322] Item 14. A method for writing information to a nucleic acid molecule, comprising: determining a symbol string representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, wherein each identifier in the plurality of identifiers corresponds to an individual symbol in the symbol string; generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, wherein each individual building block of the plurality of oligonucleotides is a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding a first volume comprising a polymerase and an amount of acyclonucleotides to a reaction volume comprising the plurality of building blocks; (c) incubating the reaction volume to add acyclonucleotides to the 3' ends of at least some of the building blocks; (d) adding reagents to the reaction volume to chemically link two or more building blocks of the plurality of building blocks, thereby generating a plurality of identifiers and a plurality of fragments; and (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification, wherein PCR amplification of any oligonucleotides that comprise acyclonucleotides is inhibited.
[0298]
[0323] Item 15. A method for writing information to a nucleic acid molecule, comprising: determining a symbol string representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, wherein each identifier in the plurality of identifiers corresponds to an individual symbol in the symbol string; generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, wherein each individual building block of the plurality of oligonucleotides is a nucleic acid molecule having a nucleic acid sequence, a 3' end, and a 5' end; (b) adding reagents to a reaction volume to chemically link two or more building blocks of the plurality of components, thereby generating a plurality of identifiers and a plurality of fragments; (c) adding a first volume comprising a polymerase and a plurality of 3'-DNA flaps to the reaction volume comprising the plurality of building blocks; (d) incubating the reaction volume to add 3'-DNA flaps to the 3' ends of at least some of the building blocks and fragments; and (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification, wherein PCR amplification of any oligonucleotides comprising a 3'-DNA flap is inhibited.
[0299]
[0324] Item 16. A method for writing information to a nucleic acid molecule, the method comprising: (a) determining a symbol string representing the information; (b) constructing a plurality of components, each individual component of the plurality of components being a nucleic acid molecule having a nucleic acid sequence, a 3' end and a 5' end; (c) chemically linking two or more components of the plurality of components, thereby generating a plurality of identifiers, each identifier of the plurality of identifiers comprising two or more components, each identifier having a first end and a second end, each component located at the first end of the identifier being a first edge component, and each component located at the second end of the identifier being a second edge component; and (d) chemically modifying each end of the first edge component, the second edge component, or both, such that the first edge component and / or the second edge component is protected from exonuclease activity, wherein each identifier of the plurality of identifiers corresponds to an individual symbol in the symbol string.
[0300]
[0325] Item 17. The method according to Item 16, wherein modifying the termini comprises constructing a hairpin loop at one or both of the termini, thereby protecting one or both termini from nuclease activity.
[0301]
[0326] Item 18. The method of any one of Items 16 to 17, wherein modifying the termini comprises adding a protelomerase recognition sequence to one or both of the termini to covalently close the one or both termini, thereby protecting the one or both termini from nuclease activity.
[0302]
[0327] Item 19. The method according to any one of Items 16 to 18, wherein modifying the termini comprises performing a phosphorothioate bond at one or both of the termini to replace a non-bridging oxygen in the phosphate backbone of the terminal oligonucleotide with a sulfur atom, thereby protecting one or both termini from nuclease activity.
[0303]
[0328] Item 20. The method according to Item 19, wherein modifying the termini comprises performing multiple phosphorothioate linkages at one or both of the termini.
[0304]
[0329] Item 21. The method according to Item 20, wherein modifying the termini comprises implementing at least three phosphorothioate bonds at one or both of the termini.
[0305]
[0331] Item 22. The method according to any one of Items 16 to 21, wherein modifying the termini comprises performing an inverse dT modification at one or both of the termini to create a 3'-3' linkage, thereby protecting one or both termini from nuclease activity.
[0306]
[0331] Item 23. A method described in any one of items 16 to 22, wherein modifying the termini includes performing a sugar residue modification at one or both of the termini, thereby protecting one or both termini from nuclease activity.
[0307]
[0332] Item 24. The method according to any one of Items 16 to 23, wherein modifying the termini comprises circularizing the identifier and linking the termini, thereby protecting one or both termini from nuclease activity.
[0308]
[0333] Item 25. The method according to any one of Items 16 to 24, wherein modifying the termini includes modifying the termini with a restriction enzyme site.
[0309]
[0334] Item 26. The method of any one of items 1 to 25, comprising targeting an overhang uniquely present on the incompletely assembled identifier by utilizing a helicase to separate two hybridized nucleic acid strands, thereby providing access to a nuclease that acts on single-stranded DNA.
[0310]
[0335] Item 27. The method according to any one of items 1 to 26, comprising treating the component with an exonuclease.
[0311]
[0336] Item 28.
[0337] 28. The method of any one of items 1 to 27, comprising selectively capturing or amplifying an identifier library comprising at least a subset of the plurality of identifiers.
[0312]
[0338] Item 29. The method of any one of Items 1 to 28, wherein each symbol in the string is one of one or more possible symbol values.
[0313]
[0339] Item 30. The method of Item 29, wherein each symbol in the string is one of two possible symbol values.
[0314]
[0340] Item 31. The method of any one of Items 29 to 30, wherein one symbol value at each position of the symbol string can be represented by the absence of a distinct identifier in an identifier library.
[0315]
[0341] Item 32. The method of Item 30, wherein the two possible symbol values are bit values of 0 and 1, and the individual symbols in the symbol string having the bit value of 0 can be represented by the absence of a distinct identifier in an identifier library, and the individual symbols in the symbol string having the bit value of 1 can be represented by the presence of the distinct identifier in the identifier library, or vice versa.
[0316]
[0342] Item 33. The method of any one of items 1 to 32, comprising chemically linking two or more components from two or more layers, each layer of the two or more layers comprising a distinct set of components.
[0317]
[0343] Item 34. The method of Item 33, wherein the individual identifiers from the identifier library include one component from each of the two or more layers.
[0318]
[0344] Item 35. The method of Item 34, wherein the two or more components are assembled in a fixed order.
[0319]
[0345] Item 36. The method of Item 34, wherein the two or more components are assembled in any order.
[0320]
[0346] Item 37. The method of Item 34, wherein the two or more components are assembled with one or more partition components disposed between two components from different layers of the two or more layers.
[0321]
[0347] Item 38. The method of Item 33, wherein the individual identifiers include one component from each layer of the subset of two or more layers.
[0322]
[0348] Item 39. The method of Item 33, wherein the individual identifiers include at least one component from each of the two or more layers.
[0323]
[0349] Item 40. The method of any one of items 1 to 39, comprising using an endonuclease to generate at least one sticky end of each component of the plurality of components.
[0324]
[0350] Item 41. The method of Item 40, wherein the at least one sticky end is at the 5' end of each of the components.
[0325]
[0351] Item 42. The method of Item 40, wherein the at least one sticky end is at the 3' end of the individual building block.
[0326]
[0352] Item 43. The method according to any one of Items 40 to 42, comprising generating two sticky ends of the individual components.
[0327]
[0353] Item 44. The method according to any one of Items 40 to 43, wherein the at least one sticky end is at least one nucleotide in length.
[0328]
[0354] Item 45. The method according to any one of Items 40 to 44, wherein the at least one sticky end is 6 nucleotides in length.
[0329]
[0355] Item 46. The method of any one of Items 1 to 45, wherein the plurality of identifiers includes a nucleic acid sequence that stores metadata of the information or conceals the information.
[0330]
[0356] Item 47. The method of any one of items 1 to 46, wherein two or more identifier libraries are combined and each identifier library of the two or more identifier libraries is tagged with a separate barcode.
[0331]
[0357] Item 48. The method of any one of items 28 to 47, wherein each individual identifier in the identifier library comprises a distinct barcode.
[0332]
[0358] Item 49. The method according to any one of Items 1 to 48, wherein the plurality of identifiers or the plurality of components containing the identifiers are selected to facilitate read, write, access, copy, and delete operations.
[0333]
[0359] Item 50. The method of any one of items 1 to 49, wherein chemically linking comprises ligating two or more components of the plurality of components together using a reagent comprising a ligase.
[0334]
[0360] Item 51. The method according to Item 50, wherein the ligase is T4 ligase, T7 ligase, T3 ligase or E. coli ligase.
[0335]
[0361] Item 52. The method according to any one of Items 50 to 51, wherein the reagent further contains an additive.
[0336]
[0362] Item 53. The method according to Item 52, wherein the additive increases the efficiency of the ligase.
[0337]
[0363] Item 54. The method according to any one of Items 52 to 53, wherein the additive comprises polyethylene glycol (PEG).
[0338]
[0364] Item 55. The method according to Item 54, wherein the PEG is PEG400, PEG6000, PEG8000, or any combination thereof.
[0339]
[0365] Item 56. The method according to any one of Items 50 to 55, wherein the ligation reaction time is at least 1 minute.
[0340]
[0366] Item 57. The method according to any one of Items 50 to 55, wherein the ligation is carried out at 30°C or higher.
[0341]
[0367] Item 58. The method according to any one of Items 50 to 57, further comprising inactivating the ligase using a buffer containing EDTA or guanidine thiocyanate.
[0342]
[0368] Item 59. The method according to any one of Items 50 to 58, wherein the final concentration of the ligase is at least about 5 CEU / μL.
[0343]
[0369] Item 60. The method according to any one of Items 50 to 59, wherein the reagent further comprises a glycerol molecule.
[0344]
[0370] Item 61. The method of any one of items 1 to 60, wherein chemically linking comprises using overlap extension polymerase chain reaction (PCR).
[0345]
[0371] Item 62. The method according to any one of Items 1 to 61, wherein the individual components are deoxyribonucleic acid (DNA) or ribonucleic acid.
[0346]
[0372] Item 63. The method according to any one of Items 1 to 62, wherein the individual components are rehydrated.
[0347]
[0373] Item 64. The method according to any one of items 1 to 63, wherein the individual components are rehydrated from dehydrated components.
[0348]
[0374] Item 65. The method of any one of Items 28 to 64, further comprising dehydrating the identifier library by dehydrating each individual identifier of at least the subset of the plurality of identifiers.
[0349]
[0375] Item 66. The method of any one of Items 28 to 65, wherein each individual identifier of at least the subset of the plurality of identifiers is dehydrated.
[0350]
[0376] Item 67. The method of any one of Items 65 to 66, further comprising rehydrating each individual identifier of at least the subset of the plurality of identifiers.
[0351]
[0377] Item 68. The method according to any one of Items 1 to 67, further comprising adding a preservative additive to the identifier library to prevent deterioration of the identifiers.
[0352]
[0378] Item 69. The method according to any one of Items 1 to 68, wherein the plurality of identifiers are copied by PCR.
[0353]
[0379] Item 70. The method of Item 69, wherein the PCR has at least 10 cycles.
[0354]
[0380] Item 71. The method of Item 69, wherein the plurality of identifiers are amplified by PCR to a concentration of 10 nanograms / microliter.
[0355]
[0381] Item 72. The method according to any one of Items 69 to 71, wherein the PCR is emulsion PCR.
[0356]
[0382] Item 73. The method according to any one of Items 1 to 72, wherein the plurality of identifiers are copied by linear amplification.
[0357]
[0383] Item 74. The method of any one of Items 69 to 73, wherein after the PCR, linear amplification is used to create more copies of the plurality of identifiers.
[0358]
[0384] Item 75. The method of any one of items 1 to 74, wherein the subset of the plurality of identifiers is accessed in one or more PCR reactions.
[0359]
[0385] Item 76. The method of any one of items 1 to 75, wherein a subset of the plurality of identifiers is accessed with one or more affinity-tagged probes.
[0360]
[0386] Item 77. The method of any one of Items 75 to 76, wherein the identifiers of the subset of the plurality of identifiers have a set of common components.
[0361]
[0387] Item 78. The method of any one of items 1 to 77, wherein the identifier is purified by gel electrophoresis.
[0362]
[0388] Item 79. The method according to any one of items 1 to 78, wherein the identifier is purified by an affinity-tagged probe.
[0363]
[0389] Item 80. The method of any one of items 1 to 79, wherein the identifier is amplified using PCR.
[0364]
[0390] Item 81. The method of any one of Items 1 to 80, wherein the identifier is designed to avoid thymine-thymine dinucleotides or cytosine-cytosine dinucleotides.
Claims
1. 1. A method for writing information into a nucleic acid molecule with reduced noise, comprising: determining a string of symbols representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, each identifier of the plurality of identifiers corresponding to a respective symbol in the string of symbols, wherein generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, each individual building block of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end and a 5' end; (b) adding a first volume containing a template-independent polymerase and an amount of dideoxynucleotides (ddNTPs) to a reaction volume containing the plurality of components; (c) incubating the reaction volume to add ddNTPs to the 3' ends of at least some of the building blocks; (d) adding a reagent to the reaction volume to chemically link two or more components of the plurality of components, thereby generating the plurality of identifiers and the plurality of fragments; (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification; wherein PCR amplification of any oligonucleotide containing a ddNTP is inhibited.
2. 2. The method of claim 1, wherein the plurality of components comprises a plurality of edge components, each edge component having a termination, and each edge component configured such that the termination of the edge component is a termination of an identifier.
3. 3. The method of claim 1, wherein the first volume comprises deoxynucleotides (dNTPs).
4. The method of any one of claims 1 to 3, wherein the polymerase comprises terminal transferase (TdT) that catalyzes the addition of a nucleotide to the 3' end of the plurality of said building blocks.
5. The method of any one of claims 1 to 4, wherein the 3' ends of the plurality of said building blocks comprise a 3' overhang.
6. The method according to any one of claims 1 to 5, wherein the ddNTP is one or more of ddATP, ddGTP, ddTTP, or ddCTP.
7. The method of any one of claims 1 to 6, wherein the polymerase comprises T4 polymerase or terminator DNA polymerase.
8. The method of any one of claims 1 to 7, comprising forming 5' overhangs on said plurality of said building blocks.
9. 9. The method of any one of claims 1 to 8, comprising removing overhangs from the 3' and / or 5' ends of the plurality of building blocks prior to the addition of the ddNTP molecules.
10. 10. The method of claim 9, wherein the overhangs are removed using a nuclease.
11. 11. The method of claim 10, wherein the nuclease comprises one or more of mung bean nuclease, nuclease P1, exonuclease I, exonuclease III, micrococcal nuclease, S1 nuclease, or a polymerase containing exonuclease activity.
12. 12. The method of any one of claims 1 to 11, wherein the addition of the ddNTP molecule comprises adding ddATP to the 3' end of the building block using a polymerase.
13. 13. The method of claim 12, wherein the polymerase comprises taq polymerase or the Klenow fragment.
14. 1. A method for writing information into a nucleic acid molecule, comprising: determining a string of symbols representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, each identifier of the plurality of identifiers corresponding to a respective symbol in the string of symbols, wherein generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, each individual building block of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end and a 5' end; (b) adding a first volume containing a polymerase and an amount of acyclonucleotides to a reaction volume containing the plurality of components; (c) incubating the reaction volume to add acyclonucleotides to the 3' ends of at least some of the building blocks; (d) adding a reagent to the reaction volume to chemically link two or more components of the plurality of components, thereby generating the plurality of identifiers and the plurality of fragments; (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification; wherein PCR amplification of any oligonucleotide containing an acyclonucleotide is inhibited.
15. 1. A method for writing information into a nucleic acid molecule, comprising: determining a string of symbols representing the information; and generating a plurality of oligonucleotides comprising a plurality of identifiers, each identifier of the plurality of identifiers corresponding to a respective symbol in the string of symbols, wherein generating the plurality of oligonucleotides comprises: (a) assembling a plurality of building blocks, each individual building block of the plurality of oligonucleotides being a nucleic acid molecule having a nucleic acid sequence, a 3' end and a 5' end; (b) adding a reagent to the reaction volume to chemically link two or more components of the plurality of components, thereby generating the plurality of identifiers and the plurality of fragments; (c) adding a first volume containing a polymerase and a plurality of 3'-DNA flaps to a reaction volume containing the plurality of components; (d) incubating the reaction volume to add 3'-DNA flaps to the 3' ends of at least some of the building blocks and fragments; (e) adding PCR primers to the reaction volume and subsequently performing PCR amplification; wherein PCR amplification of any oligonucleotide containing a 3'-DNA flap is inhibited.
16. 1. A method for writing information into a nucleic acid molecule, comprising: (a) determining a string of symbols representing the information; (b) assembling a plurality of building blocks, each individual building block of the plurality of building blocks being a nucleic acid molecule having a nucleic acid sequence, a 3' end and a 5' end; (c) chemically linking two or more components of the plurality of components, thereby generating a plurality of identifiers, each identifier of the plurality of identifiers comprising two or more components, each identifier having a first end and a second end, each component disposed at the first end of the identifier being a first edge component, and each component disposed at the second end of the identifier being a second edge component; (d) chemically modifying a terminus of each of the first edge element, the second edge element, or both, such that the first edge element and / or the second edge element are protected from exonuclease activity; wherein each identifier in the plurality of identifiers corresponds to a respective symbol in the string.
17. 17. The method of claim 16, wherein modifying the termini comprises constructing a hairpin loop at one or both of the termini, thereby protecting the one or both termini from nuclease activity.
18. 18. The method of claim 16 or 17, wherein modifying the termini comprises adding a protelomerase recognition sequence to one or both of the termini to covalently close the one or both termini, thereby protecting the one or both termini from nuclease activity.
19. 19. The method of any one of claims 16 to 18, wherein modifying the termini comprises implementing a phosphorothioate linkage at one or both of the termini to replace a non-bridging oxygen in the phosphate backbone of the terminal oligonucleotide with a sulfur atom, thereby protecting the one or both termini from nuclease activity.
20. 20. The method of claim 19, wherein modifying the termini comprises implementing multiple phosphorothioate linkages at one or both of the termini.
21. 21. The method of claim 20, wherein modifying the termini comprises implementing at least three phosphorothioate linkages at one or both of the termini.
22. 22. The method of any one of claims 16 to 21, wherein modifying the termini comprises performing an inverse dT modification at one or both of the termini to create a 3'-3' linkage, thereby protecting the one or both termini from nuclease activity.
23. 23. The method of any one of claims 16 to 22, wherein modifying the termini comprises performing a sugar residue modification at one or both of the termini, thereby protecting the one or both termini from nuclease activity.
24. 24. The method of any one of claims 16 to 23, wherein modifying the termini comprises circularizing the identifier and ligating the termini, thereby protecting the one or both termini from nuclease activity.
25. The method of any one of claims 16 to 24, wherein modifying the termini comprises modifying the termini with restriction enzyme sites.
26. 26. The method of any one of claims 1 to 25, comprising targeting an overhang that is uniquely present on the incompletely assembled identifier by utilizing a helicase to separate two hybridized nucleic acid strands, thereby providing access to a nuclease that acts on single-stranded DNA.
27. The method of any one of claims 1 to 26, comprising treating the component with an exonuclease.
28. 28. The method of any one of claims 1 to 27, comprising selectively capturing or amplifying an identifier library comprising at least a subset of said plurality of identifiers.
29. A method according to any preceding claim, wherein each symbol in the string is one of one or more possible symbol values.
30. 30. The method of claim 29, wherein each symbol in the string is one of two possible symbol values.
31. 31. The method of claim 29 or 30, wherein one symbol value at each position of the string can be represented by the absence of a distinct identifier in the identifier library.
32. 31. The method of claim 30, wherein the two possible symbol values are bit values of 0 and 1, and the individual symbols in the symbol string having the bit value of 0 can be represented by the absence of a distinct identifier in the identifier library and the individual symbols in the symbol string having the bit value of 1 can be represented by the presence of the distinct identifier in the identifier library, or vice versa.
33. 33. The method of any one of claims 1 to 32, comprising chemically linking the two or more components from two or more layers, each layer of the two or more layers comprising a distinct set of components.
34. 34. The method of claim 33, wherein the individual identifiers from the identifier library include one component from each of the two or more layers.
35. 35. The method of claim 34, wherein the two or more components are assembled in a fixed order.
36. 35. The method of claim 34, wherein the two or more components are assembled in any order.
37. 35. The method of claim 34, wherein the two or more components are assembled with one or more partition components disposed between two components from different ones of the two or more layers.
38. 34. The method of claim 33, wherein the individual identifiers include one component from each layer of the subset of two or more layers.
39. 34. The method of claim 33, wherein the individual identifier includes at least one component from each of the two or more layers.
40. 40. The method of any one of claims 1 to 39, comprising using an endonuclease to generate at least one sticky end of each component of the plurality of components.
41. 41. The method of claim 40, wherein the at least one sticky end is at the 5' end of the individual building block.
42. 41. The method of claim 40, wherein the at least one sticky end is at the 3' end of the individual building block.
43. 43. The method of any one of claims 40 to 42, comprising generating two sticky ends of the individual components.
44. 44. The method of any one of claims 40 to 43, wherein the at least one sticky end is at least one nucleotide in length.
45. 45. The method of any one of claims 40 to 44, wherein the at least one sticky end is 6 nucleotides in length.
46. 46. The method of any one of claims 1 to 45, wherein the plurality of identifiers comprises a nucleic acid sequence that stores metadata about the information or that conceals the information.
47. The method of any one of claims 1 to 46, wherein two or more identifier libraries are combined and each identifier library of the two or more identifier libraries is tagged with a separate barcode.
48. The method of any one of claims 28 to 47, wherein each individual identifier in the identifier library comprises a distinct barcode.
49. A method according to any preceding claim, wherein the plurality of identifiers or the plurality of components comprising the identifiers are selected to facilitate read, write, access, copy and delete operations.
50. 50. The method of any one of claims 1 to 49, wherein chemically linking comprises ligating two or more components of the plurality of components together using a reagent comprising a ligase.
51. 51. The method of claim 50, wherein the ligase is T4 ligase, T7 ligase, T3 ligase, or E. coli ligase.
52. 52. The method of claim 50 or 51, wherein the reagent further comprises an additive.
53. 53. The method of claim 52, wherein the additive increases the efficiency of the ligase.
54. 54. The method of claim 52 or 53, wherein the additive comprises polyethylene glycol (PEG).
55. 55. The method of claim 54, wherein the PEG is PEG 400, PEG 6000, PEG 8000, or any combination thereof.
56. The method according to any one of claims 50 to 55, wherein the reaction time for the ligation is at least 1 minute.
57. 56. The method of any one of claims 50 to 55, wherein the ligation is at 30°C or higher.
58. 58. The method of any one of claims 50 to 57, further comprising inactivating the ligase using a buffer containing EDTA or guanidine thiocyanate.
59. 59. The method of any one of claims 50 to 58, wherein the final concentration of the ligase is at least about 5 CEU / μL.
60. 60. The method of any one of claims 50 to 59, wherein the reagent further comprises a glycerol molecule.
61. 61. The method of any one of claims 1 to 60, wherein said chemically linking comprises using overlap extension polymerase chain reaction (PCR).
62. 62. The method of any one of claims 1 to 61, wherein the individual components are deoxyribonucleic acid (DNA) or ribonucleic acid.
63. 63. The method of any one of claims 1 to 62, wherein the individual components are rehydrated.
64. 64. The method of any one of claims 1 to 63, wherein the individual components are rehydrated from dehydrated components.
65. 65. The method of any one of claims 28 to 64, further comprising dehydrating the identifier library by dehydrating each individual identifier of at least the subset of the plurality of identifiers.
66. 66. The method of any one of claims 28 to 65, wherein each individual identifier of at least said subset of said plurality of identifiers is dehydrated.
67. 67. The method of claim 65 or 66, further comprising rehydrating each individual identifier of at least said subset of said plurality of identifiers.
68. 68. The method of any one of claims 1 to 67, further comprising adding a preservative additive to the identifier library to prevent degradation of the identifiers.
69. 69. The method of any one of claims 1 to 68, wherein the plurality of identifiers are copied by PCR.
70. 70. The method of claim 69, wherein the PCR has at least 10 cycles.
71. 70. The method of claim 69, wherein the plurality of identifiers are amplified by PCR to a concentration of 10 nanograms per microliter.
72. 72. The method of any one of claims 69 to 71, wherein the PCR is emulsion PCR.
73. 73. The method of any one of claims 1 to 72, wherein the plurality of identifiers are copied by linear amplification.
74. 74. The method of any one of claims 69 to 73, wherein after the PCR, linear amplification is used to create more copies of the plurality of identifiers.
75. 75. The method of any one of claims 1 to 74, wherein the subset of the plurality of identifiers is accessed in one or more PCR reactions.
76. 76. The method of any one of claims 1 to 75, wherein a subset of the plurality of identifiers is accessed with one or more affinity-tagged probes.
77. 77. A method according to claim 75 or 76, wherein the identifiers of the subset of the plurality of identifiers have a set of components in common.
78. 78. The method of any one of claims 1 to 77, wherein the identifier is purified by gel electrophoresis.
79. 79. The method of any one of claims 1 to 78, wherein the identifier is purified by an affinity-tagged probe.
80. 80. The method of any one of claims 1 to 79, wherein the identifier is amplified using PCR.
81. 81. The method of any one of claims 1 to 80, wherein the identifier is designed to avoid thymine-thymine or cytosine-cytosine dinucleotides.