Preservation of sequence-controlled polymers
Patent Information
- Application Number
- JP2023575828
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-09
- Filing Date
- 2022-06-09
- Publication Date
- 2026-01-08
AI Technical Summary
Current methods for storing biomolecules such as DNA, RNA, and proteins require cryogenic temperatures, which are energy-intensive and costly, and lack efficient room temperature storage solutions that maintain sample integrity and allow for rapid recovery and identification of large numbers of samples.
Encapsulation of biomolecules in millimeter to nanoscale capsules using molecular barcodes, allowing for ultra-high density storage at room temperature and enabling rapid recovery through orthogonal molecular barcodes and microfluidic sorting.
This method enables efficient, low-energy storage and recovery of biomolecules at room temperature, reducing the storage footprint and maintaining sample integrity for extended periods, facilitating rapid and accurate retrieval of large sample collections.
Smart Images

Figure 00000069_0000 
Figure 00000069_0001 
Figure 00000069_0002
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 208,973, filed June 9, 2021, which is hereby incorporated by reference in its entirety.
[0002] Statement Regarding Federally Sponsored Research This invention is made in part through grants received from the Office of Naval Research (ONR) under grant numbers N00014-16-1-2506, N00014-12-1-0621, N00014-18-1-2290, N00014-17-1-2609, N00014-20-1-2084, and N00014-21-1-4013, from the National Science Foundation (NSF) under grant numbers CCF1564025, 1729397, CHE1839155, OAC1940231, and CCF1956054, from the Department of Energy (DOE) under grant number DE-SC0019998, from the Army Research Office under grant number 0019999, from the Department of Energy (DOE) ... This invention was made with government support under Grant No. W911NF-13-D-0001 awarded by the Air Force Research Office (ARO), and Grant No. FA8750-19-2-1000 awarded by the Air Force Research Laboratory (AFRL). The government has certain rights in this invention.
[0003] Reference to sequence listing The Sequence Listing as a text file entitled "MIT 23164_ST25", created on May 26, 2022, and having a size of 1,614 bytes, submitted on June 9, 2022, is incorporated by reference herein pursuant to 37 C.FR § 1.52(e)(5).
[0004] FIELD OF THEINVENTION The present invention discloses a method for encapsulation of biomolecules that allows ultra-high density storage at room temperature using milli- to nano-scale capsules that can be uniquely identified using molecular barcodes. [Background technology]
[0005] 2. Background of the Invention The central dogma of biology is the progression from DNA to RNA and then finally to proteins. These biomolecules play key roles in sustaining life: DNA codes for the information for protein synthesis, and RNA executes the instructions coded in DNA. Proteins carry out most biological processes. The explosion and advancement of omics technologies has driven the demand to understand an individual's health and predisposition to disease through the collection, storage, and analysis of DNA, RNA, and proteins. Omics technologies that analyze nucleic acids, namely genomics and transcriptomics, are now scientifically advanced and commercialized on a large scale.
[0006] Large-scale storage of nucleic acid samples is crucial for basic, translational, and clinical research, synthetic biology foundries, and biodiversity conservation efforts [Ivanova and Kuzmina. Mol Ecol Resour 13, 890-898, doi:10.1111 / 1755-0998.12134 (2013);Fabre, et al. European Journal of Human Genetics 22, 379-385, doi:10.1038 / ejhg.2013.145 (2014)]. Nucleic acid storage requires robust procedures to maintain sample quality, integrity, and functionality. Currently, the required storage temperature for nucleic acids is between 4°C and -196°C [Fabre, et al. European Journal of Human Genetics 22, 379-385, doi:10.1038 / ejhg.2013.145 (2014);Muller, et al. Biopreserv Biobank 14, 89-98, doi:10.1089 / bio.2015.0022 (2016);Miernyk, et al. Biopreserv Biobank 15, 529-534, doi:10.1089 / bio.2017.0040 (2017)], at which there is negligible degradation. However, maintaining such low temperatures for long periods of time requires a great deal of energy.Large-scale cryopreservation of nucleic acid material also requires extensive robotics for access, strict cold-chain management logistics [Muller, et al. Biopreserv Biobank 14, 89-98, doi:10.1089 / bio.2015.0022 (2016);Clermont, et al. Biopreserv Biobank 12, 176-183, doi:10.1089 / bio.2013.0082 (2014);Wan, et al. Curr Issues Mol Biol 12, 135-142 (2010)] and redundant copies of samples stored in mirror storage facilities to mitigate the risk of sample loss [Muller, et al. Biopreserv Biobank 14, 89-98, doi:10.1089 / bio.2015.0022 (2016)]. Finally, cryopreservation of nucleic acids in remote or low-resource areas involves costly procedures and complex cold chain logistics to maintain the integrity and quality of isolated samples during transport [Clermont, et al. Biopreserv Biobank 12, 176-183, doi:10.1089 / bio.2013.0082 (2014)]. The transition from cryogenic to room temperature storage will reduce energy use by 40 million kilowatt-hours, eliminating 18,000 metric tons of carbon dioxide emissions per year and resulting in cost savings of $16 million over 10 years [Palmer. Nat Med 16, 1056-1057, doi:10.1038 / nm1010-1056b (2010)], and reduce the space required for cryogenic storage by 70% [Lou, et al. Clin Biochem 47, 267-273, doi:10.1016 / j.clinbiochem.2013.12.011 (2014)]. Costs and workflow complexity associated with sample processing are also reduced [Lou, et al. Clin Biochem 47, 267-273, doi:10.1016 / j.clinbiochem.2013 .12.011 (2014)].Room temperature storage of nucleic acid samples is achieved by the addition of stabilizers such as DNAstable® and RNAstable® from Biomatrica, or by using vacuum canisters such as DNAshells® and RNAshells® from Imagene. While these room temperature storage solutions can ensure nucleic acid stability for a year or longer, the space to store samples and the supporting infrastructure such as extensive robotic platforms for access and humidity control required remain significant cost considerations [Muller, et al. Biopreserv Biobank 14, 89-98, doi:10.1089 / bio.2015.0022 (2016);Lou, et al. Clin Biochem 47, 267-273, doi:10.1016 / j.clinbiochem.2013.12.011 (2014)].
[0007] Silica particles [Grass, et al. Angewandte Chemie International Edition 54, 2552-2555 (2015);Puddu, et al. Advanced healthcare materials 4, 1332-1338 (2015)], alginates [Gombotz and Wee. Advanced drug delivery reviews 31, 267-285 (1998);Machado, et al. Langmuir 29, 15926-15935 (2013)], and synthetic polymers [Gill and Ballesteros. Trends in biotechnology 18, 282-296 (2000);Zelikin, et al. ACS nano 1, 63-69 (2007)] have been used to store biomolecules at room temperature, but the ability to uniquely identify these storage materials and pool them together to realize alternative room temperature storage and retrieval platforms for biomolecules has not yet been demonstrated. Programs and functions for DNA-based data storage are described in WO2021231493A1.
[0008] There is a need for scalable biomolecule preservation that requires little energy to maintain sample integrity for periods of 10 years or longer.
[0009] Additionally, the footprint required to store biomolecular samples must be significantly reduced, allowing the rapid recovery of thousands to millions of samples.
[0010] It is therefore an object of the present invention to provide a method for preserving and recovering biomolecules collected from any source.
[0011] It is also an object of the present invention to provide methods for encapsulating biomolecules of various lengths and sizes using different chemical and biochemical preparations and different fluidic approaches.
[0012] It is also an object of the present invention to provide a method for labeling encapsulated biomolecules using a different fluidic approach.
[0013] It is also an object of this disclosure to provide a method for selecting the barcode for each particle in such a way that the encapsulated biomolecule allows for the recovery of a collection of particles related by various characteristics including, but not limited to, sample type, source, and collection date / time. Barcodes can be selected from an existing pool of sequences designed for optimal properties such as binding strength and orthogonality.
[0014] It is also an object of this disclosure to provide novel methods for designing barcode sequences that enable similarity-based retrieval by allowing probes to bind to multiple distinct barcodes of similar sequence, where these barcodes label particles whose contained biomolecules are similar under some metric of interest.
[0015] It is a further object of the presently disclosed invention to provide chemical and biochemical strategies to increase barcode sorting throughput using chemical and biochemical approaches.
[0016] It is also an object of the present invention to provide biopolymer storage structures that allow for Boolean logic calculations, which may include peptides, nucleic acids, or other sequence-controlled polymers.
[0017] It is also an object of the present invention to provide any nucleic acid origami nanostructures, as well as other nucleic acids and biopolymers, as storage blocks that can be read out using either sequencing or mass spectrometry or other analytical chemistry approaches.
[0018] It is a further object to provide nucleic acid storage blocks capable of forming stable and reconfigurable superstructures for storage block structural association and position-based storage, as well as parallel computational processing.
[0019] It is also an object to provide a nucleic acid storage object capable of accelerating decomposition in response to a specific external stimulus. [Prior art documents] [Patent documents]
[0020] [Patent Document 1] International Publication No. 2021 / 231493 [Non-patent literature]
[0021] [Non-Patent Document 1] Ivanova and Kuzmina. Mol Ecol Resour 13, 890-898, doi:10.1111 / 1755-0998.12134 (2013) [Non-Patent Document 2] Fabre, et al. European Journal of Human Genetics 22, 379-385, doi:10.1038 / ejhg.2013.145 (2014) [Non-licensed document 3] Muller, et al. Biopreserv Biobank 14, 89-98, doi:10.1089 / bio.2015.0022 (2016)
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed literature 9
Non-licensed literature 10
Non-licensed Document 11
[0022] Summary of the Invention Nucleic acids purified from any source are encapsulated in synthetic packets composed of organic or inorganic polymer networks. Encapsulation can be performed using automated liquid handling to mix the biomolecule of interest with the encapsulation reagent, or millifluidic and microfluidic approaches to confine the biomolecule and encapsulation reagent in an emulsion reaction vessel of millimeter to nanometer size. The encapsulated biomolecules are then labeled with a combination of orthogonal molecular barcodes identified from a pool of 240,000 that uniquely label and identify the contents of the sample [Xu, et al. Proceedings of the National Academy of Sciences 106, 2289-2294, doi:10.1073 / pnas.0812506106 (2009)]. The encapsulated biomolecules may also be labeled with non-orthogonal molecular barcodes that allow for similarity-based recovery, such that a collection of similar biomolecules can be simultaneously recovered because a single probe sequence can bind to any one of multiple distinct barcodes of similar sequences. Molecular barcodes may be constructed with a non-phosphate backbone to improve the stability of the strand against nucleases. The process of barcoding can be performed similarly using millifluidic or microfluidic approaches. Upon encapsulation and barcoding, all samples can be collected and pooled in a single vessel. Samples are selected from the pool using complementary probes that may contain optical, chemical, or biochemical tags that can be used as markers for downstream optical or mechanical sorting using millifluidic or microfluidic strategies. Chemical and biochemical reactions can be performed on the barcodes to improve sorting speed, sorting accuracy, and detection limits of certain sorting approaches.
[0023] Disclosed are compositions and methods for sequence-controlled storage of objects. The sequence-controlled storage objects of the present disclosure include (a) one or more different sequence-controlled polymers, and (b) a plurality of different feature tags. In some forms, the feature tags are present on the surface of the sequence-controlled storage object. In some forms, each different feature tag corresponds to a single feature that is attributed to one or more of the different sequence-controlled polymers. In some forms, the single feature that each different feature tag corresponds to is a feature that is attributed to one or more different ones of the sequence-controlled polymers. In some forms, the plurality of different feature tags collectively correspond to a plurality of features that are collectively attributed to the plurality of different sequence-controlled polymers. In some forms, each different feature tag is hybridizable and distinguishable from all of the other different feature tags.
[0024] In some forms, each of the plurality of different feature tags is a member of a different set of feature tags, each set of feature tags corresponding to a set of related features. In some forms, at least one member of the set of feature tags is a similarity-coded feature tag. In some forms, the relative hybridizability of feature tags in a set is related to the similarity of the features to which the feature tags in the set correspond, with feature tags in the set corresponding to more similar features having a closer relative hybridizability than feature tags in the set corresponding to less similar features.
[0025] In some forms, the similarity coded feature tags of a set of feature tags are similarity coded by mapping the features to which the feature tags correspond into an n-dimensional hypercube based on the similarities of the features, where n is an integer less than or equal to the number of features to which the feature tags correspond and n is a factor of the number of features to which the feature tags correspond.
[0026] In some forms, before mapping the features to which the feature tags correspond, the dimensionality of the features to which the feature tags correspond is reduced, and the reduced-dimensionality features are mapped to a hypercube based on the similarity of the reduced-dimensionality features.
[0027] In some forms, similarity coded feature tags of a set of feature tags are similarity coded by (a) reducing the dimensionality of the features to which the feature tags correspond, and (b) mapping the reduced dimensionality features to an n-dimensional hypercube based on the similarity of the reduced dimensionality features, where n is an integer less than or equal to the number of features to which the feature tags correspond, and where n is a factor of the number of features to which the feature tags correspond.
[0028] In some embodiments, the number of hypercube edges between nodes to which any two of the mapped features are mapped is proportional to the similarity of the two features. In some embodiments, at least one member of the set of feature tags is defined for hybridization, and at least one member of the set of feature tags has the same number of nucleotides.
[0029] In some forms, in at least one of the sets of feature tags, (a) members of the set of feature tags have the same number of nucleotides, and (b) each feature tag in the set differs from one or two other feature tags in the set by 1 to x mismatched nucleotides, where the mismatched nucleotides are (i) at least 2 nucleotides from either end of the feature tag and (ii) separated by at least one matching nucleotide in the feature tag, where x is the number of different nucleotide positions in the feature tags that vary within the set.
[0030] In some forms, independently for at least one of the one or more sets of feature tags, each feature tag in the set is mismatched with every other feature tag in the set by 1 to w nucleotides, where w is an integer from 2 to (y-4)÷2, and y is the number of nucleotides of the feature tag in the set, where the formula (y-4)÷2 is rounded up.
[0031] In some forms, the sequence controlled storage object further comprises a plurality of different digit tags, the digit tags being present on a surface of the storage object, the digit tags encoding a number.
[0032] In some embodiments, the array-controlled storage object further comprises a plurality of different digit tags, the digit tags being present on a surface of the storage object, each of the plurality of different digit tags corresponding to a digit value of a different place of the multi-digit number, and the number of different digit tags included in the storage object being equal to the number of places of the multi-digit number. In some embodiments, each of the plurality of different digit tags is a member of a different set of digit tags, each set of digit tags corresponding to a different place of the multi-digit number. In some embodiments, each set of digit tags has a digit tag corresponding to each possible digit value of the place of the multi-digit number to which the set of digit tags corresponds. In some embodiments, each of the different digit tags is hybridizable and distinguishable from all of the other different digit tags in all of the sets of digit tags, and each of the different digit tags is hybridizable and distinguishable from all of the different feature tags.
[0033] In some embodiments, the sequence-controlled storage object includes (a) one or more different sequence-controlled polymers, and (b) a plurality of different digit tags. In some embodiments, the digit tags are present on the surface of the storage object. In some embodiments, each of the plurality of different digit tags corresponds to a digit value of a different place of the multi-digit number, and the number of different digit tags included in the storage object is equal to the number of places of the multi-digit number. In some embodiments, each of the plurality of different digit tags is a member of a different set of digit tags, and each set of digit tags corresponds to a different place of the multi-digit number. In some embodiments, each set of digit tags has a digit tag corresponding to each possible digit value of the place of the multi-digit number to which the set of digit tags corresponds. In some embodiments, each of the different digit tags is hybridizable and distinguishable from all of the other different digit tags in all of the sets of digit tags.
[0034] In some forms, the number of the multiple digits corresponds to a feature attributable to one or more of the different sequence-controlled polymers. In some forms, the feature attributable to one or more of the different sequence-controlled polymers is a member of a set of related features, each of which has or can be associated with a different numerical value, the different numerical value corresponding to the level or intensity of the given feature relative to other features in the set of related features, and the number of the multiple digits is equal to, proportional to, or the same as a given number of digits in the numerical value of the feature attributable to one or more of the different sequence-controlled polymers.
[0035] In some forms, the difference in the numerical values that members of a set of related features have or can relate to is proportional to the similarity of the features in the set of related features. In some forms, the number of multiple digits is arbitrarily assigned to the feature that originates from one or more of the different sequence-controlled polymers to which the number of multiple digits corresponds. In some forms, the number of multiple digits is the same as the number of digits of the numerical value of the feature that originates from one or more of the different sequence-controlled polymers, starting from the most significant digit of the numerical value.
[0036] In some forms, each set of digit tags has the same number of members as the mathematical base in which the multi-digit number is expressed. In some forms, the sequence-controlled storage object further comprises one or more encapsulating agents, which coat or encapsulate the sequence-controlled polymer, and the encapsulating agent can be reversibly removed by chemical or mechanical treatment.
[0037] In some forms, the feature tag is included in one or more of the encapsulating agents. In some forms, the one or more encapsulating agents are selected from natural polymers and synthetic polymers, or combinations thereof. In some forms, the one or more encapsulating agents are selected from proteins, polysaccharides, lipids, nucleic acids, inorganic coordination polymers, metal-organic frameworks, covalent organic frameworks, inorganic coordination cages, covalent organic coordination cages, elastomers, thermoplasts, synthetic fibers, or any derivatives thereof.
[0038] In some forms, at least one of the sequence-controlled polymers is a single-stranded nucleic acid, the nucleic acid is folded into a three-dimensional polyhedral nanostructure comprising two nucleic acid helices joined by either antiparallel or parallel crossovers spanning each edge of the structure, the three-dimensional polyhedral structure is formed from single-stranded nucleic acid staple sequences hybridized to the single-stranded nucleic acid comprising the bitstream data, the single-stranded nucleic acid comprising the bitstream data is routed through Euler cycles of a network defined by the vertices and lines of the polyhedral structure, the nanostructure comprises at least one edge that comprises a double-stranded crossover or a single-stranded crossover, the location of the double-stranded crossover is determined by a spanning tree of the polyhedral structure, the staple sequences hybridize to the vertices, edges and double-stranded crossovers of the single-stranded nucleic acid comprising the bitstream data to define the shape of the nanostructure, and one or more of the staple sequences comprise one or more feature tag sequences.
[0039] In some forms, the staple strand comprises between 14 and 1,000 nucleotides, inclusive. In some forms, the single stranded nucleic acid comprises between approximately 100 and 1,000,000 nucleotides, inclusive. In some forms, the one or more staple strands comprise one or more feature tag sequences at the 5' end, the 3' end, or both the 5' and 3' ends. In some forms, the one or more feature tag sequences comprise one or more overhanging oligonucleotide sequences. In some forms, the one or more feature tag sequences comprise oligonucleotide sequences complementary to one or more feature tag sequences attached to different sequence control storage objects. In some forms, the sequence control storage object further comprises one or more additional sequence control storage objects attached thereto.
[0040] Also provided is a method for preserving a desired sequence-controlled polymer as a sequence-controlled storage object, comprising the steps of: (a) (i) one or more different sequence-controlled polymers, and (ii) a plurality of distinct feature tags; and (iii) optionally, one or more mounting media; Assembling a sequence control storage object from The feature tag is present on a surface of the sequence-controlled storage object; each distinct feature tag corresponds to a single feature attributable to one or more of the distinct sequence-controlled polymers; the single feature to which each distinct feature tag corresponds is a feature attributable to one or more distinct ones of the sequence-controlled polymers; the plurality of distinct feature tags collectively correspond to a plurality of features collectively attributable to the plurality of distinct sequence-controlled polymers; each distinct feature tag being hybridizable and distinct from all of the other distinct feature tags; (b) storing the sequence control storage object; Also disclosed is a method comprising:
[0041] In some forms, the method further comprises: (c) recovering the desired sequence-controlled polymer. In some embodiments, recovering the desired sequence-controlled polymer in step (c) comprises isolating one or more sequence-controlled stored objects from the pool of sequence-controlled stored objects. In some embodiments, the selection is determined by the sequence of one or more feature tags on the sequence-controlled stored object, the shape of the sequence-controlled stored object, affinity to a functional group bound to the sequence-controlled stored object, or a combination thereof.
[0042] In some forms, the method further comprises modifying the isolated sequence-controlled stored object by adding one or more different feature tags. In some forms, adding one or more different feature tags comprises refolding or reorganizing the sequence-controlled stored object with one or more oligonucleotides that include the different feature tags. In some forms, the one or more sequence-controlled stored objects are isolated from the pool of sequence-controlled stored objects using Boolean logic. In some forms, the one or more sequence-controlled stored objects are removed from the pool of objects using Boolean NOT logic.
[0043] In some forms, the method further comprises: (d) accessing the desired sequence-controlled polymer. In some forms, storing the sequence-controlled storage object in step (b) further comprises one or more of dehydrating, lyophilizing, or freezing the sequence-controlled storage object. In some forms, storing the sequence-controlled storage object in step (b) further comprises one or more of rehydrating or thawing the sequence-controlled storage object for processing.
[0044] In some embodiments, storing the sequence-controlled storage object comprises storing in a matrix selected from cellulose, paper, microfluidics, bulk 3D solution, on a surface using electric forces, on a surface using magnetic forces, encapsulated in inorganic or organic salts, and combinations thereof. In some embodiments, storing the sequence-controlled storage object in step (b) further comprises digitally processing the droplet containing the sequence-controlled storage object.
[0045] Also provided is a method for automating assembly of sequence-controlled storage objects, comprising using a device having a flow, the device comprising: (a) a means for adding a component of a sequence-controlled storage object to a flow; (b) a means for mixing the components, comprising: a means for mixing operably connected to the means for adding to the flow; (c) a means for annealing the components to form an assembled sequence-controlled storage object, a means for annealing operably connected to the means for mixing; and (d) a means for purifying the assembled sequence control archive object, the means for purifying is operably connected to the means for annealing, A method is also disclosed, including:
[0046] In some forms, the method further comprises: (e) means for introducing a mounting medium that preserves the sequence control entities; (f) a means for introducing a plurality of feature tags resulting from a sequence-controlled polymer; (g) means for selecting an encapsulated sequence control object from the object pool, the selecting means being operable using Boolean logic; and (h) means for removing the mounting medium to retrieve the sequence control archive object; Further includes:
[0047] In some forms, the storage block is formed by encapsulating one or more sequence-controlled polymers in one or more encapsulating agents. Exemplary encapsulating agents include proteins, lipids, sugars, polysaccharides, nucleic acids, and any derivatives thereof, as well as polystyrene, or hydrogels and synthetic polymers including silica, glass, and paramagnetic materials. These encapsulated biopolymers form discrete storage units that allow for controlled separation of the blocks. In some embodiments, the storage block comprises sequence-controlled biopolymers folded into specific nanostructure forms, such as nucleic acid nanostructures. In some forms, the storage block comprises one or more discrete units within two or more sequence-controlled biopolymers. For example, in some forms, the nucleic acid sequence is folded into a nucleic acid nanostructure that includes or is associated with one or more polypeptides or other sequence-controlled biopolymers. In some forms, the storage block comprises a nucleic acid sequence encapsulated with one or more polypeptides or other sequence-controlled biopolymers.
[0048] In some forms, the stored object may include a nucleic acid "scaffold" sequence folded into a nucleic acid nanostructure. The nucleic acid scaffold sequence may be any length, e.g., 100 to 1,000,000 nucleotides. Typically, the nucleic acid scaffold sequence is 300 to 500,000 nucleotides long, e.g., about 300 nucleotides to about 51,000 nucleotides long, inclusive. In some forms, the method provides a sequence of short single-stranded oligonucleotide staple strands, approximately 14 to 1,000 nucleotides long, e.g., about 14 to 600 nucleotides long, whereby the single-stranded nucleic acid scaffold sequence is folded into a nucleic acid nanostructure (e.g., a polyhedron or DNA brick) having any user-defined geometry. Typically, assembly of a nucleic acid nanostructure includes scaffold routing, staple strand selection, input of geometry and scaffold sequence, oligonucleotide synthesis, and folding ("nanostructuring"), as performed in either scaffolded or non-scaffolded nucleic acid origami. The staple strands, as part of the nanostructure formation, have nicks where the 5' end of the staple matches the 3' end of itself or another staple. These nicks can then have single-stranded overhanging nucleic acid sequences ("tags") of any sequence.
[0049] The method also provides nucleic acid encapsulation for storage, in which nucleic acid is encapsulated in a layer of natural or synthetic material.Any form of nucleic acid can be encapsulated, including linear, single-stranded, base-paired double-stranded, or scaffolded nucleic acid.Exemplary encapsulants include proteins, lipids, sugars, polysaccharides, nucleic acids, and any derivatives thereof, as well as polystyrene, or hydrogels and synthetic polymers, including silica, glass, and paramagnetic materials.These encapsulated nucleic acids form discrete storage units that allow controlled separation of blocks.
[0050] Thus, methods are provided for creating sequence-controlled polymeric storage objects ("SSOs"). In some forms, the storage objects are nucleic acid encapsulated units that represent nucleic acid nanostructures or nucleic acid storage objects ("NSOs"). The SSO storage "blocks" can be of variable size, reconfigurable based on external cues including buffer changes, enzymes, nucleic acid "keys", temperature, electrical signals or light, and present identity tags for physical identification and retrieval or selection. The method includes assembling SSOs together into larger superconserved blocks that spatially associate SSOs for separation and associative storage applications. The method also includes functionalizing staple chains to have tags that can be used for SSO capture, rapid purification, and computation. The method provides sequence-controlled polymers as physical, structured units of any shape and size that can be used to form supramolecular storage blocks. Nanostructuring or encapsulating storage blocks allows for natural extension of spatial separation of objects based on input signals, and associated sequence-controlled polymers can be associated into superblock storage. The address space is multiplied by the number of tags used, so 4 (k*n) (n is the number of nucleotides in the address per tag and k is the number of tags).
[0051] Selection and access of sequence-controlled polymers can be achieved by SSO capture mediated by specific and orthogonal interactions of single-stranded overhang tags. Overhang tags available in primer libraries known in the art can be included (Xu, et al., PNAS., V.106, (7) pp. 2289-2294 (2009)).
[0052] The tags from the functionalized staple chains can be modified with new addressing systems, and the sequence-controlled polymer can be refolded with a new set of tagged staples, and / or overhang sequences. This allows for a dynamic addressing system without the need to resynthesize all sequence-controlled polymer sequences. Sequence-controlled polymers encapsulated in silica or paramagnetic or sequence-controlled polymer-based nanoparticles can be reused as well, displaying tags covalently or non-covalently attached by standard chemical reactions that specify the number and stoichiometry of specific overhang sequences. Methods are also provided for accessing sequence-controlled polymers, or subsets of sequence-controlled polymers from a pool of discrete SSOs. In some forms, the sequence-controlled polymers are accessed to allow selection by Boolean logic. For example, Boolean NOT logic can be used to remove sequence-controlled polymers from the sequence-controlled polymer pool. In some forms, the removed sequence-controlled polymers are replaced, for example, with a new set of structures and addresses. In other forms, the removed sequence-controlled polymers are omitted from future calculations / selections.
[0053] In some forms, the method also includes long-term storage of the SSO, if necessary. For example, the method may include storage of scaffolded or encapsulated nucleic acids for up to 1 year, up to 10 years, up to 20 years, 30 years, or more than 30 years. Typically, the method does not include any steps or processes that are detrimental to the stability and long-term storage of the SSO. For example, only selected outputs are processed by PCR or sequencing. No new buffers and biological materials that may devalue the data are required. In some forms, the DNA is stored in a dry state to maximize its lifespan. If the DNA is stored in a dry state, suitable mechanisms and systems can be used to separate, orderly store, and rewet the dried SSO, such as lyophilization and / or freezing of the NSO. In some forms, paper-based storage is used. Paper-based storage provides a compartment that can be selected and wetted for sequencing only when necessary for separation of multiple nucleic acid storage solutions or retrieval of the storage. In further forms, the system includes droplet-based digital microfluidics, for example, on an electromagnetically actuated surface or in solution. Droplet-based digital microfluidics provides a practical means of performing the wet biochemistry required for the selection and recovery steps. Thus, in some forms, the methods include the use of droplet-based digital microfluidics to perform the selection and recovery steps.
[0054] In some forms, the storage object is a scaffold-type nucleic acid nanostructure having a desired polygonal or polyhedral shape. Thus, in some forms, the method includes providing a nucleic acid sequence, creating a nucleic acid nanostructure or a nucleic acid encapsulation unit containing the sequence, and storing the nucleic acid nanostructure or the nucleic acid encapsulation unit containing the sequence.
[0055] In some forms, the method also optionally includes organizing the sequence-controlled polymer within the storage object, such as a nucleic acid nanostructure or a nucleic acid encapsulation unit. In some forms, the method also optionally includes accessing the sequence. In further forms, the method includes retrieving the sequence from the storage object.
[0056] In some forms, the nucleic acid storage object includes a scaffold single-stranded nucleic acid of any length that is folded around the entire structure. Theoretically, there is no limit to the size of the nucleic acid scaffold strand that is folded around the entire structure, but in practice, the single-stranded nucleic acid scaffold typically includes about 100 to 1,000,000 nucleotides. In some forms, the nanostructure also includes one or more staple strands that include one or more overhanging oligonucleotide sequences. The staple strands are custom designed to anneal to the scaffold strands to form any desired three-dimensional nanostructure that contains sequence-controlled polymers. In some forms, the one or more overhanging oligonucleotide sequences are feature tags. Exemplary feature tags include barcode sequences of approximately 4 to at least 30 nucleotides in length (Xu, et al., PNAS., V.106, (7) pp. 2289-2294 (2009)). In some forms, the nucleic acid nanostructure has a regular or irregular wireframe polyhedral geometry. Typically, the geometric shape provides accessibility of the internal storage blocks by nucleic acids and enzymes. Thus, in some forms, the shape of the structure allows for the selection, or retrieval, or reconstitution of storage blocks, for example, due to the porosity of the entire supramolecular storage structure. Thus, in certain forms, the desired target structure is one that provides diffusion of small molecules throughout it, for example, to provide access to other molecules, such as enzymes and / or nucleic acids. In other forms, the desired target structure prevents access to other molecules, such as enzymes and / or nucleic acids. In some forms, the SSO comprises hydrogels, polymers, glass, silica, or paramagnetic nanoparticles with specific overhanging nucleic acid sequences or other high affinity and specificity tags that provide programmable interactions between distinct storage blocks in the SSO. Thus, in some forms, the shape of the structure itself can be used as a means to select for different or similar functionality between SSOs.
[0057] Also provided are sequence-controlled biopolymer storage objects that include nucleic acids or other sequence-controlled biopolymers encapsulated in natural or synthetic materials. In some forms, any form of nucleic acid or other biopolymer can be encapsulated. For example, in some forms, linear, single-stranded, base-paired double-stranded, or scaffolded nucleic acids are encapsulated. Exemplary encapsulating agents include proteins, lipids, sugars, polysaccharides, nucleic acids, synthetic polymers, hydrogel polymers, silica, paramagnetic materials, and metals, and any derivatives thereof. These encapsulated nucleic acids or other biopolymers are associated with one or more overhanging nucleic acid sequences that are used to add addresses and / or purification tags. In some forms, multiple layers of encapsulation and overhanging nucleic acids are designed to further sort and tag the nature of the sequence-controlled polymer.
[0058] In some forms, the stored objects have the geometry of compact brick-like user-defined structures that can also be stacked end-to-end into long ribbons or extended 2D or 3D crystal-like arrays, either through non-specific or specific stacking interactions controlled using buffers or nucleic acid overhangs or other physical associations. In some forms, one or more staple strands include "overhang" oligonucleotide sequences that are complementary to one or more staple strands or bridging oligonucleotides from different stored objects, such as different nucleic acid nanostructures. In some forms, one or more stored objects are organized into superstructures by complementarity of nucleotide sequences from one or more staple strands or to bridging nucleotides. For example, in some forms, nucleic acid nanostructures are organized into superstructures by complementarity of nucleotide sequences from one or more staple strands or to bridging nucleotides. In some forms, stored objects, such as nucleic acid nanostructures or encapsulated nucleic acids, are organized into superstructures based on user-defined associations between the above-mentioned storage blocks. The superstructured sequence-controlled polymers can then be specifically manipulated by external signals such as pH, temperature, salt, nucleic acids, enzymes, light, and microfluidic manipulations, which can be droplet-based on-chip manipulations using electrowetting or traditional two-phase flow-based microfluidics. Applying mixing and splitting operations to selective pools of SSOs, as well as other beads or reagents containing cleavage enzymes such as Cas9 or restriction enzymes, offers the ability to perform both complex and selective computations and storage manipulations and retrievals. [Brief description of the drawings]
[0059] [Figure 1]1A-1C are schematic diagrams of the objects described herein, each showing the variety of different morphologies that can occur within an addressed pool of stored objects. FIG. 1A shows a size diversity spanning several orders of magnitude of nanostructured stored objects, each with comparable morphology (shown as closed cubes), but each containing 0.5 kb to 100 kb of data. FIG. 1B is a schematic diagram showing several stored objects, each with diversity in geometry, including open wireframe polyhedra and compact brick-like geometries. FIG. 1C is a schematic diagram showing several stored objects with diversity in the number and orientation of single-stranded nucleic acid overhangs that are externally presented at predefined geometric locations as one of several means of specifically associating multiple stored blocks into a larger assembly that can be stabilized, reconstituted, or accessed in response to an exogenous cue. [Diagram 2] Figure 2 is a schematic diagram showing a connective nanostructured data framework between pools of biopolymer storage objects. Generalized storage objects (A-D), shown as cubes, can be maintained as separate individual structures or assembled into larger superstructures of AB, AC, and D, respectively, by a first signaling event. The cuboid structures can be reassembled and resorted into differentially organized larger superstructures of ABC by a second signaling event, and resorted and altered geometry to expose interior blocks, respectively, by a third signaling event, which can also be actuated exogenously / extrinsically by microfluidic or other mixing mediated by fluidics or solid-state manipulation of subpools of SSO. [Figure 3A-C]3A-3D are schematic diagrams each showing steps of a method for assembling a pool of nucleic acid storage objects. The scaffold strands of the nucleic acid origami objects may be synthesized using, for example, TDT polymerase, solid-state DNA synthesis, bacterial synthesis, PCR-based enzymatic synthesis, or template-free DNA synthesis using another approach, are multiply addressed with metadata tag overhang sequences on the staple strands (Figure 3A), a scaffold strand containing two feature tags (*) at both ends of the scaffold and a staple strand in which the overhang tags are used to code multiple addresses (A and B) in the folding data are synthesized (Figure 3B), a single-stranded nucleic acid storage scaffold is combined with staple oligonucleotides and folded into a DNA origami object (Figure 3C), and the folded, multiply addressed DNA origami object is added to the storage pool (Figure 3D). [Figure 3D] Same as above. [Figure 4A-D]4A-4D are schematic diagrams showing the encapsulation of sequence-controlled biopolymers of any form into discrete SSOs for storing sequence-controlled polymers. FIG. 4A shows single- or double-stranded DNA, RNA, PNA, LNA, or other nucleic acids or peptides, or other sequence-controlled polymers (2), either with known / characterized errors in the sequence of the polymer, or with high fidelity sequences. Sequence-controlled polymers, such as nucleic acids, are "packaged," "encapsulated," "wrapped," or "housed" (4) in gel-based beads, protein virus packages (e.g., M13, adeno-associated virus, etc.), micelles, mineralized structures, siliconized structures, metals, paramagnetic materials, or a variety of polymers and polymer types (FIG. 4B), or polymers (6) designed to encapsulate or contain one nucleic acid for multiplexed polymer storage using two or more nucleic acid objects (2, and 3) (FIG. 4C). These packaged nucleic acids (10) bear molecular identifiers, such as single-stranded tag sequences, or optional purification tags (8) that allow for selection and / or recovery of specific sequence-controlled polymers using Boolean logic (Figure 4D). Figure 4E is a schematic illustrating the workflow of multiplexed attachment and encapsulation of sequence-controlled polymers (14) and modification of molecular cores (12) for downstream molecular logic operations and sequence-controlled polymer selection. Multiple sequence-controlled polymers are attached or adsorbed by the molecular core. The molecular core is then functionalized with addressing / specificity tags (16) for multiplexed computation and selection. [Figure 4E] Same as above. [Figure 5A-B]Figures 5A-5E are schematic diagrams of how nucleic acid storage objects (NSOs) can be superstructured to spatially separate and associate storage blocks. Blocks can be associated into cohesive storage block superstructures by direct complementarity of tag sequences (Figure 5A), or by "bridging" DNA oligonucleotides complementary to two tags (Figure 5B), or by kissing loops (Figure 5C), or other secondary structure interactions including base-pair end stacking (Figure 5D). The cohesive storage block superstructures can then be used for further selection, dissociation of individual NSOs, or resorting of sequence-controlled polymers into different superstructures (Figure 5E). [Figure 5C-E] Same as above. [Figure 6] Figure 6 is a schematic diagram providing a general overview of the method used to recover a specific NSO using a single-stranded DNA sequence complementary to the tag of a specified block. An exemplary method of NSO purification and selection is based on a fixed phase complementary to the tag on the NSO: a single NSO is captured from a pool of captured NSOs using a capture support with a sequence complementary to a(a'), and then the captured NSO with overhang sequence a is released from the support. The tetrahedron is representative of any NSO that contains an encapsulated nucleic acid. [Figure 7A-B] 7A-7D are schematic diagrams showing the selection of NSOs based on both sequence and overhang geometry. Figures 7A and 7B show a tetrahedral NSO with tags a and b displayed on specific edges, Figure 7C shows a complementary geometric DNA nanostructure on a capture support with tags a' and b' displayed in positions that capture NSOs with tags a and b in the appropriate geometric positions, and Figure 7D shows that an NSO with complementary tags a and b displayed on specific edges is selected by a larger DNA nanostructure. In this way, an NSO is specifically selected based not only on the sequence of the overhang tag, but also on the geometry of the NSO. The tetrahedron is representative of any storage object, including encapsulated nucleic acids or other biological or synthetic polymers. [Figure 7C-D] Same as above. [Figure 8]FIG. 8 is a schematic diagram showing the workflow of the method used to calculate the AND logical operation on the NSO pool. A pool of differently addressed NSOs is shown, and a support (●) with a tag complementary to a (a') is used to capture NSOs with overhang sequence a, resulting in a pool of NSOs with two different configurations of feature tags (a, b and a, c, respectively), the captured NSOs with overhang sequence a are then released from the support, and a support with a tag complementary to b (b') is used to capture NSOs with further overhang sequence b released from the support, and the captured NSOs with overhang sequence b are then released from the support. Overall, the two-step capture purification results in NSOs with overhang sequences a and b. The tetrahedrons are representative of any storage object, including encapsulated nucleic acids or other biological or synthetic polymers. [Figure 9] Figure 9 is a schematic diagram showing the workflow of the method used to calculate the OR logic operation on a pool of NSOs. A pool of differently addressed NSOs is shown, where NSOs containing an overhang of sequence a or an overhang of sequence e are captured using a capture support (●) with sequences complementary to a (a') and e (e'), NSOs containing neither are washed off the capture support, and then captured NSOs with an overhang of sequence a or an overhang of sequence e are released from the support. The tetrahedron is representative of any storage object, including encapsulated nucleic acids or other biological or synthetic polymers. [Figure 10] Figure 10 is a schematic diagram showing the workflow of the method used to calculate the NOT logical operation on a pool of NSOs. A pool of differently addressed NSOs is shown, where an NSO with an overhang tag sequence of a is captured on a capture support (●) using a capture sequence complementary to a (a'), and thus all unbound objects from this capture support are objects that do not contain the overhang of a, and therefore are not a. The tetrahedron is representative of any stored object, including encapsulated nucleic acids or other biological or synthetic polymers. [Figure 11] Figure 11 is a schematic diagram showing the workflow of the method used to read out selected NSOs. The desired NSO is first selected, the NSO is denatured, the released single-stranded nucleic acid scaffold is amplified by master primer sequences flanking the DNA sequence, and the scaffold strand is sequenced. Alternatively, mass spectrometry or other analytical procedures that do not require direct polymer-based sequencing may be used to decode the sequence-controlled polymer based on mass, charge, length, or other physicochemical properties. The tetrahedron is representative of any storage object, including encapsulated nucleic acids or other biological or synthetic polymers. [Figure 12] Figure 12 is a schematic diagram showing the workflow carried out within an exemplary microfluidic device that enables automated assembly and purification of NSOs. Scaffolds and staples are provided as input to a mixing chamber ("mixer"), followed by an annealing chamber (annealer), followed by a dialysis or filter chamber (exchanger) for purifying the NSOs from the staples. If sequence-controlled polymers, or other materials, are used for preservation encapsulation in particulate form, other upstream sorting devices can be interfaced, for example, avoiding the need for annealing. [Figure 13]FIG. 13 is a schematic diagram showing the workflow carried out within an exemplary microfluidic device that allows for rapid purification of nanostructured NSOs, including the ability to "daisy-chain" devices for complex logic gating. Multiple out-ports on the capture chambers allow for the execution of AND / OR / NOT logic at the microfluidic level. A storage pool of NSOs; an exemplary signal input for selecting target NSOs based on tag overhangs; an exemplary capture chamber for capture, washing, and elution for selection based on input signal(s); an unlimited number of signal inputs and capture chambers for performing selection; further exemplary signal inputs for selecting target NSOs based on tag overhangs; further exemplary capture chambers for capture, washing, and elution for selection based on input signal(s); and a final output where the scaffold sequence is amplified, sequenced, and decoded. Electrowetting-based droplet manipulation devices such as Mondrian can be used to perform these controlled mixing and splitting operations in a rapid and controlled manner that is also fully automated. [Figure 14]FIG. 14 is a schematic diagram showing elements of an exemplary system for creating, storing, and organizing sequence-controlled polymers as reusable "storage blocks" or computational molecular elements. A structured storage block, such as a cuboctahedron, is shown as a square structured nucleic acid storage block. Storage blocks can be of many sizes, from small to large, as needed to accommodate sequence-controlled polymers. Each block can have multiple different file handles, or indexes (shown as a-d), allowing multiple addressing of sequence-controlled polymers for selection and computation. Certain modifications, such as overhang sequences, can be used to associate multiple blocks together into a larger superblock of storage, allowing rapid retrieval, re-sorting, and computation with associated or sorted sequence-controlled polymers. Modified overhangs also allow Boolean AND, OR, and NOT operations to be used on storage blocks, for example, to select one or more storage blocks to purify from a pool of storage blocks. [Figure 15A]Figures 15A and 15B are flow charts. Figure 15A shows the workflow in one system for long-term storage of sequence-controlled polymers in the form of storage blocks of DNA. To separate and later recover the sequence-controlled polymers, any number of nucleic acid storage objects (e.g., one to a dozen out of millions) are blotted and lyophilized to a long-term storage material ("paper"). The dried storage blocks are selectively re-wetted by adding water or buffer to the blot. This process can be automated to selectively pull out the correct spatially separated storage pools, and the wetted storage blocks are processed as described and sequenced, for example, by a handheld device or benchtop sequencer. Figure 15B is a flow chart that describes a general approach to molecular data storage and computation. Any digital files and folders from a computer. Digital files are coded and / or converted into molecular storage code (e.g., nucleotides, amino acids, polymers, atoms, surfaces). This code is written into the physical storage blocks used to store the data. The stored data is associated with a set of address codes to identify the storage blocks. An address is affixed to the storage block such that it can be used for subsequent reading, manipulation, selection, and computation, including physical tags, electrostatic or magnetic properties, chemical properties, or optical properties. The storage block with the address is placed in a pool of other storage blocks for storage and computation. The pool is separated based on physical properties, and some storage blocks meet the selection criteria and others do not, and are sorted as such. This and other sorting criteria can be repeated multiple times in parallel or serial. The sorted storage block of interest is purified from the pool. The sorted storage block is read and decoded into a digital format. The original digital file is retrieved from the pool. [Figure 15B] Same as above. [Figure 16]Figure 16 is a line graph showing the % of readable message population over time. Degradation of NSOs is initiated upon exposure to an external switch (▲), such as light, heat, an enzyme, a chemical reactant, or the presence of air, which activates the timed degradation of DNA, RNA, or other nucleic acids, resulting in a degraded message pool. [Figure 17A-D] Figures 17A-17D are schematic diagrams of silica encapsulation of sequence-controlled polymer storage blocks. Figure 17A shows a silica particle (18). Figure 17B shows a silica particle (20) modified to allow adsorption of DNA particles. Figure 17C shows a nucleic acid storage block (22) adsorbed to a surface-modified silica particle. Figure 17D shows a secondary silica shell (24) grown on the silica (26) to which the nucleic acid storage block is adsorbed. This shell provides environmental protection for the nucleic acid storage block. Figure 17E is a schematic diagram of an exemplary DNA assembly (double crossover or DX tile) containing a Cy3 and Cy5 energy transfer pair as a readout for monitoring the structure of the DX tile. Figure 17F is a graph showing intensity (cps) versus wavelength (nm) corresponding to the emission spectrum of the DX tile before the encapsulation process (-) and at the completion of the encapsulation step (--), respectively. [Fig. 17E-F] Same as above. [Figure 18] Figures 18A-18F show examples of the results of NSO superstructuring. Figure 18A shows a single (monomer) NSO. Figures 18B-D show exemplary "dimers" of two NSOs joined together using overhang addressing at a vertex (Figure 18B), along an edge (Figure 18C), or at a face (Figure 18D), respectively. Figures 18E-18F show NSO "tetrahedra" assembled into larger superstructures as extended tetramers addressed to assemble along an edge via complementarity (Figure 18E), and as extended tetramers with different addresses that allow for the assembly of a more compact configuration (Figure 18F), respectively. [Figure 19A-C]19A-19C are schematic diagrams illustrating the molecular shelling of stored objects. FIG. 19A is a scheme illustrating the loading of a porous core (28) with multiple sequence-controlled polymers (30), shelling (32), and the addition of a feature tag to the shelled stored object (36). FIG. 19B is a scheme illustrating the first step in the assembly of a stored object (44) from the core (38), in which a recognition site (40) is first bound, followed by complexation of a sequence-controlled polymer (42) containing one or more tags specific to the recognition site bound to the core. FIG. 19C is a scheme illustrating the final step in the assembly of the stored object (50) shown in FIG. 19B. The core (44) and associated sequence-controlled polymer are then encapsulated in a shell (46), to which a feature tag (48) is attached. [Figure 20] 20A-20B are schematic diagrams showing molecular shelling of storage objects, including modification of the shell with multiple sequence-controlled polymers and affinity tags for multiplexed molecular logic operations and sequence-controlled polymer selection. (FIG. 20A) A sequence-controlled polymer (54) attached to a molecular core (52) is further surrounded by a molecular shell (56) and functionalized with an addressing / specificity tag (58) for multiplexed computation (60), or (FIG. 20B) a sequence-controlled polymer (64) absorbed by a molecular core (62) is further surrounded by a molecular shell (68) and functionalized with an addressing / specificity tag (66) for multiplexed computation (70). The shell or core have readouts based on optical, magnetic, electrical, or physical properties of the shell / core. [Figure 21]21A-21B are schematic diagrams illustrating storage where sequence-controlled polymers are in a molecular core or shell. FIG. 21A shows a storage object formed from sequence-controlled polymers on a molecular core with readout based on optical, magnetic, electrical, or physical properties of the core. The molecular core contains an address / specificity tag for molecular logic and sequence-controlled polymer retrieval operations. FIG. 21B shows a storage object formed from sequence-controlled polymers on a molecular shell surrounding the molecular core. The shell / core has readout based on optical, magnetic, electrical, or physical properties of the shell / core. The shell is functionalized with address / specificity tags for molecular logic and sequence-controlled polymer retrieval operations. [Figure 22] Figure 22 is a schematic diagram of a proposed workflow for the storage and retrieval of biomolecules, using nucleic acids as an example. Biomolecules are extracted from samples of any origin and collected in microplates. Upon encapsulation and barcoding of the samples, the capsules are pooled together. Samples are selected using probes containing optical markers or chemical / biochemical affinity tags. The tags are used to optically or mechanically sort samples from the pool. The remainder of the pool is returned to storage until further use. [Figure 23]23A-B are schematic diagrams of data panels showing proof-of-concept storage and recovery of biomolecules using synthetic barcoded packets. Capsules containing B. taurus (containing "Eukaryote", "Animalia", "2021-01-05", and "Bos taurus" labels) and M. musculus (containing "Eukaryote", "Animalia", "2021-01-03", and "Mus musculus" labels) genomes were targeted for recovery from pools containing H. sapiens total RNA (containing "Eukaryote", "Animalia", "2021-01-03", and "Homo sapiens" labels) and SARS-CoV-2 RNA genomes (containing "Riboviria", "Orthornavirae", "2020-12-20", and "SARS-CoV-2" labels). A Boolean logic query using molecular probes matching the query strings "Eukaryote", "Animalia", and "Homo sapiens" was added to the pool (Figure 23A). Fluorescence gate selection using the different colors associated with each probe identifies the populations of interest. Selecting populations positive for "Eukaryote" and "Animalia" selects B. taurus, M. musculus, and H. sapiens. An additional "Homo sapiens" gate can be used to select populations that are negative for "Homo sapiens" or, in the Boolean logic expression, NOT Homo sapiens. Thus, the final Boolean logic search query is "Eukaryote" AND "Animalia" AND (NOT "Homo sapiens"), which selects B. taurus and M. musculus, which were verified using quantitative real-time polymerase chain reaction (Figure 23B). [Figure 24]Figures 24A-B show a proof-of-concept reaction using a barcode on the sample surface as an initiator. Figure 24A is a schematic showing hybridization-based selection, where a capsule containing a "Homo sapiens" tag (labeled "z" in the figure) is hybridized with a complementary z* tag that also contains a toehold sequence "a*" and a stem sequence "b*", triggering a hybridization chain reaction (HCR) between two hairpin structures modified with a marker, which can be a dye or a chemical / biochemical tag. Figure 24B is a graph of intensity (au) versus wavelength (nm) for HCR-modified capsules, single probe-modified capsules, and orthogonal barcode+HCR control capsules, respectively, showing the enhanced fluorescence observed for HCR-amplified capsules compared to capsules hybridized only with a complementary strand containing a single dye. [Diagram 25] Figures 25A-C are drawings of an exemplary millifluidic device that can be used to encapsulate and barcode biomolecules using an emulsion reactor. Figure 25A is a CAD design of the millifluidic device. Figure 25B shows a 3D printed millifluidic device. Figure 25C is a schematic detailing droplet formation within the device photographed in Figure 25B, where 2 mM Ca2+ and 2% (w / w) low viscosity alginate are flowed in a channel connected to a T-junction where surfactant-containing oil is flowed. [Figure 26] 26 is a schematic diagram of a process for recovering a collection of particles corresponding to a range of some numerical feature of an underlying biomolecule. Each possible digit value at each digit place of the numerical value is associated with a separate orthogonal barcode, allowing recovery of a range of values by selecting particles having specific digit values at a subset of the digit places. As an example, the numerical feature may be represented in base 3, and by selecting particles having barcodes associated with a "1" at the 27th place and a "0" at the 9th place, a collection of particles having barcodes corresponding to numerical values in the range [1000, 1100] may be recovered. [Figure 27]FIG. 27 is a schematic diagram of a barcode sequence design process that allows accurate similarity-based retrieval for features whose similarity metric is simple enough to allow accurate isometric embedding from the feature similarity space into a low-dimensional hypercube. The isometric embedding corresponds directly to the assignment of barcodes to each particle, allowing similarity-based retrieval. As an example, the schematic diagram shows a nucleic acid sequence CCCATCGTGTCATTA (SEQ ID NO: 1) with a selection of four mutations at different positions in the sequence and a simple similarity metric represented as a circular graph with eight nodes that can be embedded accurately isometrically into a four-dimensional hypercube graph. [Figure 28] Figure 28 is a schematic of a barcode sequence design process that allows approximate similarity-based recovery for features with arbitrarily complex similarity metrics. The feature similarity space is simplified and reduced to a small number of dimensions using standard dimensionality reduction. These dimensions are then further approximated by binning, which can then be directly embedded into a hypercube graph where the nodes represent mutational variants of a set of barcodes. As a proof-of-concept example, the schematic shows the process starting from a complex similarity metric derived from 4187 SARS-CoV2 genomes for which pairwise genetic similarities were calculated. This similarity metric was reduced to 18 dimensions using multidimensional scaling (MDS), and here, for visualization purposes only, the dimensionality was further reduced to 2 dimensions before plotting. After binning, linear regression showed a strong correlation between the original similarity metric and the final distance in the 54-dimensional hypercube embedding. The hypercube embedding directly corresponds to assigning 6 barcode sequences to each node of the original feature space, each with 9 mutation sites. Exemplary barcode sequences include GCCTTGTATGTGAATATCCGTGTCA (SEQ ID NO:2), and GGAGAATGATTAGCACGGAGAGTGG (SEQ ID NO:3). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0060] Detailed Description of the Invention The encapsulation chemistry is combined with the precision of DNA base pairing as a molecular barcode for the identification and retrieval of individual samples to realize a room temperature ultra-high density preservation and retrieval system for DNA, RNA, peptides, and proteins. The disclosed technology is broadly applicable to the preservation and sorting of biomolecules from any source, such as human patients, animals, and the environment.
[0061] In one implementation, biomolecules are surface adsorbed onto the surface of capsules ranging from 1 nm to 100 μm in diameter. The biomolecules are covalently or non-covalently bound to the particle surface. Encapsulation of the surface adsorbed molecules proceeds by condensation, polymerization, and crosslinking of inorganic and organic monomers with the surface adsorbed monomers. The surface of the encapsulated biomolecules is then labeled using single-stranded DNA barcodes.
[0062] In another implementation, biomolecules are encapsulated within the channels of porous particles.
[0063] In another implementation, the biomolecules and encapsulation reagents are introduced into the wells of a microplate containing the sorbent particles using an automated liquid handling device.
[0064] In another implementation, biomolecules are captured and encapsulated in an emulsion using electrically or photon-controlled microfluidic channels, with the barcode attached after encapsulation.
[0065] In another implementation, biomolecules and barcodes are combined and housed in an emulsion composed of multiple layers of aqueous and organic solvents using a microfluidic approach. Permanent encapsulation and barcoding using organic or inorganic polymers proceeds in one step.
[0066] In another implementation, the molecular barcode may include non-standard nucleotides or a non-phosphate backbone to improve the stability of the barcode.
[0067] In another implementation, the molecular barcodes can be attached using chemical synthesis or enzymes.
[0068] Selection of the encapsulated samples proceeds by hybridization of probes complementary to the barcodes of interest. The probes may contain optical, chemical, and biochemical markers for optical or mechanical sorting using millifluidic or microfluidic approaches.
[0069] In another implementation, chemical and biochemical reactions can be performed on the tags to increase sorting throughput.
[0070] This preservation and retrieval system isolates biomolecules of interest from the environment to protect their integrity for a decade or longer, eliminating the need for cryogenic storage conditions. Barcoding micron- to nanoscale capsules allows all samples to be pooled in a single container rather than millions of individual tubes, thus reducing the footprint of biomolecule preservation to desktop dimensions.
[0071] Capsules herein refer to molecules encapsulated as particles containing biomolecules and are labeled with molecular barcodes for retrieval. Encapsulating agents herein can be composed of organic and inorganic materials. Molecular barcodes herein are short primer strands of oligonucleotides derived from a pool of 240,000 [Xu, et al. Proceedings of the National Academy of Sciences 106, 2289-2294, doi:10.1073 / pnas.0812506106 (2009)]. Barcodes are taken from this pool and used with or without sequence modification to allow the retrieval of individual particles or collections of related particles. The choice of barcode allows the retrieval of collections of related particles that correspond to discrete categories, ranges of discrete numerical features (e.g., sampling date), or similarity-based retrieval for continuous or non-discrete features. The encapsulation and barcoding approach can be performed using automated liquid handling equipment or millifluidic / microfluidic devices. Samples are selected for recovery by the addition of a probe that hybridizes to the target barcode. The selected samples are sorted out of solution using optical and mechanical sorting methods, including but not limited to, using fluorescence activated sorting, magnetic sorting, electrokinetic sorting, and similar sorting approaches. Sample selection and sorting can also be performed using automated liquid handling equipment or millifluidic / microfluidic devices.
[0072] Various schemes by which barcodes can be assigned to particles to enable the selection of different collections of related particles are described below. To enable the retrieval of a collection of particles that belong to one of several discrete categories, one orthogonal barcode sequence is associated with each category, and the particle's membership in each category is indicated by the selection of the particle's corresponding barcode. To enable the retrieval of a collection of particles that belong to a range of discretized numerical features, one orthogonal barcode sequence is associated with each possible digit value for each digit of the numerical value. With this approach, a collection of particles corresponding to any numerical range of the feature can be retrieved, as long as this range can be specified by selecting specific digit values for some subset of the numerical value's digits. An example of numerical range retrieval is shown in FIG. 26.
[0073] To allow for the recovery of a collection of particles that are similar to each other with respect to continuous or non-discrete features, the barcode sequence is mutated at a small number of carefully selected sites within the sequence. The limited set of mutated variant barcode sequences is represented by a graph G, such as, but not limited to, a hypercube graph. The mutation sites are selected so that the graph G faithfully represents the binding affinity between the barcodes and the complementary sequences for the barcodes used as probes. The continuous feature similarity space is also represented by a graph H, which is then isometrically embedded in the graph G. For certain simple graphs H, an exact isometric embedding can be found using polynomial-time algorithms. For any complex graph H, an isometric embedding can be found by first performing a dimensionality reduction on the corresponding metric space represented by H. The dimensionality reduction can be performed using any standard technique that seeks to preserve distances during transformations. The low-dimensional space can then be discretized to approximate the isometric embedding in G. Examples of finding an isometric embedding for both simple and complex cases of H are shown in Figures 27 and 28.
[0074] I. Definition A "feature tag" is an oligonucleotide of defined sequence that corresponds to a feature attributed to a sequence-controlled polymer. The correspondence between a feature and a feature tag refers to a one-to-one mapping between the feature and its feature tag.
[0075] "Characteristics attributable to a sequence-controlled polymer" refers to characteristics that a sequence-controlled polymer possesses or embodies.
[0076] "Hybridically distinguishable" means orthogonal to hybridization.
[0077] "Similarity coded" means that the relative hybridization properties of feature tags are related to the similarity of the features to which they correspond, with feature tags corresponding to more similar features having closer relative hybridization properties than feature tags corresponding to less similar features. In a similarity coded set of feature tags, it is useful that the difference in hybridization energies of feature tags in the set is a monotonically increasing function of the similarity of the features to which they correspond.
[0078] "Relative hybridization" refers to the hybridization energy of a probe to a feature tag compared to the hybridization energy of the same probe to a different feature tag.
[0079] "Hybridization-defined" means that each of the feature tags in a set differs from every other feature tag in the set by 1 to x mismatched nucleotides, where the mismatched nucleotides are (i) at least 2 nucleotides from either end of the feature tag and (ii) separated by at least one matching nucleotide in the feature tag, where x is the number of different nucleotide positions in the feature tags that vary within the set.
[0080] "Number coded" means that each different digit tag corresponds to a different place digit value of a multi-digit number.
[0081] The term "payload" refers to a sequence-controlled polymer for storage. For example, in nucleic acid storage, the payload is a specified nucleotide sequence. The terms "desired polymer" or "desired nucleic acid" are used interchangeably to designate the payload contained in the sequence within a given storage object.
[0082] The term "sequence" refers to any natural or synthetic sequence-controlled polymer sequence that is stored. For example, if a nucleic acid is used to store data, the "sequence" is the nucleic acid sequence of that nucleic acid. The nucleic acid can be in the form of a linear nucleic acid sequence, a two-dimensional nucleic acid object, or a three-dimensional nucleic acid object. The nucleic acid can include a synthetic sequence or a naturally occurring sequence. The sequence of any sequence-controlled polymer can be considered to code for the data represented by the sequence of the polymer. For example, a naturally occurring nucleic acid is a sequence-controlled polymer in which the naturally occurring sequence of the nucleic acid is the data encoded by the nucleic acid.
[0083] The term "bit" is a contraction of "binary digit." In general, a "bit" refers to the basic capacity of information in computing and telecommunications. Although a "bit" conventionally represents only 1 or 0 (one or zero), other codes can be used in nucleic acids that contain four nucleotide possibilities (ATGC) at every position, and higher codecs containing consecutive 2-, 3-, 4-, etc. nucleotides can alternatively be used to represent bits, letters, or words.
[0084] The terms "nucleic acid molecule", "nucleic acid sequence", "nucleic acid fragment", "oligonucleotide", and "polynucleotide" are used interchangeably and are intended to include, but are not limited to, polymeric forms of nucleotides that may be of various lengths, either deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs or modified nucleotides thereof, including, but not limited to, locked nucleic acids (LNA) and peptide nucleic acids (PNA). Oligonucleotides are typically composed of a specific sequence of the four nucleotide bases adenine (A), cytosine (C), guanine (G), and thymine (T) (uracil (U) instead of thymine (T) if the polynucleotide is RNA). Thus, the term "oligonucleotide sequence" is an alphabetic representation of a polynucleotide molecule, or the term may be applied to the polynucleotide molecule itself. This alphabetic representation can be entered into a database of a computer having a central processing unit and used in bioinformatics applications such as functional genomics and homology searching. Oligonucleotides may contain one or more non-standard nucleotides, nucleotide analogs and / or modified nucleotides as appropriate.
[0085] The terms "staple strand" or "helper strand" are used interchangeably. When used in the context of nucleic acid nanostructure objects, "staple strand" or "helper strand" refers to an oligonucleotide that acts as a glue to hold the scaffold nucleic acid in its three-dimensional geometry.
[0086] The terms "scaffold-type origami", "origami" or "nucleic acid nanostructure" are used interchangeably. They can be one or more short single strands of nucleic acid (staple strands) (e.g., DNA) that fold a long single strand of polynucleotide (scaffold strand) into a desired shape on the order of about 10 nm to 1 micron or more. Alternatively, single stranded synthetic nucleic acid can be folded into an origami object without a helper strand, for example, using parallel crossover motifs or parallel crossover motifs. Alternatively, pure staple strands can form a finite range nucleic acid storage block. Scaffold-type origami or origami can be composed of deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs or modified nucleotides thereof, including but not limited to locked nucleic acids (LNA) and peptide nucleic acids (PNA). Scaffolds or origami composed of DNA can be referred to as, for example, scaffold-type DNA origami or DNA origami. Where the compositions, methods, and systems herein are exemplified using DNA (e.g., DNA origami), it will be recognized that other nucleic acid molecules can be substituted.
[0087] The terms "nucleic acid encapsulation" and "nucleic acid packaging" are used interchangeably. They refer to a method of encapsulating nucleic acids of any length or geometry by a material to form a discrete unit. The encapsulating material can be any suitable natural or synthetic material, such as proteins, lipids, sugars, polysaccharides, natural polymers, synthetic polymers, or derivatives thereof. Thus, the encapsulating unit is in the form of a gel-based bead, a protein virus package, a micelle, a mineralized structure, a siliconized structure, a polymeric packaging, or any combination thereof.
[0088] The term "sequence-controlled polymer" or "sequence-controlled macromolecule" refers to a macromolecule composed of two or more distinct monomer units arranged sequentially in a specific, non-random manner as a polymer "chain". That is, a sequence-controlled polymer is a polymer in which the order of the monomer units in the polymer is non-random, specified, or specifically determined. The arrangement of the two or more distinct monomer units constitutes a precise molecular "signature" or "code" within the polymer chain. A sequence-controlled polymer may be a biological polymer (i.e., a biopolymer) or a synthetic polymer. Exemplary sequence-controlled biopolymers include nucleic acids, polypeptides or proteins, linear or branched carbohydrate chains, or other sequence-controlled polymers. Exemplary sequence-controlled polymers are described in Lutz, et al., Science, 341, 1238149 (2013).
[0089] The term "sequence-controlled polymer object" refers to an object that includes a sequence-controlled polymer and one or more feature tags, digit tags, and / or barcodes.
[0090] The terms "sequence-controlled polymer storage object", or "SSO", or "storage block", or "storage object" are used interchangeably. They refer to an object that includes a sequence-controlled polymer and one or more feature tags or barcodes. The polymer includes a discrete sequence, and the feature tags allow for the selection, organization, and separation of the storage objects. In some forms, the storage object includes a sequence in the form of a continuous stretch of sequence-controlled polymer. In some forms, the storage object includes non-contiguous segments of the sequence. In some forms, the storage object includes a sequence-controlled polymer folded into a two-dimensional or three-dimensional shape. For example, the sequence-controlled polymer can be folded into a nanostructured form that is the entire SSO, such as a nanostructured nucleic acid object. In some forms, the sequence-controlled polymer is combined with one or more additional materials to form a nanoparticle. The SSO can take any form, such as a linear sequence molecule, or a two-dimensional object, or a three-dimensional object. Sometimes, storage objects are made from scaffold polymer sequences, with or without staple nucleic acid sequences, or from sequence-controlled polymers of any length / configuration encapsulated within one or more encapsulating agents.
[0091] The term "nucleic acid storage object" or "NSO" is used interchangeably to refer to an SSO that contains nucleic acid as a sequence. An NSO contains one or more segments of a nucleic acid sequence. In some forms, an NSO exists in the form of a single-stranded nucleic acid scaffold that folds on itself, or multiple single-stranded nucleic acid molecules that self-assemble into a programmed geometric block. An NSO can take any form, such as a linear nucleic acid sequence, a two-dimensional nucleic acid object, or a three-dimensional nucleic acid object. Sometimes, a nucleic acid storage object is a nucleic acid object made from a scaffold nucleic acid with or without a staple nucleic acid sequence, or from an encapsulated nucleic acid of any length / form, or any combination of these. An NSO can be composed of deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs or modified nucleotides thereof, including but not limited to locked nucleic acids (LNA) and peptide nucleic acids (PNAs). An NSO composed of DNA can be referred to as a DNA storage object ("DMO"), or the like. Where the compositions, methods, and systems herein are exemplified using DNA (eg, DMO), it will be recognized that other nucleic acid molecules can be substituted.
[0092] The terms "splint strand" and "bridge strand" are used interchangeably to refer to a nucleic acid sequence that is complementary to the strands of two or more nucleic acid sequences in separate, non-overlapping positions. For example, a first region on the splint strand is complementary to a region on the overhang tag of a first NSO, while a second region on the same splint strand is complementary to a region on the overhang tag of a second NSO. The two regions of the splint strand are positioned such that the binding of the first NSO does not sterically interfere with the binding of the second NSO. Thus, the splint strand or bridge strand serves to bring the two NSOs into close proximity at a certain, predetermined distance.
[0093] The terms "feature tag", "nucleic acid overhang", "DNA overhang tag", and "staple overhang tag" are used interchangeably to refer to nucleotides associated with an SSO that can be functionalized. In some cases, the overhang tag contains one or more nucleic acid sequences that code for the metadata of the associated SSO. In some forms, the nucleotides are added to the staple strand of the NSO. In some forms, the overhang tag contains a sequence designed to hybridize to other stationary phase objects, such as magnetic beads, surfaces, agarose or other polymer beads. In some cases, the overhang tag contains a sequence designed to hybridize to other nucleic acid sequences, such as on the tag of another SSO or on the splint strand. In other cases, the overhang contains one or more sites for conjugation to a molecule. For example, the overhang tag can be conjugated to a protein or a non-protein molecule, for example, to enable affinity binding of the SSO. Exemplary proteins for conjugating to the overhang tag include biotin and an antibody, or an antigen-binding fragment of an antibody. In some forms, overhang tags are designed and implemented within the SSO to enable programmable affinity and specificity between two interacting conserved entities, regardless of implementation, e.g., using principles of Boolean logic and computation.
[0094] The terms "encapsulation," "enveloping," "coating," "covering," and "shelling" are used interchangeably to refer to a process in which an SSO is completely or partially surrounded by an encapsulating agent. The term "encapsulating agent" refers to a molecular entity such as a polymer or other matrix.
[0095] II. Methods and Systems for Sequence-Based Storage Sequence-controlled polymers such as nucleic acid molecules (e.g., DNA) have high information density (e.g., up to 10 24It represents an excellent storage object and medium, with very high potential for high density (bits / kg), long term stability, and low energy cost to maintain.
[0096] A method for the storage of sequence-controlled polymers formed into nanostructures has been developed. The sequence-controlled polymers are folded or embedded into well-defined discrete structures that function as sequence-controlled polymer storage objects (SSOs). Thus, separate packages of sequence-controlled polymers are provided as three-dimensional structures with multiple faces that contain one or more specific sequence tags. By manipulating the SSO structure, the method allows the partitioning, association, and re-sorting of polymer sequences within each SSO. Information recovery is rapidly achieved by interpreting the sequence, structure, or other physical or chemical properties of the sequence-controlled polymers. Thus, the method allows for rapid and efficient organization and access of sequence-controlled polymers stored within SSOs.
[0097] Methods for the storage of sequence-controlled polymers of any length or any form have also been developed. Typically, sequence-controlled polymers with sequences of any desired length are packaged, encapsulated, wrapped, or housed in gel-based beads, protein virus packages, micelles, mineralized structures, siliconized structures, or polymer packages, referred to herein as "sequence-controlled polymer storage blocks". In some forms, the synthetic polymer or biopolymer comprises a single continuous polymer contained within a nanoparticle. In some forms, the synthetic polymer or biopolymer comprises multiple such polymers combined within a single nanoparticle. These discrete biopolymer "packages" function as sequence-controlled polymer storage objects (SSOs), allowing the incorporation of one or more specific tags on the surface of the structure. Some exemplary tags include nucleic acid sequence tags, protein tags, carbohydrate tags, and any affinity tags.
[0098] In some forms, the sequence controlled polymer is a biopolymer such as a nucleic acid sequence, a polypeptide amino acid sequence, a protein, a carbohydrate sequence, or a combination thereof.
[0099] A. Preservation of sequence-controlled polymers The method of storing polymers can include the assembly of sequence-controlled polymer storage objects (SSOs) that include one or more polymer sequences and one or more feature tags. The one or more polymer sequences can be present within the particle core or in association with one or more layers surrounding the core, for example embedded within an encapsulating material. The index / affinity tag is exposed and accessible. For example, the index / affinity tag is embedded within the particle or otherwise attached to the outer surface of the particle. The manner in which the index / barcode is attached to the outer surface of the core particle and / or the sequence can be varied according to the desired manner for pooling, sorting, organizing and accessing the sequence-controlled polymers.
[0100] In some forms, the "shell" that is the product of "shelling" contains a sequence-controlled polymer.
[0101] 1. Nucleic Acid Nanostructures In an exemplary embodiment, the sequence-controlled biopolymer is a nucleic acid. A method for preserving sequence-controlled polymers using nucleic acid nanostructures has been developed. Nucleic acid nanostructures formed from single-stranded nucleic acid scaffolds up to several tens of kilobases (kb) are folded into well-defined, separate structures that function as nucleic acid storage objects (NSOs). Thus, separate packages of sequence-controlled polymers are provided as three-dimensional nucleic acid structures with multiple faces that contain one or more specific sequence tags. By manipulating the NSO structure, the method allows the division, assembly, and re-sorting of sequence-controlled polymers within the NSO. Information recovery is rapidly achieved by sequencing. Thus, the method allows for rapid and efficient organization and access of sequence-controlled polymers stored within the NSO.
[0102] Methods for the storage of nucleic acids of any length or in any form have also been developed. Typically, nucleic acids of any desired length are packaged, encapsulated, enveloped, or housed in gel-based beads, protein virus packages, micelles, mineralized structures, siliconized structures, or polymeric packages, referred to herein as "nucleic acid packages". In some forms, linear nucleic acids are base-paired double strands. In other forms, linear nucleic acids include long continuous single stranded nucleic acid polymers or multiple such polymers. These discrete nucleic acid packages function as nucleic acid storage objects (NSOs) and allow the incorporation of one or more specific tags on the surface of the structure. Some exemplary tags include nucleic acid sequence tags, protein tags, carbohydrate tags, and any affinity tags.
[0103] Thus, methods for assembling sequences into sequences of single-stranded scaffolds allow for natural spatial separation of sequence-controlled polymers, tagging or addressing sequence-controlled polymers multiple times by functionalizing the staple strands used to fold the objects of interest, exchanging staple strands with different overhangs to modify the address, and associating NSOs together to further spatially separate the sequence-controlled polymers of interest. Nucleic acids can be nanostructured into a diverse set of sizes and structures and can be multiply addressed to geometrically specific locations (Figures 1A-1C). Nanostructured nucleic acids can be folded into a wide range of scaffold sizes, from just a few hundred nucleotides to hundreds of thousands of nucleotides, in user-defined, highly specific geometries that are theoretically unlimited in size. Single-stranded scaffolds can be used as scaffolds that are routed through objects folded into specific shapes by complementary single-stranded oligonucleotide staples, or by programming the single-stranded scaffold sequence to fold onto itself. These shapes can take any desired form, for example as defined by the user. In some forms, the structure is a closed, tightly packed block. In other forms, the structure has the form of an open wireframe mesh, e.g., a polyhedral structure. In each case, the geometry of the structure can be defined in any way that suits the overall storage block superstructuring and tag presentation / accessibility.
[0104] 2. Conserved access of sequence-controlled polymers A method is described for sorting, organizing, and accessing sequence-controlled polymers in a pool of different SSOs. Typically, the method selects and sorts SSOs based on intermolecular interactions between different or equally addressed SSOs in the pool. Typically, the method uses a nucleic acid label specifically bound to one or more SSOs. In some forms, each SSO contains a single tag. In other forms, each SSO contains two or more tags. Thus, in some forms, the method provides a multiply addressed SSO. The multiply addressed SSO allows for rapid selection of nucleic acids using user-defined combinations of Boolean logic, including AND, OR, and NOT logic. In some forms, the method uses a nucleic acid label to physically associate distinct SSOs with each other. Thus, in some forms, the method provides a system for rapid retrieval using prior logic, allowing physical association in ultraconserved blocks to network and spatially separate blocks of related sequence-controlled polymers. In other forms, the conserved blocks are geometrically positioned at specific locations that allow for tuning of the conserved positions.
[0105] SSOS, including nanostructured SSOs, can associate into larger superstructures based on signals to pools of stored objects (Figures 2A-2D). In some forms, pools of SSOs contained in solution are assembled based on the specific geometry of overhang sequences at precise locations. Typically, assembly occurs by complementary sequences on the overhangs, by bridging oligonucleotides (splint strands), or by protein or chemical additions to the overhangs. Superstructured SSOs can be specifically dissociated and regrouped by using external signals as desired by the user. Exemplary external signals used to control dissociation include a change in pH, a decrease in salt, an increase in temperature, application of electromagnetic radiation, displacement of the tow-hold strand, an excess of complementary strands, or enzymatic release by restriction nucleases, nickases, helicases, resolvases, release using UV-sensitive linkers, use of CRISPR / Cas9 and guide RNAs, or any combination of these.
[0106] The sequence-controlled polymer can be a biopolymer, such as DNA or a polypeptide, or a synthetic biopolymer, such as a peptidomimetic.
[0107] A non-limiting list of suitable sequence controlled polymers includes naturally occurring nucleic acids, non-naturally occurring nucleic acids, naturally occurring amino acids, non-naturally occurring amino acids, peptidomimetics such as polypeptides formed from alpha peptides, beta peptides, delta peptides, gamma peptides and combinations thereof, carbohydrates, block copolymers, and combinations thereof. Sequence-defined non-natural polymers closely resemble biopolymers, including polymers incorporating non-canonical amino acids, e.g., peptidomimetics such as β-peptides (Gellman, SH. Acc. Chem. Res., 31, 173-180 (1998)), peptide nucleic acids (PNAs), peptoids or poly-N-substituted glycines (Zuckermann, et al., J. Am. Chem. Soc., 1 14, 10646-10647(1992)), oligocarbamates (Cho, CY et al., Science, 261, 1303-1305(1993)), glycopolymers, nylon-type polyamides, and vinyl copolymers.
[0108] Enzymatic and non-enzymatic synthesis of sequence-defined non-natural polymers can be achieved by template polymerization (reviewed in Brudno Y et al., Chem Biol.; 16(3): 265-276 (2009)).
[0109] In some forms, a method is provided that includes providing a nucleic acid sequence from a pool containing a plurality of similar or different sequences. In some forms, the pool is a database of known sequences. For example, in certain forms, discrete "blocks" are contained within a pool of nucleic acid sequences ranging in size from about 100 to 1,000,000 bases, although this upper limit is theoretically unlimited. In some forms, the nucleic acid sequences within the pool of a plurality of nucleic acid sequences share one or more common sequences. When the provided nucleic acid is selected from a pool of sequences, the selection process can be performed manually, for example, by selection based on user preferences, or automatically.
[0110] B. SSO Construction In general, the purpose of generating individual SSOs is to separate blocks of sequence-controlled polymer from other blocks, to separate identification tags from the underlying sequence-controlled polymer, and to allow manipulation and selection of large packages as needed.
[0111] 1. Custom design of SSOs by inclusion of sequence-controlled polymers Sequence-controlled polymers can be formed into SSOs by encapsulation (FIGS. 4A-4E, 19A-19C, 20A-20B, and 21A-21B). For example, single-stranded and / or double-stranded DNA, or any other nucleic acid, can be used to generate NSOs by encapsulation. The encapsulated sequence-controlled polymers can take any form, such as linear DNA sequences, two- or three-dimensional DNA objects, polypeptides, proteins, etc. In some forms, the linear polymers are nucleic acids that are base-paired double strands. In other forms, the linear nucleic acids include long continuous single-stranded nucleic acid polymers or multiple such polymers. In further forms, the nucleic acids encapsulated in the same particle are a mixture of linear and non-linear nucleic acids. For example, one or more single-stranded nucleic acids and one or more scaffold-type nucleic acid nanostructures can be encapsulated in the same particle.
[0112] In some forms, the sequence-controlled polymer is packaged into discrete SSOs by encapsulation. Suitable encapsulating agents include gel-based beads, protein virus packages, micelles, mineralized structures, siliconized structures, or polymer packaging.
[0113] In some forms, the encapsulating agent is a viral capsid, or a functional part, derivative and / or analog thereof. In some forms, the encapsulating agent is a lipid that forms a micelle or liposome that surrounds the nucleic acid. In some forms, the encapsulating agent is a natural or synthetic polymer. In some forms, the encapsulating agent is mineralized, e.g., alginate beads, or polysaccharides mineralized with calcium phosphate. In other forms, the encapsulating agent is siliconized. Packaging the sequence-controlled polymer sequence into a storage block allows for selection and superstructuring by the use of molecular identifiers, i.e., "addresses." In addition to the nucleic acid overhang, other purification tags can be incorporated into the overhanging nucleic acid sequence of any SSO for purification (i.e., recovery of the sequence-controlled polymer). In some forms, the overhang contains one or more purification tags. In some forms, the overhang contains a purification tag for affinity purification. In some forms, the overhang contains one or more sites for conjugation to a nucleic acid, or a non-nucleic acid molecule. For example, the overhang tag can be conjugated to a protein or non-protein molecule, for example, to allow affinity binding of the SSO. Exemplary proteins for conjugating to the overhang tag include biotin, an antibody, or an antigen-binding fragment of an antibody.
[0114] Assembly of the stored object by encapsulation or direct assembly of sequence-controlled polymers and feature tags can generate a range of stored objects with different structures. For example, in some forms, the stored object comprises a core particle onto which one or more sequence-controlled polymers are attached. Attachment of the sequence-controlled polymer to the particle core can be achieved using covalent or non-covalent bonds. In some forms, the core molecule is coated or linked with a binding site that is recognized by a molecule that is an intermediate receptor, e.g., one or more ligands associated with the sequence-controlled polymer (see FIG. 19B). The sequence-controlled polymer can be linked or hybridized to the receptor-coated core molecule. In some forms, the polymer / core substructure is then coated with one or more encapsulating agents (i.e., "molecular shelling") to generate a coated polymer / core structure, which is linked to one or more feature tags (see FIG. 19C). Attachment of the feature tag to the coated polymer / core particle can be achieved using covalent or non-covalent bonds or hybridization of complementary nucleic acids.
[0115] In some forms, assembly of the storage object involves loading or complexing one or more sequence-controlled polymers into the interior space of a porous or otherwise accessible polymer core molecule or structure (see FIG. 19A). In some forms, assembly of the storage object involves encapsulating or shelling the polymer-loaded core to create a particle loaded with the encapsulated polymer, which is then complexed with one or more feature tags.
[0116] In some forms, the storage object includes a sequence-controlled polymer and, optionally, a core molecule and / or encapsulant coated with multiple different types of feature tags. For example, in some forms, the storage object is assembled to allow multiplexed molecular logic operations and selection of sequence-controlled polymers. For example, in some forms, one or more sequence-controlled polymer encapsulations or molecular shells containing multiple fragments of sequence-controlled polymers are labeled with multiple feature tags. The feature tags can be attached directly to or adsorbed onto the molecular core, which is further surrounded by a molecular shell and functionalized with addressing / specificity tags for multiplexed computation (Figures 20A-20B).
[0117] In some forms, the stored object includes a core molecule or encapsulant coated with sequence-controlled polymer and, optionally, a feature tag, and then coated with a shell or core that itself produces a signal or has another property that can be detected and measured to generate a readout. Thus, the outer "shell" or inner "core" of the stored particle can be used to address or label the stored object. Exemplary physical or chemical properties that can be detected and measured include optical, magnetic, electrical, or physical properties. Thus, in some forms, the outer shell or inner core of the stored object generates a readout based on the optical, magnetic, electrical, or physical properties of the shell / core. Figures 21A-21B are schematic diagrams illustrating storage where the sequence-controlled polymer is in the molecular core or shell. Thus, in some forms, the sequence-controlled polymer is placed directly on the molecular core, with a readout based on the optical, magnetic, electrical, or physical properties of the core. The molecular core also contains an address / specificity tag for molecular logic and retrieval operations of the sequence-controlled polymer. In some forms, the sequence-controlled polymer is on a molecular shell that surrounds the molecular core. The shell / core has a readout based on the optical, magnetic, electrical, or physical properties of the shell / core. The shell is functionalized with addressing / specificity tags for molecular logic and sequence-controlled polymer recovery operations. In some forms, the core structure of the particle is formed from a sequence-controlled polymer folded into a 3D polyhedral or 2D polygonal shape. For example, in some forms, the sequence-controlled polymer is a nucleic acid, which is folded into a nucleic acid nanostructure having a 2D or 3D shape and one or more feature tags are added. Thus, in some forms, the shape of the nucleic acid nanoparticle can be used to identify, sort, or select the sequence-controlled polymer in the storage object. In some forms, the nucleic acid nanoparticle contains one or more additional core or encapsulation molecules with a readout based on the optical, magnetic, electrical, or physical properties of the core.
[0118] i. Nucleic acid nanostructures Two general approaches to constructing nucleic acid storage objects (NSOs) are described below: (1) using scaffolded nucleic acids and associated staple strands; (2) using encapsulation materials to accommodate a defined amount of nucleic acid in a single NSO unit. Thus, scaffolded nucleic acid nanostructures are primarily composed of nucleic acids, but additional non-nucleic acid components can be added to the overhanging sequences, such as protein tags for purification, or nucleases for nucleic acid degradation. The encapsulated nucleic acid units can be made of any natural or synthetic material. In some forms, scaffolded nucleic acid nanostructures are also encapsulated in one or more layers of polymer for additional layers of address / metadata tags, and / or long-term stability.
[0119] a. Scaffold-type nucleic acid The method includes assembling sequence-controlled polymers into nucleic acid nanostructures. Many known methods are available for making scaffolded nucleic acids, such as DNA origami structures. Exemplary methods include Benson E et al(Benson E et al., Nature 523, 441-444 (2015)), Rothemund PW et al(Rothemund PW et al., Nature. 440, 297-302 (2006)), Douglas SM et al.,(Douglas SM et al., Nature 459, 414-418 (2009)), Ke Y et al(Ke Y et al., Science 338: 1177 (2012)), Zhang F et al(Zhang F et al., Nat. Nanotechnol. 10, 779-784 (2015)), Dietz H et al(Dietz H et al., Science, 325, 725-730 (2009)), Liu et al(Liu et al. al., Angew. Chem. Int. Ed., 50, pp. 264-267 (2011)), Zhao et al. (Zhao et al., Nano Lett., 11, pp. 2997-3002 (2011)), Woo et al. (Woo et al., Nat. Chem. 3, pp. 620-627 (2011)), and Torring et al. (Torring et al., Chem. Soc. Rev. 40, pp. 5636-5646 (2011)), which are incorporated by reference herein in their entireties.
[0120] Typically, the creation of an NSO involves one or more of the following steps: (1) Design of NSO, (2) NSO sign; (3) Assembling the NSO; and (4) Purification of the assembled NSO.
[0121] b. Custom design of nucleic acid nanostructures Nucleic acid nanostructures have a defined shape and size. Typically, one or more dimensions of the nanostructure are determined by the target sequence. The method includes designing a nanostructure that includes a target nucleic acid sequence.
[0122] Nucleic acid nanostructures for use as NSOs can be geometrically simple or geometrically complex, such as polyhedral three-dimensional structures of arbitrary geometry. Any method for manipulating, sorting, or shaping nucleic acids can be used to generate NSO nanostructures. Typically, the methods include methods for "shaping" or otherwise changing the conformation of nucleic acids, such as methods for DNA origami.
[0123] In some forms, nanostructures for nucleic acid target sequences are designed using a method to determine single-stranded oligonucleotide staple sequences that can be combined with the target sequence to form a complete three-dimensional nucleic acid nanostructure of the desired morphology and size. Thus, in some forms, the method includes automated custom design of a nucleic acid storage object (NSO) corresponding to the target nucleic acid sequence. For example, in some forms, a robust computational approach is used to generate DNA-based wireframe polyhedral structures of any scaffold sequence, symmetry and size. In certain forms, the design of the NSO corresponding to the target nucleic acid sequence includes providing geometric parameters corresponding to the desired morphology and dimensionality of the NSO, which parameters are used to generate sequences of oligonucleotide "staples" that can hybridize to the target nucleic acid "scaffold" sequence to form the desired shape. Typically, the target nucleic acid is routed throughout the Eulerian circuit of the network defined by the wireframe geometry of the nanostructure.
[0124] Thus, in some forms, the NSO is designed by a method that includes the following steps: (1) selecting a target structure which may be from a predefined set of geometries, or may further comprise the steps of: (a) determining the spatial coordinates of all vertices, the edge connectivity between the vertices, and the faces to which the vertices belong in a target structure; (b) identifying a route of a single-stranded nucleic acid scaffold sequence to follow throughout the target structure; and (2) determining the nucleic acid sequence of the single stranded nucleic acid scaffold and the nucleic acid sequence of the corresponding staple strand.
[0125] A stepwise top-down approach was demonstrated to generate arbitrary regular or irregular wireframe polyhedral DNA nanostructure origami objects composed of helices with edges that are multiples of two (i.e., 2, 4, 6, etc.) and edge lengths that are multiples of 10.5 truncated to the nearest integer.
[0126] Typically, the pathway of a scaffold nucleic acid is identified by: (i) determining the edges that form a spanning tree of the node-edge network (e.g., using Prim's algorithm); (ii) bisecting each edge that does not form a spanning tree to form two split edges; (iii) determining an Euler circuit that passes twice along each edge of the spanning tree. The orientation of successive scaffold sequences is flipped at the bisection of the node-edge network at the DX-antiparallel crossovers, and the Euler circuit defines a route for single-stranded nucleic acid scaffold sequences passing through the entire structure. In some forms, the spanning tree used to determine the location of scaffold crossovers for scaffold routing is a maximum-width spanning tree. This is important in minimizing the number of staples per object, leading to a more stable / robust structure. However, any spanning tree will lead to valid scaffold routing. In some forms, the method is implemented as a computational tool.
[0127] Given the nanoparticle geometry and scaffold sequence input, the program output is of the staple sequences required to fold the scaffold into the selected nanoparticle. The staple strands are placed at the vertices and edges of the root of the single stranded nucleic acid scaffold sequence determined in (3). In some forms, these staple oligonucleotide sequences have nick positions where either staple strand closes on itself or where two staple strands come together, and the nicked strands are placed away from the center of the object ("outside").
[0128] Exemplary methods for top-down design of nucleic acid nanostructures of arbitrary geometry are described in Venziano et al, Science, 352 (6293), 2016, the contents of which are incorporated by reference in their entirety.
[0129] In other forms, the sequence of the NSO is designed manually or using alternative computational sequence design procedures. Exemplary design strategies that can be incorporated into methods for making and using NSOs include single-stranded tile-based DNA origami (Ke Y, et al., Science 2012), e.g., brick-like DNA origami that includes a single-stranded scaffold with helper strands (Rothemund, et al., and Douglas, et al.), and pure single-stranded DNA that folds onto itself, e.g., in PX-origami, using parallel crossovers.
[0130] Alternative structured NSOs include bricks assembled using DNA duplexes, packaged in square or honeycomb lattices, bricks with holes or cavities (Douglas et al., Nature 459, 414-418 (2009); Ke Y et al., Science 338: 1177 (2012)). Parallel crossover (PX)-origami, in which the nanostructure is formed by folding one long scaffold strand onto itself, can be used instead, provided the bait sequence is still included site-specifically. Further diversity can also be introduced, such as using different edge types, including 6, 8, 10, or 12-helix bundles. Further topologies, such as ring structures, e.g., 6-helix bundle rings, can also be used.
[0131] c. Assembly of nucleic acid nanostructures The method includes the assembly of a single-stranded nucleic acid scaffold and a corresponding staple sequence into an NSO nanostructure having a desired shape and size. In some forms, the assembly is performed by hybridizing the staple to the scaffold sequence. In other forms, the NSO includes only a single-stranded DNA oligo. In further forms, the NSO includes a single-stranded DNA molecule folded onto itself. Thus, in some forms, the NSO is assembled by a DNA origami annealing reaction.
[0132] Typically, annealing can occur according to the specific parameters of the staple and / or scaffold sequences. For example, oligonucleotide staples are mixed in appropriate amounts in an appropriate reaction volume. In preferred forms, the staple strand mixture is added in an amount effective to maximize yield and correct assembly of nanostructures. For example, in some forms, the staple strand mixture is added in a molar excess of the scaffold strands. In exemplary forms, the staple strand mixture is added in a 10-20 fold molar excess of the scaffold strands. In some forms, synthetic oligonucleotide staples with and without tag overhangs are mixed with the scaffold strands and annealed by slowly lowering the temperature (annealing) over a period of 1-48 hours. This process allows the staple strands to guide the folding of the scaffolds into the final NSOs. This can be done in separate wells and added to a pool of NSOs (as in Figures 3A-3D) or added to a pool of oligonucleotides and scaffolds to generate a pool of NSOs. In Figures 3A-3D, exemplary NSOs are shown as tetrahedrons representing arbitrary conserved blocks.
[0133] By using a microfluidic automated assembly device, material usage for assembly can be minimized and assembly can be expedited (Figures 11-12). For example, in one particular configuration, oligonucleotide staples can be added to one inlet and scaffolds to a second inlet, the solutions mixed using methods known in the art, and the mixture moved through an annealing chamber where the temperature is constantly decreased over time or distance, and the output port contains the assembled NSOs for further purification or storage. A similar strategy can be applied to pure single-stranded oligo-based NSOs or single-stranded scaffold origamis using digital droplet-based microfluidics at a surface to mix and anneal solutions in the absence of helper strands.
[0134] 2. SSO Sign One or more specific labels, such as a nucleic acid sequence motif, a unique sequence identifier, or a "tag," are associated with the sequence-controlled polymer on the SSO. For example, in some forms, one or more labels are selected and then encoded into the nucleic acid sequence using a user-selected conversion method.
[0135] Typically, the label is a nucleic acid sequence motif, such as a barcode sequence. In some forms, the label includes a mechanism for direct conversion, including but not limited to a string, an integer, a date, a time, an event, a genre, metadata, a participant, a hash, or an author. In certain forms, the tag allows the user to maintain an external library of addresses, thereby allowing direct selection of sequences.
[0136] Nanostructuring the sequence-controlled polymer blocks allows a natural extension to spatially separate sequence-controlled polymers based on the input signal and associate related sequence-controlled polymers with superblock conservation. The address space is multiplied by the number of tags in use. For example, the method can be used to (k*n) It allows addressing of nucleotides with bases, where n is the number of nucleotides of address per tag, and k is the number of tags. The number of tags per nanostructure can be determined by the user. Typically, each nanostructure has at least one tag, for example, two or more tags, three or more tags, up to 10 tags, 20 tags, 100 tags, or 1,000 tags. In some forms, each edge of the polyhedron has one tag, or two or more tags. In some forms, the SSO has a number of tags that are directly proportional to the size of the polyhedron or that depend on the shape of the polyhedron.
[0137] In some forms, when the nanostructured nucleic acid object is used as an NSO, the label is a nucleic acid sequence associated with the staple sequence in the form of an overhang "tag" sequence. Exemplary overhang sequences are 4-60 nucleotides. In some forms, these overhang tag sequences are placed at the 5' end of either of the staples used to generate the wireframe DNA. In other forms, these overhang tag sequences are placed at the 3' end of either of the staples used to generate the wireframe DNA. In some forms, a combination of overhangs is used to create a logical AND / OR gate to self-assemble the SSO.
[0138] In certain embodiments, the parameters of overhang tag, including size, charge, conformation and sequence, are determined by one or more of user preferences, location on SSO, downstream purification technique, or combinations thereof.Typically, overhang tag sequence contains metadata of scaffold nucleic acid.For example, overhang tag sequence has address for identifying the location of specific sequence control polymer.In some embodiments, each overhang tag contains multiple functional elements such as address and other overhang tag sequence or region for hybridizing to crosslinked strand.
[0139] In some forms, the maximum total number of tags per individual NSO from one overhang is up to 2x (the number of staples in the NSO). For example, one staple has one tag, or two tags, two staples have one tag, two tags, three tags, or four tags, etc. These tag sequences are added to the staple sequence at user-defined positions, and the staple strands without tags are then synthesized individually or directly as a pool using any known method.
[0140] In some forms, the tag is designed to change one or more of the interactions between the tag and the scaffold nucleic acid with which it interacts. In some forms, the nucleic acid sequence of the tag is designed or engineered by adding one or more sequences that change the physical properties of the tag. Exemplary physical properties of the nucleic acid sequence that can be modified include melting temperature or nucleic acid. For example, in some forms, the melting temperature and length of the nucleic acid sequence are controlled such that 1 / 2 of the total length of the sequence, or more than 1 / 2 of the total length, is a hash value, and the remaining half of the sequence is a "homogeneous" sequence that contains a random or non-randomly generated permutation of one type of nucleotide, or two types of nucleotides, or three types of nucleotides, or more than three types of nucleotides. In an exemplary form, the melting temperature and length of the DNA sequence are controlled such that 1 / 2 of the length of the sequence is a hash value, and the remaining half of the sequence is composed of nucleotides with a GC content of 50% and a length of 18mer.
[0141] Other physical characteristics of the tag that can be altered include the secondary structure of the nucleic acid, the ratio of one or more types of nucleotides to one or more other types of nucleotides, or the length, molecular weight, or electrochemical properties of the nucleic acid sequence.
[0142] In other forms, the tag array is a category with discrete values. Exemplary discrete values include any integer value, such as a year, or a collection of integer values, such as a date. In other forms, the tag array encodes some continuous variable, such as shades of blue. In some forms, tags are used partially to store keys and partially to store values, with value-key pairs stored in the tag.
[0143] In some forms, the pool contains different sets of tag overhangs to the same target such that a single sequence-controlled polymer is addressed at many times the number of allowed functional nick positions of the target itself. In some forms, the scaffold polymer is sequence-overlapping with multiple other scaffold messages to allow bioinformatics assembly of long messages that extend beyond the size of the scaffold of a selected geometry.
[0144] 3. Purification of Assembled SSOs The method includes purifying the assembled SSO. Purification separates the assembled structure from the substrates and buffers required during the assembly process. Typically, purification is performed according to the physical properties of the nanostructure, for example, the use of filters and / or chromatographic processes (such as FPLC) according to the size and shape of the nanostructure.
[0145] In an exemplary embodiment, the SSO is purified using filtration, such as centrifugal filtration, or gravity filtration, or diffusion, such as dialysis. In some embodiments, filtration is performed using an Amicon Ultra-0.5 mL centrifugal filter (MWCO 100 kDa).
[0146] C. Storing Information as an SSO The method includes preservation of the SSO structure. The purified SSO can be placed in an appropriate buffer for storage and / or subsequent structural analysis and validation.
[0147] In some embodiments, the SSO is stored in solution. In an exemplary embodiment, the SSO is stored in an aqueous solution. Suitable aqueous storage buffers include PBS, and TAE-Mg 2+ In other forms, the SSO is stored in oil, or in an emulsion, or in other hydrophobic solutions. In some forms, the SSO is dried or dehydrated, for example, by lyophilization. In certain forms, the SSO is dried and fixed to a solid support, such as filter paper.
[0148] Storage can be at room temperature (i.e., 25° C.), 4° C., or below 4° C., e.g., −20° C., −40° C., −80° C. In some forms, the NSOs are frozen, e.g., by immersion in liquid nitrogen.
[0149] In some forms, the SSO is stored under desired long-life conditions. For example, the nucleic acid in the NSO can be maintained with high fidelity for long periods of time. For example, in some forms, the NSO is stored for up to 1 day, more than 1 day, up to 1 week, more than 1 week, up to 1 month, up to 6 months, up to 1 year, more than 1 year, up to 2 years, up to 3 years, up to 5 years, up to 10 years, more than 10 years, up to 20 years, or more than 20 years. Typically, little energy is required for maintenance (Zhirnov, V et al., Nature materials. 15, 366-370 (2016)). Typically, the NSO maintains the fidelity of the information encoded in the nanostructure or encapsulated for a longer period than tape-based storage, which has a life span of 10-30 years.
[0150] The retention of information in DNA has been improved by encapsulating the DNA in silica, with an estimated life of approximately 2,000 years at 10°C and up to approximately 2,000,000 years at -18°C (Grass, RN et al., Angew. Chem. Int. Ed. 54, 2552-2555 (2015)).
[0151] In some forms, the SSO is preserved by chemical means, e.g., encapsulation in silica (SiO2). For example, in some forms, the NSO is preserved by chemical means, e.g., encapsulation in silica (SiO2). Thus, the redundancy of sequence-controlled polymer storage can be used to ensure that copies of the NSO, which may degrade over time in a random manner in which nucleotide identity is lost, can be read out to reconstruct the entire storage. Sequencing errors can also be eliminated by reading multiple copies of the NSO and using consensus sequence mapping. The degradation of a nucleic acid storage object upon exposure to an external stimulus is shown in Figure 16.
[0152] D. Sequence-controlled polymers as SSOs The method allows the assembly of sequence-controlled polymers contained in SSOs. Typically, the assembly of sequence-controlled polymers is carried out by separating, associating, or otherwise dividing a sequence-controlled polymer with or from another sequence-controlled polymer. Thus, in some forms, the method assembles sequence-controlled polymers by associating or separating one or more SSOs. In some forms, the assembly of sequence-controlled polymers is achieved by the physical manipulation of one or more SSOs in a pool of SSOs.
[0153] 1.SSO Superstructure Assembly In some forms, the method groups or otherwise connects sequence-controlled polymers by physically associating two or more SSOs to form an SSO superstructure. Thus, the method allows for the association of a larger set of SSOs. An exemplary superstructure is shown in Figures 5D-5E, in which 10 tetrahedrons are associated together. In an exemplary form, two tetrahedron-conserved objects are associated and four tetrahedron-conserved objects are associated together in a complex SSO dimer and tetramer, respectively, by two complementary overhangs per edge. Such association techniques are not limited to tetrahedrons, i.e., any nucleic acid-conserved object with a larger or smaller set of objects in the superstructure. Associations via staple tags typically include complementary tag sequences, bridging or splint sequences, kissing loops, or hybrid interconnected staple strands, or hybrid interconnected staple strands. In some forms, association occurs based on structural complementarity and non-specific base stacking of DNA duplex ends to form larger scale 1D / 2D / 3D semi-crystalline or crystalline arrays in solution or on surfaces. Typically, buffer conditions and temperature are used to control the aggregation state of such non-specifically associated SSOs.
[0154] i. Complementary tag sequence In some forms, the SSO structures selected by the user for association are assembled such that the tag overhangs of the two objects to be associated are complementary in their nucleotide sequences. Objects with complementary sequences are brought together, the overhang sequences anneal, and the objects form a larger superstructure. Figure 5A shows an exemplary complementary tag interaction between two NSOs.
[0155] ii. Bridging or splinting arrangements In some forms, a bridging or splint oligonucleotide containing a nucleotide sequence complementary to the two overhang sequences is used to bring the two objects together at the two non-complementary tag overhang sequences. This allows for a more dynamic association since the splint strand is added after folding of the individual objects. An exemplary bridging interaction between two NSOs is shown in Figure 5B.
[0156] iii. Interconnecting Staples In a further embodiment, the two SSO structures are assembled using a hybrid staple that acts directly as a staple between the two conserved scaffolds, directly bringing the objects together during folding, in which case the SSOs are stably bound to each other.
[0157] iv. Kissing Loop In a particular embodiment, when scaffolds are mixed together, two SSO structures are assembled using a kissing loop mechanism, where a complementary loop exists in two different conserved objects, and the loop directly connects the two conserved scaffolds.In this method, the two objects are directly brought together after folding.In this case, the SSOs are stably bound to each other.An exemplary kissing loop interaction between two NSOs is shown in Figure 5C.
[0158] 2. Dissociation of SSO superstructure The method includes dissociation of SSO superstructures. Methods for dissociation of superstructure objects include multiple techniques including, but not limited to, changing pH, for example by increasing or decreasing pH, changing salt concentration, increasing temperature, displacement of toe-hold strands, enzymatic release by restriction nucleases, nickases, helicases, resolvases, UV / light sensitive linkers, or any combination thereof.
[0159] This has applications in the context of nucleic acid storage block structures, for example, in creating a superstructure of all objects related to the species H. sapiens by inserting a sequence that aggregates all objects tagged with metadata addressing the species H. sapiens. SSOs may be aggregated in this way using dendritic DNA stars that contain an array of single-stranded overhangs physically associated at a central covalent bond or on a bead.
[0160] Furthermore, re-sorting of supramolecular conserved structures can be performed using nanostructuring data. SSOs associated via splint strands, complementary tag overhangs, or kissing loop interactions can be dissociated by a variety of techniques including changing pH, lowering salt, increasing temperature, displacement of the tow-hold strand, enzymatic release by restriction nucleases, nickases, helicases, resolvases, or any combination of these. Re-association of SSOs then allows for controlled structural modification of the aggregates.
[0161] In the context of connective conservation, this allows the reassembly of new combinations of scaffolds: for example, a superstructure representing an SSO displaying a metadata tag encoding the species H. sapiens can be disassembled and a new SSO superstructure can be reassembled that assembles all NSOs displaying metadata tags encoding human neural DNA.
[0162] The tags from the functionalized staple strands can be modified with new addressing systems and the nanostructure can be refolded with a new set of tagged staples. This allows for a dynamic addressing system without the need to resynthesize all sequence-controlled polymers. Dissociation can also be used to transfer SSOs from one storage block to another based on external signals or cues as mentioned above. A schematic diagram showing a connective nanostructured data framework between a pool of nucleic acid storage objects is shown in Figure 2.
[0163] E. Access to sequence-controlled polymers within SSO The method includes a step of accessing sequence-controlled polymers. For example, nucleic acid sequences can be accessed by selecting one or more SSOs, for example, by selecting a subset of SSOs or SSO superstructures. Typically, the selection of SSOs is performed using a method that selectively captures or removes one or more sequence tags associated with one or more SSOs or a subset of SSOs. Thus, the method provides random access of information. In some forms, the selection is based on the geometry of the SSO, the size of the SSO, the sequence of the SSO, or a combination. In some forms, the nucleic acid and / or nucleic acid structure is bound to a solid phase for use in the selection and purification of SSOs. For example, the nucleic acid can be hybridized to beads such as AMPure XL SPRI beads.
[0164] In some forms, the method for recovering the encapsulated sequence storage object targets one or more populations of interest for recovery from a pool of populations. For example, in some forms, the method recovers an encapsulated sequence storage object comprising one or more populations of interest from a pool of populations, the sequence storage object comprising a molecular tag corresponding to one or more characteristics associated with the population of interest, and the recovery comprises: (i) contacting the molecular tags with a molecular probe that selectively binds to the molecular tags associated with a population of interest; (ii) isolating the sequence conserved objects bound to the probe.
[0165] To enable the recovery of a collection of particles that belong to one of several discrete categories, one orthogonal barcode sequence is associated with each category, and a particle's membership in each category is indicated by the selection of the particle's corresponding barcode. Also described below are various schemes by which barcodes can be assigned to particles to enable the selection of different collections of related particles.
[0166] 1. Geometry Selection In some forms, when nanostructured nucleic acid objects are used as NSOs, the method includes selecting the geometry of the nanostructured NSO. Thus, in some forms, an NSO with a certain geometry is selected from a pool of NSOs with different geometries (FIGS. 7A-7C). For example, in some forms, the geometry determines the location and / or accessibility of one or more tags. In some forms, NSOs that define tags in a certain orientation on the NSO can specifically capture only those NSOs. In certain forms, one or more NSOs or NSO superstructures are selected with a specific sequence and geometry that meets a specific geometric arrangement of complementary strands on a complementary object or receiving object.
[0167] For example, shown in Figures 7A-7C are nanostructured NSOs that display sequences a and b in different geometric locations, such as on two edges. These sequences are complementary to two overhangs on a complementary geometric DNA nanostructure, displaying a' and b' in ideal locations for selection of the NSO. Typically, the larger nanostructure is part of a surface or is attached to a surface or solid support by chemical, hybridization, or protein interactions. In this way, NSOs are specifically selected based not only on the sequences of the tagged overhangs, but also on the geometry of the NSO.
[0168] 2. Sequence-based selection The method includes selecting one or more components of the sequence of the SSO. The mechanism for selectively recovering only the desired portion of the pool (i.e., random access) is implemented by selecting the desired sequence tag of the SSO of interest. Methods for capturing the desired DNA sequence tag are known in the art.
[0169] In some forms, the desired sequence tag is captured by nucleic acid hybridization, where a "bait" sequence is used to select the tag region of the SSO. In some forms, the "bait" sequence is a nucleotide sequence that is complementary to the desired sequence tag. In some forms, the "bait" sequence is a DNA molecule. In other forms, the "bait" sequence is an RNA molecule. In some forms, the hybridization capture is an in-solution approach. In a preferred form, the hybridization capture is a solid-phase (immobilized) approach.
[0170] Exemplary methods for recovering NSO structures of interest from a pool of NSOs are shown in Figures 6A-6C. For example, in some forms, target SSOs in a pool of SSOs can be recovered using tag overhang sequences. In some forms, short single-stranded oligonucleotides are synthesized using known methods with sequences complementary to the tag overhang sequences of the SSOs of interest. Typically, these sequences are synthesized with a label, e.g., a biotin 5' label, that is used to capture these oligonucleotides on a stationary phase. The labeled nucleotides are bound to a fixed support. Exemplary fixed supports include streptavidin-coated beads or streptavidin-coated surfaces. When biotin is used, the nucleic acid captured with the biotin-oligonucleotide is incubated with a streptavidin support to bind (hereinafter, "capture support"). Unbound sequences are removed from the sample, e.g., by washing.
[0171] In an exemplary embodiment, specific capture is achieved by annealing SSO complementary overhang sequences to capture support. The method of specific capture of SSO by annealing includes mixing a pool of SSO with capture support and annealing by, for example, incubating at a temperature between 4°C and the melting temperature of SSO (approximately 55°C), and then cooling to allow annealing. Washing the unbound fraction from the capture support using mild conditions such as slight heating or salt reduction to remove non-specific binding allows specific capture, and then the SSO of interest can be purified from the pool.
[0172] In some forms, the capture sequence is complementary to the key-value pair such that the target address and corresponding storage block are captured, and the target address and corresponding storage block with low Hamming distance are also captured. Methods to increase or decrease this background of storage blocks with similar feature tags can be based on, for example, but not limited to, changes in temperature, pH, capture time, salt. For example, an NSO with a "sky blue" tag can be captured by selection on a "light blue" complementary capture support given certain conditions of capture.
[0173] The captured SSOs are released from the capture support by any mechanism known in the art, including, but not limited to, changing the pH, lowering the salt, increasing the temperature, displacing the toe-hold strand, enzymatic release by restriction nucleases, nickases, helicases, resolvases, or any combination thereof.
[0174] In a further embodiment, a splint strand can be generated that includes a portion of the sequence complementary to the targeted tag overhang and a second portion of the splint sequence complementary to the capture sequence on the capture support, as described for the superstructure in Figures 5A-5C.
[0175] In some forms, capture of SSOs occurs in minimal volumes, for example, using a bulk or surface-based microfluidic device. In some forms, the microfluidic device includes a surface or bead-based oligonucleotide support with sequences complementary to the tag overhang sequences of one or more SSOs. An inlet port provides an aliquot of pooled stored objects and directs them to a stationary phase capture region, allowing separation of captured objects and flow-through objects. In this manner, flow-through (i.e., unbound) objects are captured separately from captured objects (Figures 13A-13G). Prior to manipulation and capture, SSOs are stored in a dry state on paper, or other solid support matrix, for long-term storage before rewetting and manipulation prior to sequencing-based readout.
[0176] a. Fluorescence gate selection Exemplary molecular probes for use in the method for selecting and / or recovering sequence-archived objects include fluorescently labeled probes that selectively bind to the molecular tags associated with sequence-archived objects.Thus, in some forms, the method comprises fluorescent gate selection.For example, in some forms, the method for isolating the sequence-archived objects bound to probe comprises fluorescent gate selection that uses the different colors associated with each probe to identify and recover the population of interest.
[0177] In an exemplary method for recovering encapsulated sequence storage objects, capsules containing B. taurus (containing "Eukaryote", "Animalia", "2021-01-05", and "Bos taurus" labels) and M. musculus (containing "Eukaryote", "Animalia", "2021-01-03", and "Mus musculus" labels) genomes were targeted for recovery from a pool containing H. sapiens total RNA (containing "Eukaryote", "Animalia", "2021-01-03", and "Homo sapiens" labels) and SARS-CoV-2 RNA genomes (containing "Riboviria", "Orthornavirae", "2020-12-20", and "SARS-CoV-2" labels) (see Figure 23A). A Boolean logic query using molecular probes matching the query strings "Eukaryote", "Animalia", and "Homo sapiens" was added to the pool. Fluorescence gate selection using the different colors associated with each probe identifies populations of interest. Selecting populations positive with "Eukaryote" AND "Animalia" selects B. taurus, M. musculus, and H. sapiens. An additional "Homo sapiens" gate can be used to select populations negative for "Homo sapiens" or NOT Homo sapiens in the Boolean logic expression. Thus, the final Boolean logic search query is "Eukaryote" AND "Animalia" AND (NOT "Homo sapiens"), which selects B. taurus and M. musculus, which were verified using quantitative real-time polymerase chain reaction (see Figure 23B).
[0178] B hybridization chain reaction In some forms, the methods also include hybridization chain reaction (HCR). For example, in some forms, the method for isolating probe-bound sequence-stored objects includes hybridization-based selection to probes designed to have distinct hybridization properties with distinct molecular "barcode" tags on the surface of the sequence-stored objects to identify and recover the population of interest.
[0179] In some forms, at least one member of the set of feature tags is defined for hybridization, and at least one member of the set of feature tags has the same number of nucleotides. In some forms, in at least one of the sets of feature tags, (a) the members of the set of feature tags have the same number of nucleotides, and (b) each of the feature tags in the set differs from other feature tags in the set by 1 to x mismatched nucleotides, where the mismatched nucleotides are (i) at least 2 nucleotides from either end of the feature tag, and (ii) separated by at least one matching nucleotide in the feature tag, where x is the number of different nucleotide positions in the feature tag that vary in the set. In some forms, independently of one or more sets of feature tags, each feature tag in the set mismatches all other feature tags in the set by 1 to w nucleotides, where w is an integer between 2 and (y-4)÷2, and y is the number of nucleotides of the feature tags in the set, where the formula (y-4)÷2 is rounded up. In some forms, the sequence controlled storage object further comprises a plurality of different digit tags, the digit tags being present on a surface of the storage object, the digit tags encoding a number.
[0180] Thus, in some forms, the method recovers sequence storage objects containing sequences of interest by hybridization-based selection of barcodes on the sample surface as initiators. In an exemplary method, a capsule containing a "Homo sapiens" tag (e.g., labeled "z" in FIG. 24A) is coupled to a toehold sequence "a" that triggers a hybridization chain reaction (HCR) between two hairpin structures modified with a marker, which may be a dye or a chemical / biochemical tag, as shown in FIG. 24A. * " and the stem sequence "b * " and complementary z * When the marker is a fluorescent tag, enhanced fluorescence is observed in the HCR-amplified capsules compared to capsules hybridized only with the complementary strand containing a single dye, as shown in FIG.
[0181] c. Selection based on numerical range In some forms, the methods involve the selection and / or isolation of sequence archive objects based on or including molecular tags that are "barcodes," where the barcode sequence design process involves a range of some numerical characteristic of the underlying biomolecule / sequence.
[0182] In some forms, the difference in the numerical values that members of a set of related features have or can relate to is proportional to the similarity of the features in the set of related features. In some forms, the number of multiple digits is arbitrarily assigned to the feature that originates from one or more of the different sequence-controlled polymers to which the number of multiple digits corresponds. In some forms, the number of multiple digits is the same as the number of digits of the numerical value of the feature that originates from one or more of the different sequence-controlled polymers, starting from the most significant digit of the numerical value.
[0183] In some forms, each set of digit tags has as many members as the mathematical base in which the multi-digit number is represented. To enable the recovery of a collection of particles belonging to a range of discretized numerical features, one orthogonal barcode array is associated with each possible digit value for each digit of the numerical feature. With this approach, a collection of particles corresponding to any numerical range of the feature can be recovered, as long as this range can be specified by selecting specific digit values at some subset of the numerical digits. For example, in some forms, each possible digit value for each digit place of the numerical feature is associated with a separate orthogonal barcode, enabling the recovery of a range of values by selecting particles having specific digit values at some subset of the numerical digit places.
[0184] As an example, a numeric feature may be represented in base 3, and a collection of particles having barcodes corresponding to numeric values in the range [1000, 1100) may be recovered by selecting particles having barcodes associated with a "1" in the 27th digit and a "0" in the 9th digit, as shown in FIG. 26.
[0185] d. Design of sequence tags for exact and approximate similarity-based retrieval In some forms, the methods also include selection and / or isolation of sequence-stored objects based on or including molecular tags that are "barcodes," where the barcode sequence design process allows for accurate similarity-based retrieval for features whose similarity metric is simple enough to allow accurate isometric embedding from the feature similarity space into a low-dimensional hypercube. For example, in some forms, selection and / or isolation of sequence-stored objects is based on similarity determined by isometric embedding into a low-dimensional hypercube.
[0186] To allow for the recovery of a collection of particles that are similar to each other with respect to continuous or non-discrete features, the barcode sequence is mutated at a small number of carefully selected sites within the sequence. The limited set of mutated variant barcode sequences is represented by a graph G, such as, but not limited to, a hypercube graph. The mutation sites are selected so that the graph G faithfully represents the binding affinity between the barcodes and the complementary sequences for the barcodes used as probes. The similarity space of the continuous features is also represented by a graph H, which is then isometrically embedded in the graph G. For certain simple graphs H, polynomial-time algorithms can be used to find an exact isometric embedding. For any complex graph H, an isometric embedding can be found by first performing a dimensionality reduction on the corresponding metric space represented by H. The dimensionality reduction can be performed using any standard technique that seeks to preserve distances during transformations. The low-dimensional space can then be discretized to approximate the isometric embedding in G. Examples of finding an isometric embedding for both simple and complex cases of H are shown in Figures 27 and 28.
[0187] The term "hypercube", as used herein, refers to an extrapolation of a cube or square to n dimensions. For example, a four-dimensional hypercube is called a tesseract. Thus, an n-dimensional hypercube is also known as an n-cube. It is best depicted and represented in non-Euclidean geometry.
[0188] Thus, in some forms, the method for retrieving the encapsulated sequence storage objects targets one or more populations of interest for retrieval from a pool of populations based on approximate similarity-based retrieval of the target population. The method retrieves the sequence storage objects of interest from the pool of sequence storage objects, where the sequence storage objects of interest optionally include molecular tags corresponding to one or more traits associated with the complex similarity metric.
[0189] i. Barcode design with isometric embedding In some forms, the molecular "barcode" tag associated with the sequence storage object is a nucleic acid sequence that includes or encodes a sequence associated with one or more traits determined by isometric embedding, whereby the isometric embedding directly corresponds to the assignment of a barcode to each particle that enables similarity-based retrieval. Thus, in some forms, the method includes one or more steps for designing the sequence of the molecular "barcode" tag by isometric embedding.
[0190] In some forms, the method designs tags by representing the simple similarity metric as a cyclic graph with "n" nodes that can be embedded exactly isometrically into a 4-dimensional hypercube graph. In an exemplary form, the simple similarity metric is represented as a cyclic graph with 8 nodes that can be embedded exactly isometrically into a 4-dimensional hypercube graph, as shown in FIG.
[0191] A schematic diagram of an exemplary barcode sequence design process that allows approximate similarity-based recovery for features with arbitrarily complex similarity metrics is shown in Figure 28. In an exemplary form, the feature similarity space is simplified and reduced to a small number of dimensions using standard dimensionality reduction. These dimensions are then further approximated by binning, after which it can be directly embedded into a hypercube graph whose nodes represent mutational variants of a set of barcodes.
[0192] In an exemplary method, the process begins with a complex similarity metric, derived for example from 4187 SARS-CoV2 genomes for which pairwise genetic similarity was calculated. This similarity metric was reduced to 18 dimensions using multidimensional scaling (MDS), and for visualization purposes, the dimensionality was further reduced to 2 dimensions before plotting. After binning, linear regression showed a strong correlation between the original similarity metric and the final distance in a 54-dimensional hypercube embedding. The hypercube embedding corresponds to a direct assignment of six barcode sequences to each node in the original feature space.
[0193] Thus, in some forms, a method for designing a molecular barcode tag that correlates with two or more similar features includes: (a) determining a low dimensional feature similarity metric for two or more similar features by pruning a feature similarity space of the two or more similar features; (b) directly embedding the simplified features into a hypercube graph, e.g., a similarity metric correlates with distance in the hypercube embedding to provide correspondingly distinct barcode sequences; (c) generating a barcode sequence tag;
[0194] (a) Simplification of feature similarity space In some forms, the method for designing molecular barcode tags comprises one or more steps for determining the similarity metric of the complex similarity metric of two or more features.An exemplary method for providing the complex similarity between two or more pools of samples comprises determining the feature similarity metric, such as the sequence identity between each member of the pool.In an exemplary form, the collection comprises a library of genomic sequences, for example a library of distinct species, such as a library of viral genomic sequences.The similarity between members of the collection of viral genomic sequences can be evaluated, for example, by their sequence identity to each other.
[0195] In some forms, before mapping the features to which the feature tags correspond, the dimensionality of the features to which the feature tags correspond is reduced.
[0196] Thus, in some forms, the method for designing molecular barcode tags includes one or more steps for simplifying the feature similarity space by dimensionality reduction to provide a feature similarity metric. In some forms, simplifying the feature similarity space includes using standard dimensionality reduction. In certain forms, the similarity metric is reduced using multidimensional scaling (MDS). Typically, the feature similarity space is reduced to a reduced number of dimensions, such as about 2 to about 20 dimensions, inclusive. Thus, in some forms, the similarity coded feature tags of a set of feature tags are similarity coded by reducing the dimensionality of the features to which the feature tags correspond.
[0197] (b) Direct embedding into a hypercube graph In some forms, the reduced-dimensionality features are mapped to a hypercube based on similarity of the reduced-dimensionality features.
[0198] Thus, in some forms, the method includes one or more steps to further approximate the dimensionality by direct binning and embedding into an "n"-dimensional hypercube graph, where nodes represent mutational variants of the set of barcodes, where "n" is an integer less than or equal to the number of features to which the feature tags correspond, and where "n" is a factor of the number of features to which the feature tags correspond. In some forms, the method maps the dimensionality-reduced features into an n-dimensional hypercube, where n is an integer less than or equal to the number of features to which the feature tags correspond, and where n is a factor of the number of features to which the feature tags correspond, based on similarity of the dimensionality-reduced features. In some forms, the method implements a computer system to complete one or more of the steps. For example, in some forms, the mapping is implemented using a computer.
[0199] In some forms, the quality of this mapping may be evaluated by calculating the correlation between the distance in the original similarity metric and the distance in the n-dimensional hypercube after embedding. In some forms, linear regression modeling may be used to calculate this correlation. A high correlation (i.e., close to 1) indicates that the mapping well preserves the similarity between the features described by the original similarity metric. In some forms, the correlation comprises linear regression modeling. Preferably, the hypercube embedding directly corresponds to the assignment of barcode sequences to each node of the original feature space. In some forms, the number of hypercube edges between nodes to which any two of the mapped features are mapped is proportional to the similarity of the two features.
[0200] (c) Generation of molecular barcode tags In some forms, the method includes one or more steps for generating molecular barcode tags according to the assignment of barcode sequences to nodes of an n-dimensional hypercube. A restricted set of barcode sequence variants is generated by mutating at a small number of sites such that the binding affinity between the barcode and its complement (i.e., probe) is accurately represented in the n-dimensional hypercube. This hypercube determines the barcode sequence for each node of the n-dimensional hypercube in (b). Using the mapping determined in (a) and (b), this determines the barcode sequence for each node in the original feature space. The barcode sequences are then associated with the corresponding sequence-controlled polymer to generate tagged sequence-storage objects.
[0201] 3. Boolean Logic In some forms, Boolean logic of AND, OR, and NOT is applied to SSO using the tag overhang sequences described in Figures 8A-8E, 9A-9C, and 10A-10B. These logic applications are complementary. In some forms, these logic applications are applied only once. In other forms, the same logic application is applied multiple times, for example, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 20 times, 30 times, 40 times, 50 times, 100 times, or more than 100 times. Exemplary multiple applications of the same logic are a AND b AND c AND d AND e, etc. In some forms, these logic applications are used in any desired order or combination to create a large set of logical calculations. An exemplary combination is a AND b followed by NOT c. In some forms, these logical applications are used multiple times in any desired order or combination, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, or more than 100 times.
[0202] i.AND logic In some forms, AND logic is applied in the selection and purification of SSOs with two or more overhang tag sequences (Figures 8A-8E). If the targeted SSOs can be separated using AND logic, then an SSO or set of SSOs is purified from the pool of SSOs. For example, the SSO or set of SSOs of interest is purified multiple times, first using a capture support specific to one overhang of interest (i.e., capturing all SSOs with overhang sequence a). Unbound SSOs are then washed away, leaving bound SSOs attached to the capture support, as described in Figures 5A-5C. The captured SSOs are then released from the support by changing the pH, lowering the salt, increasing the temperature, displacing the tow-hold strand, enzymatic release by restriction nucleases, nickases, helicases, resolvases, UV / light-sensitive linkers, or any combination thereof. This pool of SSOs released in the first round is then applied to a second round of purification using a second, distinct set of capture sequences bound to the support. The SSOs are then captured on a second capture support with a distinct capture sequence (i.e., capturing all SSOs in the release pool with overhang sequence b), and unbound SSOs are washed as in Figures 6A-6C. The bound SSOs are then released from the support by changing the pH, lowering the salt, increasing the temperature, displacing the tow-hold strand, enzymatic release by restriction nucleases, nickases, helicases, resolvases, UV or light, or any combination of these. This results in an SSO with overhang sequences a AND b. In some forms, this AND logic purification process is repeated 2, 3, 4, 5, up to 10, or more than 10 times. In some forms, this AND logic purification process is repeated as many times as the number of occurrences of tags on a given object (2 x (number of staples)).
[0203] F. Recovery of sequence-controlled polymers from SSO The method includes recovering the sequence-controlled polymer stored within the sequence-controlled polymer object, e.g., in some forms, the method includes recovering the nucleic acid nanostructure.
[0204] 1. Recovery of sequence-controlled polymers from NSO In some forms, the method for dissociating the NSO into its single-stranded components includes denaturing the NSO. The NSO can be denatured by a change in pH, or temperature. In an exemplary form, the NSO is denatured by melting (FIGS. 11A-11D). The released single-stranded scaffold is purified and amplified by master primer sequences flanking the DNA sequence. The nucleotide sequence is read by any known sequencing method. In some forms, PCR is used to amplify the final selected message. In some forms, PCR is accomplished using a primer set specific to the NSO of interest. In some forms, PCR is performed using a set of "master primers" that are tested to be orthogonal to the sequence. Typically, the subject pool is specifically selected to narrow the pool to only those messages that meet the user's requirements. If all sequence-controlled polymers within the NSO are surrounded by a single set of master primers, then only one PCR reaction is required in the workflow. In some forms, barcode sequences are generated on the nanoparticle and / or microparticle scaffold surface using a DNA synthesizer. The barcode-modified scaffold captures the desired NSO from the target pool. In some forms, the barcode sequences generated on the chip array capture the desired NSO from the target pool for recovery and subsequent PCR amplification.
[0205] i. Sequencing Methods Any known DNA sequencing method can be used. In some forms, the nucleotide sequence is read by a sequencing method including Sanger sequencing (Sanger F et al., Proc. Natl. Acad. Sci. USA 74 (12): 5463-7 (1977)).
[0206] In some forms, the nucleotide sequence is read by Maxam&Gilbert sequencing (Maxam AM et al., Proc. Nat. Acad. Sci. USA 74,560-564 (1977)) or any other chemical method. In other forms, the sequencing is performed by PYROSEQUENCING™. In further forms, the nucleotide sequence is read by single molecule sequencing using exonucleases.
[0207] In some forms, sequencing is performed by next-generation sequencing. Some exemplary technologies include ILLUMINA®, Roche 454 sequencing, Ion torrent:Proton / PGM sequencing, SOLiD sequencing. Some exemplary commercial providers of next-generation sequencing are Pacific Biosciences, ILLUMINA®, Oxford Nanopore Technologies.
[0208] ii. Error Correction DNA synthesis introduces errors into the nucleotide sequence, with an error rate on the order of 1% per nucleotide. Furthermore, long-term storage of NSOs compromises data integrity. In some forms, the means by which NSOs are stored reduce errors by increasing data redundancy or by periodically replicating the NSOs.
[0209] iii. Data Redundancy An important aspect of DNA preservation is to devise a suitable scheme to tolerate errors by adding redundancy. In some forms, errors are tolerated by adding redundancy at the encoding stage. For example, the encoding proposed by Goldman et al. divides the input DNA nucleotides into overlapping segments, giving each segment several-fold redundancy (Goldman N et al., Nature, 494:77-80 (2013)). In some forms, the coding redundancy is built in, either exclusively or by using two payloads to form a third strand, as proposed by Bornholt J et al. (Bornholt, J et al., 21st ACM International Conference on Architectural Support for Programming Languages and Operating Systems.(2016)).
[0210] iv. Replication of NSO In long-term preservation of sequence-controlled polymers by NSO, deamination is the largest cause of information loss in ancient DNA and has the lowest energy barrier (Zhirnov V et al., Nat Mater. 23;15(4):366-70 (2016)). To combat information loss in practical preservation or archival systems, error correction codes are widely used (Kim C et al., IEEE Trans. Consum. Electron. 61, 206-214 (2015)). Fortunately, nucleic acids are easy to copy, reducing the overhead of ECC and thus making error correction the primary driver of data integrity. In some forms, nucleic acids are replicated to numerous physical copies of themselves with high fidelity and low cost.
[0211] III. Database The method can include creating a database.The database can be used to enable or support subsequent analysis of the same or different samples.For example, the database can be used to support the analysis of one or more similar types of samples with similar or different levels of heterogeneity.
[0212] For example, the method can include developing a database of sequence-controlled polymers. The database can be initiated, developed and maintained in any manner known in the art, for example, by using a data system such as a digital computer. In some forms, the sequence-controlled polymers for assembling the database can be accumulated by including a sufficiently large number of samples, for example, by creating a library of nucleic acid nanostructures and / or encapsulated nucleic acid units.
[0213] Typically, the database includes at least two different pieces of data, such as sequences or tags that can be used to identify sequence-controlled polymers or a subset of sequence-controlled polymers. In some forms, the database includes the nucleic acid sequence and / or corresponding barcode of each sequence-controlled polymer object in the pool, for example, corresponding to each SSO in the pool, or a library of SSOs. In some forms, each tag or barcode in the database corresponds to one or more sequences or other features of the sequence-controlled polymer. A database can be developed in which binary barcodes representing the sequences of different sequence-controlled polymers, such as a library of SSOs generated according to the described method, are collected. The database can store binary sequence barcodes corresponding to one or more different pools of objects. For example, the database can include tens, hundreds, thousands, or more of non-contiguous nucleic acid sequences.
[0214] In some forms, the creation of a multiply addressed pool of SSOs will serve as a database for long-term storage of sequence-controlled polymers. Multiple indexes of features allow for highly specific extraction of sequence-controlled polymers based on the features used. Thus, in some forms, the database is searched using features based on nucleic acid sequences complementary to the tags of the SSOs. In some forms, the tags are coded by a known scheme so that an external database is not required to extract SSOs based on metadata. Using direct conversion of metadata to capture sequences, sequence-controlled polymers contained within the solution database of SSOs can be mined as deeply as allowed by the number of tags allowed in a given geometry. General database queries such as PUT, GET, Delete, AND, and OR can be used against the system. Thus, the database of all sequence-controlled polymers of the SSOs can be indexed with various features of sequence-controlled polymers. After the pool of all objects has been probed to capture a particular feature of interest, the particular feature can then be extracted. The use of associative storage allows for a specific collection of records if they meet a set of user-created criteria and given the appropriate signal. For example, all sequence-controlled polymers derived from a given species can be associated into a superstructure.
[0215] IV. Composition The compositions described below include materials, compounds, and components that can be used in the disclosed methods. Various exemplary combinations, subsets, interactions, groups, etc. of these materials are described in more detail above. However, it will be recognized that each of the various other individual and collective combinations and permutations of these compounds that are not described in detail are nevertheless specifically contemplated and disclosed herein. For example, when one or more nucleic acid nanostructures are described and one or more of several permutations of structure or sequence parameters are discussed, each and every combination and permutation of possible structure or sequence parameters is specifically contemplated, unless otherwise indicated to the contrary.
[0216] These concepts apply to all aspects of this application, including, but not limited to, steps in methods of making and using the disclosed compositions. Thus, where there are various additional steps that may be performed, it is understood that each of these additional steps can be performed in any specific form or combination of forms of the disclosed methods, and each such combination should be considered specifically contemplated and disclosed.
[0217] A. Nucleic acid storage objects 1. Nucleic acid sample The nucleic acid for use in the described method can be synthetic or natural nucleic acid.In some forms, the nucleic acid sequence is not a naturally occurring nucleic acid sequence.In some forms, the nucleic acid sequence is a synthetic nucleic acid sequence.In some forms, the nucleic acid nanostructure is not a viral genomic nucleic acid.In some forms, the nucleic acid nanostructure is a virus-like particle.
[0218] Numerous other sources of nucleic acid samples are known or can be developed, and any of them can be used with the described method.In some forms, the nucleic acid used in the described method is a naturally occurring nucleic acid.Examples of nucleic acid samples suitable for use in the described method include genomic samples, RNA samples, cDNA samples, nucleic acid libraries (including cDNA and genomic libraries), whole cell samples, environmental samples, culture samples, tissue samples, body fluids, and biopsy samples.
[0219] A nucleic acid fragment is a segment of a larger nucleic acid molecule. When used in the described method, a nucleic acid fragment generally refers to a cleaved nucleic acid molecule. A nucleic acid sample incubated with a nucleic acid cleavage reagent is called a digested sample. A nucleic acid sample digested using a restriction enzyme is called a digested sample.
[0220] In certain forms, the nucleic acid sample is a fragment or portion of genomic DNA, such as human genomic DNA. Human genomic DNA is available from multiple commercial sources (e.g., Coriell No. NA23248). Thus, the nucleic acid sample can be genomic DNA, such as human genomic DNA, or any digested or cleaved sample thereof. Generally, amounts of nucleic acid between 375 bp and 1,000,000 bp per nucleic acid nanostructure are used.
[0221] 2. Nucleic Acid Nanostructures The basic technique for creating nucleic acid (e.g., DNA) origami of various shapes involves folding a long single-stranded polynucleotide, called a "scaffold strand," into a desired shape or structure using several smaller "staple strands" as glue to hold the scaffold in place. Several variants of geometry can be used to construct the NSO. For example, in some forms, an NSO can be assembled purely from short single-stranded staples, or an NSO containing a purely single-stranded scaffold can be folded onto itself, either of which can take on diverse geometries / structures including wireframe or brick-like objects.
[0222] i. Staple chain The number of staple strands varies depending on the size and shape of the scaffold strand or the complexity of the structure. For example, for relatively short scaffold strands (e.g., about 50-1,500 bases in length) and / or simple structures, the number of staple strands is low (e.g., about 5, 10, 50 or more). For longer scaffold strands (e.g., more than 1,500 bases) and / or more complex structures, the number of staple strands is hundreds to thousands (e.g., 50, 100, 300, 600, 1,000 or more helper strands).
[0223] Typically, the staple strand comprises from 10 to 600 nucleotides, for example from 14 to 600 nucleotides.
[0224] In scaffolded DNA origami, a long single-stranded DNA associates with a complementary short single-stranded oligonucleotide, bringing together two separate sequence spatial portions of the long strand and folding it into a defined shape.Historically, folding of DNA nanostructures has relied on laborious object-by-object design without the selection of a generalized scaffold sequence.
[0225] A robust computational-experimental approach is used to create DNA-based wireframe polyhedral structures of arbitrary scaffold sequence, symmetry and size. These DNA origami objects have several important properties that make them useful for DNA-based storage: 1) any number of faces or edges can be programmed to present outward-facing ssDNA tags that serve either as handles for physical association with other storage blocks or as barcodes on these storage blocks for bead-based or other physical extraction / purification; 2) unlike brick-like origami, they do not nonspecifically associate or aggregate with one another because they lack free duplex ends; 3) they are porous, so small molecules and other single-stranded nucleic acids as well as restriction enzymes and polymerases can diffuse through these storage blocks even when assembled into supramolecular storage blocks; 4) they remain stably folded under moderate ionic strength; and 5) unlike unpaired single-stranded DNA, which nonspecifically associates with itself and with other strands with partial base complementarity, these DNA nanostructure origamis encapsulate single-stranded DNA in a tightly associated, stable form that makes biochemical purification and transport practical.
[0226] ii. Geometry of NSO NSOs are nucleic acid assemblies of any geometric shape. NSOs may be two-dimensional shapes such as plates or any other 2-D shapes of any size and shape. In some forms, NSOs are simple DX-tiles, with two DNA duplexes connected by staples. DNA double crossover (DX) motifs are examples of small tiles (~4nm x ~16nm) that are programmed to generate 2D crystals (Winfree E et al., Nature.394:539-544(1998)), and these tiles often contain patterning features when two or more tiles constitute a crystallographic repeat. In some forms, NSOs are 2-D crystal arrays with parallel double helical domains with sticky ends at each connection site (Winfree E et al., Nature. 6;394(6693):539-44 (1998)). In some forms, NSOs are 2-D crystalline arrays of parallel double-helical domains held together by crossovers (Rothemund PWK et al., PLoS Biol. 2:2041-2053 (2004)). In some forms, NSOs are 2-D crystalline arrays of origami tiles whose helical axes propagate in orthogonal directions (Yan H et al., Science.301:1882-1884 (2003)).
[0227] In some forms, the NSO is a wire-frame nucleic acid (e.g., DNA) assembly of a uniform polyhedron with regular polygons as faces and is equiangular. In some forms, the NSO is a wire-frame nucleic acid (e.g., DNA) assembly of an irregular polyhedron with irregular polygons as faces. In some forms, the NSO is a wire-frame nucleic acid assembly of a convex polyhedron. In some further forms, the NSO is a brick-like square or honeycomb lattice of nucleic acid duplexes of cubes, rods, ribbons, or other linear-like geometries. The corrugated edges of these structures are used to form complementary shapes that can self-assemble by non-specific base stacking. Some exemplary superstructures of NSO include Platonic, Archimedean, Johnson, Catalan, and other polyhedral types. In some forms, the Platonic polyhedron has multiple faces, for example, 4 faces (tetrahedron), 6 faces (cube or hexahedron), 8 faces (octahedron), 12 faces (dodecahedron), 20 faces (icosahedron). In some forms, the NSO is a toroidal polyhedron and other geometries with holes. In some forms, the NSO is a wireframe nucleic acid assembly of any geometric shape. In some forms, the NSO is a wireframe nucleic acid assembly of a non-spherical topology. Some exemplary topologies include nested cubes, nested octahedra, torus, and double torus.
[0228] In a preferred form, a set of tags to be associated with sequence-controlled polymers on the NSO is selected and then encoded into a nucleic acid (such as DNA or locked nucleic acid or RNA) sequence using a user-selected conversion method. In some forms, mechanisms for direct conversion from, including but not limited to, strings, integers, dates, events, genres, metadata, participants, or authors are also included. In a further form, this further includes direct selection of sequences by the user maintaining an external library of addresses.
[0229] B. Encapsulation of sequence-controlled polymers Single-stranded and / or double-stranded DNA, or any other sequence-controlled polymer, can be encapsulated to generate SSOs. These encapsulated acidic sequence-controlled polymer units can also have one or more surface-based molecular identifiers (feature tags) for physical selection and manipulation. Typically, the encapsulated acidic sequence-controlled polymer units are designed for reversibility and recovery of intact encapsulated sequence-controlled polymers, thus allowing sequencing and readout of the sequence-controlled polymers.
[0230] The encapsulated stored objects typically contain one or more feature tags linked to the outside of the coating. The feature tags can be direct or indirect. The particles functionalized with feature tags are pooled and stored for downstream object selection and polymer recovery. In a further embodiment, the feature tags on the surface of the SSO-containing particles are used to select objects using complementary strands to isolate the desired objects from the object pool. The SSOs are released from the particles using a buffered oxide etch. The SSOs can then be processed for decoding and readout.
[0231] 1. Encapsulated sequence-controlled polymer The sequence-controlled polymer to be encapsulated can take any form, such as linear or branched polypeptide, linear or branched carbohydrate, protein, glycosylated polypeptide, linear nucleic acid sequence, two-dimensional nucleic acid object, or three-dimensional nucleic acid object. In some forms, the linear nucleic acid is a base-paired double strand. In other forms, the linear nucleic acid comprises a long continuous single-stranded nucleic acid polymer or a number of such polymers. In further forms, the sequence-controlled polymer to be encapsulated in the same particle is a mixture of one or more of linear or non-linear single-stranded or double-stranded nucleic acid molecules, polypeptides, carbohydrates, proteins, or glycosylated polypeptides. For example, in some forms, one or more single-stranded nucleic acids and one or more scaffold-type nucleic acid nanostructures are encapsulated in the same particle.
[0232] 2. Mounting medium In some forms, sequence-controlled polymers are packaged into discrete SSOs by encapsulation.For example, in some forms, nucleic acid is packaged into discrete NSOs by encapsulation.Suitable encapsulating agents include gel-based beads, protein virus packages, micelles, mineralized structures, siliconized structures, or polymer packages.
[0233] In some forms, the encapsulating agent is a viral capsid, or a functional part, derivative and / or analog thereof. In some forms, the NSO is a virus-like particle, with nucleic acid content wrapped in protein content on the surface. The viral capsid can be derived from retrovirus, human papillomavirus, M13 virus, adenovirus, adeno-associated virus, such as adenovirus 16. In a preferred form, the viral capsid used to encapsulate the NSO does not interfere with the overhang tag, i.e., the overhang tag is accessible for purification.
[0234] In some forms, the encapsulating agent is a lipid that forms a micelle or liposome that surrounds the nucleic acid. In some forms, the micelle or liposome is formed from one or more lipids that can be neutral, anionic or cationic at physiological pH. Suitable neutral and anionic lipids include, but are not limited to, sterols and lipids, such as cholesterol, phospholipids, lysolipids, lysophospholipids, sphingolipids or PEGylated lipids. Neutral and anionic lipids include, but are not limited to, 1,2-diacyl-glycero-3-phosphocholine; phosphatidylcholine (PC) such as phosphatidylserine (PS), phosphatidylglycerol, phosphatidylinositol (PI) (egg PC, soybean PC, etc.); glycolipids; sphingophospholipids such as sphingomyelin and sphingoglycolipids such as ceramide galactopyranoside, gangliosides and cerebrosides (also known as 1-ceramidylglucosides); fatty acids, carboxylic acids, and the like. Sterols containing acid groups, such as cholesterol; 1,2-diacyl-sn-glycero-3-phosphoethanolamines, including but not limited to 1,2-dioleylphosphoethanolamine (DOPE), 1,2-dihexadecylphosphoethanolamine (DHPE), 1,2-distearoylphosphatidylcholine (DSPC), 1,2-dipalmitoylphosphatidylcholine (DPPC), and 1,2-dimyristoylphosphatidylcholine (DMPC). Lipids can also include various natural (e.g., L-α-phosphatidylcholine from tissues: egg yolk, heart, brain, liver, soybean) and / or synthetic (e.g., saturated and unsaturated 1,2-diacyl-sn-glycero-3-phosphocholine, 1-acyl-2-acyl-sn-glycero-3-phosphocholine, 1,2-diheptanoyl-sn-glycero-3-phosphocholine) derivatives of lipids.
[0235] Suitable cationic lipids in micelles or liposomes include, but are not limited to, N-[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethylammonium salt, also referred to as TAP lipids, such as methyl sulfate salts.Suitable TAP lipids include, but are not limited to, DOTAP (dioleoyl-), DMTAP (dimyristoyl-), DPTAP (dipalmitoyl-), and DSTAP (distearoyl-). Suitable cationic lipids in the liposomes include, but are not limited to, dimethyldioctadecylammonium bromide (DDAB), 1,2-diacyloxy-3-trimethylammonium propane, N-[1-(2,3-dioleyloxy)propyl]-N,N-dimethylamine (DODAP), 1,2-diacyloxy-3-dimethylammonium propane, N-[1-(2,3-dioleyloxy)propyl]-N,N,N-trimethylammonium chloride (DOTMA), 1, 2-Dialkyloxy-3-dimethylammoniumpropane, dioctadecylamidoglycylspermine (DOGS), 3-[N-(N',N'-dimethylamino-ethane)carbamoyl]cholesterol (DC-Chol); 2,3-dioleoyloxy-N-(2-(sperminecarboxamido)-ethyl)-N,N-dimethyl-1-propanaminium trifluoroacetate (DOSPA), β-alanylcholesterol, cetyltrimethylammonium bromide (CTAB), diC 14-amidine, N-phenyl-butyl-N'-tetradecyl-3-tetradecylamino-propionamidine, N-(alpha-trimethylammonioacetyl)didodecyl-D-glutamic acid chloride (TMAG), ditetradecanoyl-N-(trimethylammonioacetyl)diethanolamine chloride, 1,3-dioleoyloxy-2-(6-carboxy-spermyl)-propylamide (DOSPER), and N,N,N',N'-tetramethyl-,N'-bis(2-hydroxyethyl)-2,3-dioleoyloxy-1,4-butanediammonium iodide. In one form, the cationic lipid can be a 1-[2-(acyloxy)ethyl]2-alkyl(alkenyl)-3-(2-hydroxyethyl)-imidazolinium chloride derivative, such as 1-[2-(9(Z)-octadecenoyloxy)ethyl]-2-(8(Z)-heptadecenyl-3-(2-hydroxyethyl)imidazolinium chloride (DOTIM), and 1-[2-(hexadecanoyloxy)ethyl]-2-pentadecyl-3-(2-hydroxyethyl)imidazolinium chloride (DPTIM). In one form, the cationic lipid can be a 2,3-dialkyloxypropyl quaternary ammonium compound derivative containing a hydroxyalkyl moiety on the quaternary amine, such as 1,2-dioleyl-3-dimethyl-hydroxyethyl ammonium bromide (DORI), 1,2-dioleyloxypropyl quaternary ammonium chloride (DPTIM). The dipropyl-3-dimethyl-hydroxyethyl ammonium bromide may be dipropyl-3-dimethyl-hydroxyethyl ammonium bromide (DORIE), 1,2-dioleyloxypropyl-3-dimethyl-hydroxypropyl ammonium bromide (DORIE-HP), 1,2-dioleyloxypropyl-3-dimethyl-hydroxybutyl ammonium bromide (DORIE-HB), 1,2-dioleyloxypropyl-3-dimethyl-hydroxypentyl ammonium bromide (DORIE-Hpe), 1,2-dimyristyloxypropyl-3-dimethyl-hydroxylethyl ammonium bromide (DMRIE), 1,2-dipalmityloxypropyl-3-dimethyl-hydroxyethyl ammonium bromide (DPRIE), and 1,2-disteryloxypropyl-3-dimethyl-hydroxyethyl ammonium bromide (DSRIE).
[0236] The lipid may be formed from a combination of two or more lipids, for example, a charged lipid may be combined with a lipid that is non-ionic or uncharged at physiological pH. Non-ionic lipids include, but are not limited to, cholesterol and DOPE (1,2-dioleoylglycerylphosphatidylethanolamine).
[0237] In some forms, the encapsulant is a natural or synthetic polymer.Representative natural polymers include proteins such as zein, serum albumin, gelatin, collagen, and polysaccharides such as cellulose, dextran, and alginic acid.Representative synthetic polymers include polyamides, polycarbonates, polyalkylenes, polyalkylene glycols, polyalkylene oxides, polyalkylene terephthalates, polyvinyl alcohols, polyvinyl ethers, polyvinyl esters, polyvinyl halides, polyvinyl pyrrolidones, polyglycolides, polysiloxanes, polyurethanes, alkyl celluloses, hydroxyalkyl celluloses, cellulose ethers, cellulose esters, nitrocelluloses, polymers of acrylic and methacrylic acid esters, poly[lactide-co-glycolides], polyanhydrides, polyorthoester blends, and copolymers thereof. Specific examples of these polymers include cellulose acetate, cellulose propionate, cellulose acetate butyrate, cellulose acetate phthalate, carboxymethyl cellulose, cellulose triacetate, cellulose sulfate, poly(methyl methacrylate), poly(ethyl methacrylate), poly(butyl methacrylate), poly(isobutyl methacrylate), poly(hexyl methacrylate), poly(isodecyl methacrylate), poly(lauryl methacrylate), poly(phenyl methacrylate), poly(methyl acrylate), poly(methyl meth ... Poly(isopropyl acrylate), poly(isobutyl acrylate), poly(octadecyl acrylate), polyethylene, polypropylene, poly(ethylene glycol), poly(ethylene oxide), poly(ethylene terephthalate), poly(vinyl alcohol), poly(vinyl acetate), poly(vinyl chloride), polystyrene and polyvinylpyrrolidone, polyurethanes, polylactic acid, poly(butyric acid), poly(valeric acid), poly[lactide-co-glycolide], polyanhydrides, polyorthoesters, poly(fumaric acid), and poly(maleic acid).
[0238] In some forms, the encapsulating agent is mineralized (e.g., alginate beads, or calcium phosphate mineralization of polysaccharides). In other forms, the encapsulating agent is siliconized. In one form, the nucleic acid is packaged in an inorganic structure, but has a single-stranded nucleic acid on its surface that serves as an address for association with other NSOs or for selection by Boolean logic.
[0239] In some forms, the encapsulant is a metal oxide particle. Exemplary metal oxide encapsulants include silicon dioxide (SiO2) and titanium dioxide (TiO2), and may be mesoporous, compact, or structured. In some forms, DNA is adsorbed onto the surface of modified metal oxide particles, and then coated with a polymer electrolyte, such as poly(diallyldimethylammonium chloride), poly(acrylamide-co-diallyldimethylammonium chloride), and poly(allylamine hydrochloride).
[0240] 3. Feature tags In some forms, the feature tags are synthesized directly on the encapsulated storage object. In one form, the NSO-containing particles are surface-coated with 9-O-dimethoxytrityl (DMT)-triethylene glycol, 1-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramidite. When the DNA synthesizer is used to generate the feature tags, the modified silica particles are used directly as the solid support for the DNA synthesizer. In other forms, the feature tags are synthesized separately and attached to the surface of the NSO-containing particles using chemical conjugation. For example, in some forms, feature tags are conjugated to stored objects, where the conjugation chemistry includes biotin-avidin recognition pairs, N-hydroxysuccinimide (NHS) coupling, 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) coupling, succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (SMCC)-mediated coupling, sulfo-SMCC coupling, copper-catalyzed azide-alkyne cycloaddition (CuAAC), strain-promoted azide-alkyne cycloaddition (SPAAC), or combinations thereof. The particles functionalized with feature tags are pooled and stored for downstream object selection and polymer recovery. In further forms, feature tags on the surface of SSO-containing silica particles are used to select objects using complementary strands to isolate desired data from the object pool. The SSOs are released from the silica particles using a buffered oxide etch. The SSOs can then be processed for decoding and readout.
[0241] In addition to nucleic acid overhangs, other purification tags can be incorporated into the overhang nucleic acid sequence of any SSO for purification (i.e., recovery of the target). In some forms, the overhang contains one or more purification tags. In some forms, the overhang contains a purification tag for affinity purification. In some forms, the overhang contains one or more sites for conjugation to a nucleic acid, non-nucleic acid molecule. For example, the overhang tag can be conjugated to a protein, or a non-protein molecule, for example, to enable affinity binding of the SSO. Exemplary proteins for conjugating to the overhang tag include biotin and an antibody, or an antigen-binding fragment of an antibody. Purification of antibody-tagged SSOs can be achieved, for example, by interaction with an antigen, and / or protein A, G, A / G, or L.
[0242] Further exemplary affinity tags are peptides, nucleic acids, lipids, sugars or polysaccharides.For example, the overhang contains sugars such as mannose molecules, and then mannose-containing SSOs can be selectively collected using mannose-binding lectins, and vice versa.Other overhang tags can further interact with other affinity tags, for example, any specific interaction with magnetic particles can allow purification by magnetic interaction.
[0243] 4. Nucleic Acid Overhang Tags In some embodiments, the overhang sequence is between 4 and 60 nucleotides, depending on the user's preference and downstream purification techniques. In preferred embodiments, the overhang sequence is between 4 and 25 nucleotides. In some embodiments, the overhang sequence contains 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 nucleotides in length.
[0244] In some forms, these overhang tag sequences are located at the 5' end of either of the staples used to generate the wireframe nucleic acid. In other forms, these overhang tag sequences are located at the 3' end of either of the staples used to generate the wireframe nucleic acid.
[0245] In some forms, the overhang tag sequence contains metadata of the scaffold nucleic acid or encapsulating nucleic acid. For example, the overhang tag sequence has an address for identifying the position of a particular sequence-controlled polymer. In some further forms, each overhang tag contains multiple functional elements, such as an address, and a region for hybridizing to other overhang tag sequences or crosslinking strands. These tag sequences are added to the staple sequence at user-defined positions, and the tagless staple strands are then directly synthesized individually or as a pool using any known method.
[0246] 5. Modifications to nucleotides In some forms, one or more of the nucleotides of the feature tag of the SSO are modified nucleotides. In some forms, one or more of the nucleotides of the scaffold nucleic acid sequence of the NSO are modified nucleotides. In some forms, the nucleotides of the encapsulating nucleic acid sequence of the NSO are modified. In some forms, one or more of the nucleotides of the nucleic acid staple sequence are modified nucleotides. In some forms, the nucleotides of the DNA tag sequence are modified to further diversify the addresses associated with the SSO. Examples of modified nucleotides include, but are not limited to, diaminopurine, S 2T, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5- Methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueuosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid(v), wybutoxocine, pseudouracil, queuosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, and (acp3)w,2,6-diaminopurine. Nucleic acid molecules may be modified at the base moiety (e.g., one or more atoms that are typically available to form hydrogen bonds with a complementary nucleotide, and / or one or more atoms that typically cannot form hydrogen bonds with a complementary nucleotide), the sugar moiety, or the phosphate backbone. Nucleic acid molecules may also contain amine-modifying groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexylacrylamide-dCTP (aha-dCTP), to allow for covalent attachment of amine-reactive moieties, such as N-hydroxysuccinimide ester (NHS).
[0247] Locked nucleic acids (LNAs) are a family of conformationally locked nucleotide analogs that, among other advantages, confer truly unprecedented affinity and extremely high nuclease resistance to DNA and RNA oligonucleotides (Wahlestedt C, et al., Proc. Natl Acad. Sci. USA, 975633-5638 (2000); Braasch, DA, et al., Chem. Biol. 81-7 (2001); Kurreck J, et al., Nucleic acids Res. 301911-1918 (2002)). In some forms, the scaffold DNA is a synthetic RNA-like high affinity nucleotide analog, a locked nucleic acid. In some forms, the staple strand is a synthetic locked nucleic acid.
[0248] Peptide nucleic acid (PNA) is a nucleic acid analogue in which the sugar phosphate backbone of natural nucleic acids is replaced with a synthetic peptide backbone, usually formed from N-(2-amino-ethyl)-glycine units, resulting in an achiral, uncharged mimic (Nielsen, et al., Science 254, 1497-1500 (1991)). It is chemically stable and resistant to hydrolytic (enzymatic) cleavage. In some forms, the scaffold DNA is a PNA. In some forms, the staple strand is a PNA.
[0249] In some forms, combinations of PNA, DNA, and / or LNA are used in the nucleic acid of the NSO. In other forms, combinations of PNA, DNA, and / or LNA are used in the staple strands, overhang sequences, or any nucleic acid component of the SSO.
[0250] V. Devices, Data Structures and Computer Control Data structures are described that are used in or that are created by or that are created from the described methods. A data structure is generally any form of data, information, and / or objects that are collected, organized, stored, and / or embodied in a composition or medium. For example, a set of nucleotide sequences associated with a nucleic acid nanostructure labeled with a specific sequence tag, or sequences stored in electronic form, such as on a RAM or storage disk, is a type of data structure. The described methods, or any part thereof or preparation therefor, may be controlled, managed, or otherwise assisted by computer control. Such computer control may be achieved by computer controlled processes or methods, may use and / or generate data structures, and may use computer programs. Such computer control, computer controlled processes, data structures, and computer programs are contemplated and should be understood as being described herein.
[0251] The methods and general approaches for molecular data storage and computation may be performed using a computer-based system. In some forms, one or all of the method steps are performed following input to a computer. For example, the coded data may include any digital files and folders from a computer. Digital files are coded and / or converted into molecular storage code (e.g., nucleotides, amino acids, polymers, atoms, surfaces). This code is written into physical storage blocks used to store the data. The stored data is associated with a set of address codes to identify the storage blocks. In some forms, the assembly of the storage blocks is performed by one or more automated processes, e.g., controlled by a computer. The addresses attached to the storage blocks (including physical tags, electrostatic or magnetic properties, chemical properties, or optical properties, so that they can be used for subsequent reading, manipulation, selection, and computation) are recorded in one or more databases or files written into the computer. In some forms, the physical placement of the storage blocks with addresses within a pool of other storage blocks for storage and computation may be performed through one or more automated processes, e.g., controlled by a computer. In some forms, physical separation based on physical characteristics, some storage blocks that meet the selection criteria and some that do not, and sorting are performed through one or more automated processes, such as controlled by a computer. Many cycles of this and other selection criteria can be automated or centrally controlled, such as in parallel or serial. The selections and calculations regarding these tags are recorded in one or more files or databases recorded by a computer. In some forms, physical purification and isolation of the selected storage block(s) of interest from the pool are performed through one or more automated processes, such as controlled by a computer.In some forms, the selected storage blocks are read and decoded into a digital format by one or more automated or centrally controlled processes to enable automated retrieval of the data from the pool.
[0252] A. Device In some forms, one or more devices are connected together to facilitate continuous or intermittent flow through the device as a system. In some forms, the assembly of the storage object from the components is performed using an automated device or multiple interconnected devices that are combined to generate a system. An exemplary device or system is a microfluidic device or system. In some forms, the mixing of the sequence-controlled polymer with one or more feature tags and, optionally, one or more encapsulating agents is performed in a microfluidic system.
[0253] Microfluidics, either in the form of traditional two-phase droplets or in the form of electrowetting on dielectric (EWOD) (Nelson and Kim, Journal of Adhesion Science and Technology, 26 1747-1771 (2012)), can be used to combine, separate, or otherwise manipulate specific pools of previously stored objects for computation or processing or storage / retrieval.
[0254] In some forms, automated systems are used to store and retrieve or account for stored objects.
[0255] Storage readout can be performed using either on-chip nanopore-based single molecule sequencing in the case of DNA / RNA, or PCR-based amplification and sequencing in the case of optical approaches, or other analytical chemistry approaches including mass spectrometry that utilizes the charge, size, mass, etc. of the molecules or nanoparticles to read out the information content or molecular composition of the nanoparticles, and affinity or other specific recognition tags used are also applicable to this workflow. The described methods for assembly of nucleic acid storage objects can be performed within a single device. For example, in some forms, assembly of nucleic acid storage objects is accomplished using a device that includes one or more of the following: (a) an inlet for facilitating the entry of one or more components of the nucleic acid storage object, e.g., from an external source; (b) Apparatus for mixing the components, e.g., vortex, shaker, stirring rod, turbulence coil, etc.; (c) a device for annealing the components to form an assembled nucleic acid storage object, e.g., a controllable heat source, a PCR machine, etc., and (d) Apparatus for purifying the assembled nucleic acid storage objects, for example, by affinity chromatography, high pressure liquid chromatography, filtration, etc.
[0256] The disclosed compositions and methods can be further understood through the following numbered paragraphs. 1. (a) one or more different sequence-controlled polymers, and (b) Multiple different feature tags A sequence controlled storage object comprising: The feature tag is present on a surface of the sequence-controlled storage object; each distinct feature tag corresponds to a single feature attributable to one or more of the distinct sequence-controlled polymers; the single feature to which each distinct feature tag corresponds is a feature attributable to one or more distinct ones of the sequence-controlled polymers; the plurality of distinct feature tags collectively correspond to a plurality of features collectively attributable to the plurality of distinct sequence-controlled polymers; A sequence-controlled storage object in which each distinct feature tag is hybridizable and distinct from all other distinct feature tags.
[0257] 2. The sequence-controlled storage object of paragraph 1, wherein each of the plurality of different feature tags is a member of a different set of feature tags, each set of feature tags corresponding to an associated set of features.
[0258] 3. The sequence-controlled storage object of paragraph 2, wherein at least one member of the set of feature tags is a similarity-encoding feature tag.
[0259] 4. The sequence-controlled storage object of paragraph 2, wherein the relative hybridization affinity of feature tags in a set is related to the similarity of the features to which the feature tags in the set correspond, such that feature tags in the set corresponding to more similar features have closer relative hybridization affinity than feature tags in the set corresponding to less similar features.
[0260] 5. The array-controlled storage object of paragraph 3 or 4, wherein similarity-encoded feature tags of a set of feature tags are similarity-encoded by mapping features to which the feature tags correspond into an n-dimensional hypercube based on feature similarity, where n is an integer less than or equal to the number of features to which the feature tags correspond, and n is a factor of the number of features to which the feature tags correspond.
[0261] 6. The array-controlled storage object of paragraph 5, wherein prior to mapping the feature tags to their corresponding features, the features to which the feature tags correspond are reduced in dimensionality, and the reduced-dimensional features are mapped to a hypercube based on similarity of the reduced-dimensional features.
[0262] 7. The array-controlled storage object of paragraph 3 or 4, wherein similarity-coded feature tags of a set of feature tags are (a) reduced in dimensionality to the features to which the feature tags correspond, and (b) similarity-coded by mapping the reduced-dimensionality features to an n-dimensional hypercube based on the similarity of the reduced-dimensionality features, where n is an integer less than or equal to the number of features to which the feature tags correspond, and where n is a factor of the number of features to which the feature tags correspond.
[0263] 8. An array-controlled storage object according to any one of paragraphs 5 to 7, wherein the number of hypercube edges between nodes to which any two of the mapped features are mapped is proportional to the similarity of the two features.
[0264] 9. A sequence-controlled storage object according to any one of paragraphs 2 to 8, wherein at least one member of the set of feature tags is defined for hybridization and at least one member of the set of feature tags has the same number of nucleotides.
[0265] 10. The sequence controlled storage object of any one of paragraphs 2 to 8, wherein in at least one of the sets of feature tags, (a) the members of the set of feature tags have the same number of nucleotides, and (b) each of the feature tags in the set differs from one or two other feature tags in the set by 1 to x mismatched nucleotides, where the mismatched nucleotides are (i) at least 2 nucleotides from either end of the feature tag and (ii) separated by at least one matching nucleotide in the feature tag, where x is the number of different nucleotide positions in the feature tags that vary within the set.
[0266] 11. The sequence controlled storage object of paragraph 9 or 10, wherein, independently for at least one of the one or more sets of feature tags, each feature tag in the set is mismatched with every other feature tag in the set by 1 to w nucleotides, where w is an integer between 2 and (y-4)÷2, and y is the number of nucleotides of the feature tag in the set, and where the formula (y-4)÷2 is rounded up.
[0267] 12. The array-controlled storage object of any one of paragraphs 1 to 11, further comprising a plurality of different digit tags, the digit tags being present on a surface of the storage object, and the digit tags encoding numbers.
[0268] 13. The method according to claim 1, further comprising: The digit tag is present on the surface of the stored object; each of the plurality of distinct digit tags corresponds to a digit value of a different place of the multi-digit number, the number of distinct digit tags included in the storage object being equal to the number of places of the multi-digit number; each of the plurality of different digit tags is a member of a different set of digit tags, each set of digit tags corresponding to a different digit of the plurality of digit number; each set of digit tags having a digit value corresponding to each of the possible digit values of the number of digits to which the set of digit tags corresponds; The array-controlled storage object of any one of paragraphs 1 to 11, wherein each distinct digit tag is hybridizable and distinguishable from all other distinct digit tags in all of the sets of digit tags, and each distinct digit tag is hybridizable and distinguishable from all of the distinct feature tags.
[0269] 14. (a) one or more different sequence-controlled polymers, and (b) a plurality of distinct digit tags, Digit tags present on the surface of the stored object; A sequence controlled storage object comprising: each of the plurality of distinct digit tags corresponds to a digit value of a different place of the multi-digit number, the number of distinct digit tags included in the storage object being equal to the number of places of the multi-digit number; each of the plurality of different digit tags is a member of a different set of digit tags, each set of digit tags corresponding to a different digit of the plurality of digit number; each set of digit tags corresponds to each possible digit value of the number of digits to which the set of digit tags corresponds; A sequence-controlled storage object, wherein each distinct digit tag is hybridizable and distinct from all other distinct digit tags in all of the sets of digit tags.
[0270] 15. The sequence-controlled storage object of paragraph 14, wherein the number of multiple digits corresponds to a feature attributable to one or more of the different sequence-controlled polymers.
[0271] 16. The sequence-controlled storage object of paragraph 15, wherein the feature attributable to one or more of the distinct sequence-controlled polymers is a member of a set of related features, each of the members of the set of related features having or being associated with a different numerical value, the different numerical value corresponding to the level or intensity of the given feature relative to other features in the set of related features, and the number of multiple digits is equal to, proportional to, or the same as the number of given digits of the numerical value of the feature attributable to one or more of the distinct sequence-controlled polymers.
[0272] 17. The sequence-controlled storage object of paragraph 16, wherein the numerical difference that members of the set of related features have or can relate to is proportional to the similarity of the features in the set of related features.
[0273] 18. The sequence-controlled storage object of paragraph 15, wherein multi-digit numbers are arbitrarily assigned to features attributable to one or more of the distinct sequence-controlled polymers to which the multi-digit numbers correspond.
[0274] 19. A sequence-controlled storage object according to any one of paragraphs 15 to 18, wherein the number of multiple digits is the same as a given number of digits, starting from the most significant digit of the numeric value of a feature attributable to one or more of the different sequence-controlled polymers.
[0275] 20. The array-controlled storage object of any one of paragraphs 14 to 19, wherein each set of digit tags has the same number of members as the mathematical base in which the multi-digit number is expressed.
[0276] 21. Further comprising one or more encapsulants; 21. The sequence-controlled storage object of any one of paragraphs 1 to 20, wherein an encapsulating agent coats or encapsulates the sequence-controlled polymer, and the encapsulating reagent can be reversibly removed by chemical or mechanical treatment.
[0277] 22. The sequence-controlled storage object of paragraph 21, wherein the feature tag is contained in one or more of the mounting media.
[0278] 23. The sequence-controlled storage object according to paragraph 21 or 22, wherein the one or more encapsulating agents are selected from the group comprising natural polymers and synthetic polymers, or combinations thereof.
[0279] 24. The sequence-controlled storage object of any one of paragraphs 21 to 23, wherein the one or more encapsulating agents are selected from the group including proteins, polysaccharides, lipids, nucleic acids, inorganic coordination polymers, metal-organic frameworks, covalent organic frameworks, inorganic coordination cages, covalent organic coordination cages, elastomers, thermoplastics, synthetic fibers, or any derivatives thereof.
[0280] 25. At least one of the sequence-controlled polymers is a single-stranded nucleic acid; the nucleic acid is folded into a three-dimensional polyhedral nanostructure comprising two nucleic acid helices joined by either antiparallel or parallel crossovers spanning each edge of the structure; a three-dimensional polyhedral structure is formed from single-stranded nucleic acid staple sequences hybridized to a single-stranded nucleic acid comprising the bitstream data; A single stranded nucleic acid containing bit stream data is routed through an Eulerian cycle of a network defined by the vertices and lines of a polyhedral structure; the nanostructure comprises at least one edge that comprises a double-stranded or single-stranded crossover; The positions of double-stranded crossovers are determined by a polyhedral spanning tree, A staple sequence is hybridized to the vertices, edges and double-stranded crossovers of a single-stranded nucleic acid containing bitstream data to define the shape of the nanostructure; 25. The sequence-controlled storage object of any one of paragraphs 1 to 24, wherein one or more of the staple sequences comprises one or more feature tag sequences.
[0281] 26. The sequence controlled storage object of paragraph 25, wherein the staple strand comprises between 14 and 1,000 nucleotides, inclusive.
[0282] 27. The sequence controlled storage object of paragraph 25, wherein the single-stranded nucleic acid comprises approximately 100 to 1,000,000 nucleotides, inclusive.
[0283] 28. The sequence-controlled storage object of any one of paragraphs 25 to 27, wherein one or more staple strands comprise one or more feature tag sequences at the 5' end, the 3' end, or both the 5' end and the 3' end.
[0284] 29. The sequence-controlled storage object of paragraph 28, wherein the one or more feature tag sequences comprise one or more overhanging oligonucleotide sequences.
[0285] 30. The sequence-controlled storage object of paragraph 28 or 29, wherein one or more feature tag sequences comprise oligonucleotide sequences complementary to one or more feature tag sequences bound to different sequence-controlled storage objects.
[0286] 31. The sequence control storage object of any one of paragraphs 28 to 30, further comprising one or more additional sequence control storage objects attached thereto.
[0287] 32. A method for preserving a desired sequence-controlled polymer as a sequence-controlled storage object, comprising: (a) (i) one or more different sequence-controlled polymers, and (ii) a plurality of distinct feature tags; and (iii) optionally, one or more mounting media; Assembling a sequence control storage object from The feature tag is present on a surface of the sequence-controlled storage object; each distinct feature tag corresponds to a single feature attributable to one or more of the distinct sequence-controlled polymers; the single feature to which each distinct feature tag corresponds is a feature attributable to one or more distinct ones of the sequence-controlled polymers; the plurality of distinct feature tags collectively correspond to a plurality of features collectively attributable to the plurality of distinct sequence-controlled polymers; each distinct feature tag being hybridizable and distinct from all of the other distinct feature tags; (b) storing the sequence control storage object; The method includes:
[0288] 33. 33. The method of paragraph 32, further comprising the step of (c) recovering the desired sequence controlled polymer.
[0289] 34. The method of paragraph 33, wherein recovering the desired sequence-controlled polymer in step (c) comprises isolating one or more sequence-controlled storage objects from a pool of sequence-controlled storage objects.
[0290] 35. The method of paragraph 34, wherein the selection is determined by the sequence of one or more feature tags on the sequence-controlled storage object, the shape of the sequence-controlled storage object, the affinity to a functional group bound to the sequence-controlled storage object, or a combination thereof.
[0291] 36. The method of paragraph 35, further comprising the step of modifying the isolated sequence control conserved object by adding one or more distinct feature tags.
[0292] 37. The method of paragraph 36, wherein adding one or more different feature tags comprises refolding or reassembling the sequence-controlled storage object with one or more oligonucleotides that include the different feature tags.
[0293] 38. The method of paragraph 37, wherein one or more sequence-controlled storage objects are isolated from the pool of sequence-controlled storage objects using Boolean logic.
[0294] 39. The method of paragraph 38, further comprising removing one or more sequence-controlled stored objects from the object pool using Boolean NOT logic.
[0295] 40. 40. The method of any one of paragraphs 32 to 39, further comprising the step of: (f) accessing a desired sequence controlled polymer.
[0296] 41. The method of any one of paragraphs 32 to 40, wherein the step of preserving the sequence-controlled storage object in step (b) further comprises one or more of dehydrating, lyophilizing, or freezing the sequence-controlled storage object.
[0297] 42. The method of paragraph 41, wherein the step of storing the sequence-controlled storage object in step (b) further comprises one or more of rewetting or thawing the sequence-controlled storage object for processing.
[0298] 43. The method of any one of paragraphs 32 to 42, wherein the step of storing the sequence-controlled storage object comprises storage in a matrix selected from the group including cellulose, paper, microfluidics, bulk 3D solution, on a surface using electric forces, on a surface using magnetic forces, encapsulated in inorganic or organic salts, and combinations thereof.
[0299] 44. The method of any one of paragraphs 32 to 43, wherein the step of storing the array-controlled storage object in step (b) further comprises digitally processing the droplet containing the array-controlled storage object.
[0300] 45. A method for automating the assembly of sequence-controlled storage objects according to any one of paragraphs 1 to 31, comprising using a device having a flow, the device comprising: (a) a means for adding a component of a sequence-controlled storage object to a flow; (b) a means for mixing the components, comprising: a means for mixing operably connected to the means for adding to the flow; (c) a means for annealing the components to form an assembled sequence-controlled storage object, a means for annealing operably connected to the means for mixing; and (d) a means for purifying the assembled sequence control archive object, the means for purifying is operably connected to the means for annealing, A method comprising:
[0301] 46. (e) means for introducing a mounting medium that preserves the sequence control entities; (f) a means for introducing a plurality of feature tags resulting from a sequence-controlled polymer; (g) means for selecting an encapsulated sequence control object from the object pool, the selecting means being operable using Boolean logic; and (h) means for removing the mounting medium to retrieve the sequence control archive object; 46. The method of paragraph 45, further comprising:
[0302] The invention will be further understood with reference to the following non-limiting examples. EXAMPLES
[0303] Example 1 An overview of the sample collection, nucleic acid extraction, nucleic acid encapsulation, nucleic acid storage and recovery processes is shown in Figure 22.
[0304] In one example, as shown in FIG. 23A, -1 Add a 10 µL volume of Bos taurus nucleic acid to a LoBind Eppendorf tube containing 900 µL of nuclease-free water. Then add a 10 µL volume of 50 mg mL -1 of trimethylammonium modified silica particles are added and mixed gently by inverting the tube several times. Then, trimethyl-3-trimethoxysilyl chloride and tetraethoxysilane are added and mixed on a thermal mixer at room temperature for 4 days. Upon completion of encapsulation, the mixture is centrifuged and the pellet is washed five times with ethanol. The pellet is resuspended in 900 μL of ethanol while vortexing and 50 μL of 3-(2-aminoethylamino)propyldimethoxymethylsilane is added. The mixture is mixed on a thermal mixer at room temperature for 24 hours. After surface modification of the encapsulated nucleic acid, the mixture is centrifuged and the pellet is washed five times with dimethylformamide. The pellet is resuspended in 900 μL of dimethylformamide while vortexing and 10 mg mL -1 50 μL of 2-azidoacetic acid N-hydroxysuccinimide (NHS) ester was added. The mixture was again mixed in a thermal mixer at room temperature for 24 h. After azide conversion, the mixture was centrifuged and the pellet was washed five times with dimethylformamide. The pellet was resuspended in 900 μL of dimethylformamide with vortexing and diluted to 100 mg mL -110 μL of dibenzocyclooctyne-PEG13-NHS hydroxysuccinimide ester was added. The mixture was mixed in a thermal mixer at room temperature for 4 hours. After PEG modification, the mixture was centrifuged and the pellet was washed five times with dimethylformamide. The pellet was resuspended in 200 μL of dimethylformamide while vortexing. A volume of 800 μL containing 0.050 M phosphate buffer and 6 μM of each amine-modified barcode assigned the labels "Eukaryote" (AACGATTGTTATGCCCCTAACTCAG) (SEQ ID NO: 4), "Animalia" (ATGGACGACTTGGGACGGGTATCAA) (SEQ ID NO: 5), "Bos taurus" (TAATGTGGCTTGGCTCACCGCTAGG) (SEQ ID NO: 6), and "2021-01-05" (CGATGTAGTCATCCCGATGTGCTGG) (SEQ ID NO: 7) was added. The mixture was again mixed in a thermal mixer at room temperature for 24 hours. After molecular barcode labeling, the mixture was centrifuged and the pellet was washed five times with 1× PBS containing 0.1% Tween®-20.
[0305] The exact encapsulation and barcoding procedure is repeated for additional samples, after which all encapsulated samples are pooled to form a molecular library (see, for example, Figures 23A-23B). Querying of a molecular database proceeds with the addition of probes containing chemical or biochemical markers used for downstream sorting. The sample is released by the addition of hydrofluoric acid. Desalting using spin columns removes excess salts, and the sample is now ready for a subsequent sequencing or amplification reaction.
[0306] Example 2 In another example, emulsions are used to encapsulate samples instead of synthetic or biological polymers. The aqueous phase sample, which may contain water-soluble monomers or cross-linking polymers, is made into droplets in oil containing surfactants using a microfluidic or millifluidic approach (Figures 25A-24C). Polymerization and cross-linking reactions are carried out until all monomers are used up. The emulsion is broken after polymerization and barcodes are chemically / biochemically affixed to the surface of the capsules through the non-terminal ends of the polymers.
[0307] As an example, 1 million copies of the SARS-CoV-2 RNA genome dissolved in nuclease-free water containing 2 mM Ca2+ and 2% (w / w) low-viscosity alginate were flowed through a channel connected to a T-junction in which surfactant-containing oil was flowed. Methylene blue was added to the aqueous phase and droplet formation was visualized in real time (Figure 25C).
[0308] Example 3 In another example, encapsulation and barcoding of samples are performed in a single step using multistage microfluidics (Figures 25A-25B). An aqueous phase containing nucleic acids is flowed through an oil containing surfactants and water-insoluble monomers, a crosslinker, and a polymerization initiator. The droplets are passed through another aqueous fluid stream containing barcodes labeled with chemical / biochemical handles for attachment to the non-terminal ends of the polymers. The reaction is allowed to proceed until the encapsulation polymerization is complete.
[0309] Example 4 In another sample, an isothermal chemical / biochemical amplification method can be used to select the encapsulated sample from the solution. A probe strand containing a trigger sequence or modified with a biochemical catalyst or cofactor is hybridized to the sample containing the desired barcode. Molecular labels, including but not limited to dyes and chemical / biochemical affinity tags, are amplified to improve the sorting efficiency of the proposed system.
[0310] Example 5 Design of superstructures for nucleic acid storage objects method Superstructuring by complementary overhangs was tested using two tetrahedrons. 3' single-stranded DNA overhangs of two different stapled nicks on the same edge of a tetrahedron with an edge length of 63 nucleotides were generated along with a scaffold of sequences amplified from the genomic DNA of M13 phage. Complementary sequences were created to the two overhangs on the first tetrahedron (tet-A) and placed as 3' single-stranded DNA overhangs of two different nicks on the same edge of a second tetrahedron (tet-B) whose scaffold was also amplified from M13 genomic DNA. These two structures with complementary overhangs were folded and purified separately, then pooled and slowly annealed from 43°C to 25°C for 2 hours. Validation of superstructuring was performed by gel shift mobility assay on 2% agarose and visualized under UV light with SYBR Safe DNA stain. The gel showed a shift indicative of quantitative dimer formation. This exact same procedure is used for superstructuring NSOs by using complementary strands per edge. Furthermore, a series of four tetrahedrons were structured with two overhangs per edge complementary to a second tetrahedron that has a second set of two overhangs complementary to the second set of dimers on the other side of its edge. The two tetrahedral dimers then annealed to each other to form a tetramer of tetrahedrons (shown in Figures 18B-18D). The same scaffold sequence was used to form a set of tetrahedrons of the same scaffold, but with different addresses, with curvature into a superstructure where the four tetrahedrons close back on themselves. Thus, NSOs can be assembled into elongated or closed superstructures depending on the addresses exposed.
[0311] result To demonstrate superstructuring of NSOs, NSOs were forced together at their vertices, along their edges, or at their faces using overhang addressing. Exemplary tetrahedra were shown to come together in larger superstructures by gel mobility shift assays, indicative of superstructuring, compared to monomeric, dimeric, and tetrameric NSOs, respectively. Extended tetramers were addressed to come together along their edges by complementarity, exhibiting extended configurations, as determined by transmission electron microscopy. The same tetrahedra but with different addresses were observed to form different compact configurations.
[0312] Example 6 Preservation of nucleic acid object structures on paper method Storage of NSO on paper was tested as a medium for long-term retention. Whatman paper type 42 was cut to mm scale (typically, 2 mm x 5 mm) and saturated with 15 μL of 1xTAE + 12 mM MgCl2 + 1% PEG 8000 w / v. The paper was then vacuum dried in the presence of a desiccant. After that, 15 μL of 40 nM DNA nanostructures (tetrahedrons with edge length of 63 nucleotides) were added to the paper and dried under vacuum. After at least 14 h at room temperature, the paper was transferred to another tube, washed with 15 μL of folding buffer, and the solution was separated from the paper by centrifugation. A gel mobility shift assay showed the stability of the structure. Similarly, NSO can be stored for long periods and resuspended as needed.
[0313] result The NSO was dried and stored on paper that was pretreated with 1% polyethylene glycol 8000 prior to exposure to the NSO. The NSO transferred to the paper was later rewetted and still present in assembled form as shown by gel shift assay. An exemplary paper tab containing dried NSO was stored in an Eppendorf tube.
[0314] Example 7 Preservation of nucleic acid storage target structures on metal oxides Experiments were performed to demonstrate the packaging and accessibility of nucleic acids by encapsulation or coating in non-nucleic acid polymers. Briefly, nucleic acids were housed within a polymer and addressed with one or more tags (shown in Figures 4A-4D and Figures 17A-17D).
[0315] Methods and Materials Preparation of silica particles Silica particles were prepared by mixing 800 μL of 25% w / w ammonium hydroxide, 800 μL of tetraethoxysilane, and 500 μL of distilled water in 18 mL of water. The mixture was shaken at 500 rpm for 6 hours on a platform orbital shaker at room temperature. The mixture was then centrifuged at 9,000 g for 20 minutes at room temperature and the supernatant was discarded. The silica pellets were redispersed in the solution by adding a total of 20 mL of isopropanol, followed by sonication at room temperature for 1 minute and vortexing for 5 seconds to obtain a homogenous colloidal solution. The mixture was again centrifuged at 9,000 g for 20 minutes at room temperature and the supernatant was again discarded. The pellets were redispersed in the solution by adding a total of 4 mL of isopropanol, sonication for 1 minute, and vortexing for 5 seconds until a homogenous dispersion was obtained again.
[0316] Modification of silica particles to enhance adsorption of DNA particles A 1 mL aliquot of silica particles was taken and the silica particles were immediately modified by adding 10 μL of 50% w / w N-trimethoxylsilylpropyl-N,N,N-trimethylammonium (TMAPS) chloride in methanol. The mixture was shaken at 500 rpm for 12 h at room temperature on a platform orbital shaker. The mixture was then centrifuged at 21,500 g for 4 min and the supernatant was discarded. The modified silica pellet was suspended in 1 mL of isopropanol, sonicated for 1 min, and vortexed for 5 s to obtain a homogenous solution. The mixture was again centrifuged at 21,500 g for 4 min and the supernatant was again discarded. The same washing procedure was repeated twice to remove any residual TMAPS in the solution.
[0317] Encapsulation of DNA particles 50 μg mL -1 Double crossover (DX) tiles modified with Cy3 and Cy5 energy transfer pairs as readouts were encapsulated by adding 320 μL of Cy3 and Cy5 modified DX tiles to 700 μL of water and 35 μL of functionalized silica particles (Figure 17D). The mixture was shaken in a microtube revolver for 3 min at room temperature, centrifuged at 21,500 g for 4 min, and the supernatant was discarded. The silica pellet was then suspended in 1 mL of DNAse-free water, sonicated at room temperature for 1 min, and vortexed for 5 s. The mixture was then centrifuged at 21,500 g for 4 min, and the supernatant was discarded. The silica pellet was resuspended in 500 μL of DNAse-free water, sonicated at room temperature for 1 min, and vortexed for 5 s. A volume of 0.5 μL of TMAPS was added to this mixture and mixed by vortexing for 5 s. An additional 0.5 μL of TEOS was then added. The mixture was shaken on a microtube revolver for 4 hours at room temperature, after which 4 μL of TEOS was added. The mixture was further shaken on a microtube revolver for 4 days. The mixture was centrifuged at 21,500 g for 4 minutes, and the supernatant was discarded. The silica-encapsulated DX-tile pellet was resuspended in 500 μL of DNAse-free water, sonicated at room temperature for 1 minute, and vortexed for 5 seconds. The mixture was centrifuged again at 21,500 g for 4 minutes, and the supernatant was discarded. The pellet was resuspended in 100 μL of DNAse-free water, sonicated at room temperature for 1 minute, and vortexed for 5 seconds. The DX-tile is finally encapsulated. A schematic diagram of silica encapsulation of the nucleic acid storage block is shown in Figures 17A-17D.
[0318] The protection of silica particles with DNA was tested by dropping the encapsulated particles onto paper. A volume of 10 μL was dropped onto the paper and allowed to dry at ambient temperature. A volume of 10 μL of DNA denaturant (0.1 M HCl, 0.1 M NaOH, and DNAse) was then added and allowed to dry again at room temperature.
[0319] result The surface of the silica particles is modified so that it can adsorb DNA storage objects, and the modified silica particles act as scaffolds for binding of the nucleic acid storage blocks.
[0320] The nucleic acid storage block is first adsorbed onto a surface-modified silica particle, and then a secondary silica shell is added onto the silica adsorbed with the nucleic acid storage block. A schematic diagram of an exemplary DNA assembly (double crossover or DX tile) containing an energy transfer pair of Cy3 and Cy5 as a readout for monitoring the structure of the DX tile is shown in Figure 17E. This shell provides environmental protection for the nucleic acid storage block.
[0321] The encapsulated particles were evaluated by comparing silica encapsulated particles with non-encapsulated nanoparticles under UV illumination using a long pass filter to filter out only Cy5 fluorescence. There was no change in the emission spectrum of the DX tiles upon completion of the encapsulation step, indicating that the encapsulation process does not affect the structure of the DX tiles (see Figure 17F).
[0322] To evaluate the protection of DNA storage objects by the silica encapsulation process, silica-encapsulated DX tiles were adsorbed onto paper strips and exposed to 0.1 M NaOH, 0.1 M HCl, and DNAse. The silica-coated paper was excited at 400 nm and emission was selected using a 650 nm long-pass filter.
[0323] Example 8 Microfluidic devices for the automated assembly of nucleic acid storage object structures Methods and Materials A system for automated assembly of nucleic acid storage objects was designed and assembled, which includes a 3D printed device with a size of 10 cm × 4 cm, three input ports, a mixer and annealing device on a copper plate, and three output ports, with one leg of the copper plate in a water bath at 80 °C and the other leg of the copper plate in ice water.
[0324] The input port is connected to a fluid pump, and the output is connected to a fraction collection tube, and the flow of the fluid first goes from the reagent containing the scaffold nucleic acid, the tagged staple strand and the staple into the mixer, then into the annealing device, through the annealing device into the fraction collector. In the annealing device, the fluid passes from high temperature to low temperature. The fractions are collected and purified by filtration.
[0325] The DNA nanoparticle annealing reaction in the autoassembler was realized in Tris-Acetate EDTA-MgCl2 buffer (40 mM Tris, 20 mM Acetate, 2 mM EDTA, 12 mM MgCl2, pH 8.0) with a concentration of 80 nM of ssDNA scaffold and 15-fold excess of staple strands in a reaction volume of 1.2 mL. Prior to sample injection, the device was washed with 4 mL of folding buffer at a flow rate of 100 μL / min. For sample injection, the flow rate was maintained at 10 μL / min through the autoassembler channels using a MINIPULS®3 peristaltic pump from Gilson, Inc. A temperature gradient in the autoassembler was created by connecting one end of the copper plate (denaturation region) to a water bath at 80°C and the collection end of the copper plate to a cold water bath kept at 4°C. Sample collection was monitored periodically using a nanodrop. A schematic of the automation system is shown in Figure 12. Exemplary workflows for the implementation of an automated system within an exemplary microfluidic device are also shown in FIGS.
[0326] The output from the autoassembler was gel checked on a 1% agarose gel supplemented with 12 mM MgCl2.
[0327] result The resulting nanostructure assemblies were assessed by gel electrophoresis, and folding of the assembled objects was determined by visual observation of gel bands in each lane of the gel corresponding to scaffold nucleic acid alone, scaffold mixed with staples at room temperature, scaffold mixed with staples and annealed in a thermal cycler for 3 hours, and scaffold mixed with staples and annealed in an autoassembler for 3 hours.
[0328] Folding was tested using a gel shift assay. The lane corresponding to scaffold and staple mixed and annealed in a thermal cycler for 3 hours was in the same position and intensity as the gel lane corresponding to scaffold and staple mixed and annealed in an autoassembler for 3 hours. This experiment demonstrated that the efficacy of the autoassembly system is at least as efficient as assembly using a thermal cycler.
Claims
1. A method for preserving a biomolecule, comprising: (a) assembling a storage object comprising the biomolecules, the assembling comprising encapsulating the biomolecules in one or more encapsulating agents in an emulsion; the one or more encapsulants comprise a polymer selected from the group consisting of a polysiloxane polymer, a polyurethane polymer, a polyacrylate polymer, a polyester polymer, a poly(methyl methacrylate) polymer, a polystyrene polymer, and combinations thereof; (b) storing the storage object for at least one week; A method comprising:
2. The method of claim 1, wherein the encapsulating in (a) comprises performing a polymerization reaction or a crosslinking reaction on the one or more encapsulating agents or precursors thereof in the emulsion to encapsulate the biomolecules.
3. The method of claim 1, wherein the solvent phase in the emulsion contains the one or more encapsulating agents.
4. The method of claim 1, wherein the polymer comprises a metal.
5. The method of claim 1, wherein the polymer comprises a polysiloxane polymer.
6. The method of claim 1, wherein the polymer comprises silicon.
7. The method of claim 2, wherein the encapsulating in (a) comprises polymerizing monomers of the one or more encapsulating agents in the emulsion.
8. The method of claim 7, wherein the monomer is water-soluble.
9. The method of claim 7, wherein the monomer is water-insoluble.
10. The method of claim 1, wherein the emulsion is in an emulsion reaction vessel of millimeter to nanometer size.
11. The method of claim 1, wherein the encapsulating in (a) further comprises capturing the biomolecule in the emulsion using electricity or photons.
12. The method of claim 1, further comprising generating the emulsion before or during (a).
13. The method of claim 12, wherein generating the emulsion comprises pouring an aqueous phase containing the biomolecules into an oil containing a surfactant.
14. The method of claim 13, wherein the oil further comprises a cross-linking agent or a polymerization initiator.
15. The method of claim 1, wherein (b) includes storing the object to be preserved in the emulsion.
16. The method of claim 1, further comprising breaking down the emulsion prior to said storing in (b).
17. The method of claim 1, wherein the preservation in (b) includes preserving the object to be preserved in an organic solution.
18. The method of claim 1, wherein the preservation in (b) includes storing the object to be preserved in a hydrophobic solution.
19. The method described in claim 1, wherein the preservation in (b) includes preserving the object to be preserved in an aqueous solution.
20. The method of claim 1, wherein the preservation in (b) includes dehydrating the object to be preserved.
21. The method of claim 1, wherein the preservation in (b) includes drying the object to be preserved.
22. The method of claim 1, wherein the preservation in (b) includes preserving the object to be preserved for longer than one week.
23. The method described in claim 1, wherein the storage in (b) includes storing the object to be stored at room temperature.
24. The method of claim 1, further comprising releasing the biological molecule from the preserved object after the preservation in (b).
25. The method of claim 24, further comprising sequencing the biomolecule following release from the storage object.
26. The method of claim 1, wherein the biological molecule comprises a nucleic acid, a protein, or a carbohydrate.
27. The method described in claim 26, wherein the biological molecule is a nucleic acid, and the method further comprises a step of amplifying the nucleic acid after the storage in (b).
28. The method of claim 1, wherein the biological molecule is derived from a tissue sample.
29. The method of claim 1, wherein the preserved object further comprises a cross-linking agent.
30. The method of claim 2, wherein the encapsulating in (a) includes cross-linking monomers of the one or more encapsulating agents in the emulsion.
31. The method of claim 1, further comprising: (c) dissociating the preserved object.
32. The method of claim 31, wherein the dissociating in (c) comprises a change in pH, a change in salt concentration, a change in temperature, application of electromagnetic radiation, or application of an enzymatic agent.
33. The method of claim 31, wherein the dissociating in (c) comprises application of a chemical agent.
34. The method of claim 31, wherein the dissociation in (c) occurs via the release of a UV-sensitive linker.
35. The method of claim 1, wherein the one or more encapsulants are configured to be reversibly removed by chemical or mechanical treatment.
36. The method of claim 1, wherein the encapsulation in (a) further comprises physically contacting the biomolecule with the one or more encapsulating agents.
37. The method of claim 36, wherein the one or more encapsulating agents coat the biomolecules.
38. The method of claim 36, wherein the biomolecule is coupled to a surface and the encapsulation in (a) further comprises coating the surface with the one or more encapsulating agents.
39. A biomolecule, one or more encapsulating agents that encapsulate the biomolecules, the one or more encapsulating agents comprising a polymer selected from the group consisting of a polysiloxane polymer, a polyurethane polymer, a polystyrene polymer, and combinations thereof; wherein the one or more encapsulating agents are crosslinked via one or more crosslinking agents.
40. The preserved object described in claim 39, which is in an emulsion.
41. The preserved object described in claim 39, wherein the polymer contains silicone.
42. The preserved object of claim 39, which can be dissociated via a change in pH, a change in salt concentration, a change in temperature, application of electromagnetic radiation, or enzymatic release.
43. The preserved object of claim 39, which can be dissociated through the application of a chemical agent.
44. The preserved object described in claim 39, wherein the biological material comprises a nucleic acid, a protein or a carbohydrate.
45. A preserved object as described in claim 39, wherein the one or more encapsulating agents are configured to be reversibly removed by chemical or mechanical treatment.
46. A preserved object as described in claim 39, wherein the one or more encapsulating agents coat the biological molecules.
47. A preserved object as described in claim 39, wherein the biological molecule is coupled to a surface coated with the one or more encapsulating agents.