Nucleic acid-based data storage
Patent Information
- Application Number
- JP2019545334
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-02-22
- Filing Date
- 2018-02-22
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2038-02-22
Smart Images

Figure 0007926685000022 
Figure 0007926685000023 
Figure 0007926685000024
Abstract
Description
Technical Field
[0001] <Cross Reference> This application claims the benefit of U.S. Provisional Application No. 62 / 462,284 filed on February 22, 2017, which is hereby incorporated by reference in its entirety into the present specification. <Sequence Listing> This application contains a Sequence Listing submitted electronically in ASCII format, which is hereby incorporated by reference in its entirety into the present specification. The above-mentioned ASCII copy created on February 20, 2018 has the file name 44854-738_601_SL.txt and a size of 8,636 bytes. Background Art
[0002] Biomolecule-based information storage systems (e.g., DNA-based) have large storage capacity and stability over time. However, there is a need for scalable, automated, highly accurate, and highly efficient biomolecule systems for information storage. In addition, there is a need to protect the security of such information.
[0003] <Incorporation by Reference> All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference herein to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention
[0004] A method for storing information is provided herein, comprising: (a) receiving at least one item of information in the form of at least one digital sequence; (b) receiving instructions for selecting at least one form of bioencryption, wherein the form of bioencryption is enzyme-based, electromagnetic-based, chemical-based, or affinity-based; (c) converting at least one digital sequence into a plurality of oligonucleotide sequences based on the selected form of bioencryption; (d) synthesizing a plurality of oligonucleotides encoding the oligonucleotide sequence; and (e) storing the plurality of oligonucleotides. A method for storing information is further provided herein, wherein enzyme-based bioencryption includes CRISPR / Cas-based bioencryption. A method for storing information is further provided herein, wherein enzyme-based bioencryption includes instructions for synthesizing enzyme-sensitive oligonucleotides as shown in Table 1. Methods for storing information are further provided herein, wherein electromagnetic-based bioencryption includes instructions for the synthesis of oligonucleotides sensitive to electromagnetic wavelengths from about 0.01 nm to about 400 nm. Methods for storing information are further provided herein, wherein chemical-based bioencryption includes instructions for the synthesis of oligonucleotides sensitive to the administration of gaseous ammonia or methylamine. Methods for storing information are further provided herein, wherein affinity-based bioencryption includes instructions for the synthesis of oligonucleotides sensitive to sequence tags or affinity tags. Methods for storing information are further provided herein, wherein affinity tags are biotin, digoxigenin, Ni-nitrilotriacetate, desulfurized biotin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescent tags, tandem affinity purification (TAP) tags, glutathione S-transferase (GST), polynucleotides, aptamers, antigens, or antibodies.Methods for storing information are further provided herein, in which two, three, four, or five forms of bioencryption are used. Methods for storing information are further provided herein, in which the plurality of oligonucleotides comprises at least 100,000 oligonucleotides. Methods for storing information are further provided herein, in which the plurality of oligonucleotides comprises at least 10 billion oligonucleotides.
[0005] Methods for retrieving information are provided herein, each comprising: (a) releasing a plurality of oligonucleotides from a surface; (b) applying an enzyme-based, electromagnetic-based, chemical-based, or affinity-based decryption to the plurality of oligonucleotides; (c) concentrating the plurality of oligonucleotides; (d) sequencing the oligonucleotides concentrated from the plurality of oligonucleotides to generate a nucleic acid sequence; and (e) converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence codes for at least one item of information. Further methods for retrieving information are provided herein, wherein the decryption of the plurality of oligonucleotides comprises applying a CRISPR / Cas complex to the plurality of oligonucleotides. Further methods for retrieving information are provided herein, wherein the enzyme-based decryption comprises applying an enzyme as shown in Table 1. Further methods for retrieving information are provided herein, wherein the electromagnetic-based decryption comprises applying a wavelength from about 0.01 nm to about 400 nm. Methods for retrieving information are further provided herein, wherein chemical-based deletions include applying an administration of gaseous ammonia or methylamine. Methods for retrieving information are further provided herein, wherein affinity-based deletions include applying a sequence tag or affinity tag. Methods for retrieving information are further provided herein, wherein affinity tags are biotin, digoxigenin, Ni-nitrilotriacetate, desulfurized biotin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescent tags, tandem affinity purification (TAP) tags, glutathione S-transferase (GST), polynucleotides, aptamers, antigens, or antibodies. Methods for retrieving information are further provided herein, wherein two, three, four, or five forms of deletion are used.
[0006] A system for storing information is provided herein, comprising: (a) a receiving unit that receives machine instructions for at least one item of information in the form of at least one digital sequence and machine instructions for the selection of at least one form of bioencryption, wherein the form of bioencryption is enzyme-based, electromagnetic-based, chemical-based, or affinity-based; (b) a processor unit that automatically converts at least one digital sequence into a plurality of oligonucleotide sequences based on the selected form of bioencryption; (c) a synthesizer unit that receives machine instructions from the processor unit for synthesizing a plurality of oligonucleotides encoding the oligonucleotide sequence; and (d) a storage unit for receiving the plurality of oligonucleotides deposited from the synthesizer unit. A system for storing information is further provided herein, wherein enzyme-based bioencryption includes CRISPR / Cas-based bioencryption. A system for storing information is further provided herein, wherein enzyme-based bioencryption includes machine instructions for the synthesis of enzyme-sensitive oligonucleotides as shown in Table 1. A system for storing information is further provided herein, wherein electromagnetic-based bioencryption includes machine instructions for the synthesis of oligonucleotides sensitive to electromagnetic wavelengths from about 0.01 nm to about 400 nm. A system for storing information is further provided herein, wherein chemical-based bioencryption includes machine instructions for the synthesis of oligonucleotides sensitive to the administration of gaseous ammonia or methylamine. A system for storing information is further provided herein, wherein affinity-based bioencryption includes instructions for the synthesis of oligonucleotides sensitive to sequence tags or affinity tags.A system for storing information is further provided herein, where affinity tags are biotin, digoxigenin, Ni-nitrilotriacetate, desulfurized biotin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescent tags, tandem affinity purification (TAP) tags, glutathione S-transferase (GST), polynucleotides, aptamers, antigens, or antibodies. A system for storing information is further provided herein, where a plurality of oligonucleotides comprises at least 100,000 oligonucleotides. A system for storing information is further provided herein, where a plurality of oligonucleotides comprises at least 10 billion oligonucleotides.
[0007] A system for retrieving information is provided herein, comprising: (a) a storage unit containing a plurality of oligonucleotides on a surface; (b) a deposition device for applying enzyme-based, electromagnetic-based, chemical-based, or affinity-based bioencryption to the plurality of oligonucleotides; (c) a sequencing device for sequencing the plurality of oligonucleotides to obtain a nucleic acid sequence; and (d) a processor unit for automatically converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence encodes at least one item of information. A system for retrieving information is further provided herein, wherein the deposition device applies a CRISPR / Cas complex to the plurality of oligonucleotides. A system for retrieving information is further provided herein, wherein enzyme-based bioencryption applies enzymes as shown in Table 1. A system for retrieving information is further provided herein, wherein electromagnetic-based bioencryption applies wavelengths from about 0.01 nm to about 400 nm. A system for retrieving information is further provided herein, wherein chemical-based bioencryption applies the administration of gaseous ammonia or methylamine. A system for retrieving information is further provided herein, where affinity-based bioencryption includes sequence tags or affinity tags. A system for retrieving information is further provided herein, where affinity tags are biotin, digoxigenin, Ni-nitrilotriacetate, desulfurized biotin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescent tags, tandem affinity purification (TAP) tags, glutathione S-transferase (GST), polynucleotides, aptamers, antigens, or antibodies.
[0008] A method for storing information is provided herein, the method comprising: (a) receiving information of at least one item in the form of at least one digital sequence; (b) receiving instructions for at least one form of bioencryption; (c) converting at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences; (d) synthesizing a plurality of bioencrypted oligonucleotide sequences; and (e) storing a plurality of oligonucleotides.
[0009] A method for storing information is provided herein, the method comprising: (a) receiving at least one item of information in the form of at least one digital sequence; (b) receiving instructions for enzyme-based, electromagnetic-based, chemical-based, or affinity-based bioencryption; (c) converting at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences; (d) synthesizing a plurality of bioencrypted oligonucleotide sequences; and (e) storing a plurality of oligonucleotides.
[0010] A method for storing information is provided herein, the method comprising: (a) receiving at least one item of information in the form of at least one digital sequence; (b) converting at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences, each of which comprises an additional encoded sequence to be removed by a CRISPR / Cas complex; (c) synthesizing the plurality of bioencrypted oligonucleotide sequences; and (d) storing the plurality of oligonucleotides.
[0011] A method for retrieving information is provided herein, the method comprising: (a) releasing a plurality of oligonucleotides from a surface; (b) applying at least one form of biodecryption to the plurality of oligonucleotides; (c) concentrating the plurality of oligonucleotides and thereby selecting a plurality of concentrated oligonucleotides; (d) sequencing the concentrated oligonucleotides to generate a nucleic acid sequence; and (e) converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence codes for at least one item of information.
[0012] A method for retrieving information is provided herein, the method comprising: (a) releasing a plurality of oligonucleotides from a surface; (b) applying an enzyme-based, electromagnetic-based, chemical-based, or affinity-based decryption to the plurality of oligonucleotides; (c) concentrating the plurality of oligonucleotides and thereby selecting a plurality of concentrated oligonucleotides; (d) sequencing the concentrated oligonucleotides to generate a nucleic acid sequence; and (e) converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence codes for at least one item of information.
[0013] A method for retrieving information is provided herein, the method comprising: (a) releasing a plurality of oligonucleotides from a surface; (b) applying the plurality of oligonucleotides as a CRISPR / Cas complex; (c) concentrating the plurality of oligonucleotides and thereby selecting a plurality of concentrated oligonucleotides; (d) sequencing the concentrated oligonucleotides to generate a nucleic acid sequence; and (e) converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence codes for at least one item of information.
[0014] A system for storing information is provided herein, the system comprising: (a) a receiving unit for receiving machine instructions for at least one item of information in the form of at least one digital sequence and machine instructions for at least one form of bioencryption; (b) a processor unit for converting the at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences; (c) a synthesizer unit for receiving machine instructions from the processor unit for synthesizing the plurality of bioencrypted oligonucleotide sequences; and (d) a storage unit for receiving the plurality of oligonucleotides deposited from the synthesizer unit.
[0015] A system for storing information is provided herein, the system comprising: (a) a receiving unit for receiving machine instructions for at least one item of information in the form of at least one digital sequence and machine instructions for enzyme-based, electromagnetic-based, chemical-based, or affinity-based bioencryption; (b) a processor unit for converting the at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences; (c) a synthesizer unit for receiving machine instructions from the processor unit for synthesizing the plurality of bioencrypted oligonucleotide sequences; and (d) a storage unit for receiving the plurality of oligonucleotides deposited from the synthesizer unit.
[0016] A system for storing information is provided herein, the system comprising: (a) a receiving unit for receiving machine instructions for at least one item of information in the form of at least one digital sequence and machine instructions for bioencryption by a CRISPR / Cas complex; (b) a processor unit for converting at least one digital sequence into a plurality of bioencrypted oligonucleotide sequences; (c) a synthesizer unit for receiving machine instructions from the processor unit for synthesizing the plurality of bioencrypted oligonucleotide sequences; and (d) a storage unit for receiving the plurality of oligonucleotides deposited from the synthesizer unit.
[0017] A system for retrieving information is provided herein, the system comprising: (a) a storage unit containing a plurality of oligonucleotides on a surface; (b) a deposition device for applying at least one form of biodescription to the plurality of oligonucleotides; (c) a sequencing device for sequencing the plurality of oligonucleotides to obtain a nucleic acid sequence; and (d) a processor unit for converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence encodes at least one item of information.
[0018] A system for retrieving information is provided herein, the system comprising: (a) a storage unit containing a plurality of oligonucleotides on a surface; (b) a deposition device for applying at least enzyme-based, electromagnetic-based, chemical-based, or affinity-based bioencryption to the plurality of oligonucleotides; (c) a sequencing device for sequencing the plurality of oligonucleotides to obtain a nucleic acid sequence; and (d) a processor unit for converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence encodes at least one item of information.
[0019] A system for retrieving information is provided herein, the system comprising: (a) a storage unit containing a plurality of oligonucleotides on a surface; (b) a deposition device for applying a CRISPR / Cas complex to the plurality of oligonucleotides; (c) a sequencing device for sequencing the plurality of oligonucleotides to obtain a nucleic acid sequence; and (d) a processor unit for converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence encodes at least one item of information. [Brief explanation of the drawing]
[0020] Novel features of the present invention are described in particular in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description illustrating embodiments in which the principles of the present invention are used, and to the following appended drawings. [Figure 1] This illustrates a typical workflow for nucleic acid-based data storage. [Figure 2] This illustrates a typical workflow for preserving bioencryptions. [Figure 3] This shows a typical workflow for searching for subsequent bioencryptions. [Figure 4A] This describes a bioencryption method using Cas enzymes. [Figure 4B] This describes a bioencryption method using Cas enzymes. [Figure 5] AC represents a scheme for designing various oligonucleotide sequences. [Figure 6] AC represents a scheme for designing various oligonucleotide sequences. [Figure 7] AB stands for barcode design scheme. [Figure 8] An example of a plate configured for oligonucleotide synthesis is shown, comprising 24 regions or subregions, each having an array of 256 clusters. [Figure 9] An enlarged view of the sub-region of Figure 8, which has 16 × 16 clusters, is shown as an example, with each cluster having 121 individual loci. [Figure 10] A detailed diagram of the cluster in Figure 8, which has 121 seats, is provided as an example. [Figure 11A] An example of a front view of a plate having multiple channels is provided. [Figure 11B] A cross-sectional view of a plate with multiple channels is illustrated. [Figure 12] AB represents continuous loop and open-reel configurations for flexible structures. [Figure 13] AC represents an enlarged view of a flexible structure that has planar features (sits), channels, or wells. [Figure 14A] Enlarged diagrams illustrating the structural features described herein are provided as illustrations. [Figure 14B] The markings for the structures described herein are illustrated. [Figure 14C] The markings for the structures described herein are illustrated. [Figure 15] This illustrates a material deposition apparatus for oligonucleotide synthesis. [Figure 16] This illustrates the workflow for oligonucleotide synthesis. [Figure 17] An example of a computer system is shown. [Figure 18] This is a block diagram illustrating the architecture of a computer system. [Figure 19] This diagram shows a network configured to incorporate multiple computer systems, multiple mobile phones and personal digital assistants, and network-attached storage (NAS). [Figure 20] This is a block diagram of a multiprocessor computer system using a shared virtual address memory space. [Modes for carrying out the invention]
[0021] definition
[0022] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art to which these inventions pertain.
[0023] Throughout this disclosure, various embodiments are presented in range form. It should be understood that the range form is merely for convenience and brevity and should not be interpreted as a firm limitation on the scope of any embodiment. Accordingly, unless otherwise specified in the context, range descriptions should be understood as specifically disclosing all possible subranges and the individual numerical values within those ranges to two decimal places of the lower limit. For example, a range description such as 1 to 6 should be understood as specifically disclosing subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, and the individual numerical values within those ranges, for example, 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. These upper and lower limits of the intervening ranges may be independently contained within smaller ranges and are also incorporated within the invention according to any specifically excluded limits within the defined range. If a defined range includes one or both of an upper and lower limit, the range excluding either or both of these included upper and lower limits is also included within the scope of the present invention unless the context clearly indicates otherwise.
[0024] The terms used herein are intended solely to describe specific embodiments and are not intended to limit any embodiments. Herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context explicitly indicates otherwise. The terms “comprises” and / or “comprising,” when used herein, specify the presence of a particular feature, integer, process, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, processes, operations, elements, components, or groups thereof. As used herein, the terms “and / or” include any combination of one or more of the related enumerated items.
[0025] Unless otherwise specified or the context makes clear, the term “about” in relation to a number or range of numbers, as used herein, means the explicitly stated number and its plus / minus 10%, or, for any enumerated range of values, less than or equal to 10% of the enumerated lower limit and more than or equal to 10% of the enumerated lower limit.
[0026] As used herein, the term “oligonucleotide” is interchangeable with “oligonucleotide.” The terms “oligonucleotide” and “oligonucleotide” encompass double-stranded or triple-stranded nucleic acids as well as single-stranded molecules.
[0027] Nucleic acid-based information storage
[0028] Devices, compositions, systems, and methods for storing nucleic acid-based information (data) are provided herein. A typical workflow is provided in Figure 1. In the first step, a digital sequence (i.e., digital information in binary code for computer processing) encoding an item of information is received (101). An encryption (103) scheme is applied to convert the digital sequence from binary code to a nucleic acid sequence (105). A surface material for nucleic acid elongation, a design of a site for nucleic acid elongation (also known as a placement spot), and reagents for nucleic acid synthesis are selected (107). The surface of the structure is prepared for nucleic acid synthesis (108). De novo oligonucleotide synthesis is performed (109). The synthesized oligonucleotide is stored (111) and is available in whole or in part for subsequent release (113). Once released, the whole or in part of the oligonucleotide is sequenced (115), decrypted (117), and the nucleic acid sequence is converted back to a digital sequence. Subsequently, a digital array is constructed to obtain an alignment that codes the items of the original information (119).
[0029] Methods and systems for secure DNA-based information storage are further provided herein, comprising receiving one or more digital sequences encoding at least one item of information (201), converting the one or more digital sequences into nucleic acid sequences (203), encoding the nucleic acid sequences (205), and de novo oligonucleotide synthesis of the encoded nucleic acid sequences (207). See Figure 2.
[0030] Devices, compositions, systems, and methods for nucleic acid-based information storage are provided herein, wherein machine instructions are received for conversion from digital sequences to nucleic acid sequences, bioencryption, biodecryption, or any combination thereof. Machine instructions may be received for desired items of information, for conversion, and for one or more types of bioencryption selected from an optional list, including, but not limited to, enzyme-based (e.g., CRISPR / Cas complex or restriction enzyme digest), electromagnetic radiation-based (e.g., photolysis or photodetection), chemical cleavage (e.g., treatment of gaseous ammonia or methylamine to cleave thymidine-succinyl hexamide CED phosphoramidite (ChemGenes CLP-2244)), and affinity-based (e.g., incorporation of sequence tags for hybridization, or modified nucleotides having enhanced affinity to capture reagents). Following the receipt of a specific bioencryption selection, the program module performs the steps of converting the information item into a nucleic acid sequence and applying design instructions to design a bioencrypted version of the sequence, before providing synthesis instructions to the material deposition apparatus for the de novo synthesis of oligonucleotides. In some examples, machine instructions are provided for selecting one or more species within a category of bioencryptions.
[0031] Methods and systems for secure DNA-based information retrieval are further provided herein, including the release of oligonucleotides from a surface (301), enrichment of a desired oligonucleotide (303), sequencing of the oligonucleotide (305), decryption of the nucleic acid sequence (307), and construction of one or more digital sequences encoding an item of information (309). See Figure 3.
[0032] Machine instructions as described herein may be provided for biodescription. Biodescription may include receiving machine instructions. Such instructions may include one or more formats of biodescription selected from an optional list, and for example, include, but are not limited to, enzyme-based (e.g., CRISPR / Cas complex or restriction enzyme digest), electromagnetic radiation-based (e.g., photolysis or photodetection), chemical cleavage (e.g., treatment of gaseous ammonia or methylamine for cleavage of thymidine-succinyl hexamide CED phosphoramidite (ChemGenes CLP-2244)), and affinity-based (e.g., incorporation of nucleic acid sequences for hybridization, or modified nucleotides with enhanced affinity for capturing reagents) forms of oligonucleotide biodescription. Following the receipt of a selection of a particular biodescription, the program module performs a step of releasing a regulator for oligonucleotide enrichment. Following enrichment, the oligonucleotide is sequenced, optionally aligned to a longer nucleic acid sequence, and converted into a digital sequence corresponding to an item of information. In some examples, machine instructions are provided for selecting one or more species within a biodescription category.
[0033] Information items
[0034] Optionally, an initial step in a DNA data preservation process disclosed herein includes obtaining or receiving one or more items of information in the form of an initial code (e.g., a digital sequence). Items of information include, but are not limited to, textual, auditory, and visual information. Examples of sources for items of information include, but are not limited to, books, publications, electronic databases, medical records, characters, forms, audio recordings, animal records, bioprofiles, broadcasts, films, short videos, emails, bookkeeping phone logs, internet activity logs, drawings, paintings, printed materials, photographs, pixelated graphics, and software code. Example sources of bioprofiles for items of information include, but are not limited to, gene libraries, genomes, gene expression data, and protein activity data. Example formats for items of information include, but are not limited to, .txt, .PDF, .doc, .docx, .ppt, .pptx, .xls, .xlsx, .rtf, .jpg, .gif, .psd, .bmp, .tiff, .png, and .mpeg. The size of individual files encoding items of information in digital format, or multiple files encoding items of information, is not limited, but includes up to 1024 bytes (equivalent to 1 KB), up to 1024 KB (equivalent to 1 MB), 1024 MB (equivalent to 1 GB), 1024 GB (equivalent to 1 TB), 1024 TB (equivalent to 1 PB), 1 exabyte, 1 zettabyte, 1 yottabyte, 1 xenottabyte, or more. In some examples, the amount of digital information is at least or approximately 1 gigabyte (GB). In some examples, the amount of digital information is at least or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 1000 gigabytes or more. In some examples, the amount of digital information is at least or approximately 1 terabyte (TB).In some examples, the amount of digital information is at least, or approximately, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 terabytes, or 1000 terabytes or more. In some examples, the amount of digital information is at least, or approximately 1 petabyte (PB). In some examples, the amount of digital information is at least, or approximately, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 petabytes, or 1000 petabytes or more.
[0035] Encryption
[0036] Bioencryption and biodecryption
[0037] Devices, compositions, systems, and methods, including bioencryption (also known as "biomecretion") after receiving a digital sequence that codes an item of information, are described herein. In addition to the individual forms of bioencryption and biodecryption described herein, processes for incorporating the selection of one or more classes or species to mask a biological sequence into an information storage and / or retrieval workflow are also provided herein.
[0038] Devices, compositions, systems, and methods for target enriching a nucleic acid sequence of interest from a larger population of nucleic acid sequences, including bioencryptions, are provided herein. In some examples, bioencryption is used to enhance a target signal from noise. In some examples, the target signal is the nucleic acid sequence of interest. In some examples, bioencryption involves introducing the nucleic acid sequence of interest into a larger population of nucleic acid sequences having a known sequence. The known nucleic acid sequence may be called the encryption nucleic acid sequence. In some examples, the encryption nucleic acid is decrypted. In some examples, the decryption of the known nucleic acid sequence results in an increase in the signal-to-noise ratio of the nucleic acid sequence of interest.
[0039] Devices, compositions, systems, and methods for incorporating biomolecular encryptions in information storage and / or retrieval are provided herein. Exemplary forms of bioencryption and biodecryption include, but are not limited to, enzyme-based, electromagnetic radiation-based, chemical cleavage, and affinity-based bioencryption and biodecryption.
[0040] Devices, compositions, systems, and methods, including the application of nuclease complex activity-based encryption, are provided herein. Exemplary nucleases include, but are not limited to, Cas nucleases (CRISPR-related), Zn finger nucleases (ZFNs), TAL effector nucleases, Argonaut nucleases, S1 nucleases, mangubean nucleases, or DNAse. Examples of Cas nucleases include, but are not limited to, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Cpf1, c2c1, and c2c3. In some cases, the Cas nuclease is Cas9. In some cases, the CRISPR / Cas complex results in the predetermined removal of one or more nucleic acid sequences. In some cases, the enrichment steps described herein involve the depletion of the rich sequence by hybridization (DASH). In some cases, DASH involves the application of a nuclease. For example, a nuclease such as Cas9, when bound to a CRISPR complex containing a guide RNA ("gRNA") sequence, induces strand breaks such that the longer form of the nucleic acid sequence is no longer complete. In some cases, the excised nucleic acid is unavailable for subsequent amplification following enrichment. In some cases, the gRNA guides the Cas9 enzyme to specific extension of the nucleic acid. In alternative configurations, the gRNA has multiple sites for cleavage. gRNA-based systems enable the generation of cryptographic codes with high specificity and selectivity. For example, since a CRISPR / Cas9-based system uses 20 bp to identify the sequence to cleave, at least approximately 10^12 different possibilities are available for designing a predetermined gRNA sequence for decryption using the four base systems.Following the removal of exogenous (also known as "junk") DNA, predetermined oligonucleotides encoding the target sequence undergo downstream processing (e.g., amplification and sequencing) to result in a final sequence free of extraneous (junk) sequences. In some examples, each oligonucleotide in a group of oligonucleotides is designed for modification (e.g., cleavage, transnucleotide exchange, recombination) at multiple sites. For example, each oligonucleotide in a group of oligonucleotides is synthesized with complementary regions for binding to approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more gRNA sequences. In such a configuration, each of the group of oligonucleotides undergoes cleavage, transnucleotide exchange, and recombination following nuclease (e.g., CRISPR / Cas) complex activity at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more sites.
[0041] A first process of targeted enrichment for data encryption using CRISPR / Cas9 is illustrated in Figure 4A. A population of DNA sequences (401) contains DNA information (403) and encrypted DNA (405). The DNA information (403) and encrypted DNA (405) contain adapter sequences (402) and DNA sequences (404, 406), respectively. Guide RNA (409) is added to the population of DNA base sequences (401) (407). Guide RNA (409) is used to remove encrypted DNA (405) by recognizing the cleavage sequence to be encrypted in the encrypted DNA (405). Following the addition of guide RNA (409), the encrypted DNA (405) is cleaved, resulting in a nucleic acid sequence that is no longer complete. Therefore, if the encoded DNA (405) cannot be amplified, for example, if it contains a nucleic acid sequence that is no longer complete, it is removed from the population (411), leaving behind the DNA information (403).
[0042] A second process of targeted enrichment for data encoding using CRISPR / Cas9 is illustrated in Figure 4B. A population of DNA sequences (421) includes DNA information (423) and encoded DNA (425). The DNA information (423) and encoded DNA (425) include adapter sequences (422) and DNA sequences (424, 426), respectively. Guide RNA (429) and donor DNA (431) are added to the population of DNA sequences (421) (427). The guide RNA (429) recognizes the encoded cleavage site in the encoded DNA (425) and generates a cleavage site for insertion of the donor DNA (431). Insertion of the donor DNA (431) results in an insertion or frameshift in the encoded DNA (425). In some cases, insertion of donor DNA (431) results in the introduction of a sequence tag for hybridization or incorporation of a modified nucleotide with enhanced affinity to the capture reagent. For example, donor DNA (431) is recognized by a fluorescent probe. In some cases, donor DNA (431) introduces a sequence for bioencryption and / or biodecryption based on electromagnetic radiation (e.g., photolysis or photodetection) or chemical cleavage (e.g., treatment of thymidine-succinyl hexamide CED phosphoramidite (e.g., CLP-2244 from ChemGenes) with gaseous ammonia or methylamine). In some cases, the encrypted DNA (425) is no longer recognized for amplification, and only the DNA information (423) is amplified, resulting in enrichment of the DNA information (423).
[0043] Devices, compositions, systems, and methods involving the application of nuclease complex activity-based encryption as described herein may include base exchange or sequence exchange. For example, bioencryption and biodecryption using CRISPR / Cas include base exchange or sequence exchange. In some examples, bioencryption involves CRISPR / dCas9, where inactive or “dead” Cas9 ("dCas9") no longer possesses splicing function but, with the addition of another enzymatic activity, performs modification functions of various target molecules. For example, tethering of cytidine deaminase to dCas9 converts CG DNA base pairs to TA base pairs. In other dCas9 processes, different enzymes tethered to dCas9 change base C to T or base G to A within target DNA.
[0044] Devices, compositions, systems, and methods for bioencryption and biodecryption, including the application of restriction enzymes, are provided herein. In some examples, the restriction enzyme targets an enzyme recognition site. In some examples, the enzyme recognition site is a specific nucleotide sequence. In some examples, the restriction enzyme cleaves a phosphate backbone at or near the enzyme recognition site. In some examples, cleavage of the recognition site results in a non-blunt or blunt-ended outcome. In some examples, the restriction enzyme recognizes nucleotides (e.g., A, T, G, C, U). In some examples, the restriction enzyme recognizes modifications such as methylation, hydroxylation, or glycosylation, but are not limited. In some examples, the restriction enzyme results in fragmentation. In some examples, the fragmentation produces fragments having a 5' overhang, a 3' overhang, a blunt end, or a combination thereof. In some examples, the fragments are selected, for example, based on size. In some examples, ligation is performed after fragmentation by the restriction enzyme. For example, restriction enzyme fragmentation is used to leave predictable overhangs, followed by ligation with one or more adapter oligonucleotides containing overhangs complementary to the predictable overhangs on the nucleic acid fragment. Exemplary restriction enzymes and their recognition sequences are provided in Table 1.
[0045] [Table 1-1]
[0046] [Table 1-2]
[0047] [Table 1-3]
[0048] [Table 1-4]
[0049] [Table 1-5]
[0050] [Table 1-6]
[0051] [Table 1-7]
[0052] Devices, compositions, systems, and methods for bioencryption and biodecryption, which may involve the application of repair enzymes, are provided herein. In some examples, the DNA repair enzymes are derived from specific organisms or viruses, or are non-natural variants thereof. Exemplary DNA repair enzymes include, but are not limited to, E. coli endonuclease IV, Tth endonuclease IV, human AP endonuclease, glycosylases (such as UDG, E. coli 3-methyladenine DNA glycosylase (AIkA), and human Aag), glycosylases / lyases (such as E. coli endonuclease III, E. coli endonuclease VIII, E. coli Fpg, human OGGl, and T4 PDG), and lyases. Additional exemplary DNA repair enzymes are listed in Table 2.
[0053] [Table 2-1]
[0054] [Table 2-2]
[0055] [Table 2-3]
[0056] Devices, compositions, systems, and methods for bioencryption and / or biodecryption involving nucleic acid modifications are provided herein. In some examples, nucleic acid modifications affect the activity of nucleic acid sequences in sequencing reactions. For example, nucleic acid modifications prevent the amplification of the encrypted nucleic acid sequence. In some examples, nucleic acid modifications include, but are not limited to, methylated bases, PNA (peptide nucleic acid) nucleotides, LNA (locked nucleic acid) nucleotides, and 2'-O-methyl-modified nucleotides. In some examples, nucleic acid modifications include modified nucleic acid bases other than cytosine, guanine, adenine, or thymine. Examples of non-restrictively modified nucleic acid bases include, but are not limited to, uracil, 3-meA (3-methyladenine), hypoxanthine, 8-oxoG (7,8-dihydro-8-oxoguanine), FapyG, FapyA, Tg (thymine glycol), hoU (hydroxyuracil), hmU (hydroxymethyluracil), fU (formyluracil), hoC (hydroxycytosine), fC (formylcytosine), 5-meC (5-methylcytosine), 6-meG (O6-methylguanine), 7-meG (N7-methylguanine), εC (ethenocytosine), 5-caC (5-carboxylcytosine), 2-hA, εA (ethenoadenine), 5-fU (5-fluorouracil), 3-meG (3-methylguanine), and isovaleric acid.
[0057] Devices, compositions, systems, and uses for bioencryption containing nucleic acid probe sequences are provided herein. In some examples, a nucleic acid sequence complementary to a portion of the nucleic acid probe sequence is subsequently removed by a nuclease. For example, the nuclease is a bispecific nuclease that recognizes a double-stranded nucleic acid molecule formed between the nucleic acid probe and the nucleic acid sequence. In some examples, the nucleic acid probe allows for the capture and isolation of the nucleic acid sequence. In some examples, the nucleic acid probe contains at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more than 100 nucleotide lengths.
[0058] In some cases, nucleic acid sequences are identified using nucleic acid probes that include labels such as affinity tags, including but not limited to biotin, digoxigenin, Ni-nitrilotriacetate, desulfurized biotin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescent tags, tandem affinity purification (TAP) tags, glutathione S-transferase (GST), polynucleotides, aptamers, polypeptides (e.g., antigens or antibodies), or derivatives thereof. In some cases, the labels are detected by light absorption, fluorescence, chemiluminescence, electrochemiluminescence, mass, or loading. Examples of non-fluorophore dyes include Alexa-Fluor dyes (e.g., Alexa-Fluor® 350, Alexa-Fluor® 405, Alexa-Fluor® 430, Alexa-Fluor® 488, Alexa-Fluor® 500, Alexa-Fluor® 514, Alexa-Fluor® 532, Alexa-Fluor® 546, Alexa-Fluor® 555, Alexa-Fluor® 568, Alexa-Fluor® 594, Alexa-Fluor® 610, Alexa-Fluor® 633, Alexa-Fluor® 647, Alexa-Fluor® 660, Alexa-Fluor® 680, Alexa-Fluor® 700, and Alexa-Fluor® 750), APC, Cascade Blue, Cascade Yellow. These include Yellow, R-phycoerythrin (PE), DyLight 405, DyLight 488, DyLight 550, DyLight 650, DyLight 680, DyLight 755, DyLight 800, FITC, Pacific Blue, PerCP, Rhodamine, Texas Red, Cy5, Cy5.5, and Cy7.
[0059] Devices, compositions, systems, and methods for bioencryption and / or biodecryption involving nucleic acid hybridization-based binding are provided herein. Nucleic acid probes containing affinity tags may be used. In some examples, the affinity tag allows a nucleic acid sequence to be pulled down. For example, affinity-tagged biotin is bound to a nucleic acid probe complementary to the nucleic acid sequence and pulled down using streptavidin. In some examples, the affinity tag comprises a magnetically susceptible material (e.g., a magnet, a magnetically susceptible metal). In some examples, the nucleic acid sequence is pulled down using a solid support such as streptavidin and immobilized by the solid support. In some examples, the nucleic acid sequence is pulled down in solution via beads or the like. In some examples, the nucleic acid probe allows for size-based elimination. For example, the nucleic acid probe results in a nucleic acid sequence having a different size from other nucleic acid sequences, and as a result, the nucleic acid sequence is eliminated by size-based depletion.
[0060] Devices, compositions, systems, and methods for bioencryption and / or biodecryption involving nucleic acid hybridization-based binding may include suppressed amplification. In some examples, the nucleic acid hybridization-based binding strategy targets suppressed amplification, where multiple synthetic oligonucleotides have similar regions on the forward primer to which they bind, but the reverse primer region is not readily identifiable. In such examples, a predetermined reverse primer is required. In a first exemplary workflow, a pool of reverse primers having pre-selected regions to bind to each of different synthetic oligonucleotides is generated and used in an extension amplification reaction (e.g., using DNA polymerase) to amplify the oligonucleotides for downstream processing (e.g., further amplification or DNA sequencing). Optionally, each reverse primer includes an adapter region containing a common sequence to incorporate a common reverse primer binding site by the extension amplification reaction (e.g., using DNA polymerase). In such a configuration, downstream processing is simplified because only a single forward or reverse primer is required to amplify or sequence multiple oligonucleotides. In a second exemplary workflow, multiple oligonucleotides are synthesized, each having one or two regions containing a hybridization motif, which varies but has sufficient hybridization ability to a common primer to enable downstream processing (e.g., amplification or sequencing) of the multiple oligonucleotides using a common primer on one or both of the 5' and 3' regions of each synthesized oligonucleotide. In some examples, the oligonucleotide population is designed to hybridize to approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more common primer nucleoside bases.
[0061] Devices, compositions, systems, and methods for bioencryption and / or biodecryption, including the use of electromagnetic radiation (EMR), are provided herein. In some examples, electromagnetic radiation results in detection based on image capture of cleavage or nucleic acid sequences. In some examples, EMR is applied to a surface at wavelengths of about 100 nm to about 400 nm, about 100 nm to about 300 nm, or about 100 nm to about 200 nm. In some examples, EMR is applied to a surface at wavelengths shorter than 0.01 nm. In some examples, EMR is applied to a surface at wavelengths of about 10 nm to about 400 nm, about 400 nm to about 700 nm, or about 700 nm to about 100,000 nm. For example, EMR is applied at ultraviolet (UV) wavelengths or deep UV wavelengths. In some examples, deep ultraviolet is applied to a surface at a wavelength of about 172 nm to cleave the binder from the surface. In some examples, EMR is applied using a xenon lamp. The exposure distance is a measurement between the lamp and the surface. In some cases, the exposure distance is approximately 0.1–5 cm. In some cases, the exposure distance is approximately 0.5–2 cm. In some cases, the exposure distance is approximately 0.5, 1, 2, 3, 4, or 5 cm. In some cases, EMR is applied with a laser. Examples of lasers and their wavelengths include, but are not limited to, Ar2 (126 nm), Kr2 (146 nm), F2 (157 nm), Xe2 (172 and 175 nm), and ArF (193 nm). In some cases, the nucleic acid sequence contains nucleic acid bases that are photocleavable at specific sites. In some cases, the nucleic acid sequence contains photocleavable modified nucleic acid bases. In some cases, the nucleic acid sequence is photocleavable by the application of a specific wavelength of light. In some cases, the nucleic acid sequence is photocleavable by the application of multi-wavelength light.
[0062] Devices, compositions, systems, and methods for bioencryption and / or biodecryption, including the use of chemical dissolution, are provided herein. In some examples, the nucleic acid sequence includes nucleic acid bases that are chemically cleavable at specific sites. In some examples, the nucleic acid sequence includes chemically cleavable modified nucleic acid bases. In some examples, the modified nucleic acid bases include chemically cleavable modifications. In some examples, chemical dissolution is performed using an amine reagent. In some examples, the amine reagent is a liquid, gas, aqueous reagent, or an anhydrous reagent. Non-limiting examples of amine reagents include ammonium hydroxide, ammonia gas, C1-C6 alkylamines, or methylamines.
[0063] Devices, compositions, systems, and methods for bioencryption as described herein may include the conversion of digital sequences to nucleic acid sequences. In some examples, the nucleic acid sequences are DNA sequences. In some examples, the DNA sequences are single-stranded or double-stranded. In some examples, the nucleic acid sequences are RNA sequences. In some examples, the RNA sequences are single-stranded or double-stranded. In some examples, the nucleic acid sequences are encoded in a larger population of nucleic acid sequences. In some examples, the larger population of nucleic acid sequences is a homogeneous or heterogeneous population. In some examples, the population of nucleic acid sequences includes DNA sequences. In some examples, the DNA sequences are single-stranded or double-stranded. In some examples, the population of nucleic acid sequences includes RNA sequences. In some examples, the RNA sequences are single-stranded or double-stranded.
[0064] Many nucleic acid sequences can be encoded. In some examples, the number of encoded nucleic acid sequences ranges from about 10 sequences to over 1 million sequences. In some examples, the number of nucleic acid sequences being encoded is at least approximately 10, 50, 100, 200, 500, 1,000, 2,000, 4,000, 8,000, 10,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, 55,000, 60,000, 65,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1 million, or more. In some cases, the number of nucleic acid sequences being encoded exceeds one trillion.
[0065] In some cases, the encoded nucleic acid sequence has a base length of at least 10, 25, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, or more than 300 bases. In some cases, the encoded nucleic acid sequence has a base length of 10 to 25 bases, 10 to 50 bases, 10 to 75 bases, 10 to 100 bases, 10 to 125 bases, 10 to 150 bases, 10 to 175 bases, 10 to 200 bases, 10 to 225 bases, 10 to 250 bases, 10 to 300 bases, 25 to 50 bases, 25 to 75 bases, 25 to 100 bases, 25 to 125 bases, and 25 to 150 bases. , 25-175 bases, 25-200 bases, 25-225 bases, 25-250 bases, 25-300 bases, 50-75 bases, 50-100 bases, 50-125 bases, 50-150 bases, 50-175 bases, 50-200 bases, 50-225 bases, 50-250 bases, 50-300 bases, 75-100 bases, 75-125 bases, 75-150 bases, 75-175 bases Base, 75 bases to 200 bases, 75 bases to 225 bases, 75 bases to 250 bases, 75 bases to 300 bases, 100 bases to 125 bases, 100 bases to 150 bases, 100 bases to 175 bases, 100 bases to 200 bases, 100 bases to 225 bases, 100 bases to 250 bases, 100 bases to 300 bases, 125 bases to 150 bases, 125 bases to 175 bases, 125 bases to 200 bases, 125 bases to 225 bases, 125 bases to 250 bases, 125 Includes bases ~300 bases, 150 bases ~ 175 bases, 150 bases ~ 200 bases, 150 bases ~ 225 bases, 150 bases ~ 250 bases, 150 bases ~ 300 bases, 175 bases ~ 200 bases, 175 bases ~ 225 bases, 175 bases ~ 250 bases, 175 bases ~ 300 bases, 200 bases ~ 225 bases, 200 bases ~ 250 bases, 200 bases ~ 300 bases, 225 bases ~ 250 bases, 225 bases ~ 300 bases, or 250 bases ~ 300 bases.
[0066] In some cases, the encoded nucleic acid sequence results in an enrichment of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more than 95% of the target nucleic acid sequence.
[0067] Devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may include DNA or RNA-based systems. Reference DNA is a 4-base coding system with four different nucleic acid bases available: A, T, C, or G (adenine, thymine, cytosine, and guanine). These four bases thus enable 3-base (not all are used) or 4-base coding systems. In addition, the use of uracil (U), found in RNA, enables a fifth base and 5-base coding system. Furthermore, modified nucleic acid bases may be used for coding of more than four nucleic acid bases. Nucleic acid bases that are neither reference DNA nucleic acid bases nor modified nucleic acid bases include, but are not limited to, uracil, 3-meA (3-methyladenine), hypoxanthine, 8-oxoG (7,8-dihydro-8-oxoguanine), FapyG, FapyA, and Tg (thymine). Examples include glycols, hoU (hydroxyuracil), hmU (hydroxymethyluracil), fU (formyluracil), hoC (hydroxycytosine), fC (formylcytosine), 5-meC (5-methylcytosine), 6-meG (O6-methylguanine), 7-meG (N7-methylguanine), εC (ethenocytosine), 5-caC (5-carboxylcytosine), 2-hA, εA (ethenoadenine), 5-fU (5-fluorouracil), 3-meG (3-methylguanine), hmC (hydroxymethylcytosine), and isovaleric acid. Coding schemes are further provided herein, where machine instructions provide the conversion of digital information in the form of a binary sequence into an intermediate code before it is finally converted into a final nucleic acid sequence.
[0068] In some examples, to store data in a DNA sequence, information is converted from binary codes of 1s and 0s to codes for the A, T, G, and C bases of DNA. In some examples, items of information are initially encoded in the form of digital information. In some cases, the binary code of the digital information is converted to a biomolecular-based (e.g., DNA-based) code while preserving the information that the code represents. This converted code (from digital binary code to biomolecular code) is referred to herein as resulting in a “predetermined” sequence for the deposition of biomolecules disclosed herein on the surface disclosed herein. The predetermined sequence may encode sequences of multiple oligonucleotides.
[0069] Binary code conversion
[0070] Typically, the first code is digital information, usually in the form of binary code used by computers. A general-purpose computer is an electronic device that reads "on" or "off" states represented by the numbers "0" and "1". This binary code is an application for computers to read items of multiple types of information. In binary, the number 2 is written as 10. For example, "10" represents "2 times 1". The number "3" is written as "11", meaning "2 times 1 and 1". The number "4" is written as "100", the number "5" as "101", "6" as "110", and so on. An example of the US Standard Code II (ASCII) binary code is provided in lowercase and uppercase alphabets in Table 3.
[0071] [Table 3-1]
[0072] [Table 3-2]
[0073] A method for converting information in the form of an initial code (e.g., a binary sequence) into a nucleic acid sequence is provided herein. The process may include a direct conversion from a 2-base code (i.e., binary) to a higher base code. Exemplary base codes include 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. Table 4 illustrates typical alignments between various base numbering schemes. A computer receiving machine instructions for conversion can automatically convert sequence information from one code to another.
[0074] [Table 4]
[0075] nucleic acid sequence
[0076] Methods for designing sequences of oligonucleotides described herein such that the nucleic acid sequence codes for at least a portion of an item of information are provided herein. In some examples, each oligonucleotide sequence has design features that facilitate sequence alignment and provide means for error correction during the subsequent construction process. In some configurations, oligonucleotide sequences are designed such that the exits between each oligonucleotide sequence overlap with another in the population. In some examples, each oligonucleotide sequence overlaps with a portion of just one other oligonucleotide sequence (Figure 5A). In other configurations, each oligonucleotide sequence region overlaps with two sequences such that two copies of each sequence are produced within a single oligonucleotide (Figure 5B). In yet another configuration, each oligonucleotide sequence region overlaps with two or more sequences such that three copies of each sequence are produced within a single oligonucleotide (Figure 5C). The oligonucleotide sequences described herein can code for base lengths of 10-2000, 10-500, 30-300, 50-250, or 75-200. In some examples, each oligonucleotide sequence has a base length of at least 10, 15, 20, 25, 30, 50, 100, 150, 200, 500, or more bases.
[0077] Methods, systems, and compositions are provided herein in which each oligonucleotide sequence described herein is designed to include multiple coding regions and multiple non-coding regions (Figure 6A). In such a configuration, each coding region (e.g., (601, 603, 605)) codes for at least a portion of the information item. Optionally, each coding region in the same oligonucleotide codes for the same information item, and a redundancy scheme is used as described herein (Figure 6B). In a further example, each coding region in the same oligonucleotide codes for the same sequence (Figure 6C). The oligonucleotide sequences described herein can code for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more coding regions. The sequences of oligonucleotides described herein can encode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more identical coding regions. In some examples, each of the multiple coding regions has a base length of 10-1000, 20-500, 30-300, 50-250, or 75-200. In some examples, each of the multiple coding regions has a base length of at least 10, 15, 20, 25, 30, 50, 100, 150, 200, or 200 or more. In some examples, each oligonucleotide includes a tether region (611) that links the molecule to the surface (602) of the structure.
[0078] In a configuration where multiple coding sequences exist within the same oligonucleotide, cleavage regions (607) are optionally located between each coding region. The cleavage regions (607) may be located at the junctions between each coding region or within an adapter region having a series of sequences between each coding region. Once synthesized, the cleavage regions (607) can encode sequence features that are cleaved from the chain following the application of a cleavage signal. The cleavage regions (607) may be restriction enzyme recognition sites, i.e., modified nucleic acids that are photosensitive and will be cleaved under the application of electromagnetic radiation (e.g., oligodeoxynucleotide heteropolymers carrying base-sensitive S-pivaloylthiomethyl (t-Bu-SATE) phosphotriester bonds sensitive to wavelengths of light >300 nm), or specific chemicals (e.g., thymidine-succinyl hexamide CED, which is cleaved after the application of ammonia gas). It is possible to encode modified nucleic acids that are sensitive to the application of phosphoramidite (CLP-2244 from ChemGenes). Since the design of sequences with specific cleavage schemes may not be readily apparent from the sequencing of synthetic oligonucleotides, the cleavage schemes provide a means to add a level of safety to sequences encoded by synthetic nucleic acid libraries. The sequences of oligonucleotides described herein include at least 1, 2, 3, 4, 5, 6, and 7. 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more cleavage regions can be coded. In some examples, each cleavage region codes for base lengths of 1-100, 1-50, 1-20, 1-10, 5-25, or 5-30. In some examples, each cleavage region codes for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50 It encodes 100 or more bases. In some configurations, for each oligonucleotide, each coding region is identical, and each cleavage region between coding regions is different. For example, the first cleavage region (607) is different from the second cleavage region (609). In some configurations, the cleavage region (607) near the surface (602) is identical to the next distal cleavage region (607).
[0079] A barcode is a typically known nucleic acid sequence that allows for the identification of several features of the polynucleotide to which the barcode is associated. Figure 7A provides an exemplary barcode structure. In Figure 7A, the coding regions of the first oligonucleotide (701), the second oligonucleotide (703), and the third oligonucleotide (705) have the following features (outside the surface (702)): a tether region (702), a cleavage region (707), a first primer-binding region (701), a barcode region (703), coding regions (701, 703, and 705), and a second primer-binding region (704). The oligonucleotides may be amplified using primers that recognize the first and / or second primer-binding regions. Amplification may occur for oligonucleotides bound to the surface or released from the surface (i.e., via cleavage at the cleavage region (707)). After sequencing, the barcode region (703) provides an index for identifying features associated with the coding region. In some examples, a barcode contains a nucleic acid sequence that, when ligated to a target polynucleotide, serves as an identifier for the sample from which the target polynucleotide originates. The barcode may be designed to be of a length appropriate to allow for a sufficient degree of identification, for example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, or more base lengths. Multiple barcodes (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more barcodes) may be used on the same molecule and may be optionally separated by non-barcode sequences. In some examples, barcodes are shorter than 10, 9, 8, 7, 6, 5, or 4 base lengths. In some examples, barcodes associated with some polynucleotides are of different lengths than barcodes associated with other polynucleotides. Generally, barcodes are of sufficient length and contain sequences distinct enough to allow for the identification of samples based on the barcodes they are associated with.In some configurations, a barcode and its associated sample source can be precisely identified after one or more base mutations, insertions, or deletions in the barcode sequence, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more base mutations, insertions, or deletions. In some examples, each barcode in a group of barcodes differs from all the other barcodes by at least three base positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more positions. Configurations provided herein may include a barcode sequence indicating a nucleic acid sequence that codes for a sequence in a particular region of a digital sequence. For example, a barcode sequence may indicate the location that codes for a particular oligonucleotide sequence within a larger file. In some examples, a barcode sequence may indicate which file a particular oligonucleotide sequence is associated with. In some examples, a barcode sequence includes information related to the conversion scheme of a particular sequence, adding an extra layer of safety.
[0080] Oligonucleotide sequence design schemes are provided herein, where each oligonucleotide sequence in a population of oligonucleotide sequences is designed to have at least one region common to all oligonucleotide sequences in that population. For example, all oligonucleotides in the same population may contain one or more primer regions. The design of sequence-specific primer regions allows for the selection of oligonucleotides to be amplified in selected batches from a large library of multiple oligonucleotides. Each oligonucleotide sequence may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more primer-binding sequences. A population of oligonucleotide sequences may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, 200, 500, 1000, 5000, 10000, 50000, 100000, or more non-identical binding sequences. The primer binding sequence may contain base lengths of 5-100, 10-75, 7-60, 8-60, 10-50, or 10-40.
[0081] Structure of oligonucleotide synthesis
[0082] This specification provides rigid or flexible structures for oligonucleotide synthesis, for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein. In the case of rigid structures, devices having a structure (e.g., a plate) for generating a library of oligonucleotides are provided herein. An exemplary structure (800) is illustrated in Figure 8, where structure (800) has dimensions approximately the same as a standard 96-well plate: 140 mm × 90 mm. Structure (800) includes clusters grouped into 24 regions or sub-regions (805), each sub-region (805) containing an array of 256 clusters (810). A magnified view of an exemplary sub-region (805) is shown in Figure 9. As seen in Figures 8 and 9, the structure may be essentially planar. In the enlarged view of the four clusters (Figure 9), a single cluster (910) has a Y-axis cluster pitch (distance from the center to the center of an adjacent cluster) of 1079.210 μm or 1142.694 μm, and an X-axis cluster pitch of 1125 μm. An exemplary cluster (1010) is shown in Figure 10, where the Y-axis pitch (distance from the center to the center of an adjacent locus) is 63.483 μm, and the X-axis locus pitch is 75 μm. The width of the longest locus (e.g., the diameter of a circular trajectory) is 50 μm, and the distance between locus is 24 μm. The number of locus (1005) in the exemplary cluster in Figure 10 is 121. Locus (also called “features”) may be planes, wells, or channels. An exemplary channel configuration is illustrated in Figures 11A-11B, where a plate (1105) is illustrated, comprising a main channel (1110) and a plurality of channels (1115) connected to the main channel (1110). The connection between the main channel (1110) and the plurality of channels (1115) provides fluid communication of the flow path from the main channel (1110) to each of the plurality of channels (1115). The plate (1105) described herein may include a plurality of main channels (1110). The plurality of channels (1115) together form clusters within the main channel (1110).
[0083] In the case of flexible structures, devices are provided herein that include a continuous loop (1201) wrapped around one or more fixed structures (e.g., a pair of rollers (1203)) or a discontinuous flexible structure (1207) wrapped around separate fixed structures (e.g., a pair of rollers (1205)). See Figure 12AB. Flexible structures having a surface of multiple features (seats) for oligonucleotide extension are also provided herein. Some of the features of the flexible structure (1301) may be essentially planar features (1303) (e.g., flat), channels (1305), or wells (1307). See Figure 13AC. In an exemplary configuration, each feature of the structure has a width of about 10 μm and a distance of about 21 μm between the centers of each structure. See Figure 14A. Features include, but are not limited to, circular, rectangular, tapered, or round shapes.
[0084] Structures for oligonucleotide synthesis for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may include channels. In some examples, the channels described herein have a width-to-depth (or height) ratio of 1:0.01, where width is a measurement of the width at the narrowest segment of the microchannel. In some examples, the channels described herein have a width-to-depth (or height) ratio of 0.5:0.01, where width is a measurement of the width at the narrowest segment of the microchannel. In some examples, the channels described herein have a width-to-depth (or height) ratio of about 0.01, 0.05, 0.1, 0.15, 0.16, 0.2, 0.5, or 1.
[0085] Structures for polynucleotide synthesis comprising multiple separate loci, channels, wells, or protrusions for polynucleotide synthesis are provided herein. Structures described herein comprise multiple clusters, each cluster comprising multiple wells, loci, or channels. Alternatively, structures described herein comprise a homogeneous arrangement of wells, loci, or channels. In some examples, structures described herein comprise multiple channels corresponding to multiple features (loci) within a cluster, where the height or depth of the channels is approximately 5 μm to approximately 500 μm, approximately 5 μm to approximately 400 μm, approximately 5 μm to approximately 300 μm, approximately 5 μm to approximately 200 μm, approximately 5 μm to approximately 100 μm, approximately 5 μm to approximately 50 μm, or approximately 10 μm to approximately 50 μm. In some cases, the height or depth of the channels is less than 100 μm, less than 80 μm, less than 60 μm, less than 40 μm, or less than 20 μm. In some cases, the channel height or depth is approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 μm or more. In some examples, the channel height or depth is at least 10, 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 nm, or greater than 1000 nm. In some examples, the channel height or depth is in the range of approximately 10 nm to approximately 1000 nm, approximately 25 nm to approximately 900 nm, approximately 50 nm to approximately 800 nm, approximately 75 nm to approximately 700 nm, approximately 100 nm to approximately 600 nm, or approximately 200 nm to approximately 500 nm. In some examples, the channel height or depth is in the range of approximately 50 nm to approximately 1 μm.
[0086] Structures of oligonucleotide synthesis for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may include features. In some examples, the width of a feature (e.g., an essentially planar feature, well, channel, seat, or projection) is about 0.1 μm to about 500 μm, about 0.5 μm to about 500 μm, about 1 μm to about 200 μm, about 1 μm to about 100 μm, about 5 μm to about 100 μm, or about 0.1 μm to about 100 μm, for example, about 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm. In some examples, the width of a feature (e.g., a microchannel) is less than about 100 μm, 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or 10 μm. In some examples, the feature width is at least 10, 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 nm, or greater than 1000 nm. In some examples, the feature width is in the range of approximately 10 nm to approximately 1000 nm, approximately 25 nm to approximately 900 nm, approximately 50 nm to approximately 800 nm, approximately 75 nm to approximately 700 nm, approximately 100 nm to approximately 600 nm, or approximately 200 nm to approximately 500 nm. In some examples, the feature width is in the range of approximately 50 nm to approximately 1000 nm. In some examples, the distance between the centers of two adjacent features is approximately 0.1um to 500um, 0.5um to 500um, 1um to 200um, 1um to 100um, 5um to 200um, 5um to 100um, 5um to 50um, or 5um to 30um, for example, 20um. In some examples, the total width of a feature is approximately 5um, 10um, 20um, 30um, 40um, 50um, 60um, 70um, 80um, 90um, or 100um. In some examples, the total width of a feature is approximately 1um to 100um, 30um to 100um, or 50um to 70um.In some examples, the distance between the centers of two adjacent features is approximately 0.5um to 2um, 0.5um to 2um, 0.75um to 2um, 1um to 2um, 0.2um to 1um, 0.5um to 1.5um, 0.5um to 0.8um, or 0.5um to 1um, for example, 1um. In some examples, the total width of a feature is approximately 50nm, 0.1um, 0.2um, 0.3um, 0.4um, 0.5um, 0.6um, 0.7um, 0.8um, 0.9um, 1um, 1.1um, 1.2um, 1.3um, 1.4um, or 1.5um. In some examples, the total width of a feature is approximately 0.5um to 2um, 0.75um to 1um, or 0.9um to 2um.
[0087] In some examples, each feature facilitates the synthesis of a population of oligonucleotides having a different sequence from the population of oligonucleotides grown on other features. Surfaces containing at least 10, 100, 256, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000, or more clusters are provided herein. Surfaces containing 2,000; 5,000; 10,000; 20,000; 30,000; 50,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 5,000,000; or 10,000,000 or more distinct features are provided herein. In some cases, each cluster contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 500, or more features. In some cases, each cluster contains 50–500, 50–200, 50–150, or 100–150 features. In other cases, each cluster contains 100–150 features. In the exemplary configuration, each cluster contains 109, 121, 130, or 137 features.
[0088] Provided herein are features having a longest segment width of 5 to 100 um. In some cases, the features have a longest segment width of about 30, 35, 40, 45, 50, 55, or 60 um. In some cases, the feature is a channel having a plurality of segments, wherein each segment has a center-to-center distance of 5 to 50 um. In some cases, the center-to-center distance of each segment is about 5, 10, 15, 20, or 25 um.
[0089] In some examples, the number of distinct oligonucleotides synthesized on the surface of a structure described herein depends on the number of distinct features available in the substrate. In some examples, the density of features within a cluster of a substrate is about 1 mm 2 per, at least or about 1 feature, 1 mm 2 per, at least or about 10 features, 1 mm 2 per, at least or about 25 features, 1 mm 2 per, at least or about 50 features, 1 mm 2 per, at least or about 65 features, 1 mm 2 per, at least or about 75 features, 1 mm 2 per, at least or about 100 features, 1 mm 2 per, at least or about 130 features, 1 mm 2 per, at least or about 150 features, 1 mm 2 per, at least or about 175 features, 1 mm 2 per, at least or about 200 features, 1 mm 2 per, at least or about 300 features, 1 mm 2 per, at least or about 400 features, 1 mm 2 per, at least or about 500 features, 1 mm 2 per, at least or about 1,000 features, or more. In some cases, the substrate has 1 mm 2 per 10 features to 1 mm 2 per about 500 features, 1 mm 2 per about 25 features to 1 mm 2Approximately 400 features per unit, 1 mm 2 Approximately 50 features per unit ~ 1mm 2 Approximately 500 features per unit, 1 mm 2 Approximately 100 features per unit ~ 1mm 2 Approximately 500 features per unit, 1 mm 2 Approximately 150 features per unit ~ 1mm 2 Approximately 500 features per unit, 1 mm 2 Approximately 10 characteristics per unit ~ 1mm 2 Approximately 250 features per unit, 1 mm 2 Approximately 50 features per unit ~ 1mm 2 Approximately 250 features per unit, 1 mm 2 Approximately 10 features per unit ~ 1mm 2 Approximately 200 features per unit, or 1 mm 2 Approximately 50 features per unit ~ 1mm 2 Each cluster contains approximately 200 features. In some examples, the distance between the centers of two adjacent features within a cluster is approximately 10um to 500um, 10um to 200um, or 10um to 100um. In some cases, the distance between the centers of two adjacent features is greater than approximately 10um, 20um, 30um, 40um, 50um, 60um, 70um, 80um, 90um, or 100um. In some cases, the distance between the centers of two adjacent features is less than approximately 200um, 150um, 100um, 80um, 70um, 60um, 50um, 40um, 30um, 20um, or 10um. In some cases, the distance between the centers of two adjacent features is approximately 10,000 nm, 8,000 nm, 6,000 nm, 4,000 nm, 2,000 nm, 1,000 nm, 800 nm, 600 nm, 400 nm, 200 nm, 150 nm, 100 nm, 80 μm, 70 nm, 60 nm, 50 nm, 40 nm, 30 nm, 20 nm, or 10 nm. In some examples, each square meter of the structure described herein is at least approximately 10 7 , 10 8 , 10 9 , 10 10 , 10 11 This enables the following features, where each feature supports one oligonucleotide. In some examples, 10 9The oligonucleotides are approximately 6, 5, 4, 3, 2, or 1 m of the structure described herein. 2 It is supported by those who are less than [number].
[0090] Structures for oligonucleotide synthesis for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein assist in the synthesis of a number of oligonucleotides. In some examples, the structures described herein are for 2,000; 5,000; 10,000; 20,000; 30,000; 50,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1, It provides support for the synthesis of more than 200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; and 10,000,000 non-identical oligonucleotides. In some cases, the structure encodes separate arrays: 2,000;5,000;10,000;20,000;50,000;100,000;200,000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1, It provides support for the synthesis of 200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; and 10,000,000 or more oligonucleotides. In some examples, at least a portion of the oligonucleotides have the same sequence or are configured to be synthesized with the same sequence. In some examples, the structure provides a surface environment for the growth of oligonucleotides having at least about 50, 60, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more bases.
[0091] In some examples, oligonucleotides are synthesized on distinct structural features, where each feature aids in the synthesis of a population of oligonucleotides. In some cases, each feature aids in the synthesis of a population of oligonucleotides having a different sequence from the population of oligonucleotides growing on a different locus. In some examples, structural features are located within multiple clusters. In some examples, the structure contains at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000, or more clusters. In some examples, the structure is 2,000;5,000;10,000;100,000;200,000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1,100,000;1,200,000;1,300,000;1,400,000;1,500,000;1,600,000;1,700,000;1,800,000;1,900,000;2,0 00,000;300,000;400,000;500,000;600,000;700,000;800,000;900,000;1,000,000;1,200,000;1,400,000;1,600,000;1,800,000;2,000,000;2,500,000;3,000,000;3,500,000;4,000,000;4,500,000;5,000,000; or containing 10,000,000 or more unique features. In some cases, each cluster contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, or more features (locations). In some examples, each cluster contains 50-500, 100-150, or 100-200 features. In some examples, each cluster contains 109, 121, 130, or 137 features. In some examples, each cluster contains 5, 6, 7, 8, 9, 10, 11, or 12 features.In some cases, oligonucleotides from distinct features within a single cluster, when constructed, have sequences that encode adjacent, longer oligonucleotides of a predetermined sequence.
[0092] Size of the structure
[0093] Structures for oligonucleotide synthesis used with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein include a range of sizes. In some examples, the structures described herein are roughly the size of a standard 96-well plate (e.g., about 100–200 mm × about 50–150 mm). In some examples, the structures described herein have diameters of about 1000 mm, 500 mm, 450 mm, 400 mm, 300 mm, 250 nm, 200 mm, 150 mm, 100 mm, or less than 50 mm. In some examples, the substrate diameters are about 25 mm–1000 mm, about 25 mm–800 mm, about 25 mm–600 mm, about 25 mm–500 mm, about 25 mm–400 mm, about 25 mm–300 mm, or about 25 mm–200 mm. Non-limiting examples of substrate size include approximately 300 mm, 200 mm, 150 mm, 130 mm, 100 mm, 76 mm, 51 mm, and 25 mm. In some examples, the substrate is at least approximately 100 mm. 2 ;200mm 2 500mm 2 ;1000mm 2 ;2000mm 2 5000mm 2 ;10000mm 2 ;12000mm 2 ;15000mm 2 ;20000mm 2 30,000 mm 2 40,000 mm 2 50,000mm 2It has a planar surface area of 250mm or more. In some examples, the substrate thickness is approximately 50mm to 2000mm, 50mm to 1000mm, 100mm to 1000mm, 200mm to 1000mm, and 250mm to 1000mm. Non-limiting examples of substrate thickness include 275mm, 375mm, 525mm, 625mm, 675mm, 725mm, 775mm, and 925mm. In some examples, the substrate thickness varies with diameter and depends on the composition of the substrate. For example, a structure containing a material other than silicon may have a different thickness than a silicon structure of the same diameter. The thickness of a structure can be determined by the mechanical strength of the material used, and the structure must be thick enough to support its own weight without cracking during handling. In some examples, the structure exceeds approximately 1, 2, 3, 4, 5, 10, 15, 30, 40, and 50 feet in any one dimension.
[0094] material
[0095] Structures for oligonucleotide synthesis for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may be made from a variety of materials. In some examples, the material from which the substrate / solid support of this disclosure is made exhibits low levels of oligonucleotide binding. In some situations, materials transparent to visible light and / or ultraviolet light are used. Sufficiently conductive materials, e.g., those capable of forming a uniform electric field over all or part of the substrate / solid support described herein, may be used. In some examples, such materials may be connected to an electric ground. In some cases, the substrate or solid support may be thermally conductive or thermally insulating. The materials are chemically and thermally resistant to support chemical or biochemical reactions, such as a series of oligonucleotide synthesis reactions. For flexible materials, the materials in question may include: nylon, modified and unmodified nitrocellulose, polypropylene, etc.
[0096] Regarding rigid materials, the specific materials in question include: glass; quartz glass; silicon, plastics (e.g., polytetraflouroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof); and metals (e.g., gold, platinum, etc.). Structures are made from materials selected from the group consisting of silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), and glass. The substrate / solid support or microstructure, its reactor, may be manufactured from combinations of the materials listed herein, or any other suitable materials known in the art.
[0097] The term “flexible” is used herein to refer to a structure that can be bent, folded, or similarly manipulated without breaking. In some cases, the flexible structure is bent at least 30 degrees around a roller. In some cases, the flexible structure is bent at least 180 degrees around a roller. In some cases, the flexible structure is bent at least 270 degrees around a roller. In some examples, the flexible structure is bent about 360 degrees around a roller. In some cases, the radius of the roller is less than about 10 cm, 5 cm, 3 cm, 2 cm, or 1 cm. In some examples, the flexible structure is repeatedly bent at least 100 times in either direction and straightened without breaking (e.g., cracking) or deforming at 20°C. In some examples, the flexible structures described herein have a thickness suitable for rotation. In some cases, the thickness of the flexible structures described herein is less than about 50 mm, 10 mm, 1 mm, or 0.5 mm.
[0098] Examples of flexible materials for the structures described herein include, but are not limited to, nylon (unmodified nylon, modified nylon, clear nylon), nitrocellulose, polypropylene, polycarbonate, polyethylene, polyurethane, polystyrene, acetal, acrylic, acrylonitrile, butadiene styrene (ABS), polyester films (such as polyethylene terephthalate, polymethyl methacrylate, or other acrylics), polyvinyl chloride or other vinyl resins, clear PVC foil, clear foil for printers, poly(methyl methacrylate) (PMMA), methacrylate copolymers, styrene polymers, high refractive index polymers, fluorine-containing polymers, polyethersulfones, polyimides including alicyclic structures, rubber, fabric, metal foil, and any combination thereof. Various plasticizers and modifiers may be used with the polymer substrate material to achieve the selected flexibility characteristics.
[0099] The flexible structures described herein may include plastic materials. In some examples, the flexible structures may include thermoplastic materials. Non-limiting examples of thermoplastic materials include acrylic, acrylonitrile butadiene styrene, nylon, polylactic acid, polybenzimidazole, polycarbonate, polyethersulfone, polyether ether ketone, polyetherimide, polyethylene, polyphenylene oxide, polyphenylene sulfide, polypropylene, polystyrene, polyvinyl chloride, and polytetrafluoroethylene. In some examples, the substrate includes thermoplastic materials of the polyaryl ether ketone (PEAK) family. Non-limiting examples of PEAK thermoplastics include polyether ketone (PEK), polyether ketone ketone (PEKK), poly(ether ether ketone·ketone) (PEEKK), polyether ether ketone (PEEK), and polyether ketone ether ketone ketone (PEKEKK). In some examples, the flexible structures include thermoplastic materials compatible with toluene. In some cases, the flexibility of plastic materials is increased by the addition of plasticizers. Examples of plasticizers include ester-based plasticizers such as phthalates. Phthalate ester plasticizers include bis(2-ethylhexyl)phthalate (DEHP), diisononyl phthalate (DINP), di-n-butyl phthalate (DnBP, DBP), butyl benzyl phthalate (BBzP), diisodecyl phthalate (DIDP), dioctyl phthalate (DOP, DnOP), diisooctyl phthalate (DIOP), diethyl phthalate (DEP), diisobutyl phthalate (DIBP), and di-n-hexyl phthalate. In some cases, modification of thermoplastic polymers, either through copolymerization or by adding unreactive side chains to monomers before polymerization, further increases flexibility.
[0100] Flexible structures further containing fluororubber are provided herein. Materials having about 80% fluororubber are designated as FKM. Fluororubber includes perfluoroelastomers (FFKM) and tetrafluoroethylene / propylene rubber (FEPM). There are five known types of fluororubber. Type 1 FKM consists of vinylidene fluoride (VDF) and hexafluoropropylene (HFP), and their fluorine content is typically about 66% by weight. Type 2 FKM consists of VDF, HFP, and tetrafluoroethylene (TFE), and typically has about 68%–69% fluorine. Type 3 FKM consists of VDF, TFE, and perfluoromethyl vinyl ether (PMVE), and typically has about 62%–68% fluorine. Type 4 FKM consists of propylene, TFE, and VDF, and typically has about 67% fluorine. Type 5 FKM consists of VDF, HFP, TFE, PMVE, and ethylene.
[0101] In some examples, the substrates disclosed herein include computer-readable materials. Computer-readable materials include, but are not limited to, magnetic media, open-reel tapes, cartridge tapes, cassette tapes, flexible disks, paper media, films, microfiche, continuous tapes (e.g., belts), and any medium suitable for storing electronic instructions. In some cases, the substrate includes magnetic open-reel tape or magnetic belts. In some examples, the substrate includes flexible printed circuit boards.
[0102] The structures described herein may transmit visible light and / or ultraviolet light. In some examples, the structures described herein are electrically conductive enough to form a uniform electric field throughout all or part of the structure. In some examples, the structures described herein are thermally conductive or thermally insulating. In some examples, the structures are chemically and thermally resistant to facilitate chemical reactions such as oligonucleotide synthesis reactions. In some examples, the structures are magnetic. In some examples, the structures contain metals or metallic alloys.
[0103] Structures for oligonucleotide synthesis may be 1, 2, 5, 10, 30, 50 feet, or longer in any dimension. In the case of flexible structures, they may be optionally stored in a wound state, for example, on a reel. In the case of large rigid structures (e.g., longer than 1 foot), the rigid structures may be stored vertically or horizontally.
[0104] Cryptographic key marking on the surface of the structure
[0105] Structures for oligonucleotide synthesis used with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may include cryptographic markings. Structures having markings (1401) are provided herein, where the markings provide information relating to a source item of information associated with a population of nearby oligonucleotides, a cryptographic scheme for decrypting the sequence of a population of nearby oligonucleotides, the copy number of a population of nearby oligonucleotides, or any combination thereof. See, for example, Figures 14B–14C. The markings may be visible to the naked eye or to a magnified view using a microscope. In some examples, markings on a surface may only be visible after processing conditions for exposure to the markings, such as heat, chemical, or phototreatment (e.g., UV or IR light for illuminating the markings). Exemplary inks developed by heat include, but are not limited to, cobalt chloride (which turns blue when heated). Exemplary inks developed through chemical reactions include, but are not limited to, phenolphthalein, copper sulfate, lead(II) nitrate, cobalt(II) chloride, and cerium oxalate developed with manganese sulfate and hydrogen peroxide.
[0106] Surface treatment
[0107] Structures for oligonucleotide synthesis used in conjunction with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein may include surfaces for oligonucleotide synthesis. Methods for assisting the immobilization of biomolecules on a substrate are provided herein, wherein the surface of the structure described herein includes a material and / or is coated with a material that facilitates coupling reactions with biomolecules for attachment. To prepare structures for biomolecule immobilization, surface modification can be utilized to chemically and / or physically modify the surface of the substrate by additive or subtractive methods for altering one or more chemical and / or physical properties of the surface of the substrate or selected sites or regions of the surface. For example, surface modification includes (1) altering the wettability of a surface, (2) functionalizing a surface (i.e., providing, modifying, or substituting functional groups on the surface), (3) defunctionalizing a surface (i.e., removing functional groups on the surface), (4) otherwise altering the chemical composition of a surface (e.g., by etching), (5) increasing or decreasing surface roughness, (6) providing a coating on a surface (e.g., a coating exhibiting different wettability from the surface), and / or (7) depositing particles on the surface. In some examples, the surface of a structure is selectively functionalized to generate two or more distinct regions on the structure, where at least one region has different surface or chemical properties from another region of the same structure. Such properties include, but are not limited to, surface energy, chemical termination, and chemical partial surface concentration.
[0108] In some examples, the surface of the structures disclosed herein is modified to include one or more actively functionalized surfaces configured to bind to both the substrate and the biomolecule, thereby supporting surface coupling reactions. In some examples, the surface is functionalized with a passive material that does not efficiently bind biomolecules, thereby preventing adhesion of biomolecules at the site where the passive functionalizer is bound. In some cases, the surface includes an active layer that merely defines distinct features for supporting biomolecules.
[0109] In some examples, the surface comes into contact with a mixture of functionalizing groups in any different ratios. In some examples, the mixture contains at least two, three, four, five, or more different types of functionalizing agents. In some cases, the ratio of at least two types of surface functionalizing agents in the mixture is about 1:1, 1:2, 1:5, 1:10, 2:10, 3:10, 4:10, 5:10, 6:10, 7:10, 8:10, 9:10, or any other ratio to achieve the desired surface representation of the two groups. In some examples, the desired surface tension, wettability, water contact angle, and / or contact angle of other suitable solvents are achieved by providing a substrate surface with an appropriate ratio of functionalizing agents. In some cases, the agents in the mixture are selected from appropriate reactive and inert parts, thus diluting the surface density of reactive groups to a desired level for downstream reactions. In some examples, the mixture of functionalizing reagents contains one or more reagents that bind to biomolecules and one or more reagents that do not bind to biomolecules. Therefore, reagent adjustment allows for control over the amount of biomolecular binding occurring in distinct regions of functionalization.
[0110] In some examples, the method for functionalizing a substrate involves depositing silane molecules onto the surface of the substrate. The silane molecules may be deposited on the high-energy surface of the substrate. In some examples, the high-surface-energy region includes a passive functionalizing reagent. The methods described herein provide silane groups that bind the surface, while the remaining molecules provide the distance from the surface and the free hydroxyl groups at the ends to which biomolecules attach. In some examples, the silane is an organofunctionalized alkoxysilane molecule. Non-limiting examples of organofunctionalized alkoxysilane molecules include dimethylchloro-octodecyl-silane, methyldichloro-octodecyl-silane, trichloro-octodecyl-silane, and trimethyl-octodecyl-silane, triethyl-octodecyl-silane. In some examples, the silane is an aminosilane. Examples of aminosilanes include, but are not limited to, 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, glycidyloxypropyl / trimethoxysilane, and N-(3-triethoxysilylpropyl)-4-hydroxybutylamide. In some examples, the silane includes 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, glycidyloxypropyl / trimethoxysilane, N-(3-triethoxysilylpropyl)-4-hydroxybutylamide, or any combination thereof. In some examples, the active functionalizing agent includes 11-acetoxyundecyltriethoxysilane. In some examples, the active functionalizing agent includes n-decyltriethoxysilane. In some cases, the active functionalizing agent includes glycidyloxypropyltriethoxysilane (GOPS). In some examples, the silane is a fluorosilane. In some examples, the silane is a hydrocarbon silane. In some cases, the silane is 3-iodopropyltrimethoxysilane. In some cases, the silane is octylchlorosilane.
[0111] In some cases, silane treatment is carried out on a surface via self-construction with organically functionalized alkoxysilane molecules. Organofunctionalized alkoxysilanes are classified according to their organic functional groups. Non-limiting examples of siloxane functionalizing reagents include hydroxyalkylsiloxanes (silylated surface, functionalized with diborane, and oxidized alcohol with hydrogen peroxide), diol(hydroxyalkyl)siloxanes (silylated surface and hydrolysis to diol), aminoalkylsiloxanes (amines do not require an intermediate functionalization step), glycidoxysilanes (3-glycidoxypropyl-dimethyl-ethoxysilane, glycidoxy-trimethoxysilane), mercaptosilanes (3-mercaptopropyl-trimethoxysilane, 3-4, epoxycyclohexyl-ethyltrimethoxysilane or 3-mercaptopropyl-methyl-dimethoxysilane), bicycloheptahenyl-trichlorosilane, butyl-aldehyde-trimethoxysilane, or dimeric secondary aminoalkylsiloxanes. Exemplary hydroxyalkylsiloxanes include allyltrichlorochlorosilane, which is converted to 3-hydroxypropyl, or 7-octo-1-enyltrichlorochlorosilane, which is converted to 8-hydroxyoctyl. Diol(hydroxyalkyl)siloxanes include (2,3-dihydroxypropyloxy)propyl (GOPS) derived from glycidyl, trimethoxysilane. Aminoalkylsiloxanes include 3-aminopropyltrimethoxysilane, which is converted to 3-aminopropyl (3-aminopropyl-triethoxysilane, 3-aminopropyl-diethoxy-methylsilane, 3-aminopropyl-dimethyl-ethoxysilane, or 3-aminopropyl-trimethoxysilane). In some cases, the dimeric secondary aminoalkylsiloxane is bis(3-trimethoxysilylpropyl)amine, which is converted to bis(silyloxypropyl)amine.
[0112] The actively functionalized region may contain one or more different species of silanes, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some cases, one of the silanes is present in a larger amount than the other silanes in the functionalized composition. For example, a mixed silane solution containing two silanes may include the following ratios of one silane to another: 99:1, 98:2, 97:3, 96:4, 95:5, 94:6, 93:7, 92:8, 91:9, 90:10, 89:11, 88:12, 87:13, 86:14, 85:15, 84:16, 83:17, 82:18, 81:19, 80:20, 75:25, 70:30, 65:35, 60:40, 55:45. In some cases, the active functionalizing agent contains 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane. In some cases, the active functionalizing agent has 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane in ratios of approximately 20:80 to approximately 1:99, approximately 10:90 to approximately 2:98, or approximately 5:95.
[0113] In some examples, functionalization includes the deposition of functionalizing agents onto a structure by any deposition technique, which is not limited to chemical vapor deposition (CVD), atomic layer deposition (ALD), plasma CVD (PECVD), plasma ALD (PEALD), metal-organic CVD (MOCVD), hot-wire CVD (HWCVD), latent CVD (iCVD), modification CVD (MCVD), vapor axial deposition (VAD), external vapor deposition (OVD), physical deposition (e.g., sputtering, vapor deposition), and molecular layer deposition (MLD).
[0114] Any step or component in the following functionalization process may be omitted or modified according to the desired properties of the final functionalized substrate. In some cases, additional components and / or process steps may be added to the process workflow embodied herein. In some examples, the substrate is first washed using, for example, a piranha solution. An example of a washing step includes immersing the substrate in a piranha solution (e.g., 90% H2SO4, 10% H2O2) at a high temperature (e.g., 120°C), washing (e.g., with water), and drying (e.g., with nitrogen gas). The process optionally includes a piranha post-treatment, which includes immersing the piranha-treated substrate in a basic solution (e.g., NH4OH), followed by an aqueous wash (e.g., with water). In some examples, the surface of the structure is plasma-cleaned optionally following piranha immersion and optional piranha post-treatment. An example of a plasma cleaning process includes oxygen plasma etching. In some examples, the surface is deposited in an active functionalizing agent following steam therapy. In some cases, the substrate is actively functionalized prior to washing, for example, by piranha treatment and / or plasma washing.
[0115] The process for surface functionalization optionally includes resist coating and resist stripping. In some examples, after active surface functionalization, the substrate was spin-coated with a resist (e.g., SPR(trademark) 3612 positive photoresist). In various examples, the process for surface functionalization includes lithography with patterned functionalization. In some examples, photolithography is performed following resist coating. In some examples, after lithography, the surface is visually inspected for lithographic defects. In some examples, the surface functionalization process includes a cleaning step, thereby removing substrate residue by, for example, plasma cleaning or etching. In some examples, the plasma cleaning step is performed in a step after the lithography step.
[0116] In some cases, a surface coated with a resist is treated to remove the resist, for example, after functionalization and / or lithography. In some cases, the resist is removed with a solvent (e.g., a stripping solution containing N-methyl-2-pyrrolidone). In some cases, resist stripping involves ultrasonic or sonication. In some cases, the resist is coated, stripped, and then the exposed surface is actively functionalized to create a desired different functionalization pattern.
[0117] In some examples, the methods and compositions described herein relate to the application of photoresists for the generation of modified surface properties in selective regions, where the application of the photoresist relies on the fluid properties of the surface that define the spatial distribution of the photoresist. Without being bound by theory, surface tension effects related to the applied fluid can define the flow of the photoresist. For example, surface tension and / or capillary effects may facilitate the attraction of the photoresist into the small structure in a restrained manner before the resist solvent evaporates. In some examples, the resist contacts are abutted with sharp edges to control the flow of the fluid. Substructures may be designed based on the desired flow pattern used to apply the photoresist during the manufacturing and functionalization processes. The solid organic layer left after the solvent evaporates may be used to advance the subsequent manufacturing process steps. Structures may be designed to control the flow of fluid by promoting or inhibiting wicking effects to neighboring fluid pathways. For example, a structure may be designed to avoid overlap between the upper and lower edges, which facilitates fluid retention in the superstructure and allows for a specific arrangement of the resist. In other examples, the upper and lower edges overlap, resulting in wicking of the applied fluid into the substructure. Therefore, the appropriate design may be selected depending on the desired application of the resist.
[0118] In some examples, the structures described herein have a surface comprising a material having a thickness of at least, or at least about 0.1 nm, 0.5 nm, 1 nm, 2 nm, 5 nm, 10 nm, or 25 nm, which contains reactive groups capable of binding nucleosides. Typical surfaces include, but are not limited to, glass and silicon, such as silicon dioxide and silicon nitride. In some cases, exemplary surfaces include nylon and PMMA.
[0119] In some examples, electromagnetic radiation in the form of ultraviolet light is used for surface patterning. In some examples, a lamp is used for surface patterning, and a mask mediates the location of ultraviolet light exposure to the surface. In some examples, a laser is used for surface patterning, and the opening and closing state of a shutter controls the irradiation of ultraviolet light to the surface. The laser configuration may be used in combination with a movable, flexible structure. In such a configuration, adjustment of laser exposure and the movement of the flexible structure is used to generate patterns of one or more agents having different nucleoside coupling functions.
[0120] Material deposition system
[0121] Systems and devices for the deposition and storage of biomolecules on structures described herein are provided herein. In some examples, the biomolecules are oligonucleotides that store information encoded in their sequence. In some examples, the system includes a device for applying biomolecules to the surface of a structure and / or a substrate to assist in the adhesion of the biomolecules. In one example, the device for applying biomolecules is an oligonucleotide synthesizer. In some examples, the system includes a device for processing the substrate with a fluid (e.g., a flow cell). In some examples, the system includes a device for moving the substrate between an application device and a processing device. For example, if the substrate is an open-reel tape, the system may include two or more reels that allow access to different parts of the substrate to the application device and an optional processing device at different times.
[0122] A first embodiment of an oligonucleotide material deposition system for oligonucleotide synthesis is shown in Figure 15. The system includes a material deposition device that moves in the XY direction to align with the substrate. The material deposition device can also move in the Z direction to clog with the substrate and form a decomposed reactor. The decomposed reactor is configured to allow a fluid containing oligonucleotides and / or reagents to move from the substrate to the capping element and / or vice versa. As shown in Figure 15, the fluid passes through one or both the substrate and the capping element and includes, but is not limited to, coupling reagents, capping reagents, oxidizing agents, deblocking agents, acetonitrile, and nitrogen gas. Examples of devices capable of high-resolution droplet deposition include print heads for inkjet and laser printers. Devices useful for the systems and methods described herein achieve resolutions of approximately 100 dots per inch (DPI) to approximately 50,000 DPI; approximately 100 DPI to approximately 20,000 DPI; approximately 100 DPI to approximately 10,000 DPI; approximately 100 DPI to approximately 5,000 DPI; approximately 1,000 DPI to approximately 20,000 DPI; or approximately 1,000 DPI to approximately 10,000 DPI. In some examples, the devices have resolutions of at least approximately 1,000; 2,000; 3,000; 4,000; 5,000; 10,000; 12,000 DPI, or 20,000 DPI. The high-resolution deposition performed by the devices is related to the number and density of each nozzle corresponding to the characteristics of the substrate.
[0123] Figure 16 shows an exemplary process workflow for the de novo synthesis of oligonucleotides on a substrate using an oligonucleotide synthesizer. Droplets containing oligonucleotide synthesis reagents are progressively released onto the substrate from a material deposition device, which has a piezoelectric material and electrodes for converting electrical signals into mechanical signals for releasing droplets. The droplets are released one nucleic acid base at a time at specific locations on the surface of the substrate, producing multiple synthesized oligonucleotides having a predetermined sequence that encodes data. In some cases, the synthesized oligonucleotides are stored on the substrate. The nucleic acid reagents may be deposited on the substrate surface in a non-linear or drop-on-demand manner. Examples of such methods include electromechanical transfer, electroheat transfer, and electrostatic attraction. In the electromechanical transfer method, a piezoelectric element deformed by an electrical pulse releases the droplets. In the electroheat transfer method, bubbles are generated in the chamber of the device, and the expansion force of the bubbles releases the droplets. In the electrostatic attraction method, the electrostatic force of attraction is used to release droplets onto the substrate. In some cases, the droplet frequencies are approximately 5kHz to 500kHz; approximately 5kHz to 100kHz; approximately 10kHz to 500kHz; approximately 10kHz to 100kHz; or approximately 50kHz to 500kHz. In other cases, the frequencies are less than approximately 500kHz, 200kHz, 100kHz, or 50kHz.
[0124] The size of the dispensed droplets correlates with the device's resolution. In some examples, the device deposits reagent droplets of approximately 0.01 pl to 20 pl, 0.01 pl to 10 pl, 0.01 pl to 1 pl, 0.01 pl to 0.5 pl, 0.01 pl to 0.01 pl, or 0.05 pl to 1 pl. In some examples, droplet sizes are smaller than approximately 1 pl, 0.5 pl, 0.2 pl, 0.1 pl, or 0.05 pl. The size of the droplets dispensed by the device correlates with the diameter of the deposition nozzle, where each nozzle can deposit the reagent onto the substrate features. In some examples, the oligonucleotide synthesizer in a deposition apparatus includes approximately 100 to 10,000 nozzles; approximately 100 to 5,000 nozzles; approximately 100 to 3,000 nozzles; approximately 500 to 10,000 nozzles; or approximately 100 to 5,000 nozzles. In some cases, the deposition apparatus includes 1,000; 2,000; 3,000; 4,000; 5,000; or more than 10,000 nozzles. In some examples, each material deposition apparatus includes multiple nozzles, each optionally configured to correspond to a feature on the substrate. Each nozzle deposits a different reagent component than another nozzle. In some examples, each nozzle deposits a droplet covering one or more features of the substrate. In some examples, one or more nozzles are angled. In some examples, multiple deposition apparatuses are stacked side by side to achieve a doubling of throughput. In some cases, the increase can be twofold, fourfold, eightfold, or even more. An example of a deposition device is the Samba Printhead (Fujifilm). The Samba Printhead may be used in conjunction with the Samba Web Administration Tool (SWAT).
[0125] The number of deposition sites can be increased by using the same deposition apparatus and rotating it by a specific angle or saber angle. By rotating the deposition apparatus, each nozzle is ejected with a fixed delay time corresponding to the saber angle. Asynchronous ejection creates crosstalk between nozzles. Therefore, if droplets are ejected at a specific saber angle different from 0 degrees, the volume of droplets from the nozzles may differ.
[0126] In some configurations, the oligonucleotide synthesis system configuration enables a continuous oligonucleotide synthesis process that leverages the flexibility of the substrate for movement in an open-reel process. This synthesis process operates in a continuous production line manner with the substrate passing through various steps of oligonucleotide synthesis, using one or more reels to rotate the position of the substrate. In a typical embodiment, the oligonucleotide synthesis reaction involves rotating the substrate: passing through a solvent tank, under a deposition apparatus for phospholamidite deposition, an oxidizing agent tank, an acetonitrile washing tank, and a non-blocking tank. Optionally, the tape also passes through a capping tank. The open-reel process allows the final product of the substrate, including the synthesized oligonucleotide, to be easily collected on a winding reel, where it is moved for further processing or storage.
[0127] In some configurations, oligonucleotide synthesis proceeds in a continuous production manner, as a continuous flexible tape is transported along a conveyor belt system. Similar to open-reel processes, oligonucleotide synthesis on a continuous tape operates in a production line manner, with the substrate passing through various steps of oligonucleotide synthesis during transport. However, in conveyor belt processes, the continuous tape reconsiders the oligonucleotide synthesis process without the rolling and unrolling of the tape, as in open-reel processes. In some configurations, the oligonucleotide synthesis process is divided into zones, and the continuous tape is transported through each zone one or more times per cycle. For example, an oligonucleotide synthesis reaction involves (1) transporting the substrate through a solvent tank, under a deposition apparatus for phosphoramidite deposition, an oxidizing agent tank, an acetonitrile water washing tank, and a block tank in one cycle; then (2) repeating the cycle to achieve a predetermined length of synthesized oligonucleotide. After oligonucleotide synthesis, the flexible substrate is removed from the conveyor belt system and rotated as needed for storage. Rolling may be done around the reel for storage purposes.
[0128] In a typical configuration, a flexible substrate containing thermoplastic material is coated with a nucleoside coupling reagent. The coating is patterned to the features such that each feature has a diameter of approximately 10 μm and an intercenter distance of approximately 21 μm between two adjacent features. In this example, the size of the features is sufficient to accommodate a droplet volume of 0.2 pl of deposition during the deposition process of oligonucleotide synthesis. In some cases, the density of the features is m 2 Each feature has approximately 2.2 billion characteristics (1 characteristic / 441 x 10⁻¹⁰). -12 m 2 ). In some cases, 4.5m 2 The substrate contains approximately 10 billion features, each with a diameter of 10 μm.
[0129] The deposition apparatus described herein may include approximately 2,048 nozzles, each depositing 100,000 droplets per second, with 1 nucleic acid base per droplet. For each deposition apparatus, at least approximately 1.75 × 10⁻¹⁶ per day 13 Nucleic acid bases are deposited on the substrate. In some cases, 100 to 500 nucleic acid base oligonucleotides are synthesized. In other cases, 200 nucleic acid base oligonucleotides are synthesized. Optionally, over 3 days, 1.75 × 10¹⁶ times per day. 13 In terms of the proportion of bases, at least approximately 262.5 × 10 9 The oligonucleotide is synthesized.
[0130] In some configurations, devices for applying one or more reagents to a substrate during a synthetic reaction are configured to deposit reagents and / or nucleotide monomers for nucleoside phosphoramidite-based synthesis. Reagents for oligonucleotide synthesis include reagents for oligonucleotide extension and washing buffers. In non-limiting examples, devices deposit washing reagents, coupling reagents, capping reagents, oxidizing agents, non-blocking agents, gases such as acetonitrile and nitrogen gas, and any combination thereof. In addition, devices optionally deposit reagents to prepare and / or maintain the integrity of the substrate. In some examples, oligonucleotide synthesizers deposit droplets with diameters smaller than approximately 200 μm, 100 μm, or 50 μm in volumes smaller than approximately 1000, 500, 100, 50, or 20 pl. In some cases, oligonucleotide synthesizers deposit approximately 1–10000, 1–5000, 100–5000, or 1000–5000 droplets per second.
[0131] In some configurations, during oligonucleotide synthesis, the substrate is placed in and / or sealed within a flow cell. The flow cell provides a continuous or discontinuous flow of liquids, such as those containing reagents (e.g., oxidizing agents and / or solvents) required for the reaction within the substrate. The flow cell can also provide a continuous or discontinuous flow of gases, such as nitrogen, to dry the substrate by enhanced evaporation, which is typically volatile. Various auxiliary devices are useful to improve drying and reduce residual moisture on the substrate surface. Examples of such auxiliary drying devices include, but are not limited to, vacuum sources, vacuum pumps, and vacuum tanks. In some cases, an oligonucleotide synthesis system includes one or more flow cells, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, and one or more substrates, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20. In some cases, the flow cell is configured to hold and supply reagents to the substrate during one or more steps of the synthesis reaction. In some examples, the flow cell includes a lid that slides over the top of the substrate and is secured in place to form a pressure-resistant seal around the edges of the substrate. A suitable seal includes, but is not limited to, a seal that allows for approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 atmospheres. In some cases, the lid of the flow cell can be opened to allow access to an application device, such as an oligonucleotide synthesizer. In some cases, one or more steps of an oligonucleotide synthesis method are performed on the substrate within the flow cell without substrate transport.
[0132] In some configurations, the device for processing a substrate with a fluid includes a spray bar. Nucleotide monomers are coated onto the substrate surface, and the spray bar, using its spray nozzles, sprays the substrate surface with one or more processing reagents. In some configurations, the spray nozzles are sequentially directed to correlate with different processing steps during oligonucleotide synthesis. The chemicals used in the various process steps can be modified in the spray bar to easily adapt to changes in the synthesis method or between steps of the synthesis method. In some examples, the spray bar continuously sprays a given chemical onto the substrate surface as the substrate passes through the spray bar. In some cases, the spray bar deposits over a wide area of the substrate, similar to a spray bar used in a lawn sprinkler. In some examples, the nozzles of the spray bar are positioned to provide a uniform coating of the processing material over a given area of the substrate.
[0133] In some examples, an oligonucleotide synthesis system includes one or more elements useful for downstream processing of the synthesized oligonucleotide. For example, the system includes a temperature-controlled element, such as a thermal cycling device. In some examples, the temperature-controlled element is used with multiple degraded reaction devices to carry out nucleic acid construction, such as PCA, and / or nucleic acid amplification, such as PCR.
[0134] De Novo OCR Synthesis
[0135] Systems and methods for the rapid synthesis of high-density oligonucleotides on a substrate, for use with devices, compositions, systems, and methods for bioencryption and / or biodecryption as described herein, are provided herein. In some examples, the substrate is a flexible substrate. In some examples, at least about 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , or 10 15The bases are synthesized in one day. In some cases, at least about 10 × 10 8 , 10×10 9 , 10×10 10 , 10×10 11 , or 10 x 10 12The oligonucleotides are synthesized in one day. In some cases, each synthesized oligonucleotide contains at least about 20, 50, 100, 200, 300, 400, or 500 nucleic acid bases. In some cases, these bases are synthesized with an average total error rate of less than one per 100 bases; one per 200 bases; one per 300 bases; one per 400 bases; one per 500 bases; one per 1,000 bases; one per 2,000 bases; one per 5,000 bases; one per 10,000 bases; one per 15,000 bases; and one per 20,000 bases. In some examples, these error rates are at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or higher of the synthesized oligonucleotides. In some cases, at least 90%, 95%, 98%, 99%, 99.5%, or more of these synthesized oligonucleotides are identical to the predetermined sequences they encode. In some cases, the error rate for synthetic oligonucleotides on a substrate using the methods and systems described herein is less than approximately 1 in 200. In some cases, the error rate for synthetic oligonucleotides on a substrate using the methods and systems described herein is less than approximately 1 in 1,000. In some cases, the error rate for synthetic oligonucleotides on a substrate using the methods and systems described herein is less than approximately 1 in 2,000. In some cases, the error rate for synthetic oligonucleotides on a substrate using the methods and systems described herein is less than approximately 1 in 3,000. In some cases, the error rate for synthetic oligonucleotides on a substrate using the methods and systems described herein is less than approximately 1 in 5,000. Individual types of error rates include mismatches, deletions, insertions, and / or substitutions of the synthetic oligonucleotides on the substrate. The term “error rate” refers to a comparison between the aggregate amount of synthetic oligonucleotides and the aggregate of predetermined oligonucleotide sequences. In some examples, the synthetic oligonucleotides disclosed herein include a tether of 12 to 25 bases.In some examples, the tether contains 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 43, 44, 45, 46, 47, 48, 49, 50 or more bases.
[0136] A suitable method for the synthesis of oligonucleotides on a substrate according to the present disclosure is a phosphoramidite method comprising the suppressed addition of a phosphoramidite building block (i.e., nucleoside phosphoramidite) to a growing oligonucleotide chain in a coupling step that forms a phosphytotryester bond between the phosphoramidite building block and a nucleoside bound to the substrate. In some examples, the nucleoside phosphoramidite is provided to an activated substrate. In some examples, the nucleoside phosphoramidite is provided to a substrate having an activator. In some examples, the nucleoside phosphoramidite is provided to the substrate in an excess of 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 times more than the substrate-bound nucleoside. In some examples, the addition of nucleoside phosphoramidites is carried out in an anhydrous environment, for example, in anhydrous acetonitrile. Following the addition and binding of nucleoside phosphoramidites in the coupling step, the substrate is optionally washed. In some examples, the binding step is optionally repeated one or more additional times, along with the washing step between adding nucleoside phosphoramidites to the substrate. In some examples, the oligonucleotide synthesis methods used herein include one, two, three, or more consecutive coupling steps. Prior to coupling, the binding of nucleosides to the substrate is often deprotected by the removal of a protecting group that functions to prevent polymerization. A common protecting group is 4,4'-dimethoxytrityl (DMT).
[0137] Following coupling, the phosphoramidite oligonucleotide synthesis method optionally includes a capping step. In the capping step, the grown oligonucleotide is treated with a capping agent. The capping step typically helps to block unreacted substrate-bound 5'-OH groups after coupling from further chain elongation, preventing the formation of oligonucleotides with internal base deletions. Furthermore, phosphoramidites activated with 1H-tetrazole often react to a small extent with the O6 position of guanosine. Without being bound by theory, upon oxidation with I2 / water, this byproduct undergoes depurination, possibly via O6-N7 migration. The aprinic acid moiety may be cleaved during the final deprotection of the oligonucleotide, thus reducing the yield of the full-length product. The O6 modification can be removed by treatment with a capping reagent before oxidation with I2 / water. In some examples, inclusion bodies in the capping step during oligonucleotide synthesis reduce the error rate compared to synthesis without capping. As an example, the capping step involves treating the substrate-bound oligonucleotide with a mixture of acetic anhydride and 1-methylimidazole. Following the capping process, the substrate is optionally washed.
[0138] Following the addition of nucleoside phosphoramide, and optionally after a capping step and one or more washing steps, the substrate-bound growth nucleic acid can be oxidized. The oxidation step involves oxidizing the phosphite triester to a planar four-coordinate phosphate triester, a spontaneously occurring phosphate diester-nucleoside bond-protected precursor. In some examples, oxidation of the growth oligonucleotide is optionally achieved by treatment with iodine and water in the presence of a weak base such as pyridine, lutidine, or colidine. Oxidation is sometimes carried out under anhydrous conditions using tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is carried out following oxidation. The second capping step allows for substrate drying, as residual water from the oxidation, which can persist, can inhibit subsequent coupling. Following oxidation, the substrate and growth oligonucleotide are optionally washed. In some examples, the oxidation step is replaced by a sulfidation step to obtain oligonucleotide phosphorothioates, where the capping step may be performed after sulfidation. Many reagents can efficiently transport sulfur, and these include, but are not limited to, 3-(dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also known as Beaucage's reagent), and N,N,N'N'-tetraethylthiuram disulfide (TETD).
[0139] For the subsequent cycle following nucleoside incorporation to occur via coupling, the protected 5' end of the substrate-bound growth oligonucleotide must be removed so that the primary hydroxyl group can react with the next nucleoside phosphoramidite. In some examples, the protecting group is DMT, and unblocking occurs with trichloroacetic acid in dichloromethane. Performing detritylation for extended periods or in solutions of stronger acids than recommended leads to increased depurination of the solid support-bound oligonucleotide and thus reduces the yield of the desired full-length product. The methods and compositions described herein result in controlled unblocking, limiting undesirable depurination reactions. In some examples, the substrate-bound oligonucleotide is washed after unblocking. In some cases, efficient washing after unblocking contributes to the synthesis of oligonucleotides with a low error rate.
[0140] The methods for the synthesis of oligonucleotides on a substrate described herein typically involve a series of repeated steps: applying a protected monomer to the surface of a feature of the substrate for linking with either a surface, a linker, or a pre-deprotected monomer; deprotecting the applied monomer so that it can subsequently react with the applied protected monomer; and applying another protected monomer for linking. One or more intermediate steps include oxidation and / or sulfurization. In some examples, one or more washing steps precede or follow one or all of the steps.
[0141] In some examples, oligonucleotides are synthesized with a photosensitive protecting group, where the hydroxyl groups generated on the surface are blocked by the photosensitive protecting group. When the surface is exposed to ultraviolet light via a photolithography mask, a pattern of free hydroxyl groups can be generated on the surface. These hydroxyl groups can react with photoprotected nucleoside phosphoramidites according to the phosphoramidite chemical. A second photolithography mask may be applied, and the surface is exposed to ultraviolet light to generate a second pattern of hydroxyl groups, followed by coupling with a 5'-photoprotected nucleoside phosphoramidite. Similarly, a pattern may be generated, and the oligomeric chain may be extended. Without being bound by theory, the instability of the photocleavable groups depends on the wavelength and polarity of the solvent used, and the degree of photocleavage can be affected by the exposure time and luminosity. This method can take advantage of many factors such as the precision of mask alignment, the efficiency of photoprotection group removal, and the yield of the phosphoramidite coupling step. Furthermore, unintended light leakage to neighboring areas can be minimized. The density of synthetic oligomers per spot can be monitored by adjusting the load of leader nucleosides on the synthetic surface.
[0142] The surface of a substrate that provides support for oligonucleotide synthesis may be chemically modified to allow the synthetic oligonucleotide chain to be cleaved from the surface. In some examples, the oligonucleotide chain is cleaved simultaneously with the deprotection of the oligonucleotide. In some cases, the oligonucleotide chain is cleaved after the oligonucleotide has been deprotected. In an exemplary scheme, a trialkoxysilylamine such as (CH3CH2O)3Si-(CH2)2-NH2 is reacted with the surface SiOH group of the substrate and subsequently reacted with succinic anhydride containing the amine to produce an amide bond and free OH that supports nucleic acid chain proliferation. Cleavage involves gas cleavage with ammonia or methylamine. In some examples, once released from the surface, the oligonucleotide is constructed into a larger nucleic acid that is sequenced and decrypted to extract the conserved information.
[0143] Oligonucleotides can be designed to collectively extend over a large region of a predetermined sequence that encodes information. In some examples, larger oligonucleotides are generated by a ligation reaction to ligate synthetic oligonucleotides. One example of a ligation reaction is polymerase chain construction (PCA). In some examples, at least a portion of the oligonucleotides are designed to contain an addition region that is a substrate for universal primer binding. For PCA reactions, pre-synthesized oligonucleotides contain overlaps with each other (e.g., 4, 20, 40, or more bases with overlapping sequences). During the polymerer cycle, the oligonucleotides are annealed into complementary fragments and then packed by polymerase. Thus, each cycle randomly increases the length of various fragments as the oligonucleotides find each other. The complementarity in the fragments allows for the formation of a full large span of double-stranded DNA. In some cases, after the PCA reaction is complete, an error correction step is performed using mismatch correction detection enzymes to remove mismatches in the sequence. Once larger fragments of the target sequence are generated, they can be amplified. For example, in some cases, a target sequence containing 5' and 3' terminal adapter sequences is amplified in a polymerase chain reaction (PCR) using modified primers that hybridize the adapter sequences. In some cases, the modified primers contain one or more uracil bases. The use of modified primers allows for primer removal via an enzymatic reaction that focuses on targeting the gap left by an enzyme that cleaves the modified bases and / or modified base pairs from the fragment. What remains is a double-stranded amplification product lacking the remainder of the adapter sequence. Thus, numerous amplification products can be generated in parallel with the same set of primers to produce different fragments of double-stranded DNA.
[0144] Error correction may be performed on synthesized oligonucleotides and / or constructed products. Exemplary strategies for error correction include site-directed mutagenesis by double-extension PCR to correct errors, which optionally involves two or more rounds of cloning and sequencing. In some examples, double-stranded nucleic acids with mismatches, bulges and small loops, chemically altered bases, and / or other heteroduplexes are selectively removed from a population of precisely synthesized nucleic acids. In some examples, error correction is performed using proteins / enzymes that recognize mismatched or unpaired bases in double-stranded nucleic acids and bind to, or adjacent to, the mismatched or unpaired bases in the double-stranded nucleic acids, creating single-stranded or double-stranded breaks or initiating a strand-transition event. Non-exclusive examples of proteins / enzymes for error correction include endonucleases (T7 endonuclease I, E. coli endonuclease V, T4 endonuclease VII, soybean nuclease, cell endonuclease IV, UVDE), restriction enzymes, glycosylases, ribonucleases, mismatch repair enzymes, resolvers, helicases, ligases, antibodies specific to mismatches, and their variants. Examples of specific error correction enzymes include T4 endonuclease 7, T7 endonuclease 1, S1, mung bean endonuclease, MutY, MutS, MutH, MutL, cleavase, CELI, and HINF1. In some cases, the DNA mismatch-binding protein MutS (Thermus aquaticus) is used to remove failure products from a population of synthesized products. In some examples, error correction is performed using the enzyme Correctase. In some cases, error correction is performed using SURVEYOR endonuclease (Transgenomic), mismatch-specific DNA endonucleases that scan for known and unknown mutations, and polymorphisms of heterodouble-stranded DNA.
[0145] Release, extraction, and construction
[0146] Methods and devices for reproducible information storage are provided herein. In some examples, multiple copies of the same coding region (oligonucleotide), the same cluster, the same portion of a structure containing oligonucleotides, or an entire structure containing oligonucleotides are synthesized. Multiple copies of the same oligonucleotide are synthesized, and each oligonucleotide may be attached to a separate region on a surface. The separate regions may be separated by destruction or cleavage. Alternatively, each oligonucleotide may exist in a feature in the form of a spot, well, or channel and be individually accessible. For example, contacting a feature with a cleavage reagent and then with water releases one copy of the oligonucleotide while leaving the others intact. Similarly, cleavage of an oligonucleotide in an entire region or on an entire plate allows access to a fraction of the replicated population. The replicated population may exist on separate reels, plates, belts, etc. In the case of flexible materials such as tape, the replicated region may be cleaved, and the rest of the tape may be spliced back to its original state. Alternatively, the nucleic acid information of the synthesized and stored oligonucleotides may be obtained by amplifying the oligonucleotides attached to the surface of the structure using primers and DNA polymerase.
[0147] In some examples, an aqueous or gaseous mobile medium is placed on one or more channels in a structure to move oligonucleotides from the structure to a receiving unit. For example, the mobile medium passes through the channels in the structure to adhere to and collect oligonucleotides, and move the channels in the structure to the receiving unit. In some examples, electrically conductive features and applied voltages are used to attract or repel the mobile medium to or through the channels in the structure. In some examples, slips are used to direct the mobile medium to the channels in the structure. In some cases, pressure release is used to direct the mobile medium to or through the channels in the structure. In some cases, nozzles are used to form a localized area of high pressure, which forces the mobile medium into or through the channels in the structure. In some examples, pins are used to transfer oligonucleotides from the channels in the structure to the container of the receiving unit. In such examples, the pins may contain a chemical agent to facilitate adhesion of the mobile medium. In some cases, electrically conductive features are used to attract or repel a mobile medium into or through channels in a structure by forming a potential voltage between the conductive feature and the structure. In some cases, pipette tips or other capillary flow inductive structures are used to move liquids and oligonucleotides via capillary flow. In some examples, a container comprises one or more compartments, each receiving a portion of the mobile medium and one or more oligonucleotides emitted from a single channel. In some examples, a container comprises a single compartment receiving one or more portions of the mobile medium, each containing oligonucleotides emitted from one or more structural channels.
[0148] Sequence determination
[0149] Following the extraction and / or amplification of oligonucleotides from the surface of the structure, appropriate sequencing techniques may be used to sequence the oligonucleotides. In some cases, the DNA sequence is read on the substrate or within the structural features. In some cases, oligonucleotides conserved in the substrate are extracted, optionally constructed into longer nucleic acids, and then sequenced.
[0150] Oligonucleotides synthesized and stored on the structures described herein encode data that can be interpreted by reading the sequence of the synthesized oligonucleotide and converting the sequence into a computer-readable binary code. In some cases, the sequence may require construction, and this construction step may be a nucleic acid sequencing step or a digital sequencing step.
[0151] A detection system is provided herein that includes a device capable of sequencing stored oligonucleotides either directly on the structure and / or after removal from the main structure. If the structure is an open-reel tape of a flexible material, the detection system includes a device for holding and advancing the structure through a detection position, as well as a detector positioned proximal to the detection position for detecting a signal emanating from a section of tape when the section is at the detection position. In some examples, the signal indicates the presence of an oligonucleotide. In some examples, the signal indicates the sequence of an oligonucleotide (e.g., a fluorescent signal). In some examples, information encoded within oligonucleotides on a continuous tape is read by a computer as the tape is transmitted continuously through a detector operably connected to a computer. In some examples, the detection system includes a computer system comprising an oligonucleotide sequencing device, a database for storing and retrieving data on oligonucleotide sequences, software for converting the DNA code of the oligonucleotide sequence into binary code, a computer for reading the binary code, or any combination thereof.
[0152] Computer system
[0153] In various embodiments, any of the systems described herein are operably connected to a computer and can be automated at will via a local or remote computer. In various examples, the methods and systems of this disclosure further include software programs on a computer system and their use. Accordingly, computer control for synchronizing dispensing / vacuuming / filling functions, such as synchronizing the operation, dispensing operation, and vacuum activation of a material deposition apparatus, is within the scope of this disclosure. In some examples, the computer system is programmed to interface between a user-specified nucleotide sequence and the material deposition apparatus position in order to deliver a precise reagent to a specified region of a substrate.
[0154] The computer system (1700) illustrated in Figure 17 can be understood as a logical device capable of reading instructions from a network port (1705) which may be optionally connected to a server (1709) having a medium (1711) and / or a fixed medium (1712). The system shown in Figure 17 may include a CPU (1701), a disk drive mechanism (1703), optional input devices such as a keyboard (1715) and / or a mouse (1716), and an optional monitor (1707). Data communication may be achieved to a server at a local or remote location via the indicated communication medium. The communication medium may include any means for transmitting and / or receiving data. For example, the communication medium may be a network connection, a wireless connection, or an Internet connection. Such a connection may provide communication over the World Wide Web. It is assumed that the data relating to this disclosure may be transmitted via such a network or connection for reception and / or review by a party (1722).
[0155] Figure 18 is a block diagram illustrating a first exemplary architecture of a computer system (1800) that may be used with respect to exemplary embodiments of the present disclosure. As shown in Figure 18, the exemplary computer system may include a processor (1802) for processing instructions. Examples of processors that are not limited to include: Intel Xeon® processors, AMD Opteron® processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0® processors, RM Cortex-A8 Samsung S5PC100® ARM Cortex-A8 Apple A4® processors, Marvell PXA 930® processors, or functionally equivalent processors. Multiple threads of execution may be used for parallel processing. In some examples, a processor having multiple processors or multiple cores may also be used in a single computer system, a cluster, or multiple computers distributed across a network system including mobile phones and / or personal digital assistant devices.
[0156] As illustrated in Figure 18, the high-speed cache (1804) is connected to or incorporated into the processor (1802) and provides a high-speed storage device for recently used or frequently used instructions or data by the processor (1802). The processor (1802) is connected to the northbridge (1806) by the processor bus (1808). The northbridge (1806) is connected to the random access memory (RAM) (1810) by the memory bus (1812) and manages access to the RAM (1810) by the processor (1802). The northbridge (1806) is also connected to the southbridge (1814) by the chipset bus (1816). The southbridge (1814) is, in turn, connected to the peripheral bus (1818). The peripheral bus may be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The northbridge and southbridge, often referred to as the processor chipset, manage data transfer between the processor, RAM, and peripheral components on the peripheral bus (1818). In some alternative architectures, the functionality of the northbridge may be integrated into the processor instead of using a separate northbridge chip.
[0157] In some examples, the system (1800) may include an accelerator card (1822) attached to a peripheral bus (1818). The accelerator may include a field-programmable gate array (FPGA) or other hardware for accelerating specific processing. For example, the accelerator may be used for reconstructing adaptive data or for evaluating algebraic expressions used in extended configuration processing.
[0158] Software and data are stored in external storage (1824) and may be loaded into RAM (1810) and / or cache (1804) for use by the processor. The system (1800) includes an operating system for managing system resources, and not limited to the following examples of operating systems: Linux®, Windows®, MACOS®, iOS®, and other functionally equivalent operating systems, as well as application software running on an operating system that manages data storage and optimization in accordance with the exemplary embodiments of this disclosure.
[0159] In this example, system (1800) also includes network interface cards (NICs) (1820 and 1821) connected to a peripheral bus for providing a network interface to external storage devices such as network-attached storage (NAS), and other computer systems that may be used for distributed parallel processing.
[0160] Figure 19 shows a network (1900) comprising multiple computer systems (1902a and 1902b), multiple mobile phones and personal digital assistants (1902c), and network-attached storage (NAS) (1904a and 1904b). In an exemplary embodiment, the systems (1902a, 1902b, and 1902c) can manage data storage and optimize data access to data stored in the network-attached storage (NAS) (1904a and 1904b). Mathematical models can be used with data and evaluated using distributed parallel processing across the computer systems (1902a and 1902b) and the mobile phone and personal digital assistant systems (1902c). Computer systems (1902a and 1902b), as well as mobile phones and personal information terminal systems (1902c), also provide parallel processing for adaptive data reconstruction of data stored in network-attached storage (NAS) (1904a and 1904b). Figure 19 illustrates one example, and various other computer architectures and systems may be used in conjunction with various embodiments of this disclosure. For example, blade servers can be used to provide parallel processing. Processor blades may be connected on a backplane to provide parallel processing. Storage may also be connected on a backplane or as network-attached storage (NAS) via another network interface.
[0161] In some exemplary embodiments, a processor may maintain a separate memory space and transmit data through a network interface, backplane, or other connectors for parallel processing by other processors. In other examples, some or all of the processors may use a shared virtual address memory space.
[0162] Figure 20 is a block diagram of a multiprocessor computer system (2000) using a shared virtual address memory space according to an exemplary embodiment. The system includes multiple processors (2002a)-(2002f) that can access a shared memory subsystem (2004). The system incorporates multiple programmable hardware memory algorithm processors (MAPs)(2006a)-(2006f) into the shared memory subsystem (2004). Each MAP(2006a)-(2006f) may include memory(2002a)-(2002f) and one or more field-programmable gate arrays(2010a)-(2010f). The MAPs provide configurable functional units, and certain algorithms or parts of algorithms may be provided to FPGAs(2010a)-(2010f) for processing in close cooperation with their respective processors. For example, MAPs may be used to evaluate algebraic expressions relating to a data model and, in exemplary embodiments, to perform adaptive data reconstruction. In this embodiment, each MAP is globally available through all the processors for these purposes. In one configuration, each MAP can use direct memory access (DMA) to access its associated memory (2008a)-(2008f), thereby enabling it to perform tasks independently of and asynchronously with the respective microprocessors (2002a)-(2002f). In this configuration, a MAP can directly feed results to another MAP for pipelined and parallel execution of algorithms.
[0163] The computer architectures and systems described above are merely examples, and a wide variety of other computer, mobile phone, and personal data assistant architectures and systems, including those using general-purpose processors, coprocessors, FPGAs and other programmable logic devices, systems-on-a-chip (SOCs), application-specific integrated circuits (ASICs), and any combination of other processing and logic elements, may be used in connection with the exemplary embodiments. In some examples, all or part of the computer system may run in software or hardware. All kinds of data storage media, including random-access memory, hard drives, flash memory, tape drives, disk arrays, network-attached storage (NAS), and other local or distributed data storage devices and systems, may be used in connection with the exemplary examples.
[0164] In exemplary embodiments, a computer system may run using software modules that run on any of the above or other computer architectures and systems. In other examples, the functions of the system may be partially or completely performed by firmware, programmable logic circuits such as field-programmable gate arrays (FPGAs), systems on a chip (SOCs), application-specific integrated circuits (ASICs), or other processing and logic elements. For example, set processors and optimizers may be performed with hardware acceleration through the use of hardware accelerator cards such as accelerator cards.
[0165] Methods for storing information are provided herein, the methods comprising: converting an item of information in the form of at least one digital sequence into at least one nucleic acid sequence; providing a flexible structure having a surface; synthesizing a plurality of oligonucleotides having a predetermined sequence that collectively encodes at least one nucleic acid sequence, wherein the plurality of oligonucleotides comprises at least about 100,000 oligonucleotides, and wherein the plurality of oligonucleotides extends from the surface of the flexible structure; and storing the plurality of oligonucleotides. Methods for synthesis comprising: depositing nucleosides on a surface at predetermined positions; and moving at least a portion of the flexible structure via release from a tank or spray bar are further provided herein. Methods for exposure of the surface of the structure to an oxidizing agent or non-blocking reagent via release from a tank or spray bar are further provided herein. Methods for synthesis comprising capping the deposited nucleosides on the surface are further provided herein. Methods for nucleosides comprising nucleoside phosphoramidites are further provided herein. Methods comprising open-reel tape or continuous tape are further provided herein. Methods comprising a thermoplastic material comprising a flexible structure are further provided herein. Methods comprising a polyaryletherketone comprising a thermoplastic material are further provided herein. Methods comprising a polyaryletherketone comprising a polyaryletherketone, polyetherketoneketone, poly(ether ether ketone ketone), polyether ether ketone or polyetherketone ether ketoneketone are further provided herein. Methods comprising a flexible structure comprising nylon, nitrocellulose, polypropylene, polycarbonate, polyethylene, polyurethane, polystyrene, acetal, acrylic, acrylonitrile, butadiene styrene, polyethylene terephthalate, polymethyl methacrylate, polyvinyl chloride, transparent PVC foil, poly(methyl methacrylate), styrene-based polymer, fluorine-containing polymer, polyethersulfone, or polyimide are further provided herein.Methods are further provided herein for a plurality of oligonucleotides, each oligonucleotide having a length of 50 to 500 bases. Methods are further provided herein for a plurality of oligonucleotides, each containing at least about 10 billion oligonucleotides. At least about 1.75 × 10⁻¹⁶. 13 A method for synthesizing nucleic acid bases within 24 hours is further provided herein. At least about 262.5 × 10 9 Methods for synthesizing oligonucleotides within 72 hours are further provided herein. Methods for information items being textual, auditory, or visual information are further provided herein. Methods for nucleosides including nucleoside phosphoramidites are further provided herein.
[0166] A method for storing information is provided herein, the method comprising: converting an item of information in the form of at least one digital sequence into at least one nucleic acid sequence; providing a structure having a surface; synthesizing a plurality of oligonucleotides having a predetermined sequence that collectively encodes at least one nucleic acid sequence, wherein the plurality of oligonucleotides comprises at least about 100,000 oligonucleotides, wherein the plurality of oligonucleotides extends from the surface of the structure, and wherein the synthesis comprises: washing the surface of the structure; depositing nucleosides on the surface at predetermined positions; oxidizing, deblocking, and optionally capping the nucleosides deposited on the surface; wherein washing, oxidation, deblocking, and capping include moving at least a portion of the flexible structure via discharge from a tank or spray bar; and storing the plurality of oligonucleotides. The method is further provided herein, wherein the nucleosides include nucleoside phosphoramidites.
[0167] The following examples are provided to further illustrate to those skilled in the art the principles and practices of the embodiments disclosed herein and should not be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are on a weight basis. [Examples]
[0168] Example 1: Functionalization of device surface
[0169] The device was functionalized to assist in the binding and synthesis of oligonucleotide libraries. The device surface was first washed with water for 20 minutes using a piranha solution containing 90% H2SO4 and 10% H2O2. The device was rinsed with deionized water in several beakers, held under a deionized water gooseneck tap for 5 minutes, and dried with N2. Subsequently, the device was immersed in NH4OH (1:100; 3mL:300mL) for 5 minutes, rinsed with deionized water using a hand gun, immersed in deionized water in three consecutive beakers for 1 minute each, and then rinsed again with deionized water using a hand gun. The device was then plasma-cleaned by exposing the device surface to O2. A SAMCO PC-300 instrument was used to plasma-etch O2 at 250 watts for 1 minute in downstream mode.
[0170] The cleaned device surface was actively functionalized with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutylamide using a YES-1224P deposition oven system with the following parameters: 0.5-1 Tor, 60 minutes, 70°C, vaporizer at 135°C. The device surface was a resist coated using a Brewer Science 200X spin coater. SPR(trademark) 3612 photoresist was spin-coated on the device at 2500 rpm for 40 seconds. The device was pre-baked on a Brewer heating plate at 90°C for 30 minutes. The device was exposed to photolithography using a Karl Suss MA6 mask aligner apparatus. The device was exposed for 2.2 seconds and developed with MSF 26A for 1 minute. The remaining developer was rinsed off with a handgun and the device was immersed in water for 5 minutes. The device was baked in an oven at 100°C for 30 minutes, followed by visual inspection for lithography defects using a Nikon L200. To remove the remaining resist using a SAMCO PC-300 instrument, O2 plasma etching was performed at 250 watts for 1 minute using the Descam process.
[0171] The device surface was passively functionalized with a solution of 100 μL of perfluorooctyltrichlorosilane mixed with 10 μL of diesel fuel. The device was placed in a chamber and pumped for 10 minutes, then the valve was closed to the pump and left for 10 minutes. The chamber was released to air. The device was stripped of its resist by immersing it twice in 500 mL of NMP for 5 minutes at 70°C at maximum power sonication (9 on the Crest system). Subsequently, the device was immersed in 500 mL of isopropanol for 5 minutes at room temperature at maximum power sonication. The device was immersed in 300 mL of 200 proof ethanol and blow-dried with N2. The functionalized surface was activated to function as an aid in oligonucleotide synthesis.
[0172] Example 2: Synthesis of a 50-mer sequence on an oligonucleotide synthesis device
[0173] A two-dimensional oligonucleotide synthesis device was constructed in a flow cell and connected to a flow cell (Applied Biosystems (ABI394 DNA Synthesizer)). The two-dimensional oligonucleotide synthesis device was homogenized with N-(3-triethoxysilylpropyl)-4-hydroxybutylamide (Gelest) and used to synthesize typical 50 bp ("50-mer oligonucleotide") oligonucleotides using the oligonucleotide synthesis method described herein.
[0174] The sequence of the 50-mer is as described in SEQ ID NO. 1: 5'AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3' (SEQ ID NO.: 1), where # represents thymidine-succinyl hexamide CED phosphoramidite (CLP-2244 from ChemGenes), which is a cleavable linker that allows for the release of oligonucleotides from the surface during deprotection.
[0175] The synthesis was carried out according to the protocol in Table 5, using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking) and an ABI synthesizer.
[0176] [Table 5-1]
[0177] [Table 5-2]
[0178] The phosphoramidite / activator combination was delivered via a flow cell in the same manner as the bulk reagent. Since the environment remained constantly "wet" with the reagent, no drying step was performed.
[0179] To allow for faster flow, the restrictor was removed from the ABI 394 synthesizer. Without the restrictor, the flow rates of amidite (0.1 M in ACN), activator (0.25 M benzoylthiotetrazole ("BTT"; 30-3070-xx from Glen Research in ACN), and Ox (0.02 M I2 in 20% pyridine, 10% water, and 70% THF) were approximately ~100 uL / sec, and the flow rates of acetonitrile ("ACN") and capping reagent (a 1:1 mixture of CapA and CapB, where CapA is acetic anhydride, THF / pyridine, and CapB is 16% 1-methyl acetic anhydride in THF) were approximately ~100 uL / sec. For midazole, the flow rate was approximately ~200 uL / second, and for non-blocked (3% dichloroacetic acid in toluene), it was approximately ~300 uL / second (compared to ~50 uL / second for all reagents with restrictors). The time required to completely flush out the oxidizing agent was observed, and the timing of the drug flow rate was adjusted accordingly, and additional ACN washing was introduced between different chemicals. After oligonucleotide synthesis, the tips were deprotected overnight in gaseous ammonia at 75 psi. Five drops of water were added to the surface to construct the oligonucleotides. The constructed oligonucleotides were then analyzed using a BioAnalyzer small RNA tip (not shown).
[0180] Example 3: Synthesis of a 100-base pair sequence using an oligonucleotide synthesis device
[0181] The same process for synthesizing a 50-base sequence as described in Example 2 was used for synthesizing a 100-base oligonucleotide ("100-base oligonucleotide"); 5'CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3', where # represents thymidine-succinyl hexamide CED phosphoramidite (CLP-2244 from ChemGenes); SEQ ID NO.:2) In two different silicon chips, the first silicon chip was homogeneously functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutylamide, and the second silicon chip was functionalized with a 5 / 95 mixture of 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane. Oligonucleotides extracted from the surface were analyzed using a BioAnalyzer instrument (not shown).
[0182] All 10 samples from the two chips were further PCR-amplified in 50 μL of PCR mixture (25 μL of NEB Q5 master mix, 2.5 μL of 10 μM forward primer, 2.5 μL of 10 μM reverse primer, 1 μL of oligonucleotide extracted from the surface, and up to 50 μL of water) using the following thermal cycling program: forward (5'ATGCGGGGTTCTCATCATC3'; SEQ ID NO.: 3) primer and reverse (5'CGGGATCCTTATCGTCATCG3'; SEQ ID NO.: 4) primer: 98℃, 30 seconds 98°C, 10 seconds; 63°C, 10 seconds; 72°C, 10 seconds; repeat 12 cycles. 72℃, 2 minutes
[0183] The PCR products were also run on a BioAnalyzer (not shown) and showed a sharp peak at a position of 100 base pairs. Next, the PCR-amplified samples were cloned and Sanger sequenced. Table 6 summarizes the results from Sanger sequencing for samples obtained from spots 1-5 from tip 1 and samples obtained from spots 6-10 from tip 2.
[0184] [Table 6-1]
[0185] [Table 6-2]
[0186] Therefore, the high quality and uniformity of the synthesized oligonucleotides were repeated on two chips with different surface chemistry. Overall, 89% of the sequenced oligonucleotides, corresponding to 233 out of 262 sequences, were error-free and complete.
[0187] Table 7 summarizes the sequence error characteristics from oligonucleotide samples obtained from spots 1-10.
[0188] [Table 7-1]
[0189] [Table 7-2]
[0190] Example 4: Highly accurate storage and construction of DNA-based information
[0191] Digital information was selected in the form of approximately 0.2 GB of binary data, including the content of the Universal Declaration of Human Rights in over 100 languages, the top 100 books from Project Gutenberg, and a seed database. The digital information was encrypted into nucleic acid-based sequences and split into strings. Over 10 million non-identical oligonucleotides (each corresponding to a string) were synthesized on a rigid silicon surface in a manner similar to that described in Example 2. Each non-identical oligonucleotide was 200 base pairs or less in length. The synthesized oligonucleotides were collected, sequenced, and decrypted into digital codes with 100% accuracy of the source digital information by comparing them to at least one initial digital sequence.
[0192] Example 5: Conversion of digital information to nucleic acid sequences
[0193] A computer txt file contains text information. A general-purpose computer uses a software program that includes machine instructions to convert the sequence into a 3, 4, or 5 sequence according to the received instructions. Each number in base 3 is assigned a nucleic acid (e.g., A=0, T=1, C=2). Each number in base 4 is assigned a nucleic acid (e.g., A=0, T=1, C=2, G=3). Alternatively, a 5-quinary sequence is used if each number in base 5 is assigned a nucleic acid (e.g., A=0, T=1, C=2, G=3, U=4). The sequence is generated as shown in Table 8. Machine instructions are then provided for the de novo synthesis of oligonucleotides to encode the nucleic acid sequence.
[0194] [Table 8-1]
[0195] [Table 8-2]
[0196] Example 6: Flexible surface with high density
[0197] A flexible structure containing thermoplastic material is coated with a nucleoside coupling reagent. The coating agent is patterned for high-density features. A portion of the flexible surface is illustrated in Figure 14A. Each feature has a diameter of 10 μm, and the center-to-center distance between two adjacent features is 21 μm. The size of the features is sufficient to accommodate a deposition droplet volume of 0.2 pl during the deposition process of oligonucleotide synthesis. The small size of the features allows for synthesis on the surface of a high-density oligonucleotide substrate. The density of the features is 2.2 billion features / m². 2 (1 feature / 441x10) -12 m 2 ) It has 10 billion features and is 4.5m 2 The substrate is produced, each with a diameter of 10 μm. The flexible structure is optionally placed in a continuous loop system (Figure 12A) or an open-reel system (Figure 12B) for oligonucleotide synthesis.
[0198] Example 7: Alkyl synthesis on a flexible structure
[0199] The flexible structure is prepared by incorporating several features on a thermoplastic flexible material. The structure serves to support the synthesis of oligonucleotides using oligonucleotide synthesis devices, including deposition equipment. The flexible structure is a form of flexibility mediation, similar to magnetic open-reel tapes.
[0200] De novo synthesis operates in a continuous production line manner, where the structure passes through a solvent tank, then moves under a stack of printheads, where phosphoramidite is printed onto the surface of the structure. The flexible structure, with fixed droplets deposited on its surface, is rotated into an oxidizing tank, then the tape is removed from the oxidation tank and immersed in acetonitrile washing solution, and then submerged in a non-blocking tank. Optionally, the tape is moved through a capping tank. In other workflows, in the washing step, the flexible structure is removed from the oxidation tank and sprayed with acetonitrile.
[0201] Alternatively, a spray bar is used instead of a fluid bath. In this process, nucleotides still deposit on the surface, but the flooding step is performed in a chamber equipped with spray nozzles. For example, a deposition apparatus has 2,048 nozzles, each depositing 100,000 droplets per second, with one nucleic acid base per droplet. A sequence of spray nozzles exists to mimic the sequence of flooding steps in standard phosphoramidite chemicals. This technique allows for easy modification of the chemicals packed in the spray bar to adapt to various process steps. Oligonucleotides are deprotected or cleaved in the same manner as described in Example 2.
[0202] For each deposition device, 1.75 × 10 13 A larger number of nucleic acid bases are deposited onto the structure each day. Oligonucleotides of multiple 200 nucleic acid bases are synthesized. In 3 days, 1.75 × 10¹⁶ units are produced per day. 13 The ratio of bases is 262.5 × 10 9 Synthesize oligonucleotides.
[0203] Example 8: Selective Bioencryption
[0204] The program module receives machine instructions for an item of information of a desired type for one or more categories of conversion and bioencryption, the one or more categories of bioencryption being selected from enzyme-based forms (e.g., CRISPR / Cas complexes and restriction enzyme digests), electromagnetic radiation-based forms (e.g., photolysis and photodetection), chemical cleavage forms (e.g., treatment with gaseous ammonia or methylamine to cleave thymidine-succinyl hexamide CED phosphoramidite (ChemGenes CLP-2244)), and affinity-based forms (e.g., sequence tags for hybridization, or incorporation of modified nucleotides with enhanced similarity to capture reagents). Following the receipt of a specific bioencryption selection, the program module performs the steps of converting the item of information into a nucleic acid sequence and applying design instructions to design a bioencrypted version of the sequence. It selects a specific encryption subtype within the bioencryption category. It then provides synthesis instructions to a material deposition apparatus for the de novo synthesis of oligonucleotides.
[0205] Example 9: Selected Decryption
[0206] The system provides machine instructions for the application of one or more categories of biodescriptions selected from enzyme-based (e.g., CRISPR / Cas complex or restriction enzyme digest), electromagnetic radiation-based (e.g., photolysis or photodetection), chemical cleavage-based (e.g., treatment with gaseous ammonia or methylamine for cleavage of thymidine-succinyl hexamide CED phosphoramidite (ChemGenes CLP-2244)), and affinity-based (e.g., incorporation of sequence tags for hybridization or modified nucleotides with enhanced affinity for capturing reagents). Following the receipt of a specific biodescription selection, the program module performs a step of releasing a regulator for oligonucleotide enrichment. Following enrichment, the oligonucleotide is sequenced, optionally aligned to a longer nucleic acid sequence, and converted into a digital sequence corresponding to the information item.
[0207] Example 10: Bioencryption and biodecryption of DNA sequences using CRISPR / Cas9
[0208] A digital sequence is received that encodes the information items. The digital sequence is then converted into a nucleic acid sequence. The nucleic acid sequence is then encoded into a larger population of nucleic acid sequences. The encoding process includes adding a “junk” region for detection and removal by the CRISPR / Cas9 complex. The nucleic acid sequence is synthesized as described in Examples 2-3.
[0209] A population of nucleic acid sequences containing encoded nucleic acid sequences is mixed with Cas9 and gRNA in Cas9 buffer and incubated at 37°C for 2 hours. Cas9 is then inactivated and removed by purification. The purified samples are then analyzed by next-generation sequencing.
[0210] Example 11: Bioencryption and biodecryption of DNA sequences using CRISPR / Cas9, including sequence exchange.
[0211] A digital sequence encoding information items is received, and this digital sequence is converted into a nucleic acid sequence. The nucleic acid sequence is encoded by adding specific sequences using a CRISPR / Cas9 system and a guide RNA sequence. The nucleic acid sequence is synthesized as described in Examples 2-3.
[0212] Subsequently, the nucleic acid sequence is mixed with a fluorescently labeled probe complementary to the exchanged sequence. The nucleic acid sequence recognized by the fluorescently labeled probe is removed from the population.
[0213] Example 12: Bioencryption and biodecryption of DNA sequences using restriction enzyme digests
[0214] A digital sequence encoding information items is received, and this digital sequence is converted into a nucleic acid sequence. The population of nucleic acid sequences is encoded by adding specific sequences recognized by the restriction enzyme EcoRI. The nucleic acid sequences are synthesized and stored as in Examples 2-3.
[0215] Nucleic acid sequences are incubated with EcoRI. The encoded nucleic acid sequence containing the EcoRI recognition site is cleaved. Following the cleavage of the encoded nucleic acid sequence, a sequence with a complementary overhang is hybridized and bound to the released DNA. The bound complex is then isolated, the purified sample is sequenced, and the original digital information is reconstructed.
[0216] Example 13: Bioencryption and biodecryption of DNA sequences using photolysis
[0217] A digital sequence encoding information items is received, and the digital sequence is converted into a nucleic acid sequence. A collection of nucleic acid sequences is designed to contain photocleavable nucleic acid bases. The nucleic acid sequences are synthesized and stored as in Examples 2-3.
[0218] Apply 280 nm UV-B irradiation to the nucleic acid sequence. The encrypted nucleic acid sequence containing a photocleavable site is cleaved and removed. Thereafter, the nucleic acid sequences are collected and sequenced. Alternatively, the nucleic acid sequence is released from the surface of a structure, such as by ammonia gas cleavage, and then exposed to electromagnetic radiation to effect cleavage in the nucleotide sequence. A portion of the population is enriched by, for example, pull-down assay using beads having bound complementary capture probes, PCR using primers selected only to amplify a target sequence, or size exclusion chromatography. Thereafter, the enriched nucleic acid is converted into a digital sequence, sequenced, and an item of information is obtained.
[0219] Example 14: Bioencryption and biodecryption of DNA sequences using chemical enrichment
[0220] A digital sequence encoding an item of information is received, and the digital sequence is converted into a nucleic acid sequence. A population of nucleic acid sequences is encrypted by addition of specific sequences (e.g., thymidine-succinyl hexamide CED phosphoramidite (CLP-2244 from ChemGenes), which is chemically cleavable by ammonia gas). The nucleic acid sequences are synthesized as in Examples 2-3.
[0221] Ammonia gas is applied to the nucleic acid sequence. Using the enrichment method described herein, encrypted nucleic acid sequences comprising chemically cleavable sequences are released from the population and enriched. Thereafter, the enriched nucleic acid is converted into a digital sequence, sequenced, and an item of information is obtained.
[0222] Example 15: Bioencryption and biodecryption of DNA sequences using biotin-containing nucleic acid probes
[0223] Receiving a digital array encoding an item of information, and converting the digital array into a nucleic acid sequence. Encrypting a population of nucleic acid sequences through the design of predetermined residues for containing biotin-containing nucleobases. Synthesizing the nucleic acid sequences as described in Examples 2-3.
[0224] Cleaving the nucleic acid sequences from the structure and mixing the same with streptavidin-containing beads. Thereafter, incubating the nucleic acid sequences with streptavidin magnetic beads. Pulling down the biotin-containing nucleic acid sequences by means of the magnetic beads. Thereafter, converting the concentrated nucleic acids into a digital array, sequencing the same, and obtaining the item of information.
[0225] Example 16: Bio-encryption and bio-decryption of DNA sequences using light detection
[0226] Receiving a digital array encoding an item of information, and converting the digital array into a nucleic acid sequence. Encrypting a population of nucleic acid sequences through design, so as to contain a specific sequence recognized by an Alexa488-labeled nucleic acid probe. Synthesizing the nucleic acid sequences as described in Examples 2-3.
[0227] Releasing the nucleic acid sequences from the structure and mixing the same with an Alexa488-labeled nucleic acid probe. Thereafter, fractionating the nucleic acid sequences according to fluorescence intensity. Further analyzing the nucleic acid sequences labeled with the Alexa488-labeled nucleic acid probe. Thereafter, sequencing the probe-bound nucleic acids, converting the same into a digital array, and obtaining the item of information.
[0228] Example 17: Bio-encryption and bio-decryption of DNA sequences using modified nucleotides
[0229] A digital sequence encoding information items is received, and the digital sequence is converted into a nucleic acid sequence. The population of nucleic acid sequences is encoded by designing for the addition of predetermined nucleic acid bases containing peptide nucleic acids (PNAs) at predetermined positions, and for designing for restriction enzyme recognition sizes to excise the PNA-containing sections. The nucleic acid sequences are synthesized as in Examples 2-3.
[0230] The nucleic acid sequence is released, subjected to restriction enzyme digestion, and then amplified by PCR. Nucleic acid sequences containing PNA cannot be amplified. Subsequently, the enriched and amplified nucleic acid is converted into a digital sequence, sequenced, and the information is received.
[0231] Example 18: Bioencryption and biodecryption of DNA sequences using CRISPR / Cas9 and chemical cleavage
[0232] A digital sequence encoding information items is received, and this digital sequence is converted into a nucleic acid sequence. A population of nucleic acid sequences is encoded by adding specific sequences using CRISPR / Cas9 and a guide RNA sequence. The CRISPR / Cas9 system introduces chemically cleavable sites at pre-selected positions in the nucleic acid sequence. The nucleic acid sequence is synthesized as in Examples 2-3.
[0233] Ammonia gas is applied to the nucleic acid sequence. Encrypted nucleic acid sequences containing chemically cleavable sites are cleaved and removed by size exclusion purification and analyzed by next-generation sequencing.
[0234] While preferred embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided only as examples. Many modifications, changes, and substitutions will be conceivable without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may also be used in the practice of the present invention. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents are thereby encompassed.
Claims
1. A method for storing information, wherein the method is a) A step of receiving at least one item of information in the form of at least one digital array by a computer system, b) A step of receiving instructions from the computer system for selecting at least one form of bioencryption, wherein the form of bioencryption is an enzyme-related bioencryption, and the enzyme-related bioencryption is an enzyme-related encryption of a nucleic acid sequence. c) A step of converting the at least one digital sequence into a plurality of oligonucleotide sequences based on a selected bioencryption form by the computer system, wherein the digital sequences of the plurality of oligonucleotides are configured for biodecryption by exposing the plurality of oligonucleotides to a nuclease complex. d) A step of synthesizing a plurality of oligonucleotides that encode the plurality of oligonucleotide sequences, e) A step of storing the plurality of oligonucleotides, A method that includes this.
2. The method according to claim 1, wherein the bioencryption related to the enzyme includes a bioencryption related to CRISPR / Cas.
3. The method according to claim 1, wherein the bioencryption related to the enzyme includes instructions for the synthesis of enzyme-sensitive oligonucleotides as shown in Table 1.
4. The method according to claim 1, wherein two, three, four, or five forms of bioencryption are used.
5. The method according to claim 1, wherein the plurality of oligonucleotides comprises at least 100,000 oligonucleotides.
6. The method according to claim 1, wherein the plurality of oligonucleotides comprises at least 10 billion oligonucleotides.
7. f) A step of applying enzyme-related decryptions to the plurality of oligonucleotides, g) A step of concentrating the plurality of oligonucleotides, h) A step of sequencing concentrated oligonucleotides from the plurality of oligonucleotides in order to generate a nucleic acid sequence, i) A step of converting the nucleic acid sequence into at least one digital sequence, wherein the at least one digital sequence encodes at least one item of information. The method according to claim 1, further comprising:
8. The method according to claim 7, further comprising the step of releasing a plurality of oligonucleotides from a stored surface before step f).
9. The method according to claim 1, wherein at least one item of the information includes text information, auditory information, and visual information.
10. The method according to claim 1, wherein at least one item of the information includes multiple files.
11. The method according to claim 10, wherein the plurality of files include at least 1 gigabyte.
Citation Information
Patent Citations
Method for encrypting and decrypting specific message by using nucleic acid molecule
JP2005055900A
Fusing belt for applying a protective overcoat to a photographic element
US6537741B2