Codec for DNA data storage

By designing a codec that can establish redundant during DNA data storage and sequencing, the problems of errors and ambiguity in the encoding and decoding process of DNA data in the prior art are solved, and higher reliability and efficiency are achieved.

CN119948568APending Publication Date: 2025-05-06TWIST BIOSCIENCE CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202380048529.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2023-04-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is prone to errors and ambiguity in the process of DNA data storage and sequencing, and it is difficult to effectively encode and decode DNA data.

Method used

A codec was designed to correct errors by encoding digital data into oligonucleotide pools and establishing redundancy during synthesis, storage and sequencing. The codec includes an internal codec and an external codec, which can work effectively with low sequencing coverage and high deletion, mutation, insertion rates.

Benefits of technology

It realizes efficient encoding and decoding of DNA data in the presence of errors, improving the reliability and efficiency of data storage and sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948568A_ABST
    Figure CN119948568A_ABST
Patent Text Reader

Abstract

Systems and methods for encoding digital data into oligonucleotides and decoding the oligonucleotides back to the digital data are described herein. The encoding and decoding scheme includes an internal codec for converting digital data to bases and in turn to digital data. The encoding and decoding schemes also include an external codec including an error correction scheme.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 333,305 filed on April 21, 2022, U.S. Provisional Application No. 63 / 338,760 filed on May 5, 2022, and U.S. Provisional Application No. 63 / 481,873 filed on January 27, 2023, which are incorporated by reference in their entireties.

[0003] Incorporated by Reference

[0004] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0005] background

[0006] DNA is an attractive data storage medium given its superior density, stability, energy efficiency, and longevity compared to currently used electronic media. However, errors and ambiguities may be introduced or otherwise occur at or during various stages of sequencing and sequencing-related operations and processes. Therefore, there is a need to develop methods for efficiently encoding and decoding DNA in the presence of such errors.

[0007] Overview

[0008] Provided herein are designs and implementations of various codecs for encoding digital data (e.g., binary data) into an oligonucleotide pool and decoding the pool back into digital data. The codec can include an internal codec for converting digital data into bases. The codec can also include external codes for spreading the data to be stored on many oligonucleotides and establishing redundancy to correct erasures. The codec described herein may suffer from the loss and high deletion, sudden change and insertion rate of oligonucleotides during synthesis, storage and / or sequencing. In some embodiments, the codec described herein is designed for low sequencing coverage. In some embodiments, the codec described herein is designed to optimize the synthesis of more than one polynucleotide.

[0009] Also provided herein is a method for retrieving digital information from more than one polynucleotide. The codec can include a barrel storage system that supports storing one or more objects including digital information in one or more pools. The codec can also include a storage strategy, such as an index (e.g., an index pool) and a hash (e.g., a hash module) for efficient data storage. The codec can also establish redundancy in one or more pools to correct for erasures or errors that may occur during storage or retrieval of digital information.

[0010] In one aspect, provided herein is a method for encoding data in more than one polynucleotide sequence, comprising: (a) splitting the data into more than one frame, wherein each frame in the more than one frame includes a frame index; (b) applying an external codec to each frame in the more than one frame, wherein the external codec includes an error correction scheme; (c) dividing each frame into more than one channel, wherein each channel in the more than one channel includes a channel index; (d) shuffling each channel based at least in part on the channel index; and (e) applying an internal codec to encode each channel in a polynucleotide sequence of more than one polynucleotide sequence. In some cases, the data includes more than one symbol. In some cases, the data includes binary data. In some cases, the binary data includes a byte stream or a byte array. In some cases, the shuffling in (d) includes a rotation scheme within each channel. In some cases, the shuffling in (d) includes a pseudo-random process within each channel. In some cases, the shuffling in (d) provides resistance to errors. In some cases, the error is a nucleotide synthesis error or a sequencing error. In some cases, the error includes a deletion, an insertion, or a substitution. In some cases, the error correction scheme includes Reed-Solomon (RS) code, low-density parity check (LDPC) code, turbo code, polar code or any combination thereof. In some cases, the data includes at least about 1GB to about 1TB. In some cases, more than one frame includes about 100 to about 10,000 frames. In some cases, each frame includes up to about 5000 channels. In some cases, each channel includes about 100 bits to about 300 bits. In some cases, the frame index includes about 16 bits to about 20 bits. In some cases, the channel index includes about 12 bits or about 16 bits. In some cases, the length of the polynucleotide sequence is about 100 to about 300 bases. In some cases, the frame index and / or channel index are attached to the front of each channel before (d). In some cases, applying an internal codec includes adding redundancy across more than one polynucleotide sequence. In some cases, the redundancy is about 5% to about 10%. In some cases, more than one polynucleotide sequence can be decoded in part due to redundancy across more than one polynucleotide sequence in the presence of an error. In some cases, the error includes insertion, deletion, substitution, or any combination thereof. In some cases, applying an internal codec includes: (a) combining symbols, symbol history, and symbol positions from a channel; and (b) generating candidate bases using a lookup table, hashing, or both. In some cases, the method also includes performing a base duplication check. In some cases, the symbol is a bit. In some cases, the method also includes updating symbol history, increasing channel index, increasing frame index, or any combination thereof. In some cases, the updated symbol history, increased channel index, increased frame index, or any combination thereof is combined with the symbol of a subsequent channel.In some cases, the method also includes GC filtering before synthesizing more than one polynucleotide sequence. In some cases, GC filtering includes removing about 5% to about 10% of the channels in more than one channel. In some cases, more than one polynucleotide sequence contains about 45% to about 55% GC content. In some cases, at least 90% of more than one polynucleotide sequence contains about 45% to about 55% GC content. In some cases, applying an internal codec includes: (a) using a lookup table to generate candidate bases for each symbol in the channel; and (b) selecting the next lookup table based at least in part on the symbol of the previous encoding. In some cases, applying an internal codec includes applying a coding scheme.

[0011] On the other hand, provided herein is a method for decoding more than one polynucleotide sequence to generate an output comprising data, the method comprising: (a) determining more than one polynucleotide sequence; (b) applying an internal codec to more than one polynucleotide sequence, wherein the internal codec converts each of the more than one polynucleotide sequence into a channel comprising more than one symbol, wherein the internal codec comprises a hybrid decoding algorithm, the hybrid decoding algorithm comprising a greedy algorithm and a maximum likelihood (ML) algorithm; (c) arranging the channels of data into frames based on a channel index and a frame index in each channel; and (d) applying an external codec to a frame, wherein the external codec comprises an error correction scheme, wherein the frames from the external codec are merged to generate an output comprising data. In some cases, the data comprises more than one symbol. In some cases, the data comprises binary data. In some cases, the binary data comprises a byte stream or a byte array. In some cases, the internal codec comprises a decoding scheme. In some cases, the method further comprises clustering the polynucleotide sequence before (b). In some cases, the clustering is based on an index. In some cases, the clustering comprises partially decoding a frame index, a channel index, or both. In some cases, clustering is performed using a hash function. In some cases, the method also includes comparing the polynucleotide sequence before (b). In some cases, the comparison includes analyzing the consistency of the nucleotide using an alignment algorithm. In some cases, the alignment algorithm includes a pairwise alignment algorithm, a multiple sequence alignment algorithm, or a combination thereof. In some cases, the alignment algorithm includes: (a) initializing the position of each read in more than one read, wherein the initialization includes aligning the polynucleotide sequence with position 0; (b) analyzing the consistency of the next or more bases between each read; (c) determining for each read including the following decision: whether each of the next or more bases is correct or erroneous; (d) increasing the position in the case of a given decision for each read; and (e) repeating steps (b)-(d). In some cases, the error is a deletion, substitution, or insertion. In some cases, more than one read includes about 3 to about 10 reads. In some cases, the length of each read is about 100 to about 300 bases. In some cases, the next one or more bases are about 2, 3, 4 or 5 bases. In some cases, the hybrid decoding algorithm comprises decoding based on the transition probability from one or more states. In some cases, one or more states comprise about 100 to about 1000 most probable states. In some cases, the internal codec also comprises a drift term. In some cases, the drift term comprises an integer. In some cases, the integer is related to the total number of insertions or deletions in the polynucleotide sequence. In some cases, an integer is calculated by adding the value of one or more insertions or one or more missing values ​​in the total number of insertions, deletions or both.In some cases, the value of each of the one or more insertions includes +1, and the value of each of the one or more deletions includes -1. In some cases, (c) includes de-shuffling the channel based on the channel index and grouping the channel into frames based on the frame index. In some cases, the error correction scheme includes Reed-Solomon (RS) code, low density parity check (LDPC) code, turbo code, polar code or any combination thereof. In some cases, at least one polynucleotide sequence in more than one polynucleotide sequence contains an error. In some cases, the error includes insertion, deletion, substitution or any combination thereof.

[0012] In another aspect, a device is provided herein, comprising (a) a memory; and (b) a processing device operably coupled to the memory, wherein the processing device is configured to: (i) split the data into more than one frame, wherein each of the more than one frame includes a frame index; (ii) apply an external codec to each of the more than one frame, wherein the external codec includes an error correction scheme; (iii) divide each frame into more than one channel, wherein each of the more than one channel includes a channel index; (iv) shuffle each channel based at least in part on the channel index; and (v) apply an internal codec to encode each channel in a polynucleotide sequence. In some cases, the internal codec adds redundancy so that the digital data can be decoded in the presence of errors in the polynucleotide sequence. In some cases, the internal codec includes a coding scheme. In some cases, the data includes more than one symbol. In some cases, the data includes digital data. In some cases, the device also includes a synthesizer for generating a polynucleotide sequence. In some cases, the memory, the processing device, or both are part of a computing system. In some cases, the computing system includes a cloud computing system. In some cases, the cloud computing system includes a private cloud, a public cloud, a hybrid cloud, a multi-cloud, or any combination thereof. In some cases, the cloud computing system includes infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS), or any combination thereof.

[0013] In another aspect, an apparatus is provided herein, comprising (a) a memory; (b) a sequencing device configured to determine the sequence of more than one polynucleotide; and (c) a processing device operably coupled to the memory and the sequencing device, wherein the processing device is configured to: (i) apply an internal codec to the sequence, wherein the internal codec converts each of the sequence into a channel comprising more than one symbol, wherein the internal codec comprises a hybrid decoding algorithm, wherein the hybrid decoding algorithm comprises a greedy algorithm and a maximum likelihood (ML) algorithm; (ii) arrange the channels into frames based on a channel index and a frame index in each channel; and (iii) apply an external codec to the frame, wherein the external codec comprises an error correction scheme, wherein the frames from the external codec are merged to generate an output comprising data. In some cases, the internal codec comprises a decoding scheme. In some cases, the data comprises more than one symbol. In some cases, the data comprises digital data. In some cases, the memory, the processing device, or both are part of a computing system. In some cases, the computing system comprises a cloud computing system. In some cases, the cloud computing system comprises a private cloud, a public cloud, a hybrid cloud, a multi-cloud, or any combination thereof. In some cases, cloud computing systems include infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS), or any combination thereof.

[0014] On the other hand, a method for encoding data in a polynucleotide sequence is provided herein, comprising: (a) generating an internal codec including a secret book, wherein the secret book is optimized for one or more constraints; and (b) applying the internal codec to encode the data into more than one polynucleotide sequence. In some cases, the data includes more than one symbol. In some cases, the data includes binary data. In some cases, one or more constraints are associated with nucleic acid synthesis, post-processing, storage, sequencing, or any combination thereof. In some cases, nucleic acid synthesis includes electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing, or any combination thereof. In some cases, one or more constraints associated with nucleic acid synthesis include synthesis errors. In some cases, synthesis errors include insertions, deletions, or mutations. In some cases, post-processing includes one or more of connection, cleavage, hybridization, denaturation, fixation to a solid support, extension, error correction, enrichment, separation, purification, or amplification. In some cases, storage includes cold data storage. In some cases, storage includes nucleic acid storage in a liquid phase or a solid phase. In some cases, one or more constraints associated with storage include temperature, humidity, pressure, salinity, pH, concentration, time, light, UV, O2, or any combination thereof. In some cases, temperature includes room temperature. In some cases, sequencing includes next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing, or any combination thereof. In some cases, the method further includes (c) synthesizing more than one polynucleotide comprising more than one polynucleotide sequence. In some cases, the secret book includes codewords generated in part based on base sequence. In some cases, the base sequence includes predetermined base conversions. In some cases, the internal codec includes two or more secret books. In some cases, each of the two or more secret books encodes a layer during the synthesis of more than one polynucleotide. In some cases, the layer includes each polynucleotide in more than one polynucleotide extending at least one base. In some cases, the synthesis of the layer includes one or more cycles, wherein each of the one or more cycles includes converting a mobile base according to one or more bases in the secret book. In some cases, one of the one or more cycles includes adding one or more of A, T, C, or G. In some cases, each of the two or more secret books includes a different base sequence. In some cases, the secret book includes about 12 code words. In some cases, (b) includes mapping data to more than one polynucleotide sequence based on the secret book. In some cases, the internal codec is further optimized for one or more constraints including the length, GC content, repetition, error or any combination thereof of more than one polynucleotide sequence. In some cases, 40% to 60% encoding redundancy in more than one polynucleotide sequence. In some cases, synthesis includes multiple synthesis cycles.In some cases, the number of synthesis cycles is reduced compared to the number of synthesis cycles required for synthesizing a polynucleotide sequence that is not encoded using an internal codec. In some cases, the number of synthesis cycles reduced is based in part on a flow order. In some cases, the number of synthesis cycles is reduced by at least 30%. In some cases, the number of synthesis cycles is reduced by 50%. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 300. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 155. In some cases, the polynucleotide sequence comprises one or more of A, T, C or G. In some cases, (c) includes synthesizing more than one polynucleotide on a solid support. In some cases, the solid support comprises more than one feature. In some cases, more than 25% of more than one feature is deblocked in each synthesis cycle. In some cases, at least 50% of more than one feature is deblocked in each synthesis cycle. In some cases, each of more than one polynucleotide sequence has the same length. In some cases, 80% to 100% of more than one polynucleotide sequence have the same length. In some cases, the method also includes sequencing more than one polynucleotide to generate more than one output sequence. In some cases, a greedy algorithm, a maximum likelihood (ML) algorithm or a mixed greedy ML algorithm is used to decode more than one output sequence. In some cases, more than one output sequence is decoded at least in part based on the probability of calculating an error. In some cases, errors include deletions, insertions, mutations or any combination thereof.

[0015] On the other hand, a hybrid organic-in silico platform for encoding data is provided herein, the platform comprising: (a) a computing system, the computing system comprising at least one processor and instructions executable by at least one processor to perform operations including: (i) generating an internal codec including a secret book, wherein the codebook is optimized for one or more constraints; and (ii) applying the internal codec to encode data into more than one polynucleotide sequence; and (b) a synthesizer for generating more than one polynucleotide comprising more than one polynucleotide sequence. In some cases, the data includes more than one symbol. In some cases, one or more constraints are associated with nucleic acid synthesis, post-processing, storage, sequencing, or any combination thereof. In some cases, nucleic acid synthesis includes electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing, or any combination thereof. In some cases, one or more constraints associated with nucleic acid synthesis include synthesis errors. In some cases, synthesis errors include insertions, deletions, or mutations. In some cases, post-processing includes one or more of connection, cracking, hybridization, denaturation, fixation to solid support, extension, error correction, enrichment, separation, purification and amplification. In some cases, storage includes cold data storage. In some cases, storage includes nucleic acid storage in liquid phase or solid phase. In some cases, one or more constraints associated with storage include temperature, humidity, pressure, salinity, pH, concentration, time, light, UV, O2 or any combination thereof. In some cases, temperature includes room temperature. In some cases, sequencing includes next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing or any combination thereof. In some cases, computing system includes cloud computing system. In some cases, cloud computing system includes private cloud, public cloud, hybrid cloud, multiple cloud or any combination thereof. In some cases, cloud computing system includes infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS) or any combination thereof. In some cases, the secret book includes codewords generated in part based on base sequence. In some cases, base sequence includes predetermined base conversion. In some cases, the internal codec includes two or more secret books. In some cases, each of the two or more secret books encodes a layer during the synthesis of more than one polynucleotide. In some cases, the layer comprises at least one base extension for each of the more than one polynucleotide. In some cases, the synthesis of the layer includes one or more cycles, wherein each of the one or more cycles includes converting a mobile base according to one or more bases in the secret book. In some cases, one of the one or more cycles includes adding one or more of A, T, C, or G. In some cases, each of the two or more secret books includes a different base sequence.In some cases, the instruction also causes the synthesizer to produce more than one polynucleotide. In some cases, the platform also includes a sequencer for sequencing more than one polynucleotide to produce more than one output sequence. In some cases, the instruction also causes the computing system to receive more than one output sequence. In some cases, the computing system also performs operations including the following: (iii) decoding more than one output sequence. In some cases, a greedy algorithm, a maximum likelihood (ML) algorithm or a mixed greedy ML algorithm is used to decode more than one output sequence. In some cases, more than one output sequence is decoded based at least in part on the probability of missing, inserted, suddenlyd or any combination thereof by calculation. In some cases, the platform also includes a storage unit for storing more than one polynucleotide. In some cases, operation also includes transferring more than one polynucleotide between a synthesizer, a sequencer, a storage unit or any combination thereof. In some cases, specific base conversion allows synthesis according to the flow order. In some cases, the secret book includes about 12 code words. In some cases, wherein (a) (ii) includes mapping data to more than one polynucleotide sequence based on the secret book. In some cases, the internal codec is further optimized for the constraints including the length, GC content, repetition or any combination thereof of more than one polynucleotide sequence. In some cases, 40% to 60% encoding redundancy in more than one polynucleotide sequence. In some cases, generating more than one polynucleotide includes multiple synthesis cycles. In some cases, the number of synthesis cycles is reduced compared to the number of synthesis cycles required for synthesizing a polynucleotide sequence that is not encoded using the internal codec. In some cases, the number of synthesis cycles reduced is based in part on the flow order. In some cases, the number of synthesis cycles is reduced by at least 30%. In some cases, the number of synthesis cycles is reduced by 50%. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 300. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 155. In some cases, the polynucleotide sequence comprises one or more A, T, C or G. In some cases, generating more than one polynucleotide comprises base-by-base synthesis. In some cases, the synthesizer comprises a solid support comprising more than one feature. In some cases, each of more than one feature is independently addressable by one or more electrodes of a solid support. In some cases, each of more than one feature is addressable by masking. In some cases, masking comprises a physical barrier. In some cases, masking comprises controlling the reactivity of one or more of the more than one feature. In some cases, controlling the reactivity comprises deprotecting at one or more of the more than one feature. In some cases, deprotection comprises acid generation. In some cases, deprotection comprises electrochemical deprotection. In some cases, more than 25% of the more than one feature is deblocked in each synthesis cycle.In some cases, at least 50% of more than one feature is deblocked in each synthesis cycle. In some cases, each of more than one polynucleotide sequence has the same length. In some cases, 80% to 100% of more than one polynucleotide sequence has the same length.

[0016] In one aspect, a system for storing data in DNA is provided herein, comprising: one or more processing units; a memory in communication with the one or more processing units, instructions stored in the memory and executed on the one or more processing units, the instructions causing the system to: generate more than one pool, wherein each of the more than one pools comprises a pool descriptor, a pool item comprising a payload of data, and an end descriptor; determine the first one or more hashes of the payload of each pool item; and apply a coding scheme to encode more than one pool into a sequence of more than one polynucleotide. In some embodiments, the coding scheme comprises an internal codec, an external codec, or both described herein. In some embodiments, the data comprises an information item or digital information described herein. In some embodiments, the data comprises one or more objects. In some embodiments, the one or more processing units, the memory, or both are part of a computing system. In some cases, the computing system comprises a cloud computing system. In some cases, the cloud computing system comprises a private cloud, a public cloud, a hybrid cloud, a multi-cloud, or any combination thereof. In some cases, the cloud computing system comprises infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS), or any combination thereof. In some embodiments, the instructions stored in the memory and executed on one or more processing units make the system determine the second one or more hashes of each of the one or more objects. In some embodiments, one or more objects include files or metadata associated with files. In some embodiments, the pool descriptor includes a list of versions, pool IDs, pool item descriptors, or any combination thereof. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a universal unique identifier (UUID) or a content ID. In some embodiments, the list of pool item descriptors includes the path of the object, the size of the object, the range of the pool item within the object, the offset of the pool item in the pool, or any combination thereof. In some embodiments, each of the one or more pool items also includes a hash of the pool item from the first one or more hashes. In some embodiments, the end pool descriptor includes a list of object descriptors. In some embodiments, the list of object descriptors includes the path of the object, the hash of the object from the first one or more hashes, or a combination thereof. In some embodiments, each of more than one pool is about 1GB to about 1TB. In some embodiments, more than one pool includes a redundant pool. In some embodiments, the first one or more hashes, the second one or more hashes, or both are determined using a hash module. In some embodiments, the hash module executes on one or more processing units. In some embodiments, the first one or more hashes require less memory than the one or more objects. In some embodiments, the second one or more hashes require less memory than the one or more pool items.In some embodiments, the hash module includes a hash function. In some embodiments, the hash function includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224 or SHA-512 / 256. In some embodiments, the instruction also causes the system to generate one or more index pools. In some embodiments, one or more index pools include a list of index pool descriptors and object indexes. In some embodiments, the index pool descriptor includes a version, a pool ID, the size of the pool, and a timestamp. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a UUID or a content ID. In some embodiments, the list of object indexes includes the path of the object, the hash of the object, the list of object fragments, the list of object metadata, or any combination thereof. In some embodiments, the list of object fragments includes the pool ID of the pool containing the fragment, the range of the fragment, or a combination thereof. In some embodiments, the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof. In some embodiments, the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof. In some embodiments, each of the one or more index pools is about 1 GB to about 1 TB. In some embodiments, instructions stored in memory and executed on one or more processing units cause the system to retrieve data stored in the DNA. In some embodiments, the instructions include: applying a decoding scheme to decode the sequence of more than one polynucleotide in each of more than one pool; and verifying at least the payload of each pool item using the first one or more hashes.

[0017] On the one hand, the present invention provides a device for storing information in DNA, comprising: one or more compartments, wherein each compartment comprises: (a) a library comprising more than one polynucleotide, wherein the library encodes a pool of information corresponding to one or more objects; and (b) a medium for storing more than one polynucleotide. In some embodiments, the information comprises an information item or digital information described herein. In some embodiments, the information comprises more than one symbol. In some embodiments, one or more compartments are connected. In some embodiments, one or more compartments are not connected. In some embodiments, the medium comprises a solid, a liquid, a gas, or any combination thereof. In some embodiments, the medium comprises a salt solution having a molar ratio of a salt cation to a phosphate group in DNA of less than 20:1. In some embodiments, the salt solution is dried to produce a dry product. In some embodiments, the device further comprises a solid support comprising a surface. In some embodiments, the device further comprises more than one structure located on the surface, wherein more than one polynucleotide extends from more than one structure. In some embodiments, one or more objects comprise files or metadata associated with files. In some embodiments, the pool includes a pool descriptor, one or more pool items and an end pool descriptor. In some embodiments, the pool descriptor includes a version, a pool ID, a list of pool item descriptors, or any combination thereof. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a universal unique identifier (UUID) or a content ID. In some embodiments, the list of pool item descriptors includes the path of the object, the size of the object, the range of the pool item within the object, the offset of the pool item in the pool, or any combination thereof. In some embodiments, each of the one or more pool items includes a data payload, a hash of the pool item, or a combination thereof. In some embodiments, the end pool descriptor includes a list of object descriptors. In some embodiments, the list of object descriptors includes the path of the object, the hash of the object, or a combination thereof. In some embodiments, the pool includes digital information of about 1GB to about 1TB. In some embodiments, the device also includes one or more second compartments, each of which includes a second library of encoding index pools. In some embodiments, one or more index pools include a list of index pool descriptors and object indexes. In some embodiments, the index pool descriptor includes a version, a pool ID, a size of the pool, and a timestamp. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a UUID or a content ID. In some embodiments, the list of object indexes includes the path of the object, the hash of the object, a list of object fragments, a list of object metadata, or any combination thereof. In some embodiments, the list of object fragments includes the pool ID of the pool containing the fragments, the range of the fragments, or a combination thereof. In some embodiments, the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof.In some embodiments, the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof. In some embodiments, each of the one or more index pools is about 1 GB to about 1 TB.

[0018] In another aspect, provided herein is a method for storing data in more than one polynucleotide, the method comprising: generating more than one pool, wherein each of the more than one pool comprises a pool descriptor, a pool item comprising a payload of data, and an end descriptor; determining a first one or more hashes of the payload of each pool item; and applying a coding scheme to encode the more than one pool into a sequence of more than one nucleotide. In some embodiments, the coding scheme comprises an internal codec, an external codec, or both described herein. In some embodiments, the data comprises an information item or digital information described herein. In some cases, the data comprises more than one symbol. In some embodiments, the data comprises one or more objects. In some embodiments, the method further comprises determining a second one or more hashes of each of the one or more objects. In some embodiments, the method further comprises storing more than one polynucleotide. In some embodiments, the polynucleotides in the more than one polynucleotides corresponding to each of the more than one pools are stored in separate containers of a data storage system. In some embodiments, the method further comprises generating more than one polynucleotide. In some embodiments, generating more than one polynucleotide comprises phosphoramidite-based synthesis of deoxyribonucleic acid (DNA). In some embodiments, the reagent for the synthesis based on phosphoramidite comprises nucleoside phosphoramidite, oxidant, activator or deblocking agent, or the solvent comprising acetonitrile. In some embodiments, producing more than one polynucleotide comprises enzymatic DNA synthesis. In some embodiments, the reagent for enzymatic DNA synthesis comprises terminal deoxynucleotidyl transferase (TdT) or deblocking agent, or the solvent comprising water. In some embodiments, one or more objects comprise files or metadata associated with files. In some embodiments, the pool descriptor comprises a list of versions, pool IDs, pool item descriptors, or any combination thereof. In some embodiments, the pool ID comprises a unique ID. In some embodiments, the unique ID comprises a universal unique identifier (UUID) or a content ID. In some embodiments, the list of pool item descriptors comprises the path of the object, the size of the object, the range of the pool item within the object, the offset of the pool item in the pool, or any combination thereof. In some embodiments, each of the one or more pool items also comprises a hash of the pool item from the first one or more hashes. In some embodiments, the end pool descriptor comprises a list of object descriptors. In some embodiments, the list of object descriptors includes the path of the object, the hash of the object from the first one or more hashes, or a combination thereof. In some embodiments, each of the more than one pools is about 1 GB to about 1 TB. In some embodiments, the more than one pools include redundant pools. In some embodiments, a hash module is used to determine the first one or more hashes, the second one or more hashes, or both. In some embodiments, the second one or more hashes require less memory than the one or more objects.In some embodiments, the first one or more hashes require less memory than one or more pool items. In some embodiments, the hash module includes a hash function. In some embodiments, the hash function includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224 or SHA-512 / 256. In some embodiments, the method also includes creating one or more index pools. In some embodiments, one or more index pools include a list of index pool descriptors and object indexes. In some embodiments, the index pool descriptor includes a version, a pool ID, the size of the pool, and a timestamp. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a UUID or a content ID. In some embodiments, the list of object indexes includes the path of the object, the hash of the object, the list of object fragments, the list of object metadata, or any combination thereof. In some embodiments, the list of object fragments includes the pool ID of the pool containing the fragments, the range of the fragments, or a combination thereof. In some embodiments, the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof. In some embodiments, the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof. In some embodiments, each of the one or more index pools is about 1 GB to about 1 TB.

[0019] In another aspect, provided herein is a method for retrieving data stored in more than one polynucleotide, the method comprising: determining the sequence of more than one polynucleotide, wherein the more than one polynucleotide is in more than one pool; applying a decoding scheme to decode the sequence of more than one polynucleotide of each of the more than one pools, wherein each pool comprises a pool descriptor, a pool item comprising a payload of data, and an end descriptor; and verifying at least the payload of each pool item using a first one or more hashes. In some embodiments, the decoding scheme comprises an internal codec, an external codec, or both described herein. In some embodiments, the data comprises an information item or digital information described herein. In some embodiments, the data comprises one or more objects. In some embodiments, the one or more objects comprise a file or metadata associated with a file. In some embodiments, the method further comprises verifying the one or more objects using a second one or more hashes. In some embodiments, at least verifying the payload comprises verifying the first one or more hashes using a hash function. In some embodiments, the method further comprises combining the payloads from each pool item to retrieve the data. In some embodiments, the method further comprises storing the data on a memory. In some embodiments, each of more than one pool is about 1GB to about 1TB. In some embodiments, verifying one or more objects includes verifying the second one or more hashes using a hash function. In some embodiments, the hash function includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224 or SHA-512 / 256. In some embodiments, determining the sequence includes sequencing more than one polynucleotide. In some embodiments, sequencing includes next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing or any combination thereof. In some embodiments, the method also includes accessing the index pool in one or more index pools to determine more than one pool including one or more objects. In some embodiments, the index pool includes a list of index pool descriptors and object indexes. In some embodiments, the index pool descriptor includes a version, pool ID, the size of the pool and a timestamp. In some embodiments, the pool ID includes a unique ID. In some embodiments, the unique ID includes a UUID or a content ID. In some embodiments, the list of object indexes includes the path of the object, the hash of the object, the list of object fragments, the list of object metadata, or any combination thereof. In some embodiments, the list of object fragments includes the pool ID of the pool containing the fragments, the range of the fragments, or a combination thereof. In some embodiments, the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof. In some embodiments, the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof.In some embodiments, each of the one or more index pools is about 1 GB to about 1 TB. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] A better understanding of the features and advantages of the present subject matter will be obtained by referring to the following detailed description setting forth illustrative embodiments and the accompanying drawings, in which:

[0022] Figure 1 Non-limiting examples of encoding schemes for low-level codecs according to some embodiments are shown.

[0023] Figure 2 A non-limiting example of a decoding scheme for a low-level codec according to some embodiments is shown.

[0024] Figure 3 Non-limiting examples of encoding schemes including external codecs according to some embodiments are shown.

[0025] Figure 4 A non-limiting example of an encoding scheme for a shuffled channel including data is shown in accordance with some embodiments.

[0026] Figure 5 A non-limiting example of an encoding scheme including a first inner codec according to some embodiments is shown.

[0027] Figure 6 A non-limiting example of an encoding scheme including a second inner codec according to some embodiments is shown.

[0028] Figure 7 A non-limiting example of a decoding scheme including an inner codec and an outer codec according to some embodiments is shown.

[0029] Figure 8 A non-limiting example of a greedy algorithm for decoding is shown in accordance with some embodiments.

[0030] Fig. 9 A non-limiting example of a maximum likelihood (ML) algorithm for decoding according to some embodiments is shown.

[0031] Fig.10 A non-limiting example of a computing device is shown; in this case, the device has one or more processors, memory, storage, and a network interface.

[0032] Fig.11A A non-limiting example of a "lift-off" process for making a polynucleotide synthesis surface according to some embodiments is shown.

[0033] Fig. 11BA non-limiting example of a wet etching process for making a polynucleotide synthesis surface according to some embodiments is shown. The process can also be applied to a dry etching process.

[0034] Fig.12 Non-limiting examples of encoding schemes for advanced codecs according to some embodiments are shown.

[0035] Fig.13 A non-limiting example of a decoding scheme for an advanced codec according to some embodiments is shown.

[0036] Fig.14 Non-limiting examples of digital information storage according to some embodiments are shown.

[0037] Fig.15 A non-limiting example of generating a hash according to some embodiments is shown.

[0038] Fig.16 Non-limiting examples of systems for synthesizing, storing, and sequencing more than one polynucleotide according to some embodiments are shown.

[0039] Figure 17A-17G Non-limiting examples of structures or compartments for storing more than one polynucleotide are shown according to some embodiments. Fig.17A A substantially tubular structure is shown. Fig. 17B A structure comprising a cap and a body welded flush together is shown. Fig. 17C A structure including a removable screw cap is shown. Fig.17D A structure including a diaphragm is shown. Fig.17E A structure is shown that includes two rounded, pill-shaped halves that form a seal when one half is inserted into the other. Fig.17F A structure is shown comprising a substantially flat disc-shaped container having a sealable lid. Figure 17G A structure comprising a box with an optionally attached lid is shown.

[0040] Details

[0041] Provided herein are methods and systems for storing digital information in nucleic acids. As in many storage media, synthetic DNA may have inherent errors, such as deletions, insertions, mutations or fragmentation, which may result in the erasure of complete oligonucleotides. Due to aging or sample processing, there may also be some loss of oligonucleotides. Typical techniques used in computer science and telecommunications only address erasure and / or mutations, without addressing the specific behavior of oligonucleotide pools. For example, sequencing oligonucleotide pools provide oligonucleotides in random order, while typical storage media such as hard drives provide data streams in a known and expected order created during write. In addition, many codecs for storing digital information focus on encoding digital information in nucleic acids, but may not provide a way to store and retrieve a structured list of files. Therefore, provided herein are codecs and implementations that can obtain multiple "objects" and effectively store them as one or more pools or retrieve them from one or more pools. Objects may include files or metadata associated with files. Such codec implementations can be combined with low-level codecs and / or external codecs (e.g., including error correction codes, such as but not limited to those described herein) for encoding digital information in nucleic acids.

[0042] In some cases, the method encodes data in more than one polynucleotide sequence. The data can be represented as more than one symbol. In some cases, the method includes one or more steps of: splitting the data into more than one frame; applying an external codec to each frame in more than one frame; dividing each frame into more than one channel; shuffling each channel based at least in part on a channel index; and applying an internal codec (e.g., a coding scheme) to encode each channel in a polynucleotide sequence in more than one polynucleotide sequence. In some cases, each frame in more than one frame includes a frame index. In some cases, the external codec includes an error correction scheme. In some cases, each channel in more than one channel includes a channel index.

[0043] In some cases, the method decodes more than one polynucleotide sequence to generate an output comprising data. Data can be represented as more than one symbol. In some cases, the method includes one or more steps in the following: determine more than one polynucleotide sequence; apply an internal codec (e.g., a decoding scheme) to more than one polynucleotide sequence; arrange data channels into frames based on channel indexes and frame indexes in each data channel; and apply an external codec to a frame. In some cases, the internal codec converts each of more than one polynucleotide sequence into a channel comprising more than one symbol. In some cases, the internal codec includes a hybrid decoding algorithm, which includes a greedy algorithm and a maximum likelihood (ML) algorithm. In some cases, the external codec includes an error correction scheme. In some cases, the frame from the external codec is merged to generate an output comprising data.

[0044] In some cases, the system encodes data in more than one polynucleotide sequence. In some cases, the system includes an apparatus comprising one or more of the following: a memory; and a processing device operably coupled to the memory. In some cases, the processing device is configured to perform one or more of the following steps: splitting the data into more than one frame; applying an external codec to each frame in more than one frame; dividing each frame into more than one channel; shuffling each channel based at least in part on a channel index; and applying an internal codec to encode each channel in the polynucleotide sequence. In some cases, each frame in more than one frame includes a frame index. In some cases, each channel in more than one channel includes a channel index. In some cases, the external codec includes an error correction scheme. In some cases, the internal codec adds redundancy so that the data can be decoded in the presence of errors in the polynucleotide sequence. In some cases, the internal codec includes a coding scheme.

[0045] In some cases, the system decodes more than one polynucleotide sequence to generate an output comprising data. In some cases, the system includes an apparatus comprising one or more of the following: a memory; a sequencing device configured to determine more than one polynucleotide sequence; and a processing device operably coupled to the memory. In some cases, the processing device is configured to perform one or more of the following steps: applying an internal codec to more than one polynucleotide sequence; arranging data channels into frames based on a channel index and a frame index in each data channel; and applying an external codec to a frame. In some cases, the internal codec converts each sequence into a channel comprising more than one symbol. In some cases, the internal codec includes a hybrid decoding algorithm comprising a greedy algorithm and a maximum likelihood (ML) algorithm. In some cases, the external codec includes an error correction scheme. In some cases, frames from the external codec are merged to generate an output comprising data. In some cases, the internal codec includes a decoding scheme.

[0046] In some cases, the method encodes data in a polynucleotide sequence. In some cases, the method includes one or more steps of: (a) generating an internal codec including a secret book, wherein the secret book is optimized for one or more constraints; and (b) applying the internal codec to encode the data into more than one polynucleotide sequence. In some cases, the method also includes generating more than one polynucleotide comprising more than one polynucleotide sequence.

[0047] In some cases, a hybrid organic-computer simulation platform for encoding data is provided herein. The platform includes one or more of the following: a computing system including at least one processor and instructions executable by at least one processor to perform operations; and a synthesizer for generating more than one polynucleotide comprising more than one polynucleotide sequence. In some cases, the operations include one or more of the following: generating an internal codec including a secret book, wherein the secret book is optimized for one or more constraints; and applying the internal codec to encode data into more than one polynucleotide sequence.

[0048] In some cases, the system stores information in DNA. In some cases, the system includes any one or a combination of the following: one or more processing units; a memory in communication with the one or more processing units, and instructions stored in the memory and executed on the one or more processing units. In some cases, the instructions cause the system to perform any one or a combination of the following: splitting digital information of one or more objects into more than one pool; generating a pool descriptor, one or more pool items, and an end pool descriptor in each of the more than one pool; determining a first one or more hashes of a data payload of each of the one or more pool items and a second one or more hashes of each of the one or more objects; and applying a coding scheme to encode the digital information in the more than one pool into more than one polynucleotide.

[0049] In some cases, the device is used to store information in DNA. In some cases, the device includes one or more compartments. In some cases, each compartment contains any one or a combination of: a library containing more than one polynucleotide; and a medium for storing more than one polynucleotide. In some cases, the library encodes a pool of information corresponding to one or more objects.

[0050] In some cases, the method stores the data in more than one polynucleotide. In some cases, the method includes any one or a combination of: generating more than one pool; determining a first one or more hashes of a payload of each pool item; and applying an encoding scheme to encode the more than one pool into a sequence of more than one nucleotide. In some cases, each of the more than one pool includes a pool descriptor, a pool item including the data payload, and an end descriptor.

[0051] In some cases, the method retrieves data stored in more than one polynucleotide. In some cases, the method includes any one or a combination of: determining a sequence of more than one polynucleotide; applying a decoding scheme to decode the sequence of more than one polynucleotide in each of more than one pool; and verifying at least a payload of each pool item using the first one or more hashes. In some cases, more than one polynucleotide is in more than one pool. In some cases, each pool includes a pool descriptor, a pool item including a payload of data, and an end descriptor.

[0052] Also provided herein is a method and system for optimizing polynucleotide synthesis. In some cases, synthesis is optimized using a synthesis-optimized codec (such as those provided herein). Polynucleotides can be synthesized according to the device provided herein. Electronic synthesis generally includes deblocking a specific site (e.g., a feature or site on the surface for polynucleotide synthesis) and flowing a specific base (e.g., a nucleic acid monomer), which is repeated for each base. This means that polynucleotides without a specific base sequence may need 4 cycles per layer (e.g., A, T, C, G), particularly when synthesizing millions of polynucleotides together, because the chance of the section of the polynucleotide matching in the synthesis order is very low. For example, the masking surface is shielded to protect a specific site (wherein each site contains a unique polynucleotide, and is independently addressable) from base addition, bases are coupled to unprotected sites, and then masking is changed to allow coupling bases at different sites. The layer generally includes at least one base extending each polynucleotide. For example, if the polynucleotide is M bases long, assuming that 4 cycles are needed to add a single nucleic acid to the polynucleotide each time, then 4 × M cycles may be needed for synthesis. This approach can be more expensive because it may require more time, more reagents, or both. It can also increase the chance of DNA damage because each cycle requires an oxidation step and a deblocking step, which can lead to a higher error rate.

[0053] The method, system and platform of optimizing synthesis can include the internal codec optimized to generate polynucleotide according to the specific order of base synthesis.This can allow the synthesis of polynucleotide to have less than 4×M cycles, wherein M is the number of bases of polynucleotide.The method can also provide redundancy for error correction, such as using external codec or error correction code (ECC).When the data of the same amount of polynucleotide encoding synthesized, the method can also accelerate the synthesis of polynucleotide relative to unoptimized synthesis method (for example, needing 4×M cycles).In some cases, the mixture of base (for example, two or three) flows through the surface in a single cycle.In some cases, the synthesis method is configured to be used with one or more secret books provided herein.As described herein, unoptimized synthesis method can generally refer to the synthesis of polynucleotide without base sorting.In some cases, the synthesis rate accelerates about 1.5 times, 2 times, 2.5 times, 3 times, 3.5 times or 4 times relative to unoptimized synthesis method.In some cases, the synthesis rate accelerates up to 2 times, 2.5 times, 3 times, 3.5 times or 4 times relative to unoptimized synthesis method. In some cases, the synthesis rate is accelerated by up to about 1.5 times, 2 times, 2.5 times, 3 times or 3.5 times relative to the unoptimized synthesis method. In some cases, due to the need for less oxidation steps, the synthesis rate is accelerated while improving DNA quality. In some cases, the synthesis rate is accelerated while reducing errors.

[0054] In some cases, the method encoding data provided herein.Data can be digital information or information items.Data can be represented as one or more symbols.In some instances, one or more symbols include numerical values, such as binary data.In some cases, the data represented as a symbol set are encoded into different symbol sets using a codec.In some cases, this codec is referred to as an internal codec.In some cases, different symbol sets include symbol sequences, such as polynucleotide sequences.

[0055] Method described herein may include the use or generation of an internal codec. In some cases, the method includes generating an internal codec that includes a secret book. In some cases, the secret book includes the content, structure and layout of a data set (e.g., digital information encoded in nucleic acid). In some cases, the internal codec includes two or more secret books. In some cases, the layers during the synthesis of each encoding polynucleotide in two or more secret books. In some cases, the secret book is optimized for one or more constraints. In some cases, one or more constraints are related to nucleic acid synthesis, post-processing, storage, sequencing or any combination thereof.

[0056] In some cases, a secret book is generated in a base sequence. In some cases, the secret book is optimized for one or more base conversions. In some cases, the base sequence generates one or more base conversions. Such one or more base conversions can be referred to as specific base conversions or predetermined base conversions. In some cases, each of two or more secret books includes different base sequences. In some cases, each of two or more secret books includes different one or more base conversions. In some cases, the secret book is optimized for a specific base conversion at a given layer, a loop index, a history, or any combination thereof. In some instances, the history includes one or more of the previous layers, one or more secret books encoding the previous one or more layers, the loop index of one or more previous layers, or any combination thereof. In some instances, the method includes applying an internal codec to encode data into more than one polynucleotide sequence.

[0057] Methods provided herein can be carried out on a platform. In some cases, the platform includes a hybrid organic-computer simulation platform. In some cases, the platform encodes data (e.g., binary data). In some cases, the platform includes a computing system, and the computing system includes at least one processor and instructions that can be executed by at least one processor to perform operations. In some cases, the operation includes generating an internal codec including a secret book. In some cases, a secret book is generated in base order. In some cases, the base order generates a codeword with one or more base transitions. In some cases, the operation includes applying an internal codec to encode data into more than one polynucleotide sequence. In some cases, the platform includes a synthesizer. In some cases, the platform includes a synthesizer for generating more than one polynucleotide comprising more than one polynucleotide sequence. In some cases, the synthesizer produces more than one polynucleotide sequence by synthesis, connection, assembly or any combination thereof. In some cases, the platform is integrated into one or more additional systems, such as traditional magnetic or tape storage devices.

[0058] Nucleic acid-based information storage

[0059] Provided herein are devices, compositions, systems and methods for information (data) storage based on nucleic acids.Compared with traditional binary information encoding, biomolecules such as DNA molecules provide suitable substrates (hosts) for information storage in part due to their stability over time and enhanced information encoding capabilities.In the first step, data including the first more than one symbol, for example, a digital sequence of coded information items (i.e., digital information in binary codes for computer processing) are received. An encryption scheme is applied to convert the first more than one symbol into a second more than one symbol. The second more than one symbol may include a nucleic acid sequence. For example, an encryption scheme is applied to convert the digital sequence from a binary code into a polynucleotide sequence. The surface material for nucleic acid extension, the design of the site (also referred to as, arrangementspots) for nucleic acid extension and / or the reagent for nucleic acid synthesis are selected. The surface of the prepared structure is used for nucleic acid synthesis. Then de novo polynucleotide synthesis is performed. The synthesized polynucleotides are stored in whole or in part and can be used for subsequent release.After release, all or part of the polynucleotides are sequenced, and decrypted to convert the nucleic acid sequence back to the digital sequence.Then the digital sequence is assembled to obtain the alignment coding of the original information item.

[0060] Information Items

[0061] Optionally, the early steps of the data storage process disclosed herein include obtaining or receiving data, and the data include one or more information items in the form of initial code. Information items include, but are not limited to, text, audio, and visual information. Exemplary sources of information items include, but are not limited to, books, journals, electronic databases, medical records, letters, tables, recordings, animal records, biological maps, broadcasts, films, short videos, emails, bookkeeping phone logs, Internet activity logs, drawings, paintings, prints, photos, pixelated graphics, and software codes. Exemplary biological map sources of information items include, but are not limited to, gene libraries, genomes, gene expression data, and protein activity data. Exemplary formats of information items include, but are not limited to: .txt, .PDF, .doc, .docx, .ppt, .pptx, .xls, .xlsx, .rtf, .jpg, .gif, .psd, .bmp, .tiff, .png, and .mpeg. The size of a single file or the amount of more than one file encoding an information item in a digital format includes, but is not limited to, up to 1024 bytes (equal to 1 KB), 1024 KB (equal to 1 MB), 1024 MB (equal to 1 GB), 1024 GB (equal to 1 TB), 1024 TB (equal to 1 PB), 1 exabyte, 1 zettabyte, 1 yottabyte, 1 xenottabyte, or more. In some cases, the amount of digital information is at least 1 gigabyte (GB). In some cases, the amount of digital information is at least 1 gigabyte, 2 gigabytes, 3 gigabytes, 4 gigabytes, 5 gigabytes, 6 gigabytes, 7 gigabytes, 8 gigabytes, 9 gigabytes, 10 gigabytes, 20 gigabytes, 50 gigabytes, 100 gigabytes, 200 gigabytes, 300 gigabytes, 400 gigabytes, 500 gigabytes, 600 gigabytes, 700 gigabytes, 800 gigabytes, 900 gigabytes, 1000 gigabytes, or more than 1000 gigabytes. In some cases, the amount of digital information is at least 1 terabyte (TB). In some cases, the amount of digital information is at least 1 terabyte, 2 terabytes, 3 terabytes, 4 terabytes, 5 terabytes, 6 terabytes, 7 terabytes, 8 terabytes, 9 terabytes, 10 terabytes, 20 terabytes, 50 terabytes, 100 terabytes, 200 terabytes, 300 terabytes, 400 terabytes, 500 terabytes, 600 terabytes, 700 terabytes, 800 terabytes, 900 terabytes, 1000 terabytes, or more than 1000 terabytes. In some cases, the amount of digital information is at least 1 petabyte (PB).In some cases, the amount of digital information is at least 1 beat byte, 2 beat bytes, 3 beat bytes, 4 beat bytes, 5 beat bytes, 6 beat bytes, 7 beat bytes, 8 beat bytes, 9 beat bytes, 10 beat bytes, 20 beat bytes, 50 beat bytes, 100 beat bytes, 200 beat bytes, 300 beat bytes, 400 beat bytes, 500 beat bytes, 600 beat bytes, 700 beat bytes, 800 beat bytes, 900 beat bytes, 1000 beat bytes or more than 1000 beat bytes. In some cases, digital information does not include the genome data obtained from organism. In some cases, information items are encoded. Non-limiting encoding method examples include 1 / base, 2 / base, 4 / base or other encoding methods.

[0062] Method and system for information storage

[0063] Provided herein are methods and systems for storing information (e.g., digital information). In some cases, provided herein are methods and systems for encoding. In some cases, information includes one or more objects. In some cases, one or more objects include information items, such as but not limited to those described herein. In some cases, one or more objects include files or metadata associated with the files. In some cases, methods and systems encode digital data, such as binary data. In some cases, methods and systems include internal codecs, external codecs, or a combination thereof. In some cases, binary data includes byte streams or byte arrays. In some cases, data or one or more objects are about 1GB to about 1TB. In some cases, data are about 1GB to about 1TB. In some cases, data or one or more objects are about 1 GB to about 10 GB, about 1 GB to about 50 GB, about 1 GB to about 100 GB, about 1 GB to about 500 GB, about 1 GB to about 1 TB, about 10 GB to about 50 GB, about 10 GB to about 100 GB, about 10 GB to about 500 GB, about 10 GB to about 1 TB, about 50 GB to about 100 GB, about 50 GB to about 500 GB, about 50 GB to about 1 TB, about 100 GB to about 500 GB, about 100 GB to about 1 TB, or about 500 GB to about 1 TB. In some cases, data are about 1 GB, about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB. In some cases, data or one or more objects are at least about 1 GB, about 10 GB, about 50 GB, about 100 GB, or about 500 GB. In some cases, the data or one or more objects is at most about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB.

[0064] The system storing digital information may include one or more processing units, memory communicating with one or more processing units, instructions stored in memory and executed on one or more processing units, or any combination thereof. In some cases, one or more processing units and memory are distributed across one or more physical or logical locations. In some cases, one or more processing units include a central processing unit (CPU), a graphics processing unit (GPU), a single-core processor, a multi-core processor, a processor cluster, an application-specific integrated circuit (ASIC), a programmable circuit (such as a field programmable gate array (FPGA)), an AI accelerator, and any combination thereof. In some cases, one or more processing units include a single instruction multiple data (SIMD) or a single program multiple data (SPMD) parallel architecture. As an example, one or more processing units include one or more GPUs or CPUs implementing SIMD or SPMD. In some cases, AI accelerators include Google-TPU, Graphcore, Cerebras, SambaNova, or a combination thereof. In some embodiments, in addition to hardware implementations, one or more processing units are implemented in software and / or firmware. The software or firmware implementation of the processing unit may include computer or machine executable instructions written in any suitable programming language to perform the various functions described herein. The software implementation of one or more processing units may be stored in whole or in part in the memory. Alternatively or additionally, the system may include one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc. In some cases, the memory includes removable memory, non-removable memory, local memory, and / or remote memory to provide instructions, data structures, program modules (e.g., hash modules), and storage of any other data described herein. In some cases, the memory is used to store information related to the algorithms described herein (e.g., software code, parameters, executable instructions, etc.).

[0065] The instructions stored on the memory may include one or more steps for storing digital information. Fig.12In one or more operations for storing digital information, illustrated exemplary. Dashed line operation can be performed in some embodiments, but not performed in other embodiments. In some cases, one or more steps include splitting the digital information of one or more objects into more than one pool 1205. In some cases, the object in one or more objects splits across more than one pool. In some cases, each of more than one pool is about 1GB to about 1TB. In some cases, each of more than one pool is about 1GB to about 10GB, about 1GB to about 50GB, about 1GB to about 100GB, about 1GB to about 500GB, about 1GB to about 1TB, about 10GB to about 50GB, about 10GB to about 100GB, about 10GB to about 500GB, about 10GB to about 1TB, about 50GB to about 100GB, about 50GB to about 500GB, about 50GB to about 1TB, about 100GB to about 500GB, about 50GB to about 1TB, about 100GB to about 500GB, about 100GB to about 1TB or about 500GB to about 1TB. In some cases, each of the more than one pool is about 1 GB, about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB. In some cases, each of the more than one pool is at least about 1 GB, about 10 GB, about 50 GB, about 100 GB, or about 500 GB. In some cases, each of the more than one pool is at most about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB.

[0066] In some cases, one or more objects include information items, such as files, as previously described herein. In some cases, one or more objects include metadata associated with information items (e.g., metadata associated with files). Non-limiting examples of metadata associated with objects include a list of keywords attached to the object, object size, thumbnail, text summary, ID range of a sorted keyword value database, timestamp, version, or any other data or any combination thereof providing information about one or more aspects of the object. In some instances, metadata is customizable. In some instances, metadata is used to search for objects in more than one pool.

[0067] Fig.141405. An example diagram of digital information storage is illustrated in FIG. As shown, one or more objects 1405 can be split into more than one pool 1410. In some cases, an object is split into more than one pool. In some cases, an object is split into more than one pool based on size in part. In some cases, an object is split into two, three, four, five, six, seven, eight, nine or ten pools. In some cases, more than one object is split into more than one pool. In some cases, one or more objects are in one pool. In some cases, one, two, three, four, five, six, seven, eight, nine or ten objects are in one pool. In some cases, more than one pool is repeated. In some cases, more than one pool includes redundant pools, wherein two or more pools include the same one or more objects. In some cases, two, three, four, five, six, seven, eight, nine or ten pools include the same one or more objects.

[0068] Each pool in more than one pool may include any one or combination of pool descriptors, pool items or end descriptors. In some cases, the pool includes at least one pool item. In some cases, the pool includes more than one pool item. In some cases, the pool includes at least one pool descriptor. In some cases, the pool includes more than one pool descriptor. In some cases, the pool includes at least one end descriptor. In some cases, the pool includes more than one end descriptor. As an example, each pool includes pool descriptor 1415, one or more pool items 1420 and end descriptor 1425. In some cases, the pool includes redundant pool items, pool descriptors, end pool descriptors or combinations thereof. In this case, two or more pool items, pool descriptors, end pool descriptors or combinations thereof are the same. In some cases, two, three, four, five, six, seven, eight, nine or ten pool descriptors, end pool descriptors or combinations thereof are the same.

[0069] refer to Fig.12 In some cases, one or more operations in the instruction include generating more than one pool 1210 including a pool descriptor, a pool item, and an end descriptor. In some cases, the data is divided into pools, and the instruction includes generating a pool descriptor, a pool item, an end descriptor, or any combination thereof in each of the more than one pools. In this case, the generated pool descriptor, pool item, end descriptor is added to each pool. In some cases, the pool descriptor includes a version, a pool ID, a list of pool item descriptors, or any combination thereof. In some cases, the version includes a version of the information (e.g., if the information is updated). In some cases, the version is a version of the structure of the pool. In some cases, the version can change the overall pool structure of different file systems.

[0070] In some cases, the pool ID includes a unique ID of the pool. In some instances, the unique ID includes a universal unique identifier (UUID). In some instances, the unique ID includes a content ID. In some instances, the content ID includes a digital fingerprint system that can be used to identify and / or manage the copyright or ownership of the content. In some cases, the list of pool item descriptors includes the path of the object, the size of the object (e.g., the total size of the object), the range of the pool item within the object, the offset of the pool item in the pool, or any combination thereof. In some instances, the range of the pool item within the object includes one or more positions of the payload in the pool item within the object. In some instances, one or more positions include the start and / or end range of the payload in the pool item (e.g., row 1-row 6 in pool item 1, row 7-row 13 in pool item 2, ..., etc.). In some instances, the offset of the pool item includes the payload position of the first byte of each of the one or more pool items in the payload of the pool. For example, the offset of the first pool item is 0 bytes. If the range of the first pool item is 1000-2000, then its size is 1000 bytes. In such an instance, the offset of the next pool item will be 1000 bytes. In some cases, the pool item includes a data payload and / or a hash of the pool item. In some cases, the data payload includes a stored object or a portion of an object. In some cases, the hash of the pool item includes a hash value of the stored object or a portion of an object. In some cases, the end pool descriptor includes a list of object descriptors. In some cases, the list of object descriptors includes the path of the object and / or the hash of the object. In some instances, the path of the object includes a unique path. In some instances, the path of the object includes a hierarchical structure (e.g., a directory hierarchy). In some instances, the path of the object does not include a hierarchical structure.

[0071] Systems and methods for storing digital information may include one or more hashes. In some cases, one or more hashes are determined using a hash module. In some cases, the hash module is executed on one or more processing units (such as those described herein). In some cases, the hash module includes instructions (e.g., hash functions) for determining one or more hashes. In some cases, instructions (e.g., hash functions) are stored on memory, such as those described herein. In some cases, information including an object, a part of an object, or a pool item is stored using a hash. In some cases, determine the first one or more hashes of the data payload of each of the one or more pool items and / or determine the second one or more hashes of each of the one or more objects 1215. In some cases, the data payload includes an object or a part of an object. In some cases, the hash of the pool item is attached to the data payload. In some cases, the hash of the object is attached to the end pool descriptor.

[0072] The hash can be determined using a hash function ( Fig.15 ). A hash function typically includes a function that converts an input of arbitrary length into an output with a fixed length (e.g., 224, 256, 384, 512 bits or characters). In some cases, the hash function includes a cryptographic hash function. In some cases, the hash function includes MD-5, SHA-1, SHA-2, SHA-3, RIPEMD-160, Whirlpool, BLAKE, BLAKE2, BLAKE3, or a variation thereof. In some cases, the hash function includes SHA-2. In some instances, SHA-2 includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256. The output of the hash function can be deterministic and is not feasible for reverse engineering. In addition, generating a fixed-length output can increase security because any party involved in decrypting the hash cannot tell the length of the input. In some instances, a hash is generated when an identification code, an encryption key, a password, or any variation thereof is input. In some instances, hashing allows for verification of content (eg, an information item or digital information stored in a pool) during decoding.

[0073] In some cases, input 1505 includes an object. In some instances, hash function 1510 is used to determine hash output (or hash) 1515. In some cases, input 1520 includes an object. In some instances, hash function 1525 is used to determine hash output (or hash) 1530. In some instances, hash function 1510 and hash function 1525 are the same hash function. In some instances, hash function 1510 and hash function 1525 are both SHA-256. In some instances, hash function 1510 and hash function 1525 are different hash functions. In some instances, output 1515 and output 1530 have the same length. In some instances, output 1515 and output 1530 are both 256 bits. In some instances, output 1515 and output 1530 have different lengths.

[0074] A hash function may include one or more operations that generate a hash. In some cases, one or more steps in a hash function include padding bits. In some cases, an extra bit is added to the digital information (or message) being hashed. In some instances, an extra bit is added to the message so that the length of the digital message is a modulus value less than the total number of bits. In some instances, the modulus value is 64 bits. In some instances, the number of bits is 512 bits, and the length of the digital information is 448 bits (e.g., for SHA-256). In some instances, the first extra bit includes a binary number 1. In some instances, the extra bits added subsequently include a binary number 0.

[0075] In some cases, one or more steps in the hash function include padding length. In some cases, padding length includes adding a modulus value to digital information (e.g., also referred to as a big endian (BE) integer). The modulus value or BE integer generally represents the length of the original input including binary original digital information. In some instances, the modulus value is 64 bits. In some instances, 64 bits are added to a 448-bit digital message, and the total number of bits is 512 bits (e.g., for SHA-256). In some cases, the modulus value is calculated by applying the modulus to the original digital information. As an example, if the original digital information is binary "hello world", the length of the original input is 88 bits, i.e. binary "1011000". In this way, the 0 after "1011000" is added to the end of the 448-bit digital information, so that the total number of bits is 512.

[0076] In some cases, one or more steps in the hash function include initializing one or more hash values ​​or buffers. In some cases, 8 hash values ​​or buffers are initialized. In some cases, the initialized hash value is hard-coded (e.g., a constant). In some cases, the initialized hash value represents the first 32 bits of the fractional part of the square root of the first 8 prime numbers (e.g., 2, 3, 5, 7, 11, 13, 17, 19). In some cases, one or more steps in the hash function also include initializing round constants (or keys). In some cases, 64 round constants are initialized. In some instances, each of the 64 round constants represents the first 32 bits of the fractional part of the cube root of the first 64 prime numbers (e.g., 2-311). In some cases, 64 different round constants are stored in an array.

[0077] In some cases, one or more steps in the hash function include compression. In some cases, each information block (e.g., every 512 bits) undergoes compression. During compression, each information block undergoes a fixed number of rounds. In some cases, the number of rounds is 64. In some cases, compression is performed by a one-way compression function. In some cases, the one-way compression function is a single-block length compression function. In some instances, the compression function is a Davies-Meyer, Matyas-Meyer-Oseas or Miyaguchi-Preneel compression function. In some cases, the one-way compression function is a dual-block length compression function. In some instances, the compression function is an MDC-2 / Meyer-Schilling, MDC-4 or Hirose compression function. In some cases, the output from the compression function is less than the information block. In some instances, the output has a length of 256 bits.

[0078] In some cases, one or more of the hashes (e.g., a hash of a pool item, a hash of an object) are calculated during storage of information. In some cases, all hashes (e.g., a hash of a pool item, a hash of an object) are calculated during storage of information. In some instances, this allows for stable low memory usage regardless of the size of the object. In some cases, the first one or more hashes of the data payload of each pool item require less memory than one or more objects. In some cases, the second one or more hashes of each of the one or more objects require less memory than one or more pool items. In some cases, source data (e.g., an information item) is read only once. In some cases, each pool is written once without seeking. In some instances, this minimizes data transfer and latency.

[0079] In some cases, the hashes described herein can serve one or more purposes. As non-limiting examples, the one or more purposes can include one or more of: verifying the integrity of one or more information items (e.g., objects), signature generation and verification (e.g., for digital signatures), cryptographic authentication, proof of work, or an identifier for an information item.

[0080] In some cases, encryption and / or compression can be further added. In some instances, encryption and / or compression is implemented using a streaming application programmable interface (API). In some instances, this avoids the need to store intermediate results. In some cases, the digital information to be stored has been compressed, for example, to reduce data transmission costs. In some cases, for example, for security reasons, the digital information to be stored has been encrypted.

[0081] One or more operations in the instructions stored in the memory may further include creating more than one index pool. In some cases, more than one index pool contains only indexes. In some cases, when retrieving objects stored in more than one pool encoded by more than one polynucleotide, the index pool is used. In some cases, the index pool is sorted and temporarily stored in a digital storage system (e.g., a flash drive) to search for objects. In some instances, after the pool is identified, more than one polynucleotide encoding the pool is sequenced.

[0082] In some cases, one or more index pools include a list of index pool descriptors and / or object indexes. In some cases, the index pool descriptor includes a version, a pool ID, the size of the pool, a timestamp, or a combination thereof. In some instances, the pool ID includes a unique ID for the pool. In some instances, the unique ID includes a universal unique identifier (UUID). In some instances, the unique ID includes a content ID. In some instances, the content ID includes a digital fingerprint system that can be used to identify and / or manage the copyright or ownership of the content. In some instances, the size of each index pool in more than one index pool is about 1GB to about 1TB. In some cases, the list of object indexes includes the path of the object, the hash of the object, the list of object fragments, the list of object metadata, or any combination thereof. In some instances, the path of the object includes a unique path. In some instances, the path of the object includes a hierarchical structure (e.g., a directory hierarchy). In some instances, the path of the object does not include a hierarchical structure. In some instances, the hash of the object is a hash as previously described herein (e.g., SHA-256). In some instances, the list of object fragments includes the pool ID of the pool containing the fragments, the range of the fragments, or a combination thereof. In some instances, the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof. In some instances, the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof. In some instances, the metadata is customizable. In some instances, the metadata is used to search for objects in more than one pool.

[0083] In some cases, the index pool can store information from about 1 to about 1 million pools. In some cases, the index pool can store information from about 1 pool to about 10 pools, about 1 pool to about 100 pools, about 1 pool to about 1,000 pools, about 1 pool to about 5,000 pools, about 1 pool to about 10,000 pools, about 1 pool to about 50,000 pools, about 1 pool to about 100,000 pools, about 1 pool to about 500,000 pools, about 1 pool to about 1 million pools, about 10 pools to about 100 pools, about 10 pools to about 1,000 pools, about 10 pools to about 5,000 pools, about 10 pools to about 10,000 pools pools, about 10 pools to about 50,000 pools, about 10 pools to about 100,000 pools, about 10 pools to about 500,000 pools, about 10 pools to about 1 million pools, about 100 pools to about 1,000 pools, about 100 pools to about 5,000 pools, about 100 pools to about 10,000 pools, about 100 pools to about 50,000 pools, about 100 pools to about 100,000 pools, about 100 pools to about 500,000 pools, about 100 pools to about 100,000 pools, about 100 pools to about 500,000 pools, about 100 pools to about 1 million pools, about 1,000 pools to about 5,000 pools pools, about 1,000 pools to about 10,000 pools, about 1,000 pools to about 50,000 pools, about 1,000 pools to about 100,000 pools, about 1,000 pools to about 500,000 pools, about 1,000 pools to about 1 million pools, about 5,000 pools to about 10,000 pools, about 5,000 pools to about 50,000 pools, about 5,000 pools to about 100,000 pools, about 5,000 pools to about 500,000 pools, about 5,000 pools to about 1 million pools, about 10,000 pools to about 100,000 pools, about 5,000 pools to about 500,000 pools, about 5,000 pools to about 1 million pools, about 10,000 pools to about 100,000 pools. The present invention relates to information about the number of pools from about 10,000 to about 50,000 pools, about 10,000 to about 100,000 pools, about 10,000 to about 500,000 pools, about 10,000 to about 1 million pools, about 50,000 to about 100,000 pools, about 50,000 to about 500,000 pools, about 50,000 to about 1 million pools, about 100,000 to about 500,000 pools, about 100,000 to about 1 million pools, or about 500,000 to about 1 million pools. In some cases, an index pool may store information of about 1 pool, about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, about 500,000 pools, about 100,000 pools, about 500,000 pools, or about 1 million pools.In some cases, the index pool can store information of at least about 1 pool, about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, or about 500,000 pools. In some cases, the index pool can store information of at most about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, about 500,000 pools, or about 1 million pools.

[0084] In some cases, each of one or more index pools is about 1GB to about 1TB. In some cases, each of more than one pool is about 1GB to about 1TB. In some cases, each of one or more index pools is about 1GB to about 10GB, about 1GB to about 50GB, about 1GB to about 100GB, about 1GB to about 500GB, about 1GB to about 1TB, about 10GB to about 50GB, about 10GB to about 100GB, about 10GB to about 500GB, about 10GB to about 1TB, about 50GB to about 100GB, about 50GB to about 500GB, about 50GB to about 1TB, about 100GB to about 500GB, about 100GB to about 1TB, about 100GB to about 500GB, about 100GB to about 1TB or about 500GB to about 1TB. In some cases, each of one or more index pools is about 1GB, about 10GB, about 50GB, about 100GB, about 500GB or about 1TB. In some cases, each of the one or more index pools is at least about 1 GB, about 10 GB, about 50 GB, about 100 GB, or about 500 GB. In some cases, each of the one or more index pools is at most about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB.

[0085] Coding scheme can be applied to each in more than one pool and / or index pool.In some cases, coding scheme is encoded as more than one polynucleotide 1220 by digital information in more than one pool.In some cases, coding scheme is encoded as more than one polynucleotide by digital information in index pool.In some cases, coding scheme includes the codec (for example, internal codec) for being used to encode binary data into polynucleotide sequence.In some cases, coding scheme includes error correction code (ECC).In some cases, coding scheme (for example, internal codec or low-level codec) is also designed and implemented to allow streaming read and write API access.In some cases, coding scheme (for example, internal codec or low-level codec) is also designed and implemented to match the streaming (streaming) of the system and method (for example, advanced codec) for digital storage described herein.

[0086] The encoding scheme may generally include one or more operations. The one or more operations may include one or more operations that manipulate or transform data (e.g., digital information). As non-limiting examples, the one or more operations may include splitting, shuffling, concatenating, transposing, translating, copying, marking (e.g., using an index) data or a portion of data, or any combination thereof.

[0087] A method for encoding digital information (e.g., binary data) in more than one polynucleotide sequence is schematically illustrated in Figure 1In some cases, the method for encoding numbers or data in more than one polynucleotide sequence includes splitting the data. In some cases, the data is split into more than one frame 105. In some cases, the more than one frame includes about 100 to about 10,000 frames. In some cases, the more than one frame includes about 100 frames to about 250 frames, about 100 frames to about 500 frames, about 100 frames to about 750 frames, about 100 frames to about 1,000 frames, about 100 frames to about 2,500 frames, about 100 frames to about 5,000 frames, about 100 frames to about 7,500 frames, about 100 frames to about 10,000 frames, about 250 frames to about 500 frames, about 100 frames to about 750 frames, about 100 frames to about 1,000 frames, about 100 frames to about 2,500 frames, about 100 frames to about 5,000 frames, about 100 frames to about 7,500 frames, about 100 frames to about 10,000 frames, about 250 frames to about 500 frames, about 250 frames to about 750 frames, about 250 frames to about 1,000 frames, about 250 frames to about 2,500 frames, about 250 frames to about 5,000 frames, about 250 frames to about 7,500 frames, about 250 frames to about 10,000 frames, about 500 frames to about 750 frames, about 500 frames to about 1,000 frames, about 500 frames to about 2,500 frames, about 500 frames to about 5,000 frames. 00 frames, about 500 frames to about 7,500 frames, about 500 frames to about 10,000 frames, about 750 frames to about 1,000 frames, about 750 frames to about 2,500 frames, about 750 frames to about 5,000 frames, about 750 frames to about 7,500 frames, about 750 frames to about 10,000 frames, about 1,000 frames to about 2,500 frames, about 1,000 frames to about 5,000 frames In some cases, more than one frame includes about 100 frames, about 250 frames, about 500 frames, about 750 frames, about 1,000 frames to about 10,000 frames, about 2,500 frames to about 5,000 frames, about 2,500 frames to about 7,500 frames, about 2,500 frames to about 10,000 frames, about 5,000 frames to about 7,500 frames, about 5,000 frames to about 10,000 frames, or about 7,500 frames to about 10,000 frames. In some cases, more than one frame includes about 100 frames, about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, about 7,500 frames, or about 10,000 frames. In some cases, more than one frame includes at least about 100 frames, about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, or about 7,500 frames. In some cases, more than one frame includes at most about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, about 7,500 frames, or about 10,000 frames. In some cases, each frame includes the same amount of data. In optional cases, each frame can include different amounts of data.In some cases, each frame is assigned a frame index. In some instances, the frame index increases for each frame index (e.g., 0, 1, 2, 3, 4, 5, ..., etc.). In some instances, the frame index increases monotonically for each frame index.

[0088] The method for encoding digital or binary data includes an external codec. In some cases, the method for encoding digital or binary data in more than one polynucleotide sequence includes an external codec. In some cases, the external codec is applied to data (for example, binary data). In some cases, after data are split into more than one frame, the external codec is applied to data 110. In this case, the external codec is applied to each in more than one frame. Figure 3 An example diagram of splitting a data stream into frames and applying an external codec is exemplarily illustrated in FIG.

[0089] In some cases, the external codec includes an error correction scheme or an error correction code (ECC), such as Reed-Solomon (RS) code. The external codec is used to spread digital or binary data to be stored on many oligonucleotides. In some cases, the spread data sets up redundancy, which can be used to correct erasure (for example, lost oligonucleotides). In some other cases, the spread data also sets up redundancy to correct the error from the internal codec.

[0090] In some cases, the error correction scheme includes a Reed-Solomon (RS) code. In this case, an RS encoder is used to encode binary data or more than one frame including binary data. Typically, RS codes operate on blocks of data viewed as a set of finite field elements. In some cases, the RS code includes mapping data, for example, x = (x1, ..., x k )∈F k , to the polynomial p x ,in The coded data C(x) is obtained by encoding the data in the field F at various n points a1,…,a n Evaluate p x Obtained (for example, C(x) = p x (a1),…,p x (a n )).

[0091] In some other embodiments, the RS code includes a coding scheme in which each codeword contains a message as a prefix and an error correction symbol is appended as a suffix. In some cases, the RS code is specified as RS(n,k) with m-bit symbols. In this case, the encoder takes k data symbols (each symbol is m-bit) and adds parity symbols (error correction symbols or check symbols) to form an n-symbol codeword. Here, there are nk parity symbols (or check symbols, t), each symbol is m bits. In some cases, the RS decoder corrects up to t symbols containing errors in the codeword, where 2t=nk. The codeword C(x) includes parity information CK(x), which is systematically appended to the message information M(x). The codeword C(x) can be calculated as: C(x)=x n-k M(x)+CK(x)=x n-k M(x)+x n-k M(x) mod g(x). Here, k refers to the message length (e.g., the number of symbols), t refers to the number of errors that need to be corrected, n refers to the block length (e.g., the message length n plus the correction length t), and m refers to the symbol width, where given the symbol size m, the maximum codeword length n for the RS code is n=2 m –1. In addition, x n-k refers to the displacement shift in the message, and g(x) refers to the generating polynomial, which is defined as a polynomial whose roots are sequential powers of the Galois Field (GF) primitive α (e.g., g(x) = (x-α i )(x-α i+1 )…(x-α i+n-k-1 )=g0+g1x+…+g n-k-1 x n-1-1 +x n-k ).

[0092] For example, in RS(255,223) with 8-bit symbols, the block length n is 255 codeword bytes, the message length k is 223 bytes, and the parity 2t is 32 bytes. In such an example, the RS decoder corrects up to 16 symbol errors in the codeword, which means that the decoder can correct up to 16 byte errors. The RS code can also be represented as GF(2 m ). For example, in RS GF(2 12 ) encoding scheme, such as Figure 3 As shown in FIG. 1 , n is 4095 (eg, n=2 12 –1=4096–1=4095). If k is, for example, 2499, then 2t=4095-2499=1596, and t is therefore 798.

[0093] In some cases, the error correction scheme includes a linear error correction code (or linear block code), such as a low density parity check (LDPC) code. In some cases, the error correction scheme includes a linear block error correction code, such as a polar code. In some other embodiments, the error correction scheme includes a high performance forward error correction (FEC), such as a turbo code. In some cases, the error correction scheme includes an RS code, an LDPC code, a turbo code, a polar code, or any combination thereof (e.g., an RS-based LDPC code).

[0094] In some cases, the error correction scheme includes a low-density parity-check (LDPC) code. In this case, the LDPC code is used to encode binary data or more than one frame including binary data. Typically, the structure of an LDPC code is defined by a parity check matrix that contains 0s at most entries and 1s elsewhere. For example, an (N,K) LDPC code for K information bits is a linear block code of block size N, defined by a sparse (NK)×N parity check matrix in which all elements except 1 are 0. The number of 1s in a row or column is called the degree of the row or column. In some cases, a codeword of length N is represented as a vector C, and for information bits of length K, an (N,K) code with 2K codewords is used. In some cases, a (N,K) LDPC code is defined by a (NK)×N parity check matrix H, satisfying the condition: HC T =0.

[0095] In some cases, when each row and each column of the parity check matrix has a constant degree, the LDPC code is regular, otherwise it is irregular. In some cases, irregular LDPC codes outperform regular LDPC codes. In some cases, due to the different degrees between rows and columns, irregular LDPC codes promise improved performance only when the row and column degrees are properly adjusted.

[0096] In some cases, the error correction scheme includes polar codes. In some cases, polar codes can achieve Shannon capacity through theoretical proof. In some cases, polar codes include low encoding and decoding complexity. Polar codes generally include a generator matrix G N , and the information can be obtained based on x1 N =u1 N G N Encoding, where x1 N is the encoded output bit, u1 N is the input bits before encoding, and the generator matrix is ​​defined as The code length N is defined as N = 2n, where n ≥ 0. N This includes transposed matrices such as, for example, bit-reversed matrices. Includes the Kronecker power of F, which is defined as in

[0097] In some cases, the polar code is represented by the cosecant code (N, K, A, u A c ), and the encoding process is defined as Where A is the information bit index set, G N (A) is the submatrix obtained from the row, which is in G N corresponds to the index in set A. In addition, G N (A c ) is the submatrix obtained from the row, which in G N Corresponding to the set A c The index in , and u A c are frozen bits, the number of which is (NK), where N is the code length and K is the length of the information bits. In some cases, the frozen bits are set to 0, and the above encoding process is described as x1 N =u A G N (A).

[0098] In some cases, the error correction scheme includes a turbo code. A turbo code typically includes a parallel concatenation of two or more component codes applied to different interleaved versions of the same information sequence. Typically, a recursive systematic convolutional (RSC) code is used as a component code. For example, the structure of a turbo code includes two RSC encoders (e.g., M=2) connected in parallel, and the code rate R is R=1 / 3, because R=1 / (M+1) (approximately). The input to the first RSC encoder is the original information sequence. The original information sequence d is also applied to an interleaver to produce an interleaved version d'. The interleaved version d' of the information sequence is the input to the second RSC encoder. The output from the turbo encoder includes u and a redundant portion x (1) (output from the first RSC encoder) and x (2) (output from the second encoder) is a systematic sequence. Therefore, the output of the encoder includes u1, x 1(1) 、x 1(2) 、u2、x 2(1) 、x 2(2) , where u k is the kth systematic bit (i.e., data bit), x k(1) is from the kth system bit u k The parity check output of the associated first RSC encoder; and x k(2) is from the kth system bit u kThe parity check output of the associated second RSC encoder. The decoding process of the turbo code generally includes iterative decoding. The turbo code decoding process can include two component decoders (corresponding to the two RSC encoders), an interleaver; and a deinterleaver. In some cases, the two component decoders are soft input and soft output (SISO) decoders. In some cases, the output of the two component decoders includes likelihood information about the encoded data sequence.

[0099] In some cases, once the external codec is applied, the size of the data increases. In some cases, once the external codec is applied to each of the frames including the data, the size of the frame increases. In some cases, the frame is divided into more than one channel 115. In some cases, each channel includes a channel index. In some cases, each frame includes about 1000 to about 10,000 channels. In some cases, each frame includes about 5000 channels. In some cases, each frame includes about 1,000 channels to about 2,500 channels, about 1,000 channels to about 5,000 channels, about 1,000 channels to about 7,500 channels, about 1,000 channels to about 10,000 channels, about 2,500 channels to about 5,000 channels, about 2,500 channels to about 7,500 channels, about 2,500 channels to about 10,000 channels, about 5,000 channels to about 7,500 channels, about 5,000 channels to about 10,000 channels, or about 7,500 channels to about 10,000 channels. In some cases, each frame includes about 1,000 channels, about 2,500 channels, about 5,000 channels, about 7,500 channels, or about 10,000 channels. In some cases, each frame includes at least about 1,000 channels, about 2,500 channels, about 5,000 channels or about 7,500 channels. In some cases, each frame includes at most about 2,500 channels, about 5,000 channels, about 7,500 channels or about 10,000 channels. Each channel can also include about 100 to about 300. In some cases, each channel includes about 100 to about 150, about 100 to about 200, about 100 to about 250, about 100 to about 300, about 150 to about 200, about 150 to about 250, about 150 to about 300, about 200 to about 250, about 200 to about 300 or about 250 to about 300. In some cases, each channel includes about 100 bits, 110 bits, 120 bits, 130 bits, 140 bits, 150 bits, 160 bits, 170 bits, 180 bits, 190 bits, 200 bits, 210 bits, 220 bits, 230 bits, 240 bits, 250 bits, 260 bits, 270 bits, 280 bits, 290 bits, 300 bits. In some cases, each channel includes at least about 100 bits, 110 bits, 120 bits, 130 bits, 140 bits, 150 bits, 160 bits, 170 bits, 180 bits, 190 bits, 200 bits, 210 bits, 220 bits, 230 bits, 240 bits, 250 bits, 260 bits, 270 bits, 280 bits, 290 bits, 300 bits.In some cases, each channel includes up to about 100 bits, 110 bits, 120 bits, 130 bits, 140 bits, 150 bits, 160 bits, 170 bits, 180 bits, 190 bits, 200 bits, 210 bits, 220 bits, 230 bits, 240 bits, 250 bits, 260 bits, 270 bits, 280 bits, 290 bits, 300 bits. Although the methods for encoding provided herein are illustrated using binary data, in some cases, the methods can be generally applied to data including more than one symbol.

[0100] In some cases, the method for encoding data in more than one polynucleotide sequence comprises shuffling data.In some cases, each channel is shuffled 120 at least in part based on channel index.In some cases, each channel is shuffled after external codec is applied to binary data.In some cases, shuffling each channel allows resistance to errors that may occur during synthesis or sequencing, such as those errors that affect whole oligonucleotide pool.Error can include insertion, deletion, replacement or its combination.In some cases, shuffling comprises the rotation scheme in each channel based in part on each channel index.For example, each position in the channel can be shifted according to each channel index (for example, there is no shuffling in channel 0, shifting 1 position in channel 1, shifting 2 positions in channel 2, etc.).

[0101] In some other cases, shuffling includes a pseudo-random process within each channel. In this pseudo-random shuffling process, a random seed is used to initialize a pseudo-random number generator. In some cases, the number generated by the pseudo-random number generator is determined by the random seed. Therefore, using the same seed, the pseudo-random number generator generates the same sequence of numbers. As an example, using shuffling includes a pseudo-random process to shift each bit in the channel according to the number generated by the pseudo-random number generator.

[0102] In some other cases, the channel index is used as a seed to create a permutation of some or all bits for that channel. In some cases, the permutation of some or all bits is created by sampling from a random number generator. In some cases, the permutation is stored in a precompiled form. In some cases, the use of a pseudo-random generator allows for a smaller implementation source code.

[0103] In some cases, the frame index and channel index are added to the front. In some cases, after each channel is shuffled, the frame index and channel index are added to the front of each channel. Figure 4In some cases, the frame index includes about 12 bits to about 20 bits. In some cases, the frame index includes about 12 bits to about 14 bits, about 12 bits to about 16 bits, about 12 bits to about 18 bits, about 12 bits to about 20 bits, about 14 bits to about 16 bits, about 14 bits to about 18 bits, about 14 bits to about 20 bits, about 16 bits to about 18 bits, about 16 bits to about 20 bits, about 16 bits to about 18 bits, about 16 bits to about 20 bits, or about 18 bits to about 20 bits. In some cases, the frame index includes about 12 bits, about 14 bits, about 16 bits, about 18 bits, or about 20 bits. In some cases, the frame index includes at least about 12 bits, about 14 bits, about 16 bits, or about 18 bits. In some cases, the frame index includes at most about 14 bits, about 16 bits, about 18 bits, or about 20 bits. In some cases, the channel index includes about 12 bits to about 16 bits. In some cases, the channel index includes about 12 bits to about 14 bits, about 12 bits to about 16 bits, or about 14 bits to about 16 bits. In some cases, the channel index includes about 12 bits, about 14 bits, or about 16 bits. In some cases, the channel index includes at least about 12 bits or about 14 bits. In some cases, the channel index includes at most about 14 bits or about 16 bits. Figure 4 As shown in , in some cases, the channel index is 12 bits and the frame index is 20 bits. In some cases, the channel index is the symbol width m from the RS code.

[0104] In some cases, the method for encoding data in more than one polynucleotide sequence includes an internal codec. In some cases, the internal codec is applied to data (e.g., binary data). In some cases, the internal codec is applied to the data from the external codec. In some cases, the internal codec is applied to the channel of data. In some cases, after the channel is shuffled, the internal codec is applied to the channel of data.

[0105] In some cases, the internal codec includes a coding scheme. In some cases, the internal codec including the coding scheme is applied to each channel to encode data into a polynucleotide sequence 125. The internal codec is used for converting data (for example, digital data or binary data) into nucleotide bases. In some cases, the internal codec can correct errors, such as missing, replacing or inserting errors, or any combination thereof. In some other embodiments, the internal codec is used to verify oligonucleotide and abandon any suspicious oligonucleotide to avoid contaminating external decoding. The internal codec is also encoded to index (frame index and channel index), which can allow efficient clustering during decoding.

[0106] In some cases, the coding scheme increases redundancy across more than one polynucleotide sequence. In some cases, redundancy is about 5% to about 10%. In some cases, redundancy is about 5% to about 6%, about 5% to about 7%, about 5% to about 8%, about 5% to about 9%, about 5% to about 10%, about 6% to about 7%, about 6% to about 8%, about 6% to about 9%, about 6% to about 10%, about 7% to about 8%, about 7% to about 9%, about 7% to about 10%, about 8% to about 9%, about 8% to about 10% or about 9% to about 10%. In some cases, redundancy is about 5%, about 6%, about 7%, about 8%, about 9% or about 10%. In some cases, redundancy is at least about 5%, about 6%, about 7%, about 8%, about 9% or about 10%. In some cases, redundancy is at most about 6%, about 7%, about 8%, about 9% or about 10%. In some cases, this redundancy allows the oligonucleotide pool to be decoded despite the presence of errors in individual oligonucleotides, such as insertions, deletions, substitutions, or any combination thereof.

[0107] Figure 5 An example diagram of a coding scheme is shown in . In this example diagram, the coding scheme in the internal codec combines two or more of the following: bits from each channel, bit history, and bit position (bit position). In some cases, a model (e.g., an adaptive model) is used to partition known bits into contexts (context), and each context is mapped to a bit history. In some cases, the bit history is represented by an 8-bit state. In some cases, the bit history is updated each time a context is encountered, such as by the use of a lookup table. The bit position includes a fixed number of least significant bits (LSBs). In some cases, the LSB includes a bit index of the bit to be encoded. For example, if 100 bits encode a 100-mer oligonucleotide, then "bit index" refers to an index from 0 to 99 in the bit to be encoded. The LSB includes a bit position in a binary integer representing a binary 1 position of an integer. In some cases, the LSB index has any length. In some cases, the LSB index is represented by a 2-bit state, a 3-bit state, or a 4-bit state. As an example, indices 0, 1, 2, 3, 4, 5, 6, 7, ... can be represented by 2-bit states 00, 01, 10, 11, 00, 01, 10, 11, ..., respectively. Figure 5 The encoding scheme illustrated in includes binary data, but in some cases the encoding scheme may be generally applicable to data including more than one symbol.

[0108] In some cases, the internal codec includes a candidate base for generating a position of binary data. A lookup table, a hash or a combination thereof is used to generate a candidate base for binary data. In some cases, the method previously described herein is used to determine the hash. In some cases, the binary data includes two or more of the following: a position, a position history and a position position from each channel. In some cases, the bit rate for encoding is about 1 position per base to about 2 positions per base.In some cases, the bit rate for encoding is about 1 bit per base to about 1.1 bits per base, about 1 bit per base to about 1.2 bits per base, about 1 bit per base to about 1.3 bits per base, about 1 bit per base to about 1.4 bits per base, about 1 bit per base to about 1.5 bits per base, about 1 bit per base to about 1.6 bits per base, about 1 bit per base to about 1.7 bits per base, about 1 bit per base to about 1.8 bits per base, about 1 bit per base to about 1.9 bits per base, about 1 bit per base to about 2 bits per base, about 1.1 bits per base to about 1.2 bits per base, about 1.1 bits per base to about 1.3 bits per base, about 1.1 bits per base to about 1.4 bits per base, about 1.1 bits per base to about 1.6 bits per base, about 1 bit per base to about 1.7 bits per base, about 1 bit per base to about 1.8 bits per base, about 1 bit per base to about 1.9 bits per base, about 1 bit per base to about 2 bits per base, about 1.1 bits per base to about 1.2 bits per base, about 1.1 bits per base to about 1.3 bits per base, about 1.1 bits per base to about 1.4 bits per base, about 1.1 bits per base to about 1.6 bits per base .5, about 1.1 per base to about 1.6 per base, about 1.1 per base to about 1.7 per base, about 1.1 per base to about 1.8 per base, about 1.1 per base to about 1.9 per base, about 1.1 per base to about 2 per base, about 1.2 per base to about 1.3 per base, about 1.2 per base to about 1.4 per base, about 1.2 per base to about 1.5 per base, about 1.2 per base to about 1.6 per base, about 1.2 per base to about 1.7 per base, about 1.2 per base to about 1.8 per base, about 1.2 per base to about 1.9 per base, about 1.2 per base to about 2 per base, about 1.3 per base about 1.4 per base, about 1.3 per base to about 1.5 per base, about 1.3 per base to about 1.6 per base, about 1.3 per base to about 1.7 per base, about 1.3 per base to about 1.8 per base, about 1.3 per base to about 1.9 per base, about 1.3 per base to about 2 per base, about 1.4 per base to about 1.5 per base, about 1.4 per base to about 1.6 per base, about 1.4 per base to about 1.7 per base, about 1.4 per base to about 1.8 per base, about 1.4 per base to about 1.9 per base, about 1.4 per base to about 2 per base, about 1.5 per base to about 1.6 per base, From about 1.5 per base to about 1.7 per base, from about 1.5 per base to about 1.8 per base, from about 1.5 per base to about 1.9 per base, from about 1.5 per base to about 2 per base, from about 1.6 per base to about 1.7 per base, from about 1.6 per base to about 1.8 per base, from about 1.6 per base to about 1.9 per base, from about 1.6 per base to about 2 per base, from about 1.7 per base to about 1.8 per base, from about 1.7 per base to about 1.9 per base, from about 1.7 per base to about 2 per base, from about 1.8 per base to about 1.9 per base, from about 1.8 per base to about 2 per base, or from about 1.9 per base to about 2 per base.In some cases, the bit rate used for encoding is about 1 bit per base, about 1.1 bits per base, about 1.2 bits per base, about 1.3 bits per base, about 1.4 bits per base, about 1.5 bits per base, about 1.6 bits per base, about 1.7 bits per base, about 1.8 bits per base, about 1.9 bits per base, or about 2 bits per base. In some cases, the bit rate used for encoding is at least about 1 bit per base, about 1.1 bits per base, about 1.2 bits per base, about 1.3 bits per base, about 1.4 bits per base, about 1.5 bits per base, about 1.6 bits per base, about 1.7 bits per base, about 1.8 bits per base, or about 1.9 bits per base. In some cases, the bit rate for encoding is at most about 1.1 bits per base, about 1.2 bits per base, about 1.3 bits per base, about 1.4 bits per base, about 1.5 bits per base, about 1.6 bits per base, about 1.7 bits per base, about 1.8 bits per base, about 1.9 bits per base, or about 2 bits per base. In some cases, a lookup table is used to map bits to nucleotides (e.g., A=00, T=10, C=01, G=11). In some cases, hashing includes a function that can be used to map data of any size (e.g., any number of bits) to a fixed size value (e.g., nucleotides or hash values). In some instances, hash values ​​are mapped to polynucleotide sequences.

[0109] In some cases, the internal codec includes a base duplication check. In some cases, after a candidate base is selected, a base duplication check is performed. In some cases, the base duplication check checks the duplication in two or more consecutive bases. In some cases, if there is a duplication in two or more consecutive bases, the base duplication check replaces another base with one base. In some cases, a lookup table or hash is updated based on the base updated during the base duplication check. In addition, after the base duplication check, the bit history is updated. In some cases, the frame index and / or channel index are increased. In some cases, the process is repeated until the sequence of all more than one polynucleotide sequence is determined.

[0110] In some cases, the internal codec also includes GC filtering before synthesizing more than one polynucleotide sequence. In some cases, GC filtering removes about 1% to about 10% of the channels in more than one channel. In some cases, GC filtering removes about 5% to about 10% of the channels in more than one channel. In some cases, GC filtering does not remove channels in more than one channel. In some cases, GC filtering removes about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9% or about 10%. In some cases, GC filtering removes at least about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8% or about 9%. In some cases, GC filtering removes at most about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9% or about 10%. In some cases, more than one polynucleotide sequence comprises a GC content of about 40% to about 60%. In some cases, more than one polynucleotide sequence comprises about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 50% to about 55%, about 50% to about 60% or about 55% to about 60% GC content. In some cases, more than one polynucleotide sequence comprises about 40%, about 45%, about 50%, about 55% or about 60% GC content. In some cases, more than one polynucleotide sequence comprises at least about 40%, about 45%, about 50% or about 55% GC content. In some cases, more than one polynucleotide sequence comprises up to about 45%, about 50%, about 55% or about 60% GC content. In some cases, at least 90% of more than one polynucleotide sequence comprises about 40% to about 60% GC content. In some cases, at least 90% of more than one polynucleotide sequence comprises about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 50% to about 55%, about 50% to about 60% or about 55% to about 60% GC content. In some cases, at least 90% of more than one polynucleotide sequence comprises about 40%, about 45%, about 50%, about 55% or about 60% GC content. In some cases, at least 90% of more than one polynucleotide sequence comprises at least about 40%, about 45%, about 50% or about 55% GC content. In some cases, at least 90% of more than one polynucleotide sequence comprises at most about 45%, about 50%, about 55% or about 60% GC content. In some cases, the output from the internal codec includes a final oligonucleotide pool.

[0111] Figure 6An example diagram of an optional coding scheme is shown in . In some cases, the coding scheme in the internal codec includes starting with a default lookup table. The default lookup table is used to select the word to be encoded in each channel. In some cases, the word includes more than one symbol. In some instances, the word is an 8-bit word or a byte. The application lookup table generates candidate bases for each word or byte in each channel. The next lookup table is selected based on the word or byte of the previous encoding. In some cases, the coding scheme also includes performing base duplication checks, GC filtering or its combination, as previously described herein. In some cases, the process is repeated until the sequence of all more than one polynucleotide sequence can be determined. In some cases, the output from the internal codec includes a final oligonucleotide pool or a final oligonucleotide library.

[0112] In some cases, the length of each oligonucleotide (or polynucleotide) in the library is about 20 to about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is about 20 bases to about 50 bases, about 20 bases to about 100 bases, about 20 bases to about 200 bases, about 20 bases to about 300 bases, about 20 bases to about 400 bases, about 20 bases to about 500 bases, about 50 bases to about 100 bases, about 50 bases to about 200 bases, about 50 bases to about 300 bases, about 50 bases to about 400 bases. , about 50 bases to about 500 bases, about 100 bases to about 200 bases, about 100 bases to about 300 bases, about 100 bases to about 400 bases, about 100 bases to about 500 bases, about 200 bases to about 300 bases, about 200 bases to about 400 bases, about 200 bases to about 500 bases, about 300 bases to about 400 bases, about 300 bases to about 500 bases, or about 400 bases to about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is at least about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, or about 400 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is at most about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases.

[0113] In some cases, the method for encoding data in more than one polynucleotide sequence as described herein is performed on the system. In some cases, such a system includes a device comprising a memory, a processing device operably coupled to the memory, or a combination thereof. In some cases, the memory is used to store information of binary data, polynucleotide sequences, or a combination thereof. In some cases, the information of data (e.g., binary data), polynucleotide sequences, or a combination thereof comes from one or more steps in the coding method described herein. In some cases, the memory is used to store information (e.g., software code, parameters, executable instructions, etc.) related to the algorithm described herein. In some instances, the memory may include any suitable memory described herein. In some instances, the memory may be configured according to the embodiments described herein.

[0114] In some cases, the processing device is configured to perform one or more encoding steps. In some cases, the processing device is configured to perform one or more operations including the following items: splitting the data into more than one frame; applying an external codec to each frame in more than one frame; dividing each frame into more than one channel; shuffling each channel based at least in part on a channel index; and applying an internal codec comprising a coding scheme to encode each channel in a polynucleotide sequence. In some cases, each frame in more than one frame includes a frame index. In some cases, each channel in more than one channel includes a channel index. In some cases, the external codec includes an error correction scheme. In some cases, the coding scheme increases redundancy so that binary data can be decoded in the presence of errors in the polynucleotide sequence.

[0115] Methods, systems and platforms for encoding data may include internal codecs optimized for one or more constraints. As non-limiting examples, one or more constraints may be related to nucleic acid synthesis, post-processing, storage or sequencing. In some cases, nucleic acid synthesis includes electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing or any combination thereof. In some cases, one or more constraints related to nucleic acid synthesis include synthesis errors, such as insertions, deletions or mutations. In some cases, post-processing includes connection, cracking, hybridization, denaturation, fixation to solid supports, extension, error correction, enrichment, separation, purification and amplification One or more. In some cases, storage includes cold data storage. Cold data storage can generally refer to the storage of rarely accessed data (such as data in nucleic acids). Cold data storage can be the antonym of "hot storage", and "hot storage" refers to the storage of frequently accessed data. In some instances, storage includes hot storage, wherein the data stored in nucleic acids are frequently accessed. In some cases, storage includes nucleic acid storage in liquid or solid phases. In some instances, one or more constraints associated with storage include temperature (e.g., room temperature), humidity, pressure, salinity, pH, concentration, time, light, UV, O2, or any combination thereof. In some cases, sequencing includes next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing, or any combination thereof.

[0116] Methods, systems and platforms for encoding data can include an internal codec optimized for the generation of polynucleotides. In some cases, the generation of polynucleotides includes the assembly of polynucleotides. In some cases, the generation of polynucleotides includes the synthesis of polynucleotides. Synthesis can include methods and systems described herein, or any suitable method and system known in the art. In some cases, data includes one or more symbols. In some cases, data includes a string of symbols or a sequence of symbols. In some cases, one or more symbols include binary data. In some cases, an internal codec is applied to data. In some cases, an internal codec is applied to data from an external codec (e.g., an error correction scheme), such as those provided herein. In some cases, an internal codec is applied to unencrypted data. In some cases, an internal codec is applied to encrypted data. The internal codec can be optimized to produce polynucleotides that follow a specific base sequence. In some cases, this allows more efficient polynucleotide synthesis, because the total number of synthesis cycles is reduced compared to the number of synthesis cycles required for synthesizing polynucleotides whose sequences are not encoded using the internal codec provided herein (e.g., an unoptimized synthesis method). In some cases, this allows for lower error rates because the number of oxidation and deprotection steps during the synthesis is reduced.

[0117] A method for encoding data is provided herein. In some cases, the method includes generating an internal codec including a secret book. The secret book can be optimized based on the application, manipulation, operation or use of the nucleic acid encoding the data. As described herein, the secret book can be optimized based on one or more constraints (e.g., related to nucleic acid synthesis, post-processing, storage, sequencing, etc.). The secret book can be generated in base order. In some cases, the secret book includes code words generated in part based on the base order. In some cases, the base order includes a predetermined base conversion. In some cases, the secret book generates a polynucleotide sequence by mapping data represented by one or more symbols (e.g., binary "0" and "1") to another or more symbols (such as nucleic acids (e.g., A, T, C, G)) using code words. In some cases, a specific or predetermined base conversion allows synthesis according to the base order. In some instances, pattern repetition is reduced by changing the synthesis order at each layer. Non-limiting examples of synthesis order at a given layer may include [A, G, C, T], [C, A, T, G], [T, G, A, C], or any other combination of bases A, T, G, C. In such examples, the secret book is changed for each layer. In some examples, two consecutive layers do not have the same secret book. In some examples, each layer includes a unique secret book. In some examples, two or more layers include the same secret book.

[0118] In some instances, the pattern is repeated by only allowing the specific base conversion at each base. For example, after adenine (A), only guanine (G), cytosine (C) or thymine (T) can be selected as the next base in the sequence. Alternatively, after A, base is not selected. In some instances, if G is selected, only C or T can be selected, or base is not selected alternatively. In some instances, if C is selected, only T can be selected, or base is not selected alternatively.

[0119] In some cases, the dense book comprises one, two, three, four, or five nucleotides. In some cases, the dense book comprises at least one, two, three, or four nucleotides. In some cases, the dense book comprises at most two, three, four, or five nucleotides. In some cases, the dense book comprises four nucleotides (e.g., adenine (A), thymine (T), cytosine (C), guanine (G)). For example, the specific base transitions of one or more layers include any of the following: (a) [A, T, C, G], (b) [A, T, G, C], (c) [A, G, T, C], (d) [A, G, C, T], (e) [A, C, G, T], (f) [A, C, T, G], (g) [T, C, G, A], (h) [T, C, A, G], (i) [T, G, A, C], (j) [T, G, C, A], (k) [T, A, G, C], (l) [ In some cases, the specific base conversions of one or more layers comprise natural or typical bases. In some cases, the specific base conversions of one or more layers comprise nucleotides with natural or typical bases and one or more nucleotides with non-natural or atypical bases. As an example, a secret book may include a synthesis order according to the repetition of [A, G, C, T] (e.g., A, G, C, T, A, G, C, T, ...). In such an example, the secret book may include the following codewords: A, G, C, T, AG, AC, AT, GC, GT, AGC, ACT, and AGCT. In some cases, the codewords in the secret book may be synthesized with a number of cycles equivalent to the number of nucleotides in the secret book. In some cases, the codewords in the secret book may be synthesized with 1, 2, 3, 4, or 5 synthesis cycles. In some cases, the codewords in the secret book may be synthesized with at least 1, 2, 3, 4, or 5 synthesis cycles. In some cases, the codewords in the secret book may be synthesized with at most 1, 2, 3, 4, or 5 synthesis cycles. In some cases, the transformations associated with the secret book are non-random or pseudo-non-random. In some cases, the transformations associated with the secret book are defined by a predefined mathematical algorithm or statistical algorithm.

[0120] In some cases, the synthesis order can be changed for one or more layers. A layer can generally include a stream of each base in a specific or predetermined order. For example, if the base transition is [A, T, C, G], during synthesis, the layer includes a stream of A, followed by a stream of T, a stream of C, and then a stream of G. In some cases, one or more layers can include any of the following: (a) [A, T, C, G], (b) [A, T, G, C], (c) [A, G, T, C], (d) [A, G, C, T], (e) [A, C, G, T], (f) [A, C, T, G], (g) [T, C, G, A], (h) [T, C, A, G], (i) [T, G, A, C], (j) [T, G, C, A], (k) [T, A, G, C], (l) [ In some cases, one or more specific base transitions of a layer can be repeated more than once. As an example, the synthesis order can include [A, G, C, T], [C, A, T, G], [T, G, A, C], or any combination thereof. … And the sequence can include AGCTAGCTCATGTGAC … , wherein the first layer is repeated twice. In some cases, changing one or more layers reduces pattern repetition in the sequence (e.g., repeated bases, high GC / AT, or secondary structure).

[0121] In some cases, the internal codec includes one or more secret books. In some cases, the internal codec includes one, two, three, four, five, six, seven, eight, nine or ten secret books. In some cases, the internal codec includes at least one, two, three, four, five, six, seven, eight, nine or ten secret books. In some cases, the internal codec includes at most one, two, three, four, five, six, seven, eight, nine or ten secret books. In some cases, each secret book encodes the layer during polynucleotide synthesis. In some cases, each secret book is generated with a unique base sequence. In some cases, each secret book is optimized for one or more base conversions. In some cases, the unique base sequence generates one or more unique base conversions. In some cases, each secret book is optimized for a specific base conversion at a given layer, loop index, history or any combination thereof. In some instances, history includes one or more of the previous layers, one or more secret books encoding the previous one or more layers, the loop index of one or more previous layers or any combination thereof. In some cases, each secret book is generated by a predefined mathematical algorithm or statistical algorithm.

[0122] In some cases, the secret version comprises one or more nucleotide analogs or non-natural / atypical nucleotides. Nucleotide analogs or non-natural nucleotides comprise nucleotides containing a certain type of modification. Nucleotide analogs or non-natural nucleotides comprise nucleotides containing a certain type of modification to bases, sugars or phosphate moieties. Modifications may include chemical modifications. Modifications may be, for example, modifications of 3'OH or 5'OH groups, backbones, sugar components or nucleotide bases. Modifications may include the addition of non-naturally occurring linker molecules and / or interchain or intrachain crosslinking. On the one hand, the modified nucleic acid comprises one or more of the modified 3'OH or 5'OH groups, backbones, sugar components or nucleotide bases, and / or the addition of non-naturally occurring linker molecules. On the one hand, the modified backbone comprises a backbone other than a phosphodiester backbone. On the one hand, the modified sugar comprises a sugar other than deoxyribose (in modified DNA) or a sugar other than ribose (in modified RNA). In one aspect, the modified base includes a base other than adenine, guanine, cytosine, or thymine (in a modified DNA) or a base other than adenine, guanine, cytosine, or uracil (in a modified RNA).

[0123] Nucleic acid can include at least one modified base.Modification of base moieties includes natural and synthetic modifications of A, C, G and T / U and different purine or pyrimidine bases.In some embodiments, modification is a modified form of adenine, guanine, cytosine or thymine (in modified DNA) or a modified form of adenine, guanine, cytosine or uracil (in modified RNA).Other examples of modified bases can be found in, for example, WO2019 / 014267 and US2022 / 0243244 (which are incorporated herein by reference in their entirety).

[0124] In some embodiments, the dense book comprises one or more typical nucleotides and one or more atypical nucleotides. In some cases, the typical nucleotide comprises one or more of A, T, C, G or U. In some cases, the atypical nucleotide comprises one or more nucleotide analogs or non-natural nucleotides provided herein. In some cases, the atypical nucleotide comprises one or more typical nucleotides with modification. In some cases, the dense book comprises about one, two, three, four or five typical nucleotides. In some cases, the dense book comprises about one, two, three, four or five atypical nucleotides. In some cases, the dense book comprises about at least one, two, three, four or five typical nucleotides. In some cases, the dense book comprises about at least about one, two, three, four or five atypical nucleotides. In some cases, the dense book comprises at most about one, two, three, four or five typical nucleotides. In some cases, the dense book comprises about at most about one, two, three, four or five atypical nucleotides. In some cases, the dense book comprises any combination of typical and atypical nucleotides, such as those provided herein.

[0125] In some cases, the secret book includes about 1 to about 30 codewords. In some cases, the density includes about 1 to about 5, about 1 to about 10, about 1 to about 12, about 1 to about 15, about 1 to about 18, about 1 to about 20, about 1 to about 22, about 1 to about 25, about 1 to about 28, about 1 to about 30, about 5 to about 10, about 5 to about 12, about 5 to about 15, about 5 to about 18, about 5 to about 20, about 5 to about 22, about 5 to about 25, about 5 to about 28, about 5 to about 30, about 10 to about 12, about 10 to about 15, about 10 to about 18, about 10 to about 20, about 10 to about 22, about 10 to about 25, about 10 to about 28, about 10 to about 30, about 12 to about 15, From about 12 to about 18, from about 12 to about 20, from about 12 to about 22, from about 12 to about 25, from about 12 to about 28, from about 12 to about 30, from about 15 to about 18, from about 15 to about 20, from about 15 to about 22, from about 15 to about 25, from about 15 to about 28, from about 15 to about 30, from about 18 to about 20, from about 18 to about In some cases, the cipher includes about 1, about 5, about 10, about 12, about 15, about 18, about 20, about 22, about 25, about 28, about 18 to about 30, about 20 to about 22, about 20 to about 25, about 20 to about 28, about 20 to about 30, about 22 to about 25, about 22 to about 28, about 22 to about 30, about 25 to about 28, about 25 to about 30, or about 28 to about 30 codewords. In some cases, the cipher includes about 1, about 5, about 10, about 12, about 15, about 18, about 20, about 22, about 25, about 28, or about 30 codewords. In some cases, the cipher includes at least about 1, about 5, about 10, about 12, about 15, about 18, about 20, about 22, about 25, or about 28 codewords. In some cases, the secret book includes at most about 5, about 10, about 12, about 15, about 18, about 20, about 22, about 25, about 28, or about 30 codewords.

[0126] The internal codec comprising the secret book can be applied to encode data into more than one polynucleotide sequence. In some cases, the data includes digital data. In some cases, the data includes one or more symbols. In some cases, one or more symbols are mapped to more than one polynucleotide sequence based on the secret book. For example, a numerical value such as a binary number (e.g., a sequence of 0 or 1) can be mapped to a codeword in the secret book. In some cases, the internal codec is further optimized for one or more constraints. One or more constraints may include constraints related to more than one polynucleotide sequence. In some instances, one or more constraints include the length of more than one polynucleotide sequence. In some instances, one or more constraints include the GC content of more than one polynucleotide sequence. In some instances, one or more constraints include base repetitions of more than one polynucleotide sequence. In some instances, one or more constraints include one or more errors, such as insertions, mutations or deletions. In some cases, binary data is mapped to codewords to create a conversion map. Conversion may include one or more conversions based on the value and position (e.g., index) of binary data between codewords and secret books. In some cases, one or more probabilities are calculated based on estimated deletions, insertions, and / or mutation rates during decoding. In some cases, the decoding algorithm finds one or more solutions to maximize the transition probability, as provided herein (e.g., Figure 8 and Fig. 9 ).

[0127] In some cases, a part of more than one polynucleotide sequence encodes redundancy. In some cases, the part of more than one polynucleotide sequence encoding redundancy is about 20% to about 80%. In some cases, the part of more than one polynucleotide sequence encoding redundancy is about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 60% to about 70%, about 60% to about 80% or about 70% to about 80%. In some cases, the portion of the coding redundancy in more than one polynucleotide sequence is about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%. In some cases, the portion of the coding redundancy in more than one polynucleotide sequence is at least about 20%, about 30%, about 40%, about 50%, about 60%, or about 70%. In some cases, the portion of the coding redundancy in more than one polynucleotide sequence is at most about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%.

[0128] In some cases, more than one polynucleotide sequence has the same length. In some cases, about 70% to about 100% of the more than one polynucleotide sequence has the same length. In some cases, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 95%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 95%, about 80% to about 100%, about 85% to about 90%, about 85% to about 95%, about 85% to about 100%, about 90% to about 95%, about 90% to about 100%, or about 95% to about 100% of more than one polynucleotide sequence are the same length. In some cases, about 70%, about 75%, about 80%, about 85%, about 90%, about 95% or about 100% in more than one polynucleotide sequence have the same length. In some cases, at least about 70%, about 75%, about 80%, about 85%, about 90% or about 95% in more than one polynucleotide sequence have the same length. In some cases, at most about 75%, about 80%, about 85%, about 90%, about 95% or about 100% in more than one polynucleotide sequence have the same length. In some cases, more than one polynucleotide sequence has different lengths. In some cases, more than one polynucleotide sequence differs by 1% to about 30%. In some cases, more than one polynucleotide sequence differs by about 1% to about 5%, about 1% to about 10%, about 1% to about 15%, about 1% to about 20%, about 1% to about 25%, about 1% to about 30%, about 5% to about 10%, about 5% to about 15%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 20% to about 25%, about 20% to about 30%, or about 25% to about 30%. In some cases, more than one polynucleotide sequence differs by about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%. In some cases, more than one polynucleotide sequence differs by at least about 1%, about 5%, about 10%, about 15%, about 20%, or about 25%. In some cases, more than one polynucleotide sequence differs by at most about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%.

[0129] More than one polynucleotide comprising more than one polynucleotide sequence can be produced. In some cases, more than one polynucleotide is synthesized. In some cases, synthesis includes base-by-base synthesis. In some cases, synthesis includes synthesis cycles. Synthesis cycles generally refer to one or more steps carried out to achieve nucleotide coupling. Synthesis cycles may include one or more of the following: deblocking (or deprotection), coupling, oxidation and capping. In some cases, synthesis includes multiple synthesis cycles. The internal codec can allow more efficient synthesis by reducing the number of required synthesis cycles. In some cases, compared with the number of synthesis cycles required for synthesizing more than one polynucleotide with a sequence not encoded by the internal codec, the number of synthesis cycles required for synthesizing more than one polynucleotide comprising more than one polynucleotide sequence encoded by the internal codec is reduced. In some cases, the number of synthesis cycles has been reduced by about 5% to about 80%. In some cases, the number of synthesis cycles is reduced by about 5% to about 10%, about 5% to about 20%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 5% to about 60%, about 5% to about 70%, about 5% to about 80%, about 10% to about 20%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 20% to about 30%, about 20% to about 40%, about 20% to about In some cases, the number of synthesis cycles is reduced by about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 150%, about 200%, about 400%, about 500%, about 600%, about 700%, about 800%, about 1000%, about 2000%, about 6000%, about 7000%, about 8000%, about 1500%, about 2 ... In some cases, the number of synthesis cycles is reduced by at least about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60% or about 70%. In some cases, the number of synthesis cycles is reduced by at most about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70% or about 80%. As an example, an internal codec with 12 code words using 4 nucleotides encodes more than one polynucleotide sequence with about 50% redundancy. Therefore, in this example, 6 values ​​of binary data are mapped to 12 code words, which is equivalent to log2 (12) = 3.58 bits of information. However, considering redundancy, such as 2x redundancy, corresponding to 3.58 / 2 = 1.79 bits of information for each code word.If the payload in each of more than one polynucleotide sequence is about 100 bits, this requires about 100 bits / 1.79 bits per codeword=55.8 codewords. Using optimized internal codec and circular sorting, as described herein, a codeword requires 4 cycles, producing about 55.8 codewords × 4 cycles per codeword=about 224 synthesis cycles. However, in the absence of internal codec, synthesis will require 400 cycles (for example, 4 × 100). As another example, if the payload in each of more than one polynucleotide sequence is about 300 bits, this requires about 447 synthesis cycles (for example, 200 / 1.79 × 4). However, in the absence of internal codec, synthesis will require 800 cycles (for example, 4 × 200). As another example, if the payload in each of more than one polynucleotide sequence is about 300 bits, this requires about 670 synthesis cycles (for example, 300 / 1.79 × 4). However, without the internal codec, the synthesis would require 1200 cycles (eg, 4×300).

[0130] In some cases, more than one polynucleotide has the same length.In some cases, about 70% to about 100% in more than one polynucleotide has the same length.In some cases, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 95%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 95%, about 80% to about 100%, about 85% to about 90%, about 85% to about 95%, about 85% to about 100%, about 90% to about 95%, about 90% to about 100% or about 95% to about 100% have the same length. In some cases, about 70%, about 75%, about 80%, about 85%, about 90%, about 95% or about 100% in more than one polynucleotide have the same length. In some cases, at least about 70%, about 75%, about 80%, about 85%, about 90% or about 95% in more than one polynucleotide have the same length. In some cases, at most about 75%, about 80%, about 85%, about 90%, about 95% or about 100% in more than one polynucleotide have the same length. In some cases, more than one polynucleotide has different lengths. In some cases, more than one polynucleotide differs by 1% to about 30%. In some cases, more than one polynucleotide differs by about 1% to about 5%, about 1% to about 10%, about 1% to about 15%, about 1% to about 20%, about 1% to about 25%, about 1% to about 30%, about 5% to about 10%, about 5% to about 15%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 20% to about 25%, about 20% to about 30%, or about 25% to about 30%. In some cases, more than one polynucleotide differs by about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%. In some cases, more than one polynucleotide differs by at least about 1%, about 5%, about 10%, about 15%, about 20% or about 25%. In some cases, more than one polynucleotide differs by at most about 5%, about 10%, about 15%, about 20%, about 25% or about 30%. In some cases, the efficiency of PCR is related to the amount of polynucleotides with the same length. In some cases, more than one polynucleotide with the same length ensures that PCR does not change the distribution of polynucleotides. In some cases, 90% or more of more than one polynucleotide have the same length to ensure that PCR does not change the distribution of polynucleotides.

[0131] In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 400. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 300. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 200. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 300. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is about 200. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 224. In some cases, for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 100. In some cases, for a polynucleotide sequence comprising about 200 bases, the number of synthesis cycles is less than 800. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is less than 600. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is less than 500. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is less than 400. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is less than 300. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is about 500. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is about 400. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is about 300. In some cases, for a polynucleotide sequence comprising 200 bases, the number of synthesis cycles is about 200. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is less than 1200. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is less than 1000. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is less than 800. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is less than 600. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is less than 400. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is about 600. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is about 500. In some cases, for a polynucleotide sequence comprising 300 bases, the number of synthesis cycles is about 450. In some cases, the polynucleotide sequence comprises four nucleotides. In some cases, the polynucleotide sequence comprises one or more of A, T, C, and G. In some cases, the polynucleotide sequence comprises one, two, three, four, or five nucleotides.In some cases, the polynucleotide sequence comprises at least one, two, three, four or five nucleotides. In some cases, the polynucleotide sequence comprises at most one, two, three, four or five nucleotides. In some cases, about 10%, 20%, 25%, 30%, 33%, 40%, 50%, 60%, 66%, 70%, 75%, 80% or 90% of the polynucleotide sequence encoding redundancy. In some cases, up to about 10%, 20%, 25%, 30%, 33%, 40%, 50%, 60%, 66%, 70%, 75%, 80% or 90% of the polynucleotide sequence encoding redundancy. In some cases, at most about 10%, 20%, 25%, 30%, 33%, 40%, 50%, 60%, 66%, 70%, 75%, 80% or 90% of the polynucleotide sequence encoding redundancy. In some examples, the polynucleotide sequence comprises about 1.5x, 2x, 2.5x, 3x, 3.5x, or 4x redundancy.

[0132] In some cases, more than one polynucleotide is synthesized on a solid support, such as those provided herein. A solid support can be a substrate as provided herein. In some cases, a solid support comprises more than one feature (or site). More than one polynucleotide can be synthesized on more than one feature. In some cases, more than one feature of each synthesis cycle of about 25% to about 80% is deblocked. In some cases, each synthesis cycle of about 25% to about 30%, about 25% to about 35%, about 25% to about 40%, about 25% to about 45%, about 25% to about 50%, about 25% to about 55%, about 25% to about 60%, about 25% to about 65%, about 25% to about 70%, about 25% to about 75%, about 25% to about 80%, about 30% to about 35%, about 30% to about 40%, about 30% to about 45%, about 30% to about 50%, about 30% to about 55%, about 25% to about 60%, about 25% to about 65%, about 25% to about 70%, about 25% to about 75%, about 25% to about 80%, about 30% to about 35%, about 30% to about 40%, about 30% to about 45%, about 30% to about 50%, about 30% to about 55 %, about 30% to about 60%, about 30% to about 65%, about 30% to about 70%, about 30% to about 75%, about 30% to about 80%, about 35% to about 40%, about 35% to about 45%, about 35% to about 50%, about 35% to about 55%, about 35% to about 60%, about 35% to about 65%, about 35% to about 70%, about 35% to about 75%, about 35% to about 80%, about 40% to about 45%, about 40% to about 50%, about 40% to about 55% , about 40% to about 60%, about 40% to about 65%, about 40% to about 70%, about 40% to about 75%, about 40% to about 80%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 45% to about 65%, about 45% to about 70%, about 45% to about 75%, about 45% to about 80%, about 50% to about 55%, about 50% to about 60%, about 50% to about 65%, about 50% to about 70%, about 50% to about 75%, About 50% to about 80%, about 55% to about 60%, about 55% to about 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 65% to about 70%, about 65% to about 75%, about 65% to about 80%, about 70% to about 75%, about 70% to about 80%, or about 75% to about 80% of more than one feature is deblocked. In some cases, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, or about 80% of more than one feature is deblocked per synthesis cycle. In some cases, at least about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, or about 75% of more than one feature is deblocked per synthesis cycle.In some cases, up to about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, or about 80% of more than one feature is deblocked per synthesis cycle.

[0133] More than one feature on a solid support can be independently addressable. In some cases, more than one feature is independently addressable by controlling the proximity of reagents to certain parts. In some cases, more than one feature is independently addressable by controlling the reactivity of polynucleotides at each feature in more than one feature. In some cases, more than one feature is independently addressable by one or more electrodes of a solid support. In U.S. Patent No. 10936953 or U.S. Patent No. 9267213 (which are incorporated herein by reference in their entirety), an example of a device comprising a solid support containing an addressable site (e.g., feature) is described. In some cases, more than one feature is addressable by masking a specific region. In some instances, a specific region is chemically functionalized, such as, for example, by modifying the surface with a hydrophobic or hydrophilic chemical group. As an example, more than one feature can be masked using the method described in U.S. Patent No. 10894242, U.S. Patent No. 10195580 or WO2022 / 047076 (which is incorporated herein by reference in its entirety). In some cases, more than one feature is addressable by electrochemical deblocking. In some cases, more than one feature is addressable by acid generation. In some cases, one or more electrodes can be used to produce one or more chemical reactions (for example, electrochemically generated acid (EGA) for nucleotide deprotection). In some cases, electrochemical deblocking includes a solution based on an organic solvent, for deblocking during the synthesis of any one of a variety of oligomers (for example, oligonucleotides). In this case, the removal of the blocking portion on the molecule is related to the chemistry based on acid, and the covalent bonding of the next nucleotide can be allowed. Electrochemical deblocking includes applying a voltage or current to one or more features via one or more electrodes on a solid support (e.g., an electrode microarray) to locally generate an acid or base (depending on whether the electrode is an anode or a cathode), which can affect the removal of an acid or base labile protecting group (part) bound to a chemical substance. In some cases, masking techniques using photogenerated acid-addressable are used in combination with a photosensitizer for deblocking. In some cases, more than one feature is addressable by metal-catalyzed deprotection (e.g., palladium-catalyzed deprotection).

[0134] In some cases, more than one feature may be addressable by masking methods. In some cases, lift-off fabrication methods ( Fig.11AIn some cases, the lift-off method includes adding a sacrificial layer (e.g., photoresist or "PR") to the base layer coated with the oxide layer, adding a conductive layer, and removing the sacrificial layer. In some cases, a dry etching manufacturing method ( Fig. 11B ). In some cases, the dry etching method includes adding one or more layers to the base layer, such as an oxide layer, a first intermediate layer (e.g., TiN or other material), a conductive layer (e.g., platinum), a second intermediate layer (e.g., TiN or other material), and a sacrificial layer (e.g., photoresist); partially removing the second intermediate layer to expose the conductive layer; partially removing the conductive layer to expose the first intermediate layer; partially removing the first conductive layer to expose the first intermediate layer; and partially removing the first intermediate layer to expose the oxide layer. As an example, the surface of the base layer including silicon and the top layer including oxide can be patterned with a removable masking material (such as photoresist) ( Fig.11A ). The entire surface including the masking can be plated with platinum, and then the masking layer can be removed. The previously masked areas are then exposed to oxide, while the unmasked areas include platinum on top of the oxide layer. As another example, a surface including a base layer of silicon, a first layer including oxide, a second layer of titanium nitride, a third layer including platinum, and a fourth layer including titanium nitride (from bottom to top) can be patterned with a removable masking material such as a photoresist ( Fig. 11B ). The unmasked fourth layer can be removed to expose the third layer, and the photoresist can be removed to expose the masked fourth layer. Removing all remaining second and fourth layers can produce a surface including a base layer of silicon, and a top layer of oxide and "islands" of platinum patterned on top of titanium nitride.

[0135] In some cases, by way of non-limiting example, one or more electrodes for generating electrochemical reagents may include metals such as iridium and / or platinum, and other metals such as palladium, gold, silver, copper, mercury, nickel, zinc, titanium, tungsten, aluminum, and alloys of various metals, and other conductive materials such as carbon, including glassy carbon, meshed glassy carbon, substrate plane graphite, edge plane graphite or graphite. In some cases, doped oxides such as indium tin oxide and semiconductors such as silicon oxide and gallium arsenide may also be used. Additionally, the electrode may be composed of conductive polymers, metal-doped polymers, conductive ceramics and conductive clays. In some cases, platinum and palladium contain favorable properties related to their ability to absorb hydrogen (e.g., their ability to "preload" hydrogen before use). In some cases, one or more electrodes may be connected to a power supply. In some cases, the electrode is connected to a power supply by a CMOS (complementary metal oxide semiconductor) switching circuit, a radio and microwave frequency addressable switch, an optical addressable switch, a direct connection from an electrode to a bonding pad on the periphery of a semiconductor chip, or any combination thereof. The CMOS switching circuit may include connecting each electrode to a CMOS transistor switch. The switches can be accessed by sending an electronic address signal along a common bus to an SRAM (static random access memory) circuit associated with each electrode. When the switch is "on", the electrode can be connected to a power source. Radio and microwave frequency addressable switches can involve electrodes that are switched by RF or microwave signals. This can allow the switch to be thrown with and / or without the use of switching logic. The switch can be tuned to receive a specific frequency or modulation frequency and switch without switching logic. Optically addressable switches can be switched by light. In some cases, one or more electrodes can also be switched with and without switching logic. In some instances, the optical signal can be spatially positioned to provide switching without switching logic, for example, by scanning a laser beam across an array of electrodes, where the electrode is switched each time the laser illuminates the electrode.

[0136] The sequence of more than one polynucleotide can be determined. In some cases, more than one polynucleotide can be sequenced according to the system and method provided herein. By way of non-limiting example, sequencing can include next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing or any combination thereof. In some cases, more than one polynucleotide can be sequenced via a sequencing instrument. In some cases, more than one polynucleotide is sequenced to produce more than one output sequence. In some cases, more than one output sequence overlaps with more than one polynucleotide sequence. In some cases, the overlap is about 50% to 100%. In some cases, the overlap is about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some cases, the overlap is about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some cases, the overlap is at least about 50%, about 60%, about 70%, about 80%, about 90%, or about 90%. In some cases, the overlap is at most about 60%, about 70%, about 80%, about 90%, or about 100%. In some cases, more than one output sequence is decoded using the methods described herein. For example, more than one output sequence is decoded using a greedy algorithm, a maximum likelihood (ML) algorithm, or a hybrid greedy ML algorithm. In some cases, more than one output sequence is decoded based at least in part on the probability of a deletion, insertion, mutation, or any combination thereof being calculated.

[0137] A platform for encoding data is also provided herein. In some cases, the platform includes a hybrid organic-computer simulation platform. In some cases, the platform includes a computing system, a synthesizer, or a combination thereof. In some cases, the computing system includes at least one processor and instructions that can be executed by at least one processor to perform operations. The computing system or at least one processor can be those provided herein. In some cases, the computing system includes a distributed computing system. In some cases, the computing system includes a cloud computing system. The cloud computing system may include a private cloud, a public cloud, a hybrid cloud, multiple clouds, or any combination thereof. The cloud computing system may include infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS), or any combination thereof. In some cases, the operation includes generating an internal codec including a secret book, such as those provided herein. In some cases, the secret book is optimized for one or more constraints, such as one or more constraints related to nucleic acid synthesis, post-processing, storage, or sequencing. In some cases, nucleic acid synthesis includes electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing, or any combination thereof. In some cases, one or more constraints related to nucleic acid synthesis include synthesis errors, such as insertions, deletions, or mutations. In some cases, post-processing includes one or more of connection, lysis, hybridization, denaturation, fixation to solid support, extension, error correction, enrichment, separation, purification and amplification. In some cases, storage includes cold data storage. Cold data storage can generally refer to the storage of rarely accessed data (such as data in nucleic acid). Cold data storage can be the antonym of "hot storage", and the "hot storage" refers to the storage of frequently accessed data. In some instances, storage includes hot storage, wherein the data stored in nucleic acid is frequently accessed. In some cases, storage includes nucleic acid storage in liquid or solid phase. In some instances, one or more constraints associated with storage include temperature (for example, room temperature), humidity, pressure, salinity, pH, concentration, time, light, UV, O2 or any combination thereof. In some cases, sequencing includes next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing or any combination thereof.

[0138] In some cases, the secret book is generated in a base sequence (e.g., [A, T, C, G], etc.). In some cases, the secret book includes codewords generated based on the base sequence. In some cases, the base sequence includes predetermined base transitions. In some cases, the operation includes applying an internal codec to encode binary data into more than one polynucleotide sequence using the methods provided herein.

[0139] In some cases, the synthesizer produces more than one polynucleotide comprising more than one polynucleotide sequence. In some cases, the synthesizer produces more than one polynucleotide sequence by synthesis, connection, assembly or any combination thereof. The synthesis method can be those provided herein (e.g., phosphoramidite, enzymatic, etc.). In some cases, the instruction from the computing system also causes the synthesizer to generate more than one polynucleotide. In some cases, the synthesizer is used to synthesize polynucleotides. In some cases, the synthesizer is used to assemble polynucleotides. In some cases, an optional assembly module is used to assemble polynucleotides. In some instances, assembly includes overlap extension polymerase chain reaction (PCR), polymerase cycle assembly (PCA), sticky end connection, bio-brick assembly, Golden Gate assembly, Gibson assembly, recombinase assembly, ligase cycle reaction, template-guided connection or any combination thereof. In some cases, the synthesizer and assembly module are in fluid communication, electronic communication or its combination.

[0140] The platform can also include a sequencer. The sequencer can include systems and devices for performing sequencing methods provided herein, or those known in the art. In some cases, the sequencer sequences more than one polynucleotide to generate more than one output sequence. The sequencing method can be those provided herein. In some cases, the instruction also causes the computing system to receive more than one output sequence. In some cases, the computing system further performs operations including decoding more than one output sequence. The computing system can decode more than one output sequence or any other polynucleotide sequence using a decoding scheme provided herein. In some cases, a greedy algorithm, a maximum likelihood (ML) algorithm, or a mixed greedy ML algorithm is used to decode more than one output sequence. In some cases, more than one output sequence is decoded based at least in part on the probability of missing, inserted, mutated, or any combination thereof calculated. The platform can also include a storage unit. In some cases, the storage unit stores more than one polynucleotide. The polynucleotide can be stored in a solution as a liquid, or dried as a solid. The polynucleotide can be stored on a substrate, such as those provided herein. In some cases, the instruction of the computing system causes more than one polynucleotide to be transferred between a synthesizer, a sequencer, a storage unit, or any combination thereof.

[0141] De novo polynucleotide synthesis

[0142] Provided herein are systems and methods for synthesizing polynucleotides on a substrate. In some cases, a final oligonucleotide pool from an internal codec is synthesized. In some cases, a library 1225 (e.g., Fig.12In some instances, a library comprising more than one polynucleotide from an encoding scheme encodes a pool in more than one pool. In some instances, a library comprising more than one polynucleotide from an encoding scheme encodes an index pool. In some cases, the method comprises using electrochemical deprotection. In some cases, the substrate is a flexible substrate. In some cases, at least 10 10 10 11 10 12 10 13 10 14 or 10 15 In some cases, at least 10 × 10 8 pcs, 10×10 9 pcs, 10×10 10 pcs, 10×10 11 or 10×10 12polynucleotides. In some cases, each polynucleotide synthesized comprises at least 20, 50, 100, 200, 300, 400, or 500 nucleobases. In some cases, the total average error rate for synthesizing these bases is less than about 1 in 100; 200; 300; 400; 500; 1000; 2000; 5000; 10000; 15000; 20000 bases. In some cases, these error rates are for at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5% or more of the synthesized polynucleotides. In some cases, these at least 90%, 95%, 98%, 99%, 99.5% or more of the synthesized polynucleotides are indistinguishable from the predetermined sequences they encode. In some cases, the error rate of the polynucleotides synthesized on the substrate using the methods and systems described herein is less than about 1 / 200, less than about 1 / 1,000, less than about 1 / 2,000, less than about 1 / 3,000 or less than about 1 / 5,000. Individual types of error rates include mismatches, deletions, insertions and / or substitutions of the polynucleotides synthesized on the substrate. The term "error rate" refers to the total amount of the synthesized polynucleotides compared to the total amount of the predetermined polynucleotide sequence. In some cases, the synthetic methods provided herein (e.g., based on inkjet synthetic methods) have a result better than about 0.1% deletion rate, 0.1% mutation rate (or substitution rate), 0.05% insertion rate or any combination thereof. For example, the synthesized polynucleotides can have a deletion rate of less than or about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, a mutation rate of less than or about 0.001%, 0.005%, 0.01%, 0.05%, or 0.1%, an insertion rate of less than or about 0.001%, 0.005%, 0.01%, or 0.05%, or any combination thereof. In some cases, the synthesized polynucleotides disclosed herein comprise a tether of 12 to 25 bases. In some cases, the tether comprises 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more bases.

[0143] Methods, systems, devices and compositions are described herein, wherein the chemical reactions used in electrochemically controlling polynucleotide synthesis are used. In some cases, the electrochemical reaction is controlled by any energy source such as light, heat, radiation or electricity. For example, an electrode is used to control the chemical reaction at all or a portion of discrete sites on the surface. In some cases, the electrode is charged by applying an electric potential to the electrode to control one or more chemical steps in the synthesis of polynucleotides. In some cases, these electrodes are addressable. In some cases, any number of chemical steps described herein are controlled by one or more electrodes. The electrochemical reaction can include oxidation, reduction, acid / base chemistry or other reactions controlled by the electrode. In some cases, the electrode produces electrons or protons, which are used as reagents for chemical conversion. In some cases, the electrode directly produces reagents such as acids. In some cases, the acid is a proton. In some cases, the electrode directly produces reagents such as alkalis. Acids or alkalis are generally used to cleave protective groups, or affect the kinetics of a variety of polynucleotide synthesis reactions, such as by adjusting the pH of the reaction solution. In some cases, the polynucleotide synthesis reactions controlled by electrochemistry include redox active metals or other redox active organic materials. In some cases, metal or organic catalysts are used for these electrochemical reactions. In some cases, the acid is produced by oxidation of the quinone.

[0144] The control of chemical reaction is not limited to the electrochemical generation of reagents; chemical reactivity can be indirectly affected by the electric field (or gradient) generated by the electrode to the biophysical changes of the substrate or reagent. In some cases, the substrate includes but is not limited to nucleic acids. In some cases, an electric field that repels or attracts a specific reagent or substrate toward or away from an electrode or surface is generated. In some cases, such a field is generated by applying an electric potential to one or more electrodes. For example, negatively charged nucleic acids are repelled from the negatively charged electrode surface. In some cases, such repulsion or attraction of polynucleotides or other reagents caused by a local electric field provides the movement of polynucleotides or other reagents within or outside the region of a synthetic device or structure. In some cases, the electrode generates an electric field, which repels polynucleotides away from the synthetic surface, structure or device. In some cases, the electrode generates an electric field, which attracts polynucleotides toward the synthetic surface, structure or device. In some cases, protons are repelled from the positively charged surface to limit the contact of protons with the substrate or its part. In some cases, a repulsive force or an attractive force is used to allow or prevent a reagent or substrate from entering a specific area of ​​the synthetic surface. In some cases, nucleoside monomers are prevented from contacting polynucleotide chains by applying an electric field near one or both components. Such an arrangement allows the gating of specific reagents, which can eliminate the need for blocking groups when controlling the concentration or contact rate between reagents and / or substrates. In some cases, unprotected nucleoside monomers are used for polynucleotide synthesis. Alternatively, applying a field near one or both components promotes the contact of nucleoside monomers with polynucleotide chains. In addition, applying an electric field to a substrate can change the reactivity or conformation of a substrate. In exemplary applications, the electric field generated by an electrode is used to prevent the polynucleotides at adjacent sites from interacting. In some cases, substrates are polynucleotides optionally attached to a surface. In some cases, the application of an electric field changes the three-dimensional structure of polynucleotides. Such changes include the folding or unfolding of multiple structures such as spirals, hairpins, loops or other three-dimensional nucleic acid structures. Such changes are useful for manipulating nucleic acids inside a well, channel or other structures. In some cases, an electric field is applied to a nucleic acid substrate to prevent secondary structure. In some cases, an electric field eliminates the need for a joint or attachment to a solid support during polynucleotide synthesis.

[0145] The suitable method for synthesizing polynucleotide on the substrate of the present disclosure is based on the DNA synthesis of phosphoramidite.In some cases, the reagent for the synthesis based on phosphoramidite comprises any one or combination in the following: nucleoside phosphoramidite, oxidant, activator or remove blocking agent, or comprises the solvent of acetonitrile.In some cases, the synthetic method based on phosphoramidite is included in the polynucleotide chain that phosphoramidite structural unit (i.e. nucleoside phosphoramidite) is controlledly added to growth in the coupling step, and the coupling step forms phosphite triester bond between the phosphoramidite structural unit and the nucleoside bonded with substrate.In some cases, nucleoside phosphoramidite is provided to activated substrate.In some cases, nucleoside phosphoramidite is provided to substrate together with activator. In some cases, nucleoside phosphoramidites are provided to substrates in an excess of 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100 times or more of the nucleoside bound by substrate. In some cases, the interpolation of nucleoside phosphoramidites is carried out in anhydrous environment, for example, in anhydrous acetonitrile. After adding and connecting nucleoside phosphoramidites in the coupling step, the substrate is optionally washed. In some cases, the coupling step is repeated once or more times in addition, and washing steps are optionally carried out between the addition of the nucleoside phosphoramidites added to the substrate. In some cases, the polynucleotide synthesis method used herein includes 1, 2, 3 or more sequential coupling steps. In many cases, before coupling, the nucleoside bonded to the substrate is deprotected by removing the protecting group, and wherein the protecting group plays the role of preventing polymerization. The protecting group can include any chemical group that prevents the polynucleotide chain from extending. In some cases, the protecting group is cleaved (or removed) in the presence of an acid. In some cases, the protecting group is cleaved in the presence of a base. In some cases, the protecting group is removed with electromagnetic radiation such as light, heat or other energy sources. In some cases, the protecting group is removed by oxidation or reduction reaction. In some cases, the protecting group includes a triarylmethyl group. In some cases, the protecting group includes an aryl ether. In some cases, the protecting group includes a disulfide. In some cases, the protecting group includes an acid-unstable silane. In some cases, the protecting group includes an acetal. In some cases, the protecting group includes a ketal. In some cases, the protecting group includes an enol ether. In some cases, the protecting group includes a methoxybenzyl group. In some cases, the protecting group includes an azide. In some cases, the protecting group is 4,4'-dimethoxytrityl (DMT). In some cases, the protecting group is tert-butyl carbonate. In some cases, the protecting group is tert-butyl ester. In some cases, the protecting group includes a base-unstable group.

[0146] After coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step. In the capping step, the growing polynucleotide is treated with a capping agent. The capping step is generally used to prevent further chain extension of the unreacted substrate-bound 5'-OH group after coupling to prevent the formation of a polynucleotide with an internal base deletion. In addition, phosphoramidites activated with 1H-tetrazole usually react with the O6 position of guanosine to a small extent. Without being bound by theory, when oxidized with I2 / water, this byproduct (possibly migrating via O6-N7) undergoes depurination. In the final deprotection process of the polynucleotide, the apurinic site can eventually be cleaved, thereby reducing the yield of the full-length product. The O6 modification can be removed by treating with a capping agent before oxidation with I2 / water. In some cases, including a capping step during polynucleotide synthesis reduces the error rate compared to synthesis without capping. As an example, the capping step includes treating the polynucleotide bound to the substrate with a mixture of acetic anhydride and 1-methylimidazole. After the capping step, the substrate is optionally washed.

[0147] After adding nucleoside phosphoramidites, and optionally after capping and one or more washing steps, substrates described herein include nucleic acids of combined growth that can be oxidized. The oxidation step includes oxidation of phosphite triester to tetracoordinate phosphotriester, which is a protected precursor of naturally occurring phosphodiester internucleoside bonds. In some cases, phosphite triester is electrochemically oxidized. In some cases, the oxidation of the polynucleotides of growth is achieved by optionally treating with iodine and water in the presence of a weak base such as pyridine, lutidine or collidine. Oxidation is sometimes carried out under anhydrous conditions using tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, capping steps are performed after oxidation. The second capping step allows the substrate to dry, because the water remaining from oxidation that may persist may inhibit subsequent coupling. After oxidation, the substrate and the polynucleotides of growth are optionally washed. In some cases, the oxidation step is replaced by a sulfurization step to obtain a polynucleotide phosphorothioate, wherein any capping step can be performed after sulfurization. Many reagents are capable of efficient sulfur transfer, including but not limited to 3-(dimethylaminomethylene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also known as Beaucage reagent) and N,N,N'N'-tetraethylthiuram disulfide (TETD).

[0148] In order to be incorporated into the subsequent circulation of the nucleoside that occurs by coupling, the protected 5 ' end (or 3 ' end, if the synthesis is carried out in 5 ' to 3 ' direction) of the polynucleotide of the growth of substrate binding is removed so that the primary hydroxyl group can react with the next nucleoside phosphoramidite. In some cases, the blocking group is DMT, and deblocking is carried out with trichloroacetic acid in dichloromethane. In some cases, the blocking group is DMT, and deblocking is carried out with the proton produced by electrochemistry. The depurination of the polynucleotide of solid support binding may be increased by continuing the extended time or detritylation with a stronger acid solution than recommended, and thereby reduce the productive rate of the desired full-length product. Method and composition described herein provide the controlled deblocking condition that limits the undesirable depurination reaction. In some cases, the polynucleotide of substrate binding is washed after deblocking. In some cases, the effective washing after deblocking helps to synthesize the polynucleotide with low error rate.

[0149] The methods described herein for synthesizing polynucleotides on substrates can involve a repetitive sequence of the following steps: applying a protected monomer to the surface of a substrate feature to connect to the surface, a joint, or to a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and applying another protected monomer for connection. One or more intermediate steps include oxidation and / or sulfurization. In some cases, one or more washing steps are present before or after one or all of the steps.

[0150] The methods described herein for synthesizing polynucleotides on a substrate may include an oxidation step. For example, the method involves a repetitive sequence of the following steps: applying a protected monomer to a surface of a substrate feature to connect to the surface, a linker, or to a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; applying another protected monomer for connection, and oxidation and / or sulfurization. In some cases, there is one or more washing steps before or after one or all of the steps.

[0151] The methods described herein for synthesizing polynucleotides on substrates may also include an iterative sequence of the following steps: applying a protected monomer to a surface of a substrate feature to attach to the surface, a linker, or to a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and oxidizing and / or sulfiding. In some cases, one or more washing steps are performed before or after one or all of the steps.

[0152] The method for synthesizing polynucleotides on substrates described herein can also include an iterative sequence of the following steps: applying a protected monomer to the surface of a substrate feature to connect to the surface, a joint or to a previously deprotected monomer; and oxidation and / or sulfurization. In some cases, one or more washing steps are performed before or after one or all of the steps.

[0153] The methods described herein for synthesizing polynucleotides on substrates may also include an iterative sequence of the following steps: applying a protected monomer to a surface of a substrate feature to attach to the surface, a linker, or to a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and oxidizing and / or sulfiding. In some cases, one or more washing steps are performed before or after one or all of the steps.

[0154] In some cases, polynucleotides are synthesized with light-labile protective groups, wherein the hydroxyl groups produced on the surface are blocked by light-labile protective groups. When the surface is exposed to UV light such as by a photolithography mask, a pattern of free hydroxyl groups can be produced on the surface. According to phosphoramidite chemistry, these hydroxyl groups can react with the nucleoside phosphoramidites of light protection. A second photolithography mask can be applied, and the surface can be exposed to UV light to produce a second pattern of hydroxyl groups, followed by coupling with the nucleoside phosphoramidites of 5'-light protection. Similarly, a pattern can be produced, and an oligomer chain can be extended. Not bound by theory, the instability of the photocleavable group depends on the wavelength and the polarity of the solvent used, and the rate of photocleavable may be affected by exposure duration and light intensity. The method can utilize many factors, such as the alignment accuracy of the mask, the removal efficiency of the photoprotective group and the yield of the phosphoramidite coupling step. In addition, the accidental leakage of light leaking to adjacent sites can be minimized. The density of the synthesized oligomer of each point can be monitored by adjusting the load of the leading nucleoside on the synthetic surface.

[0155] The surface of the substrate providing support for polynucleotide synthesis described herein can be chemically modified to allow the synthesized polynucleotide chain to be cracked from the surface. In some cases, the polynucleotide chain is cracked while the polynucleotide is deprotected. In some cases, the polynucleotide chain is cracked after the polynucleotide is deprotected. In an exemplary scheme, a trialkoxysilylamine such as (CH3CH2O)3Si-(CH2)2-NH2 reacts with the surface SiOH group of the substrate, and then reacts with succinic anhydride and amine to produce amide bonds and free OH, supporting nucleic acid chain growth on free OH. Cracking includes the gas cracking using ammonia or methylamine. In some cases, cracking includes the joint cracking using an electrically generated reagent such as an acid or base. In some cases, after being released from the surface, the polynucleotide is assembled into larger nucleic acids, which are sequenced and decoded to extract the stored information.

[0156] In some cases, synthesis includes enzymatic synthesis. Enzymatic synthesis can be carried out on the surface described herein. In some cases, enzymatic synthesis includes chain extension enzyme. In some cases, chain extension enzyme is polymerase. In some cases, polymerase is non-template dependent polymerase. In some cases, polymerase is RNA polymerase or DNA polymerase. In some cases, polymerase is DNA polymerase. In some cases, enzymatic DNA synthesis uses water as solvent, and reagent is enzyme terminal deoxynucleotidyl transferase (TdT) or deblocking agent. In some cases, the enzymatic synthesis of DNA uses non-template dependent DNA polymerase terminal deoxynucleotidyl transferase (TdT), which is a protein that has evolved to rapidly catalyze the connection of naturally occurring dNTPs. TdT adds nucleotides indiscriminately, and therefore prevents it from continuing unregulated synthesis by various techniques, such as tethering TdT, producing variant enzymes and using nucleotides including reversible terminators to prevent chain extension. TdT activity reaches maximum at about 37 ° C, and enzymatic reaction is carried out in an aqueous environment. Examples of DNA polymerases include, but are not limited to, polA, polB, polC, polD, polY, polX, reverse transcriptase (RT), and high-fidelity polymerases. In some cases, the polymerase is a modified polymerase. In some embodiments, the polymerase comprises Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, 9°Nm TM 、Therminator TM DNA polymerase, Tne, Tma, TfI, Tth, TIi, Stoffel fragment, Vent TM and Deep Vent TM DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, Escherichia coli (E.coli) DNA polymerase I, Escherichia coli DNA polymerase III, Archaea DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase and III reverse transcriptase. In some embodiments, the polymerase is DNA polymerase 1-Klenow fragment, Vent polymerase, DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7RNA polymerase, Therminator TM DNA polymerase, POLB polymerase, SP6 RNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase or III reverse transcriptase. The polymerase molecule used in the methods described herein can be polymerase theta, DNA polymerase, or any enzyme that can extend a nucleotide chain. In some embodiments, the polymerase is tri29. In some embodiments, the polymerase is a protein with a pocket that works around a terminal phosphate group (e.g., a triphosphate group).

[0157] In some embodiments, enzymatic synthesis uses TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, the method uses TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid mutations of amino acid residues accessible to the surface. In some embodiments, TdT is a variant of TdT. In some embodiments, the variant of TdT comprises a cysteine ​​mutation (e.g., NTT-1). In some embodiments, the variant of TdT is NTT-1, NTT-2 or NTT-3. In some cases, the variant TdT comprises at least 70%, 80%, 90% or 95% sequence identity with wild-type TdT. In some embodiments, enzymatic synthesis can use polymerase θ with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, enzymatic synthesis can use polymerase θ with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations of amino acid residues accessible to the surface. In some embodiments, polymerase θ is a variant of polymerase θ. In some cases, variant polymerase θ comprises at least 70%, 80%, 90%, or 95% sequence identity to wild-type polymerase θ. In some embodiments, polymerase θ is encoded by POLQ.

[0158] In some embodiments, the enzyme described herein (e.g., TdT) comprises one or more non-natural amino acids. In some cases, the non-natural amino acid comprises: a lysine analog; an aromatic side chain; an azido group; an alkyne group; or an aldehyde or ketone group. In some cases, the non-natural amino acid does not comprise an aromatic side chain. In some embodiments, the non-natural amino acid is selected from N6-azidoethoxycarbonyl-L-lysine (AzK), N6-propargylethoxycarbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyl tetrazine lysine, allyloxycarbonyl lysine, 2-amino -8-Oxonanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargylphenylalanine, 3-methylphenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl-tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2- In some embodiments, the enzymes described herein are fused to one or more other enzymes. For example, TdT is fused to other enzymes such as helicases.

[0159] Various joints can be used for conjugating enzymes or other nucleic acid (e.g., polymerase) binding moieties with one or more base pairing moieties (e.g., modified nucleotides) during the enzymatic synthesis of polynucleotides. The conjugation of nucleotides or other base pairing moieties with joints can be achieved by any means known in the field of chemical conjugation methods. For example, base-modified nucleotides comprising the addition of free amine groups are considered to be conjugated with joints as described herein. For example, primary amines can be connected to bases in a manner such that they can react with heterobifunctional polyethylene glycol (PEG) joints, to produce nucleotides containing variable length PEG joints, which will still be appropriately bound to the enzyme active site. Examples of such amine-containing nucleotides include 5-propargylamino-dNTP, 5-propargylamino-NTP, aminoallyl-dNTP and aminoallyl-NTP. In some embodiments, amine-containing nucleotides are suitable for conjugation with joints based on PEG. The length of the PEG linker can vary, for example, from 1-1000, from 1-500, from 1-11, from 1-100, from 1-50 or from 1-10 subunits. Non-limiting examples of other suitable linkers may include, but are not limited to, poly-T and poly-A oligonucleotide chains (e.g., ranging in length from about 1 base to about 1,000 bases), peptide linkers (e.g., ranging in length from about 1 residue to about 1,000 residues of poly-glycine or poly-alanine) or carbon chain linkers (e.g., C6, C12, C18, C24, etc.). In some embodiments, the linker comprises an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the linker comprises a maleimide group. The connection of nucleotides can be achieved by disulfide bond formation (forming an easily cleavable connection), amide formation, ester formation, protein-ligand bond (e.g., biotin-streptavidin bond), by alkylation (e.g., using substituted iodoacetamide reagents), or adduct formation using aldehydes and amines or hydrazines. In some embodiments, the linker comprises, for example, a maltose group, a biotin group, an O2-benzylcytosine group or an O2-benzylcytosine derivative, an O6-benzylguanine group or an O6-benzylguanine derivative. The length of the linker can vary depending on the type of nucleotide (or other base pairing moiety) and enzyme (or other nucleic acid binding moiety).In some cases, the linker used to connect the nucleotide to the enzyme can have a diameter of about 0.1 nm-1,000 nm, 0.5 nm-500 nm, 0.5 nm-400 nm, 0.5 nm-300 nm, 0.5 nm-200 nm, 0.5 nm-100 nm, 0.5 nm-50 nm, 0.6 nm-500 nm, 0.6 nm-400 nm, 0.6 nm-300 nm, 0.6 nm-200 nm, 0.6 nm-100 nm, 0.6 nm-50 nm, 1 nm-500 nm, In some embodiments, the chemical linker is an acid cleavable linker. In some embodiments, the chemical linker is an alkali cleavable linker. In some embodiments, the chemical linker is a photocleavable linker. In some embodiments, the chemical linker is selected from the group consisting of a silyl linker, an alkyl linker, a polyether linker, a polysulfonyl linker, a polysulfoxide linker, and any combination thereof. In some embodiments, the linker is cleaved by an enzyme. In some embodiments, the enzyme is a protease, esterase, glycosylase or peptidase. In some embodiments, the lyase destroys the bond in the polymerase. In some embodiments, the lyase directly cleaves the connected nucleosides.

[0160] Surface described herein can be reused after polynucleotide cleavage, to support other polynucleotide synthesis cycle.For example, joint can be used again without additional processing / chemical modification.In some cases, joint is non-covalently bound to substrate surface or polynucleotide.In some embodiments, joint remains attached to polynucleotide after from surface cleavage.In some embodiments, joint comprises reversible covalent bond, such as ester, amide, ketal, β substituted ketone, heterocyclic compound or other groups capable of reversible cleavage.In some cases, such reversible cleavage reaction is by adding or removing reagent, or is controlled by the electrochemical process controlled by electrode.Optionally, after multiple cycles, chemical joint or surface-bound chemical group is regenerated to recover reactivity and remove the unwanted by-product formed on such joint or surface-bound chemical group.

[0161] Device for polynucleotide storage

[0162] The polynucleotide library of synthesis can be stored in the device. In some cases, the device includes a polynucleotide data storage system. In some cases, the library of the coding pool (for example, more than one pool or index pool) is stored in a compartment. In some cases, by way of non-limiting example, the compartment includes an active surface (for example, a site), a tube or any other physical storage solution. In some instances, the compartment is labeled. In some instances, the label includes a barcode, a name (for example, a customer name, a sample type, etc.), a timestamp, a list of stored objects or any combination thereof.

[0163] In some cases, the device for storing digital information in DNA includes one or more compartments. In some cases, each of the one or more compartments includes a library, and the library includes more than one polynucleotide. In some instances, the library encoding includes a pool (for example, a pool in more than one pool described herein) of digital information corresponding to one or more objects. In some instances, the pool includes a pool descriptor, one or more pool items, an end pool descriptor, such as those described herein. In some instances, the pool includes about 1GB to about 1TB of digital information, as previously described herein.

[0164] The compartment or structure for storing more than one polynucleotide can be of any shape or size. In some cases, the compartment is substantially spherical, tubular ( Fig.17A ), ovate, cone, cube, cuboid, cylinder, wedge, hexagonal prism, square-bottomed pyramid, triangular-bottomed pyramid, triangular prism, torus, hemisphere, spiral, heart-shaped or other shapes. In some cases, the shape is configured to allow the structure to be opened or closed to the external environment. In some cases, such closure is facilitated by welding, seals, diaphragms, or other mechanisms for limiting the movement of gases or other substances into or out of the structure. In some cases, the compartments include holes, grooves, diaphragms, valves, or ports for adding or removing nucleic acids, fluids, gases, or other materials from the structure. In some cases, the structure for storing more than one polynucleotide includes a cap and a body ( Fig. 17B In some cases, the compartment for storing more than one polynucleotide comprises a removable screw cap ( Fig. 17C In some cases, the structure includes a diaphragm ( Fig.17D In some cases, the structure includes two rounded pill-shaped halves that form a seal when one half is inserted into the other half ( Fig.17E In some cases, the structure includes a substantially flat disk-shaped container having a sealable lid ( Fig.17F ). In some cases, the compartment comprises a box with an optionally attached lid ( Figure 17G). In some instances, the shape is a cylinder or a disk. In some instances, a cylinder or a disk shape is preferred for automated handling and / or filling of the compartments.

[0165] In some cases, each of the one or more compartments comprises a medium for storing more than one polynucleotide. In some instances, the medium includes a solid, a liquid, a gas or any combination thereof. In some instances, the medium comprises a saline solution. In some instances, the mol ratio of salt to DNA can be in the range of about 20:1 to about 2:1. In some instances, the mol ratio depends on the molecular weight of the salt used and depends on the relative amount of the salt and DNA of the combination. In some instances, the mol ratio between the cation of the salt and the negatively charged phosphate group of the DNA is calculated. In some instances, the saline solution comprises a mol ratio of the salt cation less than 20:1 to the phosphate group in the DNA. In some instances, the saline solution is dried to produce a dry product. In some cases, by way of non-limiting example, the saline solution comprises calcium chloride, calcium nitrate, calcium carbonate, calcium phosphate, magnesium chloride, magnesium sulfate, magnesium nitrate, magnesium carbonate, lanthanum chloride, lanthanum nitrate, lanthanum carbonate, lanthanum bromide or its mixture. In some cases, the salt solution comprises barium (II) chloride dihydrate, calcium chloride dihydrate, anhydrous copper (II) chloride, lanthanum trichloride, magnesium dichloride hexahydrate, sodium chloride, or strontium chloride hexahydrate. In some cases, the concentration of the salt solution is about 0.01 nM to about 0.1 nM.

[0166] In some cases, each of the one or more compartments is connected. In some cases, each of the one or more compartments is connected through a medium. In some cases, each of the one or more compartments is not connected. In some cases, each of the one or more compartments is not connected through a medium.

[0167] In some cases, the device also includes one or more second compartments. In some cases, each of the one or more second compartments includes a second library. In some instances, the second library encoding index pool, such as those described herein. In some cases, one or more second compartments include a medium as previously described herein. In some cases, one or more second compartments include the same medium as one or more compartments. In some cases, one or more second compartments include a medium different from one or more compartments. In some cases, each of the one or more second compartments is communicated with each other and / or is communicated with one or more compartments (for example, by a medium). In some cases, each of the one or more second compartments is not communicated with each other and / or is not communicated with one or more compartments.

[0168] In some cases, the device also includes a solid support comprising a surface. Thus, devices for nucleic acid synthesis and storage based on solid supports are described herein, wherein the solid supports have different sizes. In some cases, the size of the solid support is between about 40mm and 120mm multiplied by between about 25mm and 100mm. In some cases, the size of the solid support is about 80mm multiplied by about 50mm. In some cases, the width of the solid support is at least or about 10mm, 20mm, 40mm, 60mm, 80mm, 100mm, 150mm, 200mm, 300mm, 400mm, 500mm or greater than 500mm. In some cases, the height of the solid support is at least or about 10mm, 20mm, 40mm, 60mm, 80mm, 100mm, 150mm, 200mm, 300mm, 400mm, 500mm or greater than 500mm. In some cases, the solid support has a width of at least or about 10mm, 20mm, 40mm, 60mm, 80mm, 100mm, 150mm, 200mm, 300mm, 400mm, 500mm or greater than 500mm. 2 ; 200mm 2 ; 500mm 2 ; 1,000mm 2 ; 2,000mm 2 ; 4,500mm 2 ; 5,000mm 2 ; 10,000mm 2 ; 12,000mm 2 ; 15,000mm 2 ; 20,000mm 2 ; 30,000mm 2 ; 40,000mm 2 ; 50,000mm 2 In some cases, the thickness of the solid support is between about 50mm and about 2000mm, between about 50mm and about 1000mm, between about 100mm and about 1000mm, between about 200mm and about 1000mm, or between about 250mm and about 1000mm. Non-limiting examples of solid support thickness include 275mm, 375mm, 525mm, 625mm, 675mm, 725mm, 775mm and 925mm. In some cases, the thickness of the solid support is at least or about 0.5mm, 1.0mm, 1.5mm, 2.0mm, 2.5mm, 3.0mm, 3.5mm, 4.0mm or greater than 4.0mm.

[0169] Devices are described herein, wherein two or more solid supports are assembled. In some cases, solid supports are joined together on a larger unit. Engagement can include other exchange media between fluid exchange, electrical signals or solid supports. The unit can be joined with any number of servers, computers or networking devices. For example, more than one solid support is integrated into a rack unit, which is conveniently inserted into or removed from a server rack. A rack unit can include any number of solid supports. In some cases, a rack unit includes at least 1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 50,000, 100,000 or more than 100,000 solid supports. In some cases, two or more solid supports do not engage each other. Nucleic acids (and information stored therein) present on a solid support can be accessed from a rack unit. Access includes removing polynucleotides from solid supports, directly analyzing polynucleotides on solid supports, or allowing manipulation or identifying any other method of information stored in nucleic acids. In some cases, information is accessed from a single site on more than one rack, a single rack, a single solid support in a rack, a part of a solid support, or a solid support. In many cases, access includes engaging nucleic acid with another device, another device such as a mass spectrometer, HPLC, sequencing instrument, PCR thermal cycler, or other devices for manipulating nucleic acids. In some cases, access to nucleic acid information is achieved by cracking polynucleotides from all or part of solid supports. In some cases, cracking includes being exposed to chemical reagents (ammonia or other reagents), electric potential, radiation, heat, light, acoustics, or other forms of energy that can manipulate chemical bonds. In some cases, cracking occurs by charging one or more electrodes near polynucleotides. In some cases, electromagnetic radiation in the form of UV light is used for cracking polynucleotides. In some cases, lamps are used for cracking polynucleotides, and mask mediates UV light exposure positions on surfaces. In some cases, a laser is used to lyse polynucleotides, and a shutter open / closed state controls exposure of the surface to UV light. In some cases, access to nucleic acid information (including removal / addition of racks, solid supports, reagents, nucleic acids or other components) is fully automated.

[0170] The solid support described herein comprises an active region. In some cases, the active region comprises a region or site for nucleic acid synthesis. In some cases, the active region comprises a region or site for nucleic acid storage. In some instances, the region or site comprises one or more compartments. In some instances, the region or site comprises a second one or more compartments. In some cases, the region is addressable. In some instances, the region is addressable by an electrode.

[0171] In some cases, the active area includes a width of at least or about 0.5mm, 1mm, 1.5mm, 2mm, 2.5mm, 3mm, 5mm, 7mm, 10mm, 12mm, 14mm, 16mm, 18mm, 20mm, 25mm, 30mm, 35mm, 40mm, 45mm, 50mm, 60mm, 70mm, 80mm or more than 80mm. In some cases, the active area includes a height of at least or about 0.5mm, 1mm, 1.5mm, 2mm, 2.5mm, 3mm, 5mm, 7mm, 10mm, 12mm, 14mm, 16mm, 18mm, 20mm, 25mm, 30mm, 35mm, 40mm, 45mm, 50mm, 60mm, 70mm, 80mm or more than 80mm.

[0172] Described herein are devices for nucleic acid synthesis and storage based on solid supports, wherein the solid support has multiple sites (e.g., points) or positions for synthesis or storage. In some cases, the solid support comprises up to or about 10,000 times 10,000 positions in one area. In some cases, the solid support comprises between about 1000 and 20,000 times between about 1000 and 20,000 positions in one area. In some cases, the solid support comprises at least or about 10, 30, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 12,000, 14,000, 16,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, 41,000, 42,000, 43,000, 44,000, 45,000, 46,000, 47,000, 48,000, 49,000, 5 In some cases, the area is up to 0.25 square inches, 0.5 square inches, 0.75 square inches, 1.0 square inches, 1.25 square inches, 1.5 square inches, or 2.0 square inches. In some cases, solid support comprises a site with at least or about 0.1um, 0.2um, 0.25um, 0.3um, 0.4um, 0.5um, 1.0um, 1.5um, 2.0um, 2.5um, 3.0um, 3.5um, 4.0um, 4.5um, 5um, 6um, 7um, 8um, 9um, 10um or a spacing greater than 10um. In some cases, solid support comprises a site with a spacing of about 5um. In some cases, solid support comprises a site with a spacing of about 2um. In some cases, solid support comprises a site with a spacing of about 1um. In some cases, solid support comprises a site with a spacing of about 0.2um. In some cases, the solid support comprises sites having a spacing of about 0.2um to about 10um, about 0.2um to about 8um, about 0.5um to about 10um, about 1um to about 10um, about 2um to about 8um, about 3um to about 5um, about 1um to about 3um, or about 0.5um to about 3um. In some cases, the solid support comprises sites having a spacing of about 0.1um to about 3um.

[0173] Solid supports for nucleic acid synthesis or storage as described herein include high capacities for storing data. For example, the capacity of the solid support is at least or about 1 megabyte, 2 megabytes, 3 megabytes, 4 megabytes, 5 megabytes, 6 megabytes, 7 megabytes, 8 megabytes, 9 megabytes, 10 megabytes, 20 megabytes, 50 megabytes, 100 megabytes, 200 megabytes, 300 megabytes, 400 megabytes, 500 megabytes, 600 megabytes, 700 megabytes, 800 megabytes, 900 megabytes, 1000 megabytes or more than 1000 megabytes. In some cases, the capacity of the solid support is between about 1 megabyte to 10 megabytes, 1 megabyte to 50 megabytes, 1 megabyte to 100 megabytes, 1 megabyte to 500 megabytes, 1 megabyte to 1000 megabytes, 10 megabytes to 50 megabytes, 10 megabytes to 100 megabytes, 10 megabytes to 500 megabytes, 10 megabytes to 1000 megabytes, 50 megabytes to 100 megabytes, 50 megabytes to 500 megabytes, 50 megabytes to 1000 megabytes, 100 megabytes to 500 megabytes, 100 megabytes to 1000 megabytes, 200 megabytes to 500 megabytes, 200 megabytes to 1000 megabytes, 500 megabytes to 1000 megabytes, or about 800 megabytes to 1000 megabytes. For example, the solid support can have a capacity of at least or about 1 Gbyte, 2 Gbyte, 3 Gbyte, 4 Gbyte, 5 Gbyte, 6 Gbyte, 7 Gbyte, 8 Gbyte, 9 Gbyte, 10 Gbyte, 20 Gbyte, 50 Gbyte, 100 Gbyte, 200 Gbyte, 300 Gbyte, 400 Gbyte, 500 Gbyte, 600 Gbyte, 700 Gbyte, 800 Gbyte, 900 Gbyte, 1000 Gbyte, or more than 1000 Gbyte. In some cases, the capacity of the solid support is between about 1 Gbyte to 10 Gbytes, 1 Gbyte to 50 Gbytes, 1 Gbyte to 100 Gbytes, 1 Gbyte to 500 Gbytes, 1 Gbyte to 1000 Gbytes, 10 Gbyte to 50 Gbytes, 10 Gbyte to 100 Gbytes, 10 Gbyte to 500 Gbytes, 10 Gbyte to 1000 Gbytes, 50 Gbyte to 100 Gbytes, 50 Gbyte to 500 Gbytes, 50 Gbyte to 1000 Gbytes, 100 Gbyte to 500 Gbytes, 100 Gbyte to 1000 Gbytes, 200 Gbyte to 500 Gbytes, 200 Gbyte to 1000 Gbytes, 500 Gbyte to 1000 Gbytes, or about 800 Gbyte to 1000 Gbytes.For example, the solid support has a capacity of at least or about 1 terabyte, 2 terabytes, 3 terabytes, 4 terabytes, 5 terabytes, 6 terabytes, 7 terabytes, 8 terabytes, 9 terabytes, 10 terabytes, 20 terabytes, 50 terabytes, 100 terabytes, 200 terabytes, 300 terabytes, 400 terabytes, 500 terabytes, 600 terabytes, 700 terabytes, 800 terabytes, 900 terabytes, 1000 terabytes, or more than 1000 terabytes. In some cases, the capacity of the solid support is between about 1 terabyte to 10 terabytes, 1 terabyte to 50 terabytes, 1 terabyte to 100 terabytes, 1 terabyte to 500 terabytes, 1 terabyte to 1000 terabytes, 10 terabyte to 50 terabytes, 10 terabyte to 100 terabytes, 10 terabyte to 500 terabytes, 10 terabyte to 1000 terabytes, 50 terabyte to 1000 terabytes, 50 terabyte to 500 terabytes, 50 terabyte to 1000 terabytes, 100 terabyte to 500 terabytes, 100 terabyte to 1000 terabytes, 200 terabytes to 500 terabytes, 200 terabytes to 1000 terabytes, 500 terabytes to 1000 terabytes, or about 800 terabytes to 1000 terabytes. For example, the solid support has a capacity of at least or about 1 petabyte, 2 petabytes, 3 petabytes, 4 petabytes, 5 petabytes, 6 petabytes, 7 petabytes, 8 petabytes, 9 petabytes, 10 petabytes, 20 petabytes, 50 petabytes, 100 petabytes, 200 petabytes, 300 petabytes, 400 petabytes, 500 petabytes, 600 petabytes, 700 petabytes, 800 petabytes, 900 petabytes, 1000 petabytes, or more than 1000 petabytes. In some cases, the capacity of the solid support is between about 1 petabyte to 10 petabytes, 1 petabyte to 50 petabytes, 1 petabyte to 100 petabytes, 1 petabyte to 500 petabytes, 1 petabyte to 1000 petabytes, 10 petabyte to 50 petabytes, 10 petabyte to 100 petabytes, 10 petabyte to 500 petabytes, 10 petabyte to 1000 petabytes, 50 petabyte to 100 petabytes, 50 petabyte to 500 petabytes, 50 petabyte to 1000 petabytes, 100 petabyte to 500 petabytes, 100 petabyte to 1000 petabytes, 200 petabyte to 500 petabytes, 200 petabyte to 1000 petabytes, 500 petabyte to 1000 petabytes, or about 800 petabytes to 1000 petabytes. In some cases, the capacity of the solid support is about 100 petabytes.

[0174] In some cases, data is stored as a data packet array as a droplet. In some instances, the data packet array is an addressable data packet. In some instances, the data packet is addressable using electrodes. In some cases, data is stored as a data packet array as a droplet on a point. In some cases, data is stored as a data packet array as a dry trap. In some cases, the array contains at least or about 1 gigabyte, 2 gigabytes, 3 gigabytes, 4 gigabytes, 5 gigabytes, 6 gigabytes, 7 gigabytes, 8 gigabytes, 9 gigabytes, 10 gigabytes, 20 gigabytes, 50 gigabytes, 100 gigabytes, 200 gigabytes, or more than 200 gigabytes of data. In some cases, the array contains at least or about 1 terabyte, 2 terabytes, 3 terabytes, 4 terabytes, 5 terabytes, 6 terabytes, 7 terabytes, 8 terabytes, 9 terabytes, 10 terabytes, 20 terabytes, 50 terabytes, 100 terabytes, 200 terabytes, or more than 200 terabytes of data. In some cases, the information items are stored in a data background. For example, an information item encodes about 10 megabytes to about 100 megabytes of data and is stored in 1 petabyte of background data. In some cases, an information item encodes at least or about 1 megabyte, 10 megabytes, 20 megabytes, 30 megabytes, 40 megabytes, 50 megabytes, 60 megabytes, 70 megabytes, 80 megabytes, 90 megabytes, 100 megabytes, 150 megabytes, 200 megabytes, 300 megabytes, 400 megabytes, 500 megabytes, or more than 500 megabytes of data and is stored in 1 petabyte, 10 petabytes, 20 petabytes, 30 petabytes, 40 petabytes, 50 petabytes, 60 petabytes, 70 petabytes, 80 petabytes, 90 petabytes, 100 petabytes, 150 petabytes, 200 petabytes, 300 petabytes, 400 petabytes, 500 petabytes, or more than 500 petabytes of background data.

[0175] Provided herein are devices for nucleic acid synthesis and storage based on solid supports, wherein after synthesis, the polynucleotides are collected in a data packet as one or more droplets. In some cases, the polynucleotides are collected in a data packet and stored as one or more droplets. In some cases, the number of droplets is at least or about 1, 10, 20, 50, 100, 200, 300, 500, 1000, 2500, 5000, 7500, 10,000, 25,000, 50,000, 75,000, 100,000, 1 million, 5 million, 10 million, 25 million, 50 million, 75 million, 100 million, 250 million, 50 million, 75 million, 100 million, 250 million, 500 million, 750 million, or more than 750 million droplets. In some cases, the droplet volume comprises a diameter of 5um, 10um, 15um, 20um, 25um, 30um, 35um, 40um, 45um, 50um, 55um, 60um, 65um, 70um, 75um, 80um, 85um, 90um, 95um, 100um, or greater than 100um (micrometers). In some cases, the droplet volume comprises a diameter of 1um-100um, 10um-90um, 20um-80um, 30um-70um, or 40um-50um.

[0176] In some cases, the polynucleotides collected in the data packets include similar sequences. In some cases, the polynucleotides also include non-identical sequences to be used as labels or barcodes. For example, non-identical sequences are used to index polynucleotides stored on solid supports, and then search for specific polynucleotides based on non-identical sequences. Exemplary labels or barcode lengths include but are not limited to the barcode sequences of the length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 or more bases. In some cases, labels or barcodes include at least or about 10, 50, 75, 100, 200, 300, 400 or more than 400 base pairs.

[0177] Provided herein are devices for nucleic acid synthesis and storage based on solid supports, wherein polynucleotides are collected in data packets comprising redundancy. For example, the data packet comprises about 100 to about 1000 copies of every kind of polynucleotide. In some cases, the data packet comprises at least or about 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000 or more than 2000 copies of every kind of polynucleotide. In some cases, the data packet comprises about 1000X to about 5000X synthetic redundancy. In some cases, the synthetic redundancy is at least or about 500X, 1000X, 1500X, 2000X, 2500X, 3000X, 3500X, 4000X, 5000X, 6000X, 7000X, 8000X or greater than 8000X. The polynucleotides synthesized using the method based on the solid support as described herein comprise various lengths. In some cases, the synthetic polynucleotides are further stored on a solid support. In some cases, the polynucleotide length is between about 100 to about 1000 bases. In some cases, the polynucleotide comprises at least or about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, or more bases in length.

[0178] Sequencing

[0179] Extraction and / or amplification of polynucleotide from the surface of synthesis or storage polynucleotide.After extraction and / or amplification of polynucleotide from the surface of structure, suitable sequencing technology can be adopted to order-check polynucleotide.In some cases, DNA sequence is read on substrate or in the feature of structure.In some cases, extract the polynucleotide stored on substrate, and optionally assemble into longer nucleic acid and then order-check.Can use system and method described herein to extract polynucleotide from substrate.

[0180] The polynucleotides synthesized and stored on the structures described herein encode data that can be retrieved or interpreted by reading the sequence of the synthesized polynucleotides and converting the sequence into a computer-readable representation (e.g., a string of symbols, such as a binary code). In some cases, the sequence needs to be assembled, and the assembly step may need to be performed at the polynucleotide sequence stage or at the digital sequence stage.

[0181] Provided herein is a detection system, which includes a device that can sequence the stored polynucleotides directly on the structure and / or after being removed from a main structure (e.g., a synthetic structure, a storage structure, etc.). In the case where the structure is a reel-type tape of a flexible material, the detection system includes a device for holding the structure and making the structure advance through a detection position and a detector arranged near the detection position, and the detector is used to detect a signal derived from the part when a part of the belt is located at the detection position. In some cases, the signal indicates the presence of polynucleotides. In some cases, the signal indicates the sequence (e.g., fluorescent signal) of polynucleotides. In some cases, the information encoded in the polynucleotides on a continuous tape is read by a computer when the tape is continuously transmitted through a detector that is operably connected to a computer. In some cases, the detection system includes a computer system, which includes a polynucleotide sequencing device, a database for storing and retrieving data related to a polynucleotide sequence, a software for converting the DNA code of a polynucleotide sequence into a symbol string such as a binary code, a computer for reading a binary code, or any combination thereof.

[0182] Provided herein is a sequencing system that can be integrated into the device described herein. A variety of sequencing methods are well known in the art and include "base calling", wherein the identity of the base in the target polynucleotide is identified. In some cases, the polynucleotides synthesized using the methods, devices, compositions and systems described herein are sequenced after cleavage from the synthesis surface. In some cases, sequencing occurs during polynucleotide synthesis or occurs simultaneously with polynucleotide synthesis, wherein base calling occurs immediately after or before the nucleoside monomer extends to the growing polynucleotide chain. The method for base calling includes measuring the current / voltage generated by the addition of bases to the template chain by polymerase catalysis. In some cases, the synthesis surface includes an enzyme such as a polymerase. In some cases, such an enzyme is tethered to an electrode or synthesis surface. In some cases, the enzyme includes terminal deoxynucleotidyl transferase or a variant thereof.

[0183] In some cases, the polynucleotides cleaved from the substrate surface or the amplified polynucleotides can be processed by techniques such as conventional or massively parallel sequencing. Sequencing can be performed by various methods available in the art, for example, methods involving the incorporation of one or more chain terminating nucleotides, for example, can be performed by, for example, the DNA sequencing kit from Applied Biosystems. In other embodiments, sequencing may include performing a next generation sequencing (NGS) method, e.g., primer extension followed by semiconductor-based detection (e.g., Ion Torrent from ThermoFisher Scientific). TM systems) or via fluorescence detection (e.g., Illumina systems).

[0184] Method and system for information retrieval

[0185] Provided herein are methods and systems for retrieving information (e.g., digital information). In some cases, provided herein are methods and systems for decoding. In some cases, the method and system decode polynucleotide sequences (e.g., polynucleotides, oligonucleotides, more than one polynucleotide, etc.). In some cases, polynucleotide sequences are encoded using methods described herein. In some cases, methods and systems include internal codecs, external codecs, or a combination thereof. In some cases, information includes one or more objects, as previously described herein. In some cases, each of one or more objects is about 1GB to about 1TB, as previously described herein. In some cases, one or more objects include information items, such as but not limited to those described herein.

[0186] In some cases, the systems and methods decode polynucleotide sequences (e.g., polynucleotides, oligonucleotides, more than one polynucleotide, etc.). Exemplary methods for retrieving digital information stored in more than one polynucleotide are described in Fig.13 In this case, Fig.12 Following the general operations illustrated in FIG. , more than one polynucleotide may have been split into more than one pool. In some cases, methods for retrieving digital information stored in more than one polynucleotide include Fig.13 One or more operations illustrated in FIG.

[0187] In some cases, retrieving digital information stored in more than one polynucleotide includes accessing index pool 1300. In some cases, accessing index pool includes sequencing the library encoding index pool completely or partially. In some instances, the index pool is encoded in the library using the system and method described herein. In some instances, the polynucleotides in the library encoding index pool are sequenced using the system and method described herein. In some cases, more than one index pool is accessed. In some cases, the polynucleotides in more than one library are sequenced. In some cases, the sequencing library is temporarily stored in a memory storage system (e.g., a flash drive). In some cases, the sequenced library is converted into digital information to retrieve the index pool. In some cases, the index pool is temporarily stored in a memory storage system (e.g., a flash drive). In some cases, the digital information in the index pool is used to search for one or more objects of interest. In some instances, one or more objects of interest are stored in a library containing more than one polynucleotide encoding one or more objects. In some instances, metadata associated with one or more objects is used to search for one or more objects of interest. In some cases, the index pool is accessed to determine more than one pool corresponding to the one or more objects. However, in some cases, one or more objects in one or more of the more than one pools may be known, and the index pool may not need to be accessed.

[0188] In some cases, one or more objects of interest are retrieved, for example, from a compartment in a storage device. In some cases, retrieving digital information stored in more than one polynucleotide includes sequencing more than one polynucleotide corresponding to one or more objects in more than one pool 1305. In some cases, more than one polynucleotide is in a library. In some cases, the library is in a compartment of a device, as previously described herein. In some cases, more than one polynucleotide in a library encoding a pool is sequenced using the systems and methods described herein. In some cases, a pool is encoded in a library using the systems and methods described herein. In some cases, more than one polynucleotide in more than one compartment is sequenced to retrieve one or more objects.

[0189] In some cases, retrieving the digital information stored in more than one polynucleotide also includes applying a decoding scheme 210. In some cases, the decoding scheme decodes the digital information in more than one pool. In some cases, the decoding scheme is applied to a library comprising the sequencing of more than one polynucleotide. In some cases, the decoding scheme includes an internal codec, ECC or a combination thereof. In some cases, the decoding scheme decodes more than one polynucleotide sequence to generate an output comprising digital information (e.g., an object). In some cases, the decoding scheme includes undoing the operation in the coding scheme. In some instances, the operation includes splitting, shuffling, splicing, transposing, translating, duplicating, marking (e.g., using an index) data or a part of data, or any combination thereof.

[0190] A method of decoding more than one polynucleotide sequence to generate an output comprising data (e.g., binary data) is described, for example, in Figure 2 Schematically illustrated in FIG. In some cases, the method for decoding more than one polynucleotide sequence can include determining more than one polynucleotide sequence 205. In some cases, determining more than one polynucleotide sequence includes sequencing the nucleotides. In some cases, the nucleotides are sequenced using the method described herein.

[0191] After sequencing more than one nucleotide, the encoded data (e.g., one or more objects) is decoded. In some cases, by way of non-limiting example, using Figure 7 The schematic diagram illustrated in FIG. decodes more than one nucleotide. The output from sequencing includes an unordered list of reads (e.g., polynucleotide sequences), such as Figure 7 As shown in .

[0192] In some cases, sequencing and / or unsorted reads are clustered after sequencing. In some cases, clustering is performed before applying an internal codec. In some cases, reads are clustered based on an index such as a frame index, a channel index, or a combination thereof. In such cases, reads are partially decoded to obtain a frame index, a channel index, or a combination thereof. In some cases, clustering is performed using a hash function, as previously described herein. In some cases, if a base in a polynucleotide sequence is determined using a hash in a coding scheme, a hash function is used, as previously described herein.

[0193] In some cases, the reads of sequencing are compared. In some cases, the polynucleotides of sequencing are compared after they are clustered. In some cases, the reads of clustering are compared. In some cases, the reads are compared before applying the internal codec. In some cases, the comparison includes using an alignment algorithm to analyze the consistency of the reads (e.g., nucleic acid or polynucleotide sequences). In some instances, the alignment algorithm includes a pairwise alignment algorithm, a multiple sequence alignment algorithm, or a combination thereof.

[0194] In some cases, the pairwise alignment algorithm includes initializing the position of each read. Initialization includes aligning the polynucleotide sequence to position 0. The consistency of the next one or more bases is analyzed between the reads. In some cases, the consistency of about 3 to about 10 reads is analyzed. In some cases, about 3 to about 4, about 3 to about 5, about 3 to about 6, about 3 to about 7, about 3 to about 8, about 3 to about 9, about 3 to about 10, about 4 to about 5, about 4 to about 6, about 4 to about 7, about 4 to about 8, about 4 to about 9, about 4 to about 10, about 5 to about 6, about 5 to about 7, about 5 to about 8, about 5 to about 9, about 5 to about 10, about 6 to about 7, about 6 to about 8, about 6 to about 9, about 6 to about 10, about 7 to about 8, about 7 to about 9, about 7 to about 10, about 8 to about 9, about 8 to about 10, or about 9 to about 10 reads are analyzed for identity. In some cases, the consistency of about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 reads is analyzed. In some cases, the consistency of at least about 3, about 4, about 5, about 6, about 7, about 8, or about 9 reads is analyzed. In some cases, the consistency of up to about 4, about 5, about 6, about 7, about 8, about 9, or about 10 reads is analyzed. In some cases, the next one or more bases include the next 2 to 10 bases. In some cases, the next one or more bases are about 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases. In some cases, the next one or more bases are at least about 2, 3, 4, 5, 6, 7, 8, or 9 bases. In some cases, the next one or more bases are at most about 3, 4, 5, 6, 7, 8, 9 or 10 bases. In some cases, the next one or more bases are about 2, 3, 4 or 5 bases. The consistency between the reads is analyzed, and it is determined whether the next one or more bases are correct. If there is consistency between the bases at the position (e.g., x) between all reads, the subsequent bases (e.g., x+1) can be analyzed subsequently. If there is an inconsistency in the base at the position (e.g., x) in the read, it is determined whether the read containing the inconsistency has an error. If there is an error), the position is increased, for example, x+1. In some cases, these steps are repeated until the end of the read is reached.

[0195] In some cases, the method (or decoding scheme) for decoding more than one polynucleotide sequence includes an internal codec. In some cases, the internal codec is applied to more than one nucleic acid (or polynucleotide) sequence. In some cases, the internal codec includes a decoding scheme. The internal codec is used to convert the polynucleotide sequence into data (e.g., digital or binary data). In some cases, the internal codec can correct deletions, substitutions or insertion errors or any combination thereof. For example, the internal codec provided herein can correct errors up to 12% deletions, 6% mutations or 2% insertions or any combination thereof. In some instances, the internal codec can correct errors up to 6% deletions, 3% mutations or 1% insertions or any combination thereof. In some instances, the internal codec can correct about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11% or 12% deletions; about 1%, 2%, 3%, 4%, 5% or 6% mutations; or about 1% or 2% insertions; or the error of any combination thereof. In some instances, the internal codec can correct up to about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11% or 12% deletion; about 1%, 2%, 3%, 4%, 5% or 6% mutation; or about 1% or 2% insertion; or errors in any combination thereof. In some other embodiments, the internal codec is used to verify the oligonucleotides and discard any suspicious oligonucleotides to avoid contaminating the external decoding. In some cases, the internal codec allows efficient decoding using indexes (frame index and channel index).

[0196] An internal codec including a decoding scheme is applied to more than one polynucleotide sequence 210. The decoding scheme of the internal codec can convert each of more than one polynucleotide sequence into a data channel. In some cases, the internal codec is applied to more than one nucleotide that has been sequenced. In some cases, the internal codec is applied to unordered reads. In some cases, after the reads or more than one nucleotide are clustered, the internal codec is applied to the reads or more than one nucleotide, as described herein. In some cases, after the reads or more than one nucleotide are aligned, the internal codec is applied to the reads or more than one nucleotide, as described herein.

[0197] In some cases, the decoding scheme can decode reads at a rate of at least about 50,000, 100,000, 150,000, or 200,000 reads per second if the software is running on an 8-core processing chip (e.g., an 8-core processor). In some examples, the decoding scheme decodes reads at a rate of at least about 100,000 reads per second (e.g., about 500 million reads per hour corresponds to about 138,000 reads per second) provided that the software is running on an 8-core processing chip (e.g., 8 cores). 9i). However, those skilled in the art will appreciate that the decoding rate may be accelerated by changing one or more hardware parameters, one or more software methods, or both. In some cases, the decoding scheme may be scaled horizontally or vertically. By way of non-limiting example, the one or more hardware parameters may include clock speed, cores, cache size, RAM size, CPU, Fig.10 Components in, or any other hardware parameters, or combinations of parameters known in the art. By way of non-limiting example, one or more software methods may include implementation using concurrent, parallel, distributed methods, or any other methods known in the art.

[0198] In some cases, the inner codec includes a decoding scheme comprising a greedy algorithm. In some cases, the inner codec includes a decoding scheme comprising a maximum likelihood (ML) algorithm. In some cases, the inner codec includes a decoding scheme comprising a hybrid greedy ML algorithm.

[0199] Figure 8 In the example, a decoding scheme including a greedy algorithm (e.g., a greedy decoder) is illustrated. As shown, the greedy algorithm considers only the transition from the most likely state when it decodes each bit position in the sequence. In some cases, each bit is guessed one at a time using the greedy algorithm. In some cases, more than one bit is guessed using the greedy algorithm at a given time. In some cases, the x-axis includes the bit position and the y-axis includes the state. In some cases, the state includes one or more valid coding states S analyzed at each bit position. In some cases, each state S is assigned a probability. In some cases, the state S is defined as the coded bit, bit history, and bit position from each channel. In some cases, the state S is defined as a bit history and a bit word. The greedy algorithm repeatedly searches for the highest possible state at each position until the highest possible end state is reached. In some cases, the decoded bit is traced back by following the highest possible state at each bit position. In some cases, this produces a fully decoded bit. In some cases, the greedy decoder finds a local optimal solution. In some cases, the local optimal solution is an approximation of the global optimal solution. Compared to other decoding schemes (such as those described herein), the greedy decoder provides a solution (or ending state) in a reasonable amount of time.

[0200] In some instances, the greedy decoder can correct errors of up to 6% deletions, 4% mutations, or 1% insertions, or any combination thereof. In some instances, the greedy decoder can correct errors of up to 3% deletions, 2% mutations, or 0.5% insertions, or any combination thereof. In some instances, the greedy decoder can correct errors of about 1%, 2%, 3%, 4%, 5%, or 6% deletions; about 1%, 2%, 3%, or 4% mutations; or about 0.5% or 1% insertions; or any combination thereof.

[0201] In some cases, the performance of the decoding scheme is improved by knowing where the polynucleotide sequence ends. In some cases, the oligonucleotide length is determined during sequencing, for example, by pair-end sequencing. In some cases, a drift term is introduced into the greedy algorithm. The drift term includes an integer associated with the total number of insertions and deletions. Each insertion is represented by a +1 value, and each deletion is represented by a -1 value. For example, if there is no insertion and there are 2 deletions, the total drift is -2. In such an example, the greedy algorithm considers all the end decoding states that do not match the oligonucleotide length as invalid and discards them. Therefore, the drift term allows the greedy algorithm to know which end decoding states are valid, and can further improve performance. Like this, in some cases, such as Figure 8 and Fig. 9 As shown in , the decoding scheme also includes a z-axis corresponding to the drift.

[0202] Fig. 9 A decoding scheme (or internal codec) including an ML algorithm (or ML decoder) is exemplarily illustrated in FIG. As shown, the ML algorithm considers transitions from all states when decoding each bit position in the sequence. The states are defined as previously described herein. In some cases, each bit is guessed one at a time using the ML algorithm. In some cases, more than one bit is guessed using the ML algorithm at a given time. In some cases, the ML algorithm repeatedly finds all transition states at each position until an end candidate state is determined. In some cases, the x-axis includes the bit position and the y-axis includes the state, as previously described herein. In some cases, as previously described herein, a drift term is used to filter the end candidate state. In some cases, the ML algorithm provides a global optimal solution by tracking all state transitions. In some cases, compared to other decoding schemes (such as those described herein), the ML algorithm is computationally intensive.

[0203] In some instances, the ML decoder can correct errors of up to 12% deletions, 6% mutations, or 2% insertions, or any combination thereof. In some instances, the ML decoder can correct errors of up to 6% deletions, 3% mutations, or 1% insertions, or any combination thereof. In some instances, the ML decoder can correct errors of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, or 12% deletions; about 1%, 2%, 3%, 4%, 5%, or 6% mutations; or about 0.5%, 1%, 1.5%, or 2% insertions; or any combination thereof. In some instances, the ML decoder can correct errors of up to about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, or 12% deletions; about 1%, 2%, 3%, 4%, 5%, or 6% mutations; or about 0.5%, 1%, 1.5%, or 2% insertions; or any combination thereof.

[0204] In some cases, the decoding scheme of the inner codec includes a hybrid greedy ML algorithm. In some cases, the hybrid greedy ML algorithm includes a greedy algorithm and an ML algorithm. The hybrid greedy ML algorithm considers transitions from more than one state when it decodes each bit position in the sequence. In some cases, the more than one state is about 100 to about 1000 states when decoding each bit position in the sequence. In some cases, the more than one state is about 100 to about 200, about 100 to about 300, about 100 to about 400, about 100 to about 500, about 100 to about 600, about 100 to about 700, about 100 to about 800, about 100 to about 900, about 100 to about 1,000, about 200 to about 300, about 200 to about 400, about 200 to about 500, about 200 to about 600, about 200 to about 700, about 200 to about 800, about 200 to about 900, about 200 to about 1,000, about 300 to about 400, about 300 to about 500, about 300 to about 600, about 300 to about 700, about 300 to about 800, about 300 to about 900, about 200 to about 1,000, about 300 to about 400, about 300 to about 500, about 300 to about 600, about 300 to about 700, about 300 to about 800, about 300 From about 400 to about 500, from about 400 to about 600, from about 400 to about 700, from about 400 to about 800, from about 400 to about 900, from about 400 to about 1,000, from about 500 to about 600, from about 500 to about 700, from about 500 to about 800, from about 400 to about 900, from about 400 to about 1,000, from about 500 to about 600, from about 500 to about 700, from about 500 to about 800, from about 500 to about 900, About 500 to about 1,000, about 600 to about 700, about 600 to about 800, about 600 to about 900, about 600 to about 1,000, about 700 to about 800, about 700 to about 900, about 700 to about 1,000, about 800 to about 900, about 800 to about 1,000, or about 900 to about 1,000 states. In some cases, the more than one state is about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, or about 1,000 states. In some cases, the more than one state is at least about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, or about 900 states. In some cases, the more than one state is at most about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, or about 1,000 states. The states are defined as previously described herein. In some cases, each bit is guessed one at a time using a hybrid greedy ML algorithm.In some cases, more than one bit is guessed at a given time using the hybrid greedy ML algorithm. In some cases, the hybrid greedy ML algorithm repeatedly finds about 100 to about 1000 transition states at each position until an end candidate state is determined. In some cases, a drift term is used to filter the end candidate states, as previously described herein. In some cases, the hybrid greedy ML algorithm provides a global optimal solution while being computationally less expensive relative to other decoding schemes (such as the ML algorithm described herein).

[0205] In some instances, the hybrid greedy ML decoder can correct errors of up to 15% deletions, 10% mutations, or 5% insertions, or any combination thereof. In some instances, the hybrid greedy ML decoder can correct errors of up to 12% deletions, 6% mutations, or 2% insertions, or any combination thereof. In some instances, the hybrid greedy ML decoder can correct errors of up to 6% deletions, 3% mutations, or 1% insertions, or any combination thereof. In some instances, the hybrid greedy ML decoder can correct errors of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15% deletions; about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% mutations; or about 1%, 2%, 3%, or 4% insertions; or any combination thereof. In some examples, the hybrid greedy ML decoder can correct errors of up to about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15% deletions; about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% mutations; or about 1%, 2%, 3%, or 4% insertions; or any combination thereof.

[0206] In some cases, the decoding scheme in the inner codec includes a beam search decoder or a random sampling decoder (e.g., a pure sampling decoder, a top-K sampling decoder, etc.). In some cases, the beam search decoder or the random sampling decoder provides a diversity of candidate states compared to a greedy decoder.

[0207] In some cases, the internal codec also includes a checksum. In some cases, the checksum is used to verify data integrity, detect errors or combinations thereof. In some cases, a checksum function or a checksum algorithm (e.g., parity check bytes or parity check work (longitudinal parity check), and complement code, position correlation, fuzzy checksum, etc.) is used to generate a checksum. Examples of checksum functions or algorithms include, but are not limited to, BSD checksums (Unix), SYSV checksums (Unix), sum4, sum8, sum16, sum32, fletcher-4, fletcher-8, fletcher-16, fletcher-32, Adler-32, xor8, Luhn algorithm, Verhoeff algorithm or Damm algorithm. In some cases, instead of only taking the highest possible path, some best possible paths are considered and tested for checksum. In some cases, the checksum includes an RS code (e.g., a small RS code). In this case, the decoder provides a list of possibilities (e.g., "list decoding"), assuming that the user can decide which one it actually is.

[0208] In some cases, the method and system for decoding includes arranging channels into frames. In some cases, the decoded channels from the internal codec are arranged into frames based on the channel index and the frame index 215. In some cases, one or more channels are missing from the frame, such as Figure 7 In some cases, the channel is missing due to the error that occurs during the synthesis or sequencing of nucleotides. In some cases, about 1% to about 10% of the channel is missing from the frame. In some cases, about 1% to about 2%, about 1% to about 4%, about 1% to about 6%, about 1% to about 8%, about 1% to about 10%, about 2% to about 4%, about 2% to about 6%, about 2% to about 8%, about 2% to about 10%, about 4% to about 6%, about 4% to about 8%, about 4% to about 10%, about 6% to about 8%, about 6% to about 10% or about 8% to about 10% of the channel is missing from the frame. In some cases, about 1%, about 2%, about 4%, about 6%, about 8% or about 10% of the channel is missing from the frame. In some cases, at least about 1%, about 2%, about 4%, about 6% or about 8% of the channel is missing from the frame. In some cases, up to about 2%, about 4%, about 6%, about 8%, or about 10% of the channels are missing from the frame.

[0209] In some cases, the internal codec includes a "format". In some cases, there is no a priori information about the size of the data (e.g., binary data) during decoding. Therefore, in some cases, frame index 0 includes the size of the data. In some cases, after arranging the channels into frames and / or sorting the frames, frame 0 is decoded first. Data is then extracted from frame 0 to reject frames that exceed the expected data size (e.g., from an erroneously decoded oligo).

[0210] In some cases, the internal codec includes a hash (e.g., SHA-256). In some cases, the hash verifies that the data was decoded correctly. In some cases, encoding and decoding are performed as a stream by using a hash at the end (after the external codec or ECC). In some cases, this can limit memory usage to temporary buffers only.

[0211] The method for decoding more than one polynucleotide sequence includes an external codec or an error correction code (ECC). In some cases, more than one polynucleotide sequence is decoded into data (e.g., binary data). In some cases, external codec or ECC are applied to each 220 in the frame. In some cases, external codec or ECC are applied to the channel from the internal codec. In some cases, external codec or ECC is applied after the channel from the internal codec is arranged into a frame.

[0212] In some cases, the outer codec includes an error correction scheme, or the code is based on an error correction scheme for encoding data (e.g., binary data). In some cases, the error correction scheme includes Reed-Solomon (RS) codes, LDPC codes, polar codes, turbo codes, or any combination thereof.

[0213] In some cases, the error correction scheme of the outer codec includes a Reed-Solomon (RS) code. In this case, the RS decoder receives a codeword r(x), which is the original codeword c(x) plus an error e(x) (e.g., r(x)=c(x)+e(x)). In some cases, the error e(x) is 0. In some cases, the RS decoder attempts to identify the location and size of up to t errors (or 2t erasures). The RS code then attempts to correct these identified errors and / or erasures.

[0214] In some cases, the RS decoder includes a syndrome calculation. In some cases, the syndrome calculation includes receiving incoming symbols and dividing them into a generator polynomial g(x), as previously described herein. In some cases, the syndrome is calculated by substituting the 2t roots (or syndrome of the RS codeword c(x)) of the generator polynomial g(x) into r(x). In some cases, the generator polynomial g(x) is a known parameter of the decoder. In some cases, the RS codeword c(x) has a 2t syndrome that depends on the error.

[0215] In some cases, the RS decoder includes finding the symbol error position. In some cases, parity check or check symbol t makes the syndrome calculation zero in the absence of error. In some cases, parity check or check symbol t includes the remainder in the RS decoder. If there is an error, the resulting polynomial g (x) is passed to the Euclidean algorithm. In some cases, the factor of the remainder is found using the Euclidean algorithm. In some cases, the result is evaluated on the iteration for each incoming symbol. In some cases, errors are found and errors are corrected. In some cases, the correction codeword c (x) is output from the RS decoder. In some cases, there are more errors (e.g., e (x)>2t) that can be corrected than the RS code in the codeword. In this case, the codeword r (x) received is output from the RS decoder. In some cases, the received codeword r (x) is output together with the indication (e.g., sign) of error correction failure. In some cases, the received codeword r (x) is discarded (e.g., a channel or frame including binary data as described herein).

[0216] In some cases, frames from an external codec (or ECC) are merged to generate an output including data 225. In some cases, the data includes binary data, which can be a byte stream or a byte array, as previously described herein.

[0217] The decoding methods described herein (e.g., internal codecs, external codecs, or both) can be used to recover data when there is an error in at least one polynucleotide sequence in more than one stored polynucleotide sequence. In some cases, the error includes an insertion, a deletion, a substitution, or any combination thereof. In some cases, data is recovered when there is an error (e.g., an error rate) in about 0.001% to about 30% of the polynucleotide sequences in more than one nucleotide. In some cases, in the presence of about 0.001% to about 0.01%, about 0.001% to about 0.1%, about 0.001% to about 0.5%, about 0.001% to about 1%, about 0.001% to about 2%, about 0.001% to about 5%, about 0.001% to about 10%, about 0.001% to about 15%, about 0.001% to about 20%, about 0.001% to about 25%, about 0.001% to about 30%, about 0.01% to about 0.1%, about 0.01% to about 0.5%, about 0.01% to about 1%, about 0.01% to about 2%, about 0.01% to about 5%, about 0.01% to about 10%, about 0.01% to about 15%, about 0.01% to about 20%, about 0.01% to about 25%, about 0.01% to about 30%, about 0.1% to about 0.5%, about 0.1% to about 1%, about 0.1% to about 2%, about 0.1% to about 5%, about 0.1% to about 10%, about 0.1% to about 15%, about 0.1% to about 20%, about 0.1% to about 25%, about 0.1% to about 30%. % to about 30%, about 0.5% to about 1%, about 0.5% to about 2%, about 0.5% to about 5%, about 0.5% to about 10%, about 0.5% to about 15%, about 0.5% to about 20%, about 0.5% to about 25%, about 0.5% to about 30%, about 1% to about 2%, about 1% to about 5%, about 1% to about 10%, about 1% to about 15%, about 1% to about 20%, about 1% to about 25%, about 1% to about 30%, about 2% to about 5%, about 2% to about 10%, about 2% to about 15%, about 2 In some cases, data is recovered in the presence of an error rate of about 0.001%, about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 10% to 15%, about 10% to 20%, about 10% to 25%, about 10% to 30%, about 15% to 20%, about 15% to 25%, about 15% to 30%, about 20% to 25%, about 20% to 30%, or about 25% to 30%. In some cases, data is recovered in the presence of an error rate of about 0.001%, about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%.In some cases, data is recovered in the presence of an error rate of at least about 0.001%, about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, or about 25%. In some cases, data is recovered in the presence of an error rate of at most about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%.

[0218] In some cases, decoding schemes (e.g., external decoding and internal decoding) are used together with soft decoding. Soft decoding generally refers to decoding by considering the range of possible values ​​(e.g., using probability estimation). As an example, sequencing can carry the quality of each base, which can be considered during probability calculation. In such an example, each state includes a final probability, which can be used as, for example, log-likelihood in the external decoder if the external decoder supports soft decoding. In addition, clustering and alignment can provide soft information about alignment confidence. As another example, the LDPC external codec includes an iterative decoder. This provides the possibility of going back and forth between internal and external decoders in an iterative manner rather than a single pass. However, in some cases, this is accompanied by the cost of higher computational requirements.

[0219] Decoding can be run on at least one logic element, programmable logic or processor. Non-limiting examples of at least one logic element, programmable logic or processor include programmable logic controller (PLC), programmable logic array (PLA), programmable array logic (PAL), general logic array (GLA), complex programmable logic decision (CPLD), field programmable gate array (FPGA) or application specific integrated circuit (ASIC), GPU, CPU, AI accelerator or any combination thereof. In some cases, AI accelerators include Google-TPU, Graphcore, Cerebras, SambaNova or a combination thereof. In some cases, decoding is based on compute-on-memory technology (such as but not limited to UpMem) operation.

[0220] The hash of the present disclosure can allow the verification of digital information during retrieval. In some cases, retrieving the digital information stored in more than one polynucleotide also includes verifying at least one or more objects 1315. In some cases, the first one or more hashes in more than one pool are used to verify one or more objects. In some cases, retrieving the digital information stored in more than one polynucleotide also includes verifying one or more pool items. In some cases, the second one or more hashes in more than one pool are used to verify one or more pool items. In some instances, if an object is stored across more than one pool in more than one pool, more than one pool item is assembled into the object. In such instances, the first one or more hashes of the data payload of each of the pool items, the second one or more hashes of one or more objects, or a combination thereof enable appropriate assembly verification.

[0221] Verifying a hash typically includes generating a hash (e.g., a cryptographic hash). Verification may also include comparing the generated hash to a previously determined hash. In some cases, the same hash function is used to determine the previous hash and the new hash. In some cases, the hash function includes a cryptographic hash function. In some cases, the hash function includes MD-5, SHA-1, SHA-2, SHA-3, RIPEMD-160, Whirlpool, BLAKE, BLAKE2, BLAKE3, or a variant thereof. In some cases, the hash function includes SHA-2. In some instances, SHA-2 includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256. In some cases, if the new hash matches the previous hash, the integrity of the information item (e.g., the object) is verified. In some cases, if the new hash does not match the previous hash, the verification fails. In some cases, if the verification fails, the integrity of the information item is not verified. In some cases, if the verification fails, the information item has been modified and / or corrupted.

[0222] Retrieving digital information may include combining information stored across pool items and / or more than one pool. In some cases, retrieving digital information stored in more than one polynucleotide also includes combining digital information in more than one pool 1320. In some cases, data payloads in one or more pool items are combined. In some cases, data payloads in one or more pool items across more than one pool are combined. In some cases, the combined data payloads include digital information. In some cases, the retrieved data or digital information is stored on a memory 1325.

[0223] In some cases, the retrieved digital information is presented to the user. In some cases, the information is presented to the user on an interface. In some cases, the interface is an interface of an electronic device (e.g., a personal electronic device). In some cases, the electronic device includes an application configured to communicate with the system described herein via a computer network to access the information.

[0224] In some cases, the methods described herein for decoding more than one polynucleotide sequence to generate an output comprising digital data (e.g., binary data) are performed on a system. In some cases, the system performs Figure 1 , Figure 2 or both of the operations generally illustrated in the figure. In some cases, such a system includes an apparatus comprising a memory, a sequencing device, a processing device operably coupled to the memory, or a combination thereof. In some cases, the sequencing device is operably coupled to the memory, the processing device, or a combination thereof. In some cases, the memory is used to store information of binary data, polynucleotide sequences, or a combination thereof. In some cases, the information of binary data, polynucleotide sequences, or a combination thereof comes from one or more steps in the encoding method described herein. In some cases, the memory is used to store information (e.g., software code, parameters, executable instructions, etc.) related to the algorithm described herein. In some instances, the memory may include any suitable memory described herein. In some instances, the memory may be configured according to the embodiments described herein. In some instances, the sequencing device is configured to determine more than one polynucleotide sequence using the method described herein.

[0225] In some cases, the processing device is configured to perform one or more decoding steps. In some cases, the processing device is configured to perform one or more steps including the following: the internal codec including the decoding scheme is applied to more than one polynucleotide sequence; the channel of the binary data is arranged into a frame based on the channel index and the frame index in each channel of the binary data; and the external codec is applied to the frame. In some cases, the decoding scheme converts each of more than one polynucleotide sequence into a channel of binary data. In some cases, the decoding scheme includes a hybrid decoding algorithm, and the hybrid decoding algorithm includes a greedy algorithm and a maximum likelihood (ML) algorithm. In some cases, the external codec includes an error correction scheme. In some cases, the frame from the external codec is merged to generate an output including binary data.

[0226] Methods for retrieving digital information in DNA (or polynucleotides) can be performed on the system. In some cases, the system performs Fig.12 , Fig.13or both. In some cases, such a system includes an apparatus that includes one or more processing units, memory, instructions, sequencing devices, or a combination thereof. In some cases, the memory communicates with the one or more processing units. In some cases, the instructions are stored on the memory. In some cases, the sequencing device communicates with the memory, one or more processing units, or a combination thereof. In some cases, the one or more processing units and the memory are distributed across one or more physical or logical locations.

[0227] In some cases, the memory is used to store data or digital information, polynucleotide sequences (e.g., partially or completely decoded sequences), or a combination thereof. In some cases, the memory is used to store information related to the algorithms described herein (e.g., software code, parameters, executable instructions, etc.). In some instances, the memory may include any suitable memory described herein. In some instances, the memory may be configured according to the embodiments described herein. In some instances, the sequencing device is configured to determine more than one polynucleotide sequence using the methods described herein.

[0228] In some cases, one or more processing units include a central processing unit (CPU), a graphics processing unit (GPU), a single-core processor, a multi-core processor, a processor cluster, an application-specific integrated circuit (ASIC), a programmable circuit (such as a field programmable gate array (FPGA)), an AI accelerator, and any combination of its variants. In some cases, one or more processing units include a single instruction multiple data (SIMD) or a single program multiple data (SPMD) parallel architecture. As an example, one or more processing units include one or more GPUs or CPUs that implement SIMD or SPMD. In some cases, AI accelerators include Google-TPU, Graphcore, Cerebras, SambaNova, or a combination thereof. In some embodiments, in addition to hardware implementations, one or more of the processing units are implemented in software and / or firmware. The software or firmware implementation of the processing unit may include computer or machine executable instructions written in any suitable programming language to perform the various functions described herein. The software implementation of one or more processing units may be stored in whole or in part in memory. Alternatively or additionally, the system may include one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc. In some cases, decoding is performed based on compute-on-memory technology (such as, but not limited to, UpMem).

[0229] In some cases, one or more processing units are configured to perform one or more decoding steps. In some cases, the processing device is configured to perform one or more steps including: applying a decoding scheme to decode digital information in more than one pool; using the first one or more hashes in more than one pool to at least verify the data payload in the pool item; combining the digital information in more than one pool to retrieve one or more objects; and storing the digital information on the memory. In some cases, one or more processing units are configured to perform one or more steps including: applying an internal codec to more than one polynucleotide; or applying ECC to more than one polynucleotide. In some cases, the internal codec converts each of more than one polynucleotide into digital information. In some cases, the internal codec includes a hybrid decoding algorithm, which includes a greedy algorithm and a maximum likelihood (ML) algorithm. In some cases, the output from the ECC is merged to generate an output including digital information.

[0230] DNA Data Storage

[0231] The polynucleotides encoding the information described herein can be stored in a data storage system. The system for data storage may include one or more modules. In some cases, some or all of the one or more modules are in communication. In some instances, some or all of the one or more modules are in communication to allow polynucleotides to be transferred between them. In some instances, some or all of the one or more modules are fluidically coupled. In some instances, some or all of the one or more modules are fluidically coupled with one or more tubes. Fluids can generally refer to one or more liquids used in various processes involving the treatment of polynucleotides, including but not limited to synthesis, amplification, sequencing preparation and sequencing. In some instances, some or all of the modules communicate to allow control commands to be transmitted between the modules of the system. In some instances, some or all of the one or more modules are electrically coupled. The modules in the system may include but are not limited to synthesizer units, amplification chambers, sequencer units, storage units, controllers, robotic systems, or any combination thereof. In some instances, the modules may also include a fluid source, a database or a file system, or both. In some instances, the database or file system tracks the storage capacity of the system. In some embodiments, the database or file system can be used to track the rack (or tray), slot (for capsule) or both. In some instances, the database or file system is used to determine the layout of the rack in the storage system. In some cases, the movement of polynucleotides between one or more modules of the system is completed by one or more pipes or robotic systems. In some instances, the database or file system is used to guide the robotic system to the correct position in the storage system. In some cases, the system is automated.

[0232] Fig.16 16. A non-limiting example of a system for data storage is illustrated in FIG. A system for data storage may include a synthesizer unit 1610. The synthesizer unit may be used to synthesize more than one polynucleotide encoding digital information. In some cases, the system includes more than one synthesizer unit 1610. Polynucleotides may be synthesized using the method provided herein or any other suitable synthesis method known in the art. The fluid and / or electronic control of the polynucleotide synthesis in the synthesizer unit 1610 may be performed by a controller 1635. In some cases, the electronic device in the synthesizer unit 1610 communicates with the controller 1635. In some cases, the synthesizer unit 1610 has an input for receiving a DNA sequence. In some cases, the synthesizer unit 1610 has an input for receiving a fluid for polynucleotide synthesis. In some cases, the synthesizer unit 1610 has an output for eluting the synthesized polynucleotides. In some cases, the synthesized polynucleotides are transferred to another component of the system, such as, by way of non-limiting example, a storage unit, an amplification chamber, or a sequencing unit.

[0233] The synthesizer unit may include a solid support. The solid support may include a device for polynucleotide storage described herein. The solid support may include a surface for polynucleotide synthesis. In some cases, the solid support, the surface, or both include materials described herein. In some cases, the material includes a metal or an organic polymer. In some cases, the material includes steel (e.g., stainless steel) or other metal alloys. In some cases, the material includes polyethylene, polypropylene, or other polymers. In some cases, the structure includes a flexible material, such as those provided herein. Exemplary flexible materials include, but are not limited to, modified nylon, unmodified nylon, nitrocellulose, and polypropylene. In some cases, the material includes a rigid material, such as those provided herein. Exemplary rigid materials include, but are not limited to, glass, fused silica, silicon, silicon dioxide, silicon nitride, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and blends thereof) and metals (e.g., steel, gold, platinum). In some cases, the material disclosed herein may be made of a material comprising silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), glass, or any combination thereof. In some examples, the materials disclosed herein are made with combinations of the materials listed herein or any other suitable materials known in the art.

[0234] In some cases, polynucleotides are deprotected, cracked and / or eluted from synthesizer unit 1610, and transferred to another module in the system. In some cases, robot system 1630 or fluid pipe is used to transport polynucleotides to another module in the system. Robot system 1630 can be controlled by controller 1635. Robot system generally includes a system for manipulating more than one polynucleotide. In some cases, robot system is used to manipulate a structure comprising more than one polynucleotide, such as those described herein. By way of non-limiting example, manipulation can include moving, storing, retrieving, processing, transferring or any combination thereof. Robot system can be similar to those used for moving wafers and trays of chips between processing devices in semiconductor processing. Robot system 1630 can be used to select and transfer polynucleotides between modules of the system. For example, robot system 1635 can include a label reader to verify the structure in storage unit 1615. In some cases, robot system 1635 includes a label reader, and the structure in storage unit 1615 includes a label (e.g., a barcode or RFID tag). After verification, robot system 1630 can transfer the structure to the component of the system. Additionally, the robotic system 1630 can transfer structures to precise locations in components of the system. In some cases, the robotic system can allow polynucleotides to be added and / or removed from modules in the data storage system. In some cases, the robotic system allows structures containing more than one polynucleotide to be placed in and / or retrieved from locations in an identifiable layout in the storage unit 1615. As further described herein, the controller 1635 can be used to control the robotic system 1630.

[0235] In some cases, one or more droplets containing polynucleotides are transferred from synthesizer unit 1610 to storage unit 1615. In some cases, some or all of the polynucleotides synthesized on the solid support are transferred to a structure for storage. The structure or compartment can have various shapes and sizes. The structure can also include a label (e.g., a barcode or RFID tag). In some cases, more than one polynucleotide is transferred to a structure in synthesizer unit 1610. In some cases, more than one polynucleotide is transferred to a structure in storage unit 1615. The fluid and / or electronic control of the polynucleotide synthesis in storage unit 1615 can be performed by controller 1635. In some cases, the electronic device in storage unit 1615 communicates with controller 1635. In some cases, the polynucleotides are stored in storage unit 1615 at room temperature. In some cases, the system includes a database or file system for tracking the storage capacity in storage unit 1615. In some instances, the database includes a control application database. In some cases, the database or file system is part of controller 1635.

[0236] Structures containing more than one polynucleotide may be stored in storage unit 1615 in an identifiable layout. The identifiable layout may include a rack or more than one rack or a variation thereof. A rack may be used to hold one or more structures containing more than one polynucleotide. In some cases, each structure is stored at a fixed position in the identifiable layout. In some cases, the tag includes information about the location of the structure in the identifiable layout. As an example, the tag may encode metadata including the location of the structure in the identifiable layout. In some cases, the rack may be located in a data center. In some cases, the rack uses a mechanical structure that is commonly used to install conventional computing and data storage resources in a rack unit. For example, a rack may include an opening suitable for supporting a disk drive, a processing blade, and / or other computer devices. In some cases, the rack includes a tag. In some instances, the tag includes information about the structure stored in / on the rack. In some instances, the tag includes a list of structures stored in / on the rack.

[0237] In some cases, the storage unit 1615 can be accessed using a robotic system 1630. In some cases, the identifiable layout in the storage unit 1615 includes robotically addressable slots. Each slot can hold a structure comprising more than one polynucleotide. In some cases, each slot comprises a width, depth, length, or any combination thereof for accommodating a structure comprising more than one polynucleotide. In some cases, a rack comprises more than one slot, wherein each slot holds a structure comprising more than one polynucleotide.

[0238] The system for storing polynucleotides can also include an amplification chamber 1620. The amplification unit can be used to amplify more than one polynucleotide. In some cases, the system includes more than one amplification chamber 1620. In some cases, a structure is selected from a storage unit 1615, and the polynucleotide from the structure is transferred to the amplification chamber 1620. In some cases, the polynucleotide from the synthesizer unit 1610 is transferred to the amplification chamber 1620 for size selection, PCR or other types of amplification or preparation for storage. Size selection generally involves selecting DNA of a target size and rejecting much shorter or much longer chains. In some cases, a filter is tuned to capture DNA of a specific size range. In some cases, other methods include PCR, electrophoresis, primer capture by a solid phase bond complementary to the terminal sequence of the synthesized oligonucleotide, or use isothermal polymerase. The fluid and / or electronic control of the polynucleotide synthesis in the amplification chamber 1620 can be performed by a controller 1635. In some cases, the electronic device in the amplification chamber 1620 communicates with the controller 1635.

[0239] The system for storing polynucleotides can also include a sequencing unit 1625. The sequencing unit 1625 can be used to sequence more than one polynucleotide. In some cases, more than one polynucleotide is transferred from the amplification chamber 1620 to the sequencing unit 1625. In some cases, the system can include additional modules for performing additional sequencing preparation steps. In some instances, more than one polynucleotide is transferred from the amplification chamber 1620 to the sequencing unit 1625 using one or more tubes or a robotic system 1630. In some cases, the amplification chamber 1620 and the sequencing unit 1625 are fluidically coupled. The fluid and / or electronic control of the polynucleotide synthesis in the sequencing unit 1625 can be performed by a controller 1635. In some cases, the electronic device in the sequencing unit 1625 communicates with the controller 1635.

[0240] In some cases, the system includes large-scale sequencing of polynucleotides. In some cases, large-scale sequencing includes dense and highly parallel sequencers. In some cases, the system includes more than one sequencing unit 1625. In some cases, sequencing unit 1625 uses centrifugal force and / or vacuum / pressure to add or empty reagents from sequencing unit 1625 to sequencing unit 1625. In some cases, sequencing unit 1625 is based on light (e.g., with light source and sensor on chip), based on nanopore (e.g., Oxford Nanopore Technologies (ONT)) or relates to other operations (e.g., based on light methods, such as PacBio or other sequencing technologies). In some cases, sequencing unit 1625 adopts sequencing methods provided herein. In some cases, sequencing unit 1625 uses nanopore or benefits from other electrical sequencing technologies of bulk fluids provided by semiconductor manufacturing equipment. In some cases, one or more modules described herein include camera. Camera can be used for capturing one or more optical features of polynucleotides in module. As an example, a camera can be used in a synthesizer unit, a sequencing unit, or both, to capture optical characteristics of polynucleotides attached to a surface on a solid support as described herein.

[0241] The system for storing polynucleotides may include a robotic system 1630 as described herein. The robotic system can generally be used to manipulate the polynucleotides in the system. Manipulation may include, but is not limited to, moving, storing, retrieving, processing, transferring, or any combination thereof. In some cases, the robotic system transfers more than one polynucleotide between modules in the system. In some instances, the robotic system manipulates (e.g., transfers) more than one polynucleotide in a structure for storage as described herein. In some cases, the robotic system manipulates (e.g., transfers) more than one polynucleotide in a rack. In some instances, the rack includes more than one structure, each structure including a label. In some instances, the rack includes more than one solid support for synthesis and / or sequencing. In some cases, the robotic system includes a robotic hand or a robotic picker. In some cases, the robotic system 1630 is fully integrated with the storage system control software and / or firmware in the controller 1635. In some cases, the robotic system 1630 is fully integrated with an external host application. In some cases, the robotic system 1630 is fully automated.

[0242] The system for storing polynucleotides can include controller 1635. Controllers can be used to control modules, components, fluids, robots or any combination thereof generally. Modules, components, fluids, electronic devices, robots or any combination thereof can be used for synthesis, storage, retrieval, sequencing and / or amplification of polynucleotides. In some cases, controller 1635 can catalog all storage structures loaded, unloaded and / or stored in a rack. Polynucleotides can encode digital information as described herein. Modules, components, fluids, electronic devices, robots or any combination thereof can be used for execution methods, models or algorithms, such as encoding or decoding polynucleotides.

[0243] In some cases, controller 1635 controls the physical position of more than one polynucleotide.In some cases, controller 1635 provides command to one or more modules of system.In some instances, controller 1635 controls robot (for example, robot system 1630), actuator and fluid valve or any other device of system.In some cases, controller 1635 allows synchronization and control for processing and / or transfer polynucleotide module.In some instances, via fluid processing and / or transfer polynucleotide.In some instances, via electronic device processing and / or transfer polynucleotide.In some cases, controller 1635 controls physical parameter in one or more modules, such as but not limited to pressure, vacuum, temperature, volume (for example, fluid) or its any combination.

[0244] In some cases, the controller 1635 calls an encoder module or a decoder module. In some cases, the encoder module encodes the digital information into more than one polynucleotide. In some cases, the encoder module converts one or more codecs (such as those described herein (e.g., Figure 1 , Figure 3-Figure 6 , Fig.12 , Figure 13-Figure 14 )) is applied to digital information. In some cases, the decoder module decodes the sequence of more than one polynucleotide to retrieve the digital information. In some cases, the decoder module combines one or more codecs (such as those described herein (e.g., Figure 2 , Figure 7-Figure 9 , Fig.13)) is applied to the sequence of more than one polynucleotide. In some cases, the decoding module performs reorganization, error correction and outputs digital information (e.g., binary data). In some cases, the output comprising digital information is transmitted to an operating system and / or a file system. Output can be provided on a display (such as a graphical user interface (GUI)) or any other suitable display (such as those described herein) for providing digital information. In some cases, controller 1635 is implemented in one or more software modules, such as those described herein. In some cases, controller 1635 responds to commands from an operating system, such as those described herein.

[0245] Some definitions

[0246] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter belongs.

[0247] Throughout the present disclosure, numerical features are presented in range format. It should be understood that the description in range form is only for convenience and simplicity, and should not be interpreted as an unchangeable limitation to the scope of any embodiment. Therefore, unless the context clearly stipulates otherwise, the description of the range should be considered to have specifically disclosed all possible sub-ranges and the individual numerical values ​​in the range, up to one tenth of the lower limit unit. For example, the description of the range should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual numerals in the range, such as 1.1, 2, 2.3, 5 and 5.9. Regardless of the width of the range, this is applicable. The upper and lower limits of these intermediate ranges can be included independently in a smaller range, and are also included in the present invention, subject to any specifically excluded limits in the range. In the case where the range includes one or two of the limits, unless the context clearly stipulates otherwise, the scope of any one or two excluded in those included limits is also included in the present invention.

[0248] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit any embodiment. Unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an" and "the" are intended to also include the plural forms. It should also be understood that the terms "comprises" and / or "comprising" when used in this specification specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the listed related items.

[0249] References throughout this specification to "some instances," "other instances," or "a particular instance" mean that a particular feature, structure, or characteristic described in conjunction with that instance is included in at least one instance. Thus, the phrases "in some instances" or "in other instances" or "in a particular instance" that appear in various places throughout this specification do not necessarily all refer to the same instance. Furthermore, in one or more instances, the particular features, structures, or characteristics may be combined in any suitable manner.

[0250] Unless specifically stated or obvious from the context, as used herein, the term "about" referring to a number or a range of numbers should be understood to mean the stated number and a number + / - 10% thereof, or a number that is 10% below the listed lower limit and 10% above the listed upper limit for the value listed for the range.

[0251] As used herein, the terms "preselected sequence", "predefined sequence" or "predetermined sequence" are used interchangeably. These terms mean that the sequence of a polymer is known and selected prior to the synthesis or assembly of the polymer. In particular, various aspects of the invention are described herein, primarily with respect to the preparation of nucleic acid molecules, in which the sequence of a polynucleotide is known and selected prior to the synthesis or assembly of the nucleic acid molecules.

[0252] As used herein, the term "hash" or "hashes" may generally refer to a fixed-length string output from a hash function. A hash function may generally include a function that receives an input of arbitrary length into an output having a fixed length. In some cases, the input may be one or more bits, which may be passed through a hash function to generate a hash. In some cases, a hash function may be deterministic, and it may not be feasible to reverse engineer the input from the hash output. The action of feeding an input into a hash function may be referred to as "hashing."

[0253] As used herein, the term "symbol" generally refers to a representation of a digital information unit. Digital information can be segmented or converted into one or more symbols. In an example, a symbol can be a bit, and the bit can have a numerical value. In some examples, a symbol can have a value of '0' or '1'. In some examples, digital information can be represented as a sequence of symbols or a string of symbols. In some examples, a sequence of symbols or a string of symbols can include binary data.

[0254] Unless otherwise indicated, the polynucleotide sequence described herein can include DNA or RNA or its analogs or derivatives. As used herein, the term nucleic acid, nucleotide, polynucleotide, oligonucleotide (oligonucleotides), oligonucleotide (oligos), oligonucleic acid (oligonucleic acids) are used synonymously in the full text to represent the polymer of nucleoside monomers. As used herein, the term nucleic acid sequence, polynucleotide sequence (polynucleotide sequences), oligonucleotide sequence (oligonucleotides sequences), oligonucleotide sequence (oligo sequences) or oligonucleic acid sequence (oligonucleic acid sequences) are also used synonymously in the full text to represent the sequence of the polymer of nucleoside monomers. In some cases, nucleic acid is connected via phosphate or sulfur-containing bonds. In some cases, nucleic acid includes DNA, RNA, atypical nucleic acid, non-natural nucleic acid or other nucleosides. In some cases, nucleotides include atypical bases, sugar or other parts. In some cases, nucleotides include terminators configured to prevent extension reactions. In some cases, such terminators are removed before subsequent nucleotides are added to the growing chain.

[0255] Computing System

[0256] refer to Fig.10, a block diagram depicting an exemplary machine including a computer system 1000 (e.g., a processing or computing system) is shown within which a set of instructions may be executed to cause an apparatus to perform or implement any one or more of the aspects and / or methods for static code scheduling of the present disclosure. Fig.10 The components in the description are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or combination of two or more such components to implement a particular embodiment.

[0257] Including Fig.10 The computer system platform shown can be used to encode data represented as a set of symbols into another set of symbols. In some cases, the computer system uses a program to convert a first string of symbols into a second string of symbols. In some cases, the computer system executes a program to convert data into more than one polynucleotide sequence, convert more than one polynucleotide sequence into data, or both. As an example, Fig.10 The computing system generally illustrated in FIG. 1 can be used to execute one or more software programs for encoding a string of symbols (representing an item of information) into a polynucleotide sequence (e.g., Figure 1 , Figure 3-Figure 6 , Fig.12 or Figure 14-15 ), decoding the polynucleotide sequence back into a symbol string (e.g., Figure 2 , Figure 7-Figure 9 or Fig.13 More specifically, the data can be represented as digital symbols, such as binary values ​​"0" and "1", and the computer system 1000 can execute a computer program (e.g., an internal codec, an external codec, or both) to convert the data into more than one polynucleotide sequence. In some cases, the computer system executes a program to convert a first one or more polynucleotide sequences into a second one or more polynucleotide sequences.

[0258] The platform for encoding data may also include one or more components, such as a synthesizer, a sequencer, a storage unit, or any combination thereof. In such a platform, in some cases, the computer system 1000 communicates electronically with any of the one or more components, such as a synthesizer, a sequencer, a storage unit, or any combination thereof. In some cases, one or more components are operably linked to the computer system and are optionally automated locally or remotely by the computer. In various cases, the methods and systems described herein also include software programs for the operation of one or more components of the platform on the computer system and their use. Therefore, synchronization of dispensing / vacuum / refilling functions such as coordinating and synchronizing material deposition, device movement, dispensing actions, and vacuum actuation are all within the scope of the disclosure provided herein. In some cases, the computer system is programmed to engage between a user-specified base sequence and the position of the material deposition device to deliver the correct building blocks and / or reagents to a specified area (e.g., a specific site) of the substrate. In addition, such as Fig.10 The computer system of the system shown in can be used for monitoring one or more components in the platform. For example, the computer system can be used for monitoring one or more sensor data from the sensor integrated in the component or connected to the component. In some cases, the computer system uses a program to monitor and detect irregularities of one or more parameters, such as pressure, volume, flow, temperature, vacuum, orientation angle, humidity or any other physical parameter that can be measured in the system and platform described herein. The computer system including the program can analyze the pattern in one or more sensor data, and if any irregularity is detected or if any data or data combination falls outside a threshold value (e.g., predetermined or dynamic threshold), the user is optionally warned by the HMI.

[0259] The program can be executed on a computer system provided herein. In some cases, the program includes a statistical algorithm or a machine learning algorithm. In some cases, the algorithm including machine learning (ML) i...

Claims

1. A method for encoding data represented by more than one symbol in more than one polynucleotide sequence, comprising: (a) splitting data into more than one frames, wherein each of the more than one frames comprises a frame index; (b) applying an outer codec to each of the more than one frames, wherein the outer codec includes an error correction scheme; (c) dividing each frame into more than one channel, wherein each channel of the more than one channel comprises a channel index; (d) shuffling each channel based at least in part on the channel index; and (e) applying an inner codec to encode each channel in the polynucleotide sequences of the more than one polynucleotide sequences. 2 . The method of claim 1 , wherein the data comprises binary data, wherein the binary data comprises a byte stream or a byte array.

3. The method of claim 1, wherein the shuffling in (d) comprises a rotation scheme within each channel.

4. The method of claim 1, wherein the shuffling in (d) comprises a pseudo-random process within each channel.

5. The method of claim 1, wherein the shuffling in (d) provides resistance to errors. The method according to claim 5 , wherein the error is a nucleotide synthesis error or a sequencing error. The method of claim 5 , wherein the error comprises a deletion, an insertion or a substitution.

8. The method of claim 1, wherein the error correction scheme comprises a Reed-Solomon (RS) code, a low-density parity-check (LDPC) code, a turbo code, a polar code, or any combination thereof.

9. The method of claim 1, wherein the data comprises at least about 1 GB to about 1 TB of data.

10. The method of claim 1, wherein the more than one frame comprises about 100 to about 10,000 frames.

11. The method of claim 1, wherein each frame includes up to approximately 5000 channels.

12. The method of claim 1, wherein each channel comprises about 100 bits to about 300 bits.

13. The method of claim 1, wherein the frame index comprises about 16 bits to about 20 bits.

14. The method of claim 1, wherein the channel index comprises approximately 12 bits or approximately 16 bits.

15. The method of claim 1, wherein the polynucleotide sequence is about 100 to about 300 bases in length.

16. The method of claim 1, wherein the frame index and / or the channel index is appended to the front of each channel before (d).

17. The method of claim 1, wherein applying the internal codec comprises adding redundancy across the more than one polynucleotide sequences.

18. The method of claim 17, wherein the redundancy is about 5% to about 10%.

19. The method of claim 17, wherein the more than one polynucleotide sequences are capable of being decoded in the presence of errors due in part to the redundancy across the more than one polynucleotide sequences.

20. The method of claim 19, wherein the error comprises an insertion, a deletion, a substitution, or any combination thereof.

21. The method of claim 1, wherein applying the inner codec comprises: (a) combining symbols from the channel, symbol history, and symbol positions; and (b) Generate candidate bases using a lookup table, hashing, or both.

22. The method of claim 21, further comprising performing a base duplication check.

23. The method of claim 21, further comprising updating the symbol history, increasing the channel index, increasing the frame index, or any combination thereof.

24. The method of claim 23, wherein the updated symbol history, the incremented channel index, the incremented frame index, or a combination thereof is combined with symbols of a subsequent channel.

25. The method of claim 21, further comprising performing GC filtering prior to synthesizing said more than one said polynucleotide sequences.

26. The method of claim 25, wherein the GC filtering comprises removing about 5% to about 10% of the more than one channels.

27. The method of claim 1, wherein the more than one polynucleotide sequences comprise a GC content of about 45% to about 55%.

28. The method of claim 1, wherein at least 90% of the more than one polynucleotide sequences comprise a GC content of about 45% to about 55%.

29. The method of claim 1, wherein applying the inner codec comprises: (a) Generate candidate bases for each symbol in the channel using a lookup table; and (b) selecting a next lookup table based at least in part on a previously encoded symbol.

30. A method for decoding more than one polynucleotide sequence to generate an output comprising data represented by more than one symbol, comprising: (a) determining said more than one polynucleotide sequences; (b) applying an inner codec to the more than one polynucleotide sequences, wherein the inner codec converts each of the more than one polynucleotide sequences into a channel comprising more than one symbol, wherein the inner codec comprises a hybrid decoding algorithm comprising a greedy algorithm and a maximum likelihood (ML) algorithm; (c) arranging channels of data into frames based on a channel index and a frame index for each channel; and (d) applying an outer codec to the frame, wherein the outer codec includes an error correction scheme, Wherein frames from said outer codec are merged to generate an output comprising said data.

31. The method of claim 30, further comprising clustering the polynucleotide sequences before (b).

32. The method of claim 31 , wherein the clustering is based on an index.

33. The method of claim 32, wherein clustering comprises partially decoding the frame index, the channel index, or both.

34. The method of claim 31 , wherein the clustering is performed using a hash function.

35. The method of claim 30, further comprising aligning the polynucleotide sequences before (b).

36. The method of claim 35, wherein aligning comprises analyzing the identity of the nucleotides using an alignment algorithm.

37. The method of claim 36, wherein the alignment algorithm comprises a pairwise alignment algorithm, a multiple sequence alignment algorithm, or a combination thereof.

38. The method of claim 36, wherein the alignment algorithm comprises: (a) initializing the position of each of the more than one reads, wherein the initialization comprises aligning the polynucleotide sequence to position 0; (b) analyzing the identity of the next one or more bases between each read segment; (c) determining for each read whether each of the next one or more bases is correct or erroneous; (d) incrementing the position given the determination for each read; and (e) Repeat steps (b)-(d).

39. The method of claim 38, wherein the more than one read comprises about 3 to about 10 reads.

40. The method of claim 38, wherein each read is about 100 to about 300 bases in length.

41. The method of claim 38, wherein the next one or more bases are about 2, 3, 4, or 5 bases.

42. The method of claim 30, wherein the hybrid decoding algorithm comprises decoding based on transition probabilities from one or more states.

43. The method of claim 42, wherein the one or more states include about 100 to about 1000 most likely states.

44. The method of claim 42, wherein the inner codec further comprises a drift term.

45. The method of claim 44, wherein the drift term comprises an integer.

46. ​​The method of claim 45, wherein the integer is related to the total number of insertions or deletions in the polynucleotide sequence.

47. The method of claim 46, wherein the integer is calculated by adding the values ​​of insertions and / or deletions in the total number of insertions or deletions.

48. The method of claim 47, wherein the value of the insertion comprises +1 and the value of the deletion comprises -1.

49. The method of claim 30, wherein (c) comprises deshuffling the channels based on the channel index and grouping the channels into frames based on the frame index.

50. The method of claim 30, wherein the error correction scheme comprises a Reed-Solomon (RS) code, a low-density parity-check (LDPC) code, a turbo code, a polar code, or any combination thereof.

51. The method of claim 30, wherein at least one of the more than one polynucleotide sequences comprises an error.

52. The method of claim 51, wherein the error comprises an insertion, a deletion, a substitution, or any combination thereof.

53. An apparatus comprising: (a) Memory; (b) a processing device operatively coupled to the memory, wherein the processing device is configured to: (i) splitting the data into more than one frames, wherein each of the more than one frames comprises a frame index; (ii) applying an outer codec to each of the more than one frames, wherein the outer codec includes an error correction scheme; (iii) dividing each frame into more than one channel, wherein each channel of the more than one channel comprises a channel index; (iv) shuffling each channel based at least in part on the channel index; and (v) Applying an internal codec to encode each channel in a polynucleotide sequence.

54. An apparatus comprising: (a) Memory; (b) a sequencing apparatus configured to determine the sequence of more than one nucleotide; and (c) a processing device, the processing device being operatively coupled to the memory and the sequencing device, wherein the processing device is configured to: (i) applying an inner codec to the sequences, wherein the inner codec converts each of the sequences into a channel comprising more than one symbol, wherein the inner codec comprises a hybrid decoding algorithm, the hybrid decoding algorithm comprising a greedy algorithm and a maximum likelihood (ML) algorithm; (ii) arranging the channels into frames based on the channel index and the frame index in each channel; and (iii) applying an outer codec to the frame, wherein the outer codec includes an error correction scheme, Wherein frames from said outer codec are merged to generate an output comprising said data.

55. A method for encoding data in a polynucleotide sequence, comprising: (a) generating an inner codec including a secret book, wherein the secret book is optimized with respect to one or more constraints; and (b) applying the inner codec to encode the data into more than one polynucleotide sequence.

56. The method of claim 55, wherein the one or more constraints are related to nucleic acid synthesis, post-processing, storage, sequencing, or any combination thereof.

57. The method of claim 56, wherein the nucleic acid synthesis comprises electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing, or any combination thereof.

58. The method of claim 56, wherein the one or more constraints associated with nucleic acid synthesis include synthesis errors.

59. The method of claim 58, wherein the synthesis error comprises an insertion, a deletion, or a mutation.

60. The method of claim 56, wherein post-processing comprises one or more of ligation, cleavage, hybridization, denaturation, immobilization to a solid support, extension, error correction, enrichment, separation, purification, or amplification.

61. The method of claim 56, wherein the storing comprises cold data storage.

62. The method of claim 56, wherein the storage comprises storage of the nucleic acid in a liquid phase or a solid phase.

63. The method of claim 56, wherein one or more constraints associated with storage include temperature, humidity, pressure, salinity, pH, concentration, time, light, UV, O2, or any combination thereof.

64. The method of claim 63, wherein the temperature comprises room temperature.

65. The method of claim 56, wherein sequencing comprises next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, sequencing by synthesis, Sanger sequencing, or any combination thereof.

66. The method of claim 55, further comprising (c) synthesizing more than one polynucleotide comprising the more than one polynucleotide sequences.

67. The method of claim 55, wherein the secret book comprises codewords generated based in part on a sequence of bases.

68. The method of claim 67, wherein the base sequence includes predetermined base transitions.

69. The method of claim 55, wherein the inner codec comprises two or more codebooks.

70. The method of claim 69, wherein each of the two or more secret books encodes a layer during synthesis of the more than one polynucleotide.

71. The method of claim 70, wherein the layer comprises each of the more than one polynucleotides extended by at least one base.

72. The method of claim 71, wherein synthesis of the layer comprises one or more cycles, wherein each of the one or more cycles comprises flowing bases according to one or more base transitions of the codebook.

73. The method of claim 72, wherein a cycle in the one or more cycles comprises adding one or more of A, T, C, or G.

74. The method of claim 69, wherein each of the two or more secret books comprises a different sequence of bases.

75. The method of claim 55, wherein the secret book includes approximately 12 codewords.

76. The method of claim 55, wherein (b) comprises mapping the data to more than one polynucleotide sequence based on the secret book.

77. The method of claim 55, wherein the internal codec is further optimized for one or more constraints including length, GC content, repeats, errors, or any combination thereof of the more than one polynucleotide sequences.

78. The method of claim 55, wherein 40% to 60% of the more than one polynucleotide sequences encode redundancy.

79. The method of claim 55, wherein synthesizing comprises a plurality of cycles of synthesis.

80. The method of claim 79, wherein the number of synthesis cycles is reduced compared to the number of synthesis cycles required to synthesize a polynucleotide sequence not encoded using the internal codec.

81. The method of claim 80, wherein the reduced number of synthesis cycles is based in part on a flow sequence.

82. The method of claim 80, wherein the number of synthesis cycles is reduced by at least 30%.

83. The method of claim 80, wherein the number of synthesis cycles is reduced by 50%.

84. The method of claim 80, wherein for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 300.

85. The method of claim 80, wherein for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is about 155.

86. The method of claim 84, wherein the polynucleotide sequence comprises one or more of A, T, C, or G.

87. The method of claim 66, wherein (c) comprises synthesizing the more than one polynucleotides on a solid support.

88. The method of claim 87, wherein the solid support comprises more than one feature.

89. The method of claim 88, wherein more than 25% of the more than one features are deblocked in each synthesis cycle.

90. The method of claim 88, wherein at least 50% of the more than one features are deblocked in each synthesis cycle.

91. The method of claim 55, wherein each of the more than one polynucleotide sequences has the same length.

92. The method of claim 55, wherein 80% to 100% of the more than one polynucleotide sequences are of the same length.

93. The method of claim 55, further comprising sequencing the more than one polynucleotides to produce more than one output sequence.

94. The method of claim 55, wherein the more than one output sequences are decoded using a greedy algorithm, a maximum likelihood (ML) algorithm, or a hybrid greedy ML algorithm.

95. The method of claim 55, wherein the more than one output sequences are decoded based at least in part on calculating a probability of error.

96. The method of claim 95, wherein the error comprises a deletion, an insertion, a mutation, or any combination thereof.

97. A hybrid organic-computer simulation platform for encoding data, the platform comprising: (a) A computing system comprising at least one processor and instructions executable by the at least one processor to perform operations comprising: (i) generating an inner codec comprising a codebook, wherein the codebook is optimized with respect to one or more constraints; and (ii) applying said inner codec to encode said data into more than one polynucleotide sequence; and (b) a synthesizer for producing more than one polynucleotide comprising the more than one polynucleotide sequences.

98. The platform of claim 97, wherein the one or more constraints are related to nucleic acid synthesis, post-processing, storage, sequencing, or any combination thereof.

99. The platform of claim 98, wherein the nucleic acid synthesis comprises electrochemical synthesis, enzymatic synthesis, phosphoramidite synthesis, inkjet printing, or any combination thereof.

100. The platform of claim 98, wherein the one or more constraints associated with nucleic acid synthesis include synthesis errors.

101. The platform of claim 100, wherein the synthesis error comprises an insertion, a deletion, or a mutation.

102. The platform of claim 98, wherein post-processing comprises one or more of ligation, cleavage, hybridization, denaturation, fixation to a solid support, extension, error correction, enrichment, separation, purification, and amplification.

103. The platform of claim 98, wherein the storage comprises cold data storage.

104. The platform of claim 98, wherein storage comprises storage of nucleic acids in a liquid phase or a solid phase.

105. The platform of claim 98, wherein one or more constraints associated with storage include temperature, humidity, pressure, salinity, pH, concentration, time, light, UV, O2, or any combination thereof.

106. The platform of claim 105, wherein the temperature comprises room temperature.

107. The platform of claim 98, wherein sequencing comprises next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, sequencing by synthesis, Sanger sequencing, or any combination thereof.

108. The platform of claim 97, wherein the computing system comprises a cloud computing system.

109. The platform of claim 108, wherein the cloud computing system comprises a private cloud, a public cloud, a hybrid cloud, a multi-cloud, or any combination thereof.

110. The platform of claim 108, wherein the cloud computing system comprises infrastructure as a service (IaaS), platform as a service (PaaS), software as a service (SaaS), or any combination thereof.

111. The platform of claim 97, wherein the secret book comprises codewords generated based in part on a sequence of bases.

112. The platform of claim 97, wherein the base sequence includes predetermined base transitions.

113. The platform of claim 97, wherein the internal codec comprises two or more codebooks.

114. The platform of claim 113, wherein each of the two or more secret books encodes a layer during synthesis of the more than one polynucleotide.

115. The platform of claim 114, wherein the layer comprises each of the more than one polynucleotides extended by at least one base.

116. The platform of claim 115, wherein synthesis of the layer comprises one or more cycles, wherein each of the one or more cycles comprises flowing bases according to the one or more base transitions of the compact book.

117. The platform of claim 116, wherein a cycle in the one or more cycles comprises adding one or more of A, T, C, or G.

118. The platform of claim 113, wherein each of the two or more secret books comprises a different sequence of bases.

119. The platform of claim 97, wherein the instructions further cause the synthesizer to generate the more than one polynucleotides.

120. The platform of claim 97, further comprising a sequencer for sequencing the more than one polynucleotides to generate more than one output sequence.

121. The platform of claim 120, wherein the instructions further cause the computing system to receive the more than one output sequences.

122. The platform of claim 120, wherein the computing system further performs operations comprising: (iii) decoding the more than one output sequences.

123. The platform of claim 122, wherein the more than one output sequences are decoded using a greedy algorithm, a maximum likelihood (ML) algorithm, or a hybrid greedy ML algorithm.

124. The platform of claim 122, wherein decoding the more than one output sequences is based at least in part on calculating probabilities of deletions, insertions, mutations, or any combination thereof.

125. The platform of claim 97, further comprising a storage unit for storing the more than one polynucleotide.

126. The platform of claim 125, wherein the operations further comprise transferring the more than one polynucleotides between the synthesizer, the sequencer, the storage unit, or any combination thereof.

127. The platform of claim 97, wherein specific base transitions allow synthesis to proceed according to a flow order.

128. The platform of claim 97, wherein the secret book comprises 12 codewords.

129. The platform of claim 97, wherein (a)(ii) comprises mapping binary data to more than one polynucleotide sequence based on the secret book.

130. The platform of claim 97, wherein the internal codec is further optimized for constraints including length, GC content, repeats, errors, or any combination thereof of the more than one polynucleotide sequences.

131. The platform of claim 97, wherein 40% to 60% of the more than one polynucleotide sequences encode redundancy.

132. The platform of claim 97, wherein producing the more than one polynucleotide comprises multiple cycles of synthesis.

133. The platform of claim 132, wherein the number of synthesis cycles is reduced compared to the number of synthesis cycles required to synthesize a polynucleotide sequence not encoded using the internal codec.

134. The platform of claim 133, wherein the reduced number of synthesis cycles is based in part on a flow sequence.

135. The platform of claim 133, wherein the number of synthesis cycles is reduced by at least 30%.

136. The platform of claim 133, wherein the number of synthesis cycles is reduced by 50%.

137. The platform of claim 133, wherein for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is less than 300.

138. The platform of claim 133, wherein for a polynucleotide sequence comprising 100 bases, the number of synthesis cycles is 155.

139. The platform of claim 138, wherein the polynucleotide sequence comprises one or more A, T, C, or G.

140. The platform of claim 97, wherein generating the more than one polynucleotide comprises base-by-base synthesis.

141. The platform of claim 97, wherein the synthesizer comprises a solid support comprising more than one feature.

142. The platform of claim 141, wherein each of the more than one features is independently addressable by one or more electrodes of the solid support.

143. The platform of claim 141, wherein each of the more than one features is addressable by masking.

144. The platform of claim 143, wherein the masking comprises a physical barrier.

145. The platform of claim 143, wherein the masking comprises controlling reactivity at one or more of the more than one features.

146. The platform of claim 145, wherein controlling reactivity comprises deprotection at one or more of the more than one features.

147. The platform of claim 146, wherein the deprotection comprises acid generation.

148. The platform of claim 146, wherein the deprotection comprises electrochemical deprotection.

149. The platform of claim 141, wherein greater than 25% of the more than one features are deblocked in each synthesis cycle.

150. The platform of claim 141, wherein at least 50% of the more than one features are deblocked in each synthesis cycle.

151. The platform of claim 97, wherein each of the more than one polynucleotide sequences has the same length.

152. The platform of claim 97, wherein 80% to 100% of the more than one polynucleotide sequences have the same length.

153. A system for storing data in DNA, the system comprising: one or more processing units; a memory in communication with the one or more processing units, instructions stored in the memory and executed on the one or more processing units, the instructions causing the system to: generating more than one pool, wherein each of the more than one pool comprises a pool descriptor, a pool entry including a payload of the data, and a peer descriptor; determining a first one or more hashes of the payload of each pool entry; and An encoding scheme is applied to encode the more than one pools as sequences of more than one polynucleotide.

154. A system according to claim 153, wherein the data includes one or more objects.

155. The system of claim 154, wherein instructions stored in the memory and executed on the one or more processing units cause the system to determine a second one or more hashes for each of the one or more objects.

156. The system of claim 153, wherein the one or more objects comprise a file or metadata associated with the file.

157. The system of claim 153, wherein the pool descriptor comprises a version, a pool ID, a list of pool entry descriptors, or any combination thereof.

158. The system of claim 157, wherein the pool ID comprises a unique ID.

159. The system of claim 158, wherein the unique ID comprises a universally unique identifier (UUID) or a content ID.

160. The system of claim 157, wherein the list of pool item descriptors comprises a path to an object, a size of an object, a scope of the pool item within an object, an offset of the pool item in a pool, or any combination thereof.

161. The system of claim 153, wherein each of the one or more pool entries further comprises a hash of a pool entry from the first one or more hashes, or a combination thereof.

162. A system according to claim 153, wherein the end pool descriptor includes a list of object descriptors.

163. The system of claim 162, wherein the list of object descriptors comprises a path to an object, a hash of the object from the first one or more hashes, or a combination thereof.

164. The system of claim 153, wherein each of the more than one pools is about 1 GB to about 1 TB.

165. The system of claim 153, wherein the more than one pools include redundant pools.

166. The system of claim 155, wherein the first one or more hashes, the second one or more hashes, or both are determined using a hash module.

167. The system of claim 165, wherein the hashing module executes on the one or more processing units.

168. The system of claim 153, wherein the first one or more hashes require less memory than the one or more objects.

169. The system of claim 153, wherein the second one or more hashes require less memory than the one or more pool entries.

170. The system of claim 165, wherein the hash module comprises a hash function.

171. The system of claim 170, wherein the hash function comprises SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256.

172. The system of claim 153, wherein the instructions further cause the system to generate one or more index pools.

173. The system of claim 172, wherein the one or more index pools comprise an index pool descriptor and a list of object indexes.

174. The system of claim 173, wherein the index pool descriptor comprises a version, a pool ID, a size of the pool, and a timestamp.

175. The system of claim 174, wherein the pool ID comprises a unique ID.

176. The system of claim 175, wherein the unique ID comprises a UUID or a content ID.

177. The system of claim 173, wherein the list of object indexes comprises a path to an object, a hash of an object, a list of object fragments, a list of object metadata, or any combination thereof.

178. The system of claim 177, wherein the list of object fragments comprises a pool ID of a pool containing fragments, a range of fragments, or a combination thereof.

179. The system of claim 177, wherein the list of object metadata comprises a type of metadata, a metadata payload, or a combination thereof.

180. A system according to claim 179, wherein the types of metadata include a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof.

181. The system of claim 172, wherein each of the one or more index pools is about 1 GB to about 1 TB.

182. A system according to claim 153, wherein the instructions stored in the memory and executed on the one or more processing units cause the system to retrieve the data stored in the DNA.

183. The system of claim 182, wherein the instructions include: applying a decoding scheme to decode the sequence of the more than one polynucleotides in each of the more than one pools; and At least the payload of each pool entry is verified using the first one or more hashes.

184. A device for storing information in DNA, comprising: One or more compartments, wherein each compartment comprises: (a) a library comprising more than one polynucleotide, wherein the library encodes a pool comprising information corresponding to one or more objects; and (b) a medium for storing said more than one polynucleotide.

185. The device of claim 184, wherein the one or more compartments are connected.

186. The device of claim 184, wherein the one or more compartments are not interconnected.

187. The device of claim 184, wherein the medium comprises a solid, a liquid, a gas, or any combination thereof.

188. An apparatus according to claim 184, wherein the medium comprises a salt solution having a molar ratio of salt cations to phosphate groups in the DNA of less than 20:

1.

189. The apparatus of claim 188, wherein the salt solution is dried to produce a dry product.

190. The device of claim 184, further comprising a solid support comprising a surface.

191. The device of claim 184, further comprising more than one structure located on the surface, wherein the more than one polynucleotides extend from the more than one structure.

192. The apparatus of claim 184, wherein the one or more objects comprise a file or metadata associated with the file.

193. The apparatus of claim 184, wherein the pool comprises a pool descriptor, one or more pool items, and an end pool descriptor.

194. The apparatus of claim 193, wherein the pool descriptor comprises a version, a pool ID, a list of pool entry descriptors, or any combination thereof.

195. The device of claim 194, wherein the pool ID comprises a unique ID.

196. The device of claim 195, wherein the unique ID comprises a universally unique identifier (UUID) or a content ID.

197. An apparatus according to claim 194, wherein the list of pool item descriptors includes a path to an object, a size of an object, a scope of the pool item within an object, an offset of the pool item in a pool, or any combination thereof.

198. The apparatus of claim 184, wherein each of the one or more pool entries comprises a data payload, a hash of the pool entry, or a combination thereof.

199. An apparatus according to claim 184, wherein the end pool descriptor comprises a list of object descriptors.

200. The apparatus of claim 199, wherein the list of object descriptors comprises a path to an object, a hash of an object, or a combination thereof.

201. The apparatus of claim 184, wherein the pool comprises between about 1 GB and about 1 TB of digital information.

202. The device of claim 184, further comprising one or more second compartments, wherein each of the one or more second compartments comprises a second library encoding an index pool.

203. An apparatus according to claim 202, wherein the index pool comprises an index pool descriptor and a list of object indexes.

204. The apparatus of claim 203, wherein the index pool descriptor comprises a version, a pool ID, a size of the pool, and a timestamp.

205. The device of claim 204, wherein the pool ID comprises a unique ID.

206. The device of claim 205, wherein the unique ID comprises a UUID or a content ID.

207. An apparatus according to claim 203, wherein the list of object indexes includes a path to the object, a hash of the object, a list of object fragments, a list of object metadata, or any combination thereof.

208. The apparatus of claim 207, wherein the list of object fragments comprises a pool ID of a pool containing fragments, a range of fragments, or a combination thereof.

209. The apparatus of claim 207, wherein the list of object metadata comprises a type of metadata, a metadata payload, or a combination thereof.

210. An apparatus according to claim 209, wherein the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof.

211. The device of claim 203, wherein the index pool is about 1 GB to about 1 TB.

212. A method for storing data in more than one polynucleotide, comprising: generating more than one pool, wherein each of the more than one pool comprises a pool descriptor, a pool entry comprising a payload of the data, and an end descriptor; determining a first one or more hashes of the payload of each pool entry; and An encoding scheme is applied to encode the more than one pools as a sequence of more than one nucleotide.

213. A method according to claim 212, wherein the data includes one or more objects.

214. The method of claim 213, wherein the method further comprises determining a second one or more hashes for each of the one or more objects.

215. The method of claim 212, further comprising storing the more than one polynucleotide.

216. The method of claim 213, wherein the polynucleotides of the more than one polynucleotides corresponding to each of the more than one pools are stored in separate containers of a data storage system.

217. The method according to claim 212 also includes producing more than one polynucleotide.

218. The method of claim 217, wherein generating the more than one polynucleotides comprises phosphoramidite-based synthesis of deoxyribonucleic acid (DNA).

219. The method of claim 217, wherein the reagents used in the phosphoramidite-based synthesis comprise a nucleoside phosphoramidite, an oxidizing agent, an activating agent, or a deblocking agent, or a solvent comprising acetonitrile.

220. The method of claim 217, wherein producing the more than one polynucleotides comprises enzymatic DNA synthesis.

221. The method of claim 220, wherein the reagents for enzymatic DNA synthesis comprise terminal deoxynucleotidyl transferase (TdT) or a deblocking agent, or a solvent comprising water.

222. The method of claim 212, wherein the one or more objects comprise a file or metadata associated with the file.

223. The method of claim 212, wherein the pool descriptor comprises a version, a pool ID, a list of pool entry descriptors, or any combination thereof.

224. The method of claim 223, wherein the pool ID comprises a unique ID.

225. The method of claim 224, wherein the unique ID comprises a universally unique identifier (UUID) or a content ID.

226. A method according to claim 223, wherein the list of pool item descriptors includes a path to an object, a size of an object, a scope of the pool item within an object, an offset of the pool item in a pool, or any combination thereof.

227. The method of claim 212, wherein each of the one or more pool entries further comprises a hash of a pool entry from the first one or more hashes.

228. A method according to claim 227, wherein the end pool descriptor includes a list of object descriptors.

229. The method of claim 228, wherein the list of object descriptors comprises a path to an object, a hash of the object from the first one or more hashes, or a combination thereof.

230. The method of claim 212, wherein each of the more than one pools is about 1 GB to about 1 TB.

231. The method of claim 212, wherein the more than one pools include redundant pools.

232. The method of claim 213, wherein the first one or more hashes, the second one or more hashes, or both are determined using a hash module.

233. The method of claim 213, wherein the second one or more hashes require less memory than the one or more objects.

234. The method of claim 212, wherein the first one or more hashes require less memory than the one or more pool entries.

235. The method of claim 232, wherein the hash module comprises a hash function.

236. The method of claim 235, wherein the hash function comprises SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256.

237. The method according to claim 212 also includes creating one or more index pools.

238. A method according to claim 237, wherein the one or more index pools include an index pool descriptor and a list of object indexes.

239. The method of claim 238, wherein the index pool descriptor comprises a version, a pool ID, a size of the pool, and a timestamp.

240. The method of claim 239, wherein the pool ID comprises a unique ID.

241. The method of claim 240, wherein the unique ID comprises a UUID or a content ID.

242. The method of claim 238, wherein the list of object indexes comprises a path to an object, a hash of an object, a list of object fragments, a list of object metadata, or any combination thereof.

243. The method of claim 242, wherein the list of object fragments comprises a pool ID of a pool containing fragments, a range of fragments, or a combination thereof.

244. The method of claim 242, wherein the list of object metadata comprises a type of metadata, a metadata payload, or a combination thereof.

245. A method according to claim 244, wherein the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof.

246. The method of claim 237, wherein each of the one or more index pools is about 1 GB to about 1 TB.

247. A method for retrieving data stored in more than one polynucleotide, comprising: determining the sequences of the more than one polynucleotides, wherein the more than one polynucleotides are in more than one pool; applying a decoding scheme to decode the sequences of the more than one polynucleotides in each of the more than one pools, wherein each pool comprises a pool descriptor, a pool entry including a payload of the data, and an end descriptor; and At least the payload of each pool entry is verified using the first one or more hashes.

248. A method according to claim 247, wherein the data includes one or more objects.

249. A method according to claim 248, wherein the one or more objects include a file or metadata associated with the file.

250. The method of claim 248, wherein the method further comprises verifying the one or more objects using a second one or more hashes.

251. The method of claim 247, wherein verifying at least the payload comprises verifying the first one or more hashes using a hash function.

252. The method of claim 247, further comprising combining the payload from each pool item to retrieve the data.

253. The method according to claim 247 also includes storing the data in a memory.

254. The method of claim 247, wherein each of the more than one pools is about 1 GB to about 1 TB.

255. The method of claim 250, wherein verifying the one or more objects comprises verifying the second one or more hashes using a hash function.

256. A method according to claim 251, wherein the hash function comprises SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224 or SHA-512 / 256.

257. The method of claim 247, wherein determining the sequence comprises sequencing the more than one polynucleotides.

258. The method of claim 257, wherein sequencing comprises next generation sequencing, parallel sequencing, single molecule real-time sequencing, nanopore sequencing, synthetic sequencing, Sanger sequencing, or any combination thereof.

259. The method of claim 248, further comprising accessing an index pool in the one or more index pools to determine more than one pool that includes the one or more objects.

260. A method according to claim 259, wherein the index pool includes an index pool descriptor and a list of object indexes.

261. The method of claim 260, wherein the index pool descriptor comprises a version, a pool ID, a size of the pool, and a timestamp.

262. The method of claim 261, wherein the pool ID comprises a unique ID.

263. The method of claim 262, wherein the unique ID comprises a UUID or a content ID.

264. A method according to claim 260, wherein the list of object indexes includes a path to the object, a hash of the object, a list of object fragments, a list of object metadata, or any combination thereof.

265. The method of claim 264, wherein the list of object fragments comprises a pool ID of a pool containing fragments, a range of fragments, or a combination thereof.

266. A method according to claim 264, wherein the list of object metadata includes the type of metadata, the metadata payload, or a combination thereof.

267. A method according to claim 266, wherein the type of metadata includes a list of keywords attached to the object, a thumbnail, a text summary, an ID range of a sorted keyword value database, a timestamp, a version, or any combination thereof.

268. The method of claim 259, wherein each of the one or more index pools is about 1 GB to about 1 TB.

Citation Information

Patent Citations

  • Method of selectively masking one or more sites on a surface and a method of synthesising an array of molecules

    US10195580B2

  • Heated nanowells for polynucleotide synthesis

    US10894242B2

  • DNA-based digital information storage with sidewall electrodes

    US10936953B2

  • Compositions and methods for in vivo synthesis of unnatural polypeptides

    US20220243244A1

  • Electrochemical deblocking solution for electrochemical oligomer synthesis on an electrode array

    US9267213B1