Multi-layer error correction code for DNA data storage

By employing a multi-layered error correction code configuration in DNA data storage, the problem of insufficient robustness of error correction codes in existing technologies is solved, storage efficiency and decoding efficiency are improved, and robust protection of DNA data is achieved.

CN121368754APending Publication Date: 2026-01-20WESTERN DIGITAL TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480040689.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-11
Filing Date
2024-05-24
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing DNA data storage technologies are not robust and effective enough in terms of error correction codes, especially when storing and retrieving data in oligonucleotide groups, where it is difficult to effectively correct errors.

Method used

A multi-layer error correction code configuration is adopted to encode data units into multiple symbols and allocate these symbols in an oligonucleotide set. Multi-layer codewords, including redundant data and cyclic redundancy check values, are generated using error correction codes. Encoding and decoding are performed by an encoder and a decoder to achieve robust protection of the data.

Benefits of technology

It improves the storage efficiency of oligonucleotide pools, reduces decoding complexity, and enhances decoding efficiency through selective and iterative decoding, thereby improving the reliability and stability of DNA data storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121368754A_ABST
    Figure CN121368754A_ABST
Patent Text Reader

Abstract

Example systems and methods for DNA data storage using multi-layer error correction codes distributed among oligonucleotides are described. The data unit may be encoded as a set of codewords, where each codeword is distributed as a symbol on a different oligonucleotide. The codeword may include: a first layer codeword set, the first layer codeword set including CRC and ECC redundant data; and one or more additional layer codewords, the one or more additional layer codewords including permutation data and corresponding ECC redundancy data. The decoding may include a series of decoding iterations between the first codeword layer and the additional codeword layer.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims the benefit of the entire contents of U.S. Non-Provisional Application No. 18 / 464,494, entitled “Multi-Tier Error Correction Codes for DNA Data Storage,” and filed on September 11, 2023, in the United States Patent and Trademark Office, and which is hereby incorporated by reference in its entirety for all purposes. TECHNICAL FIELD

[0002] The present disclosure relates to deoxyribonucleic acid (DNA) data storage. In particular, the present disclosure relates to error correction of data stored as a set of synthetic DNA oligonucleotides. BACKGROUND

[0003] DNA is a promising technology for information storage. It has the potential for ultra-high density 3D storage, large storage capacity, and long service life. Currently, DNA synthesis technology provides tools for synthesizing and manipulating relatively short synthetic DNA strands (oligonucleotides). For example, some oligonucleotides can include 40 to 350 bases, encoding two times the number of bits in a configuration that maps to four DNA nucleotides or sequences thereof.

[0004] Similar to other data storage technologies, binary data can be encoded using various techniques before being stored in oligonucleotides, and various decoding and error correction techniques can be applied after the data stored in oligonucleotides is read back into binary data. Because the payload capacity of oligonucleotides is relatively short, Reed-Solomon error correction codes have been applied to individual oligonucleotides to enable error correction on a per-oligonucleotide basis. Other schemes have been proposed for applying larger and more complex error correction codes to data from a group of oligonucleotides, such as a group of oligonucleotides storing a particular data object.

[0005] DNA decoding can be divided into two stages. In the first stage, deletion and insertion errors are eliminated using correlation analysis within individual oligonucleotides. In the second stage, error correction codes can be applied to regular data decoding. For example, Reed-Solomon decoding can be applied to each oligonucleotide or larger block sizes based on multiple oligonucleotides.

[0006] There is a need for techniques to apply more robust and effective error correction codes to DNA data storage and retrieval. BRIEF DESCRIPTION OF DRAWINGS

[0007] The technology presented herein is illustrated in the accompanying drawings, which are by way of example and not limitation.

[0008] Figure 1A is a block diagram of a prior art DNA data storage process.

[0009] Figure 1B is a block diagram of a prior art DNA data storage decoding process for oligonucleotides encoded with binary data.

[0010] Figure 2 is a block diagram of an example encoding system and an example decoding system for DNA data storage using multi-layer error-correcting codes (ECCs).

[0011] Figure 3 is a diagram of an example matrix for constructing codewords in a pool of oligonucleotides for multi-layer error-correcting codes.

[0012] Figure 4 is a diagram of an example configuration for oligonucleotide pool encoding.

[0013] Figure 5 is a diagram of an example configuration for oligonucleotide pool decoding.

[0014] Figure 6 is a block diagram of an example method for encoding data units in a pool of oligonucleotides using multi-layer ECCs.

[0015] Figure 7 is a block diagram of an example method for decoding data units from a pool of oligonucleotides using multi-layer ECCs.

[0016] Figure 8 is a block diagram of an example method for storing data units distributed among oligonucleotides in a pool of DNA data storage oligonucleotides. SUMMARY

[0017] Various aspects are described for DNA data storage using a multi-layer error-correcting code configuration distributed among oligonucleotides in a pool of oligonucleotides.

[0018] One general aspect includes a system including an encoder configured to: determine a set of oligonucleotides for encoding a data unit, where each oligonucleotide in the set of oligonucleotides encodes a plurality of symbols; determine a first codeword of an error-correcting code, where the first codeword can include a first set of symbols for encoding the data unit; assign symbols from the first set of symbols among a plurality of oligonucleotides from the set of oligonucleotides, where each oligonucleotide in the plurality of oligonucleotides receives one symbol from the first set of symbols; and output write data for the set of oligonucleotides to a synthesis interface for synthesis of the set of oligonucleotides.

[0019] Implementations can include one or more of the following features. The plurality of symbols can be encoded in sequential positions along a length of each oligonucleotide; and the first set of symbols can occupy the same sequential positions in the plurality of oligonucleotides. The first codeword can include a plurality of symbols corresponding to user data in the data unit and at least one symbol corresponding to redundancy data of the error correction code. The first codeword can further include at least one symbol corresponding to a cyclic redundancy check value. The encoder can be further configured to: determine a first set of codewords corresponding to the data unit and a first set of redundancy data of the data unit, wherein the first set of codewords includes the first codeword; determine at least one set of permuted data based on the data unit and the first set of redundancy data; determine a second set of codewords including the at least one set of permuted data and a second set of redundancy data for the at least one set of permuted data; and assign symbols for the first set of codewords and the second set of codewords to the set of oligonucleotides. The encoder can be further configured to, in response to determining the second set of redundancy data, add a codeword to the first set of codewords, the codeword including the second set of redundancy data and a third set of redundancy data for the second set of redundancy data. The set of oligonucleotides can store an aggregate number of symbols; the at least one set of permuted data can include a plurality of sets of permuted data based on the data unit and the first set of redundancy data; and the aggregate number of symbols can be substantially equal to a number of symbols in a combination of the first set of codewords and the second set of codewords. The system can include a decoder configured to: receive read data determined from sequencing the set of oligonucleotides; determine the first set of symbols of the first codeword from the read data; assemble the first codeword; decode the first codeword using the error correction code; and output the data unit based on the decoded first codeword. The decoder can be further configured to: determine the first set of codewords from the read data, the first set of codewords including a plurality of codewords corresponding to the data unit and a first set of redundancy data of the data unit, wherein the first set of codewords includes the first codeword; determine a second set of codewords from the read data, the second set of codewords including at least one set of permuted data based on the data unit and the first set of redundancy data and a second set of redundancy data for the at least one set of permuted data; decode the data unit using the first set of codewords; and selectively decode the second set of codewords in response to failing to decode at least one codeword in the first set of codewords. The decoder can be further configured to: determine a cyclic redundancy check value for each codeword in the first set of codewords; determine a validation mask by evaluating the cyclic redundancy check value for each codeword in the first set of codewords; and use the validation mask to determine a target codeword for selective decoding of the second set of codewords.

[0020] Another general aspect includes a method comprising: receiving read data determined from sequencing a set of oligonucleotides, wherein each oligonucleotide of the set of oligonucleotides encodes a plurality of symbols; determining a first set of symbols of a first codeword from the read data, wherein the first set of symbols encodes a portion of a data unit using an error correcting code; assembling the first codeword from the first set of symbols, wherein the first set of symbols is distributed among a plurality of oligonucleotides from the set of oligonucleotides, and each oligonucleotide of the plurality of oligonucleotides receives one symbol from the first set of symbols; decoding the first codeword using the error correcting code; and outputting the data unit based on the decoded first codeword.

[0021] Implementations can include one or more of the following features. The plurality of symbols can be encoded in sequential positions along a length of each oligonucleotide; and the first set of symbols occupies a same sequential position in the plurality of oligonucleotides. The first codeword can include a plurality of symbols corresponding to user data in the data unit and at least one symbol corresponding to redundancy data of the error correction code. The first codeword can further include at least one symbol corresponding to a cyclic redundancy check value. The method can include determining a first set of codewords from the read data, the first set of codewords including a plurality of codewords corresponding to the data unit and a first set of redundancy data for the data unit, wherein the first set of codewords includes the first codeword; determining a second set of codewords from the read data, the second set of codewords including at least one permuted set of data based on the data unit and the first set of redundancy data and a second set of redundancy data for the at least one permuted set of data; decoding the data unit using the first set of codewords; and selectively decoding the second set of codewords in response to failing to decode at least one codeword in the first set of codewords. The method can include determining a cyclic redundancy check value for each codeword in the first set of codewords; determining a verification mask by evaluating the cyclic redundancy check value for each codeword in the first set of codewords; and using the verification mask to determine a target codeword for selective decoding of the second set of codewords. The method can include iteratively decoding the data unit by alternating between decoding using the first set of codewords and decoding using the second set of codewords, wherein: the first set of codewords corresponds to a first codeword layer that encoded the data unit; the at least one permuted set of data can include a plurality of permuted sets of data based on the data unit and the first set of redundancy data; and the second set of codewords corresponds to a plurality of additional codeword layers that encoded the plurality of permuted sets of data. The method can include receiving the data unit; determining a first set of codewords including a plurality of codewords corresponding to the data unit and a first set of redundancy data for the data unit, wherein the first set of codewords includes the first codeword; determining at least one permuted set of data based on the data unit and the first set of redundancy data; determining a second set of codewords including the at least one permuted set of data and a second set of redundancy data for the at least one permuted set of data; assigning symbols for the first set of codewords and the second set of codewords to a set of oligonucleotides; and outputting write data for the set of oligonucleotides to a synthesis interface for synthesis of the set of oligonucleotides. The method can include adding a codeword to the first set of codewords in response to determining the second set of redundancy data, the codeword including the second set of redundancy data and a third set of redundancy data for the second set of redundancy data.

[0022] Another general aspect includes a system comprising: means for receiving read data determined from sequencing a set of oligonucleotides, wherein each oligonucleotide in the set of oligonucleotides encodes a plurality of symbols; means for determining a first set of symbols for a first codeword from the read data, wherein the first set of symbols encodes a portion of a data unit using error correction codes; means for assembling a first codeword from the first set of symbols, wherein the first set of symbols is distributed among a plurality of oligonucleotides from the set of oligonucleotides, and each of the plurality of oligonucleotides receives a symbol from the first set of symbols; means for decoding the first codeword using error correction codes; and means for outputting a data unit based on the decoded first codeword.

[0023] This disclosure describes various aspects of an innovative technique capable of encoding and decoding user data stored in a DNA oligonucleotide pool using multilayer error correction codes. The configuration of the multilayer error correction codes provided by the present invention is applicable to a variety of computer systems for storing or retrieving data stored as oligonucleotide sets in DNA storage media. The configuration can be applied to various DNA synthesis and sequencing technologies to generate write data for storage as base pairs and to process read data read from those base pairs. The novel techniques described herein include numerous innovative technical features and advantages over existing solutions, including but not limited to: (1) improved storage efficiency of the oligonucleotide pool, (2) reduced decoding complexity using relatively short codewords, and (3) improved decoding efficiency based on selective decoding and iteration in the multilayer error correction code configuration. Detailed Implementation

[0024] Novel data storage technologies are being developed for long-term data storage using synthetic DNA encoded with binary data. While current methods may be limited by the time spent synthesizing and sequencing DNA, the speed of these systems is increasing, and the density and durability of DNA as a data storage medium are remarkable. Figure 1A In the example configuration, method 100 can be used to store and retrieve data from DNA.

[0025] At box 110, the binary data to be stored in the DNA medium can be identified. For example, any conventional computer data source can be stored as a target in the DNA medium, such as files, databases, objects, data blocks from block storage devices, software code, etc. Due to the high storage density and durability of the DNA medium, the data targeted for storage can include very large data stores with archival value, such as collections of images, videos, scientific data, software, and other archival data.

[0026] At block 112, binary data can be converted to DNA code. For example, conventional computer data objects or data files can be encoded according to a DNA symbol index, such as: A or T = 1 and C or G = 0; A = 00, T = 01, C = 10, and G = 11; or more complex symbol indexes that map DNA base sequences to predetermined binary data patterns. In some configurations, source data can be encoded according to an oligonucleotide length format that includes addressing and redundancy data for recovery and reconstruction of the source data during retrieval, prior to conversion to DNA code.

[0027] At block 114, DNA can be synthesized to embody the DNA code determined at block 112. For example, the DNA code can be used as a template for generating a plurality of synthetic DNA oligonucleotides embodying the DNA code using various DNA synthesis techniques. In some configurations, large data units are divided into segments that match the payload capacity of the oligonucleotide length used, and each segment is synthesized in a corresponding DNA oligonucleotide. In some configurations, solid-phase DNA synthesis can be used to produce the desired oligonucleotides. For example, each desired oligonucleotide can be built up one base at a time on a solid support matrix to match the desired DNA sequence, such as using phosphoramidite synthesis chemistry in a four-step chain extension cycle. In some configurations, column-based or microarray-based oligonucleotide synthesizers can be used.

[0028] At block 116, the DNA medium can be stored. For example, the resulting set of DNA oligonucleotides for a data unit can be placed in a fluid or solid carrier medium. The resulting DNA medium of the set of oligonucleotides and its carrier can then be stored for any length of time with high levels of stability (e.g., DNA has been successfully sequenced for tens of thousands of years). In some configurations, the DNA medium can include a set of DNA oligonucleotides in pores or a solid matrix suspended in a carrier fluid that can itself be stored or attached to another object. In some configurations, the set of oligonucleotides associated with a particular data unit can be referred to as an oligonucleotide pool.

[0029] At block 118, the DNA oligonucleotides can be recovered from the stored medium. For example, the oligonucleotides can be separated from the carrier fluid or solid matrix for processing. The resulting set of DNA oligonucleotides can be transferred to a new solution for a sequencing process, or can be stored in a solution that can receive other polymerase chain reaction (PCR) reagents.

[0030] At block 120, the DNA oligonucleotides can be sequenced and read into a DNA data signal corresponding to the base sequence in the oligonucleotides. For example, the set of oligonucleotides can be processed by PCR to amplify a variable number of copies of the oligonucleotides from the stored set of oligonucleotides. In some configurations, the PCR amplification can result in a variable number of copies of each oligonucleotide.

[0031] At block 122, the data signal can be read from the sequenced DNA oligonucleotides. For example, the sequenced oligonucleotides can be read by a nanopore reader to generate an electrical signal corresponding to the base sequence. In some configurations, each oligonucleotide can be threaded through a nanopore, and a voltage across the nanopore can generate a differential signal with an amplitude corresponding to the different resistances of the bases. The analog DNA data signal can then be converted back to digital data based on one or more decoding steps, as further described with respect to method 130 in Figure 1B

[0032] In Figure 1B method 130 can be used to convert an analog read signal corresponding to a DNA base sequence back to the digital data units that were the original target of the DNA storage process. In the illustrated example, the original digital data units, such as a data file, are divided into data subunits corresponding to the payload size of an oligonucleotide, and the set of oligonucleotides corresponding to the subunits of the data unit can be reassembled into the original data unit. The example oligonucleotide format 140, including primers 142 and 148 that can be added to support PCR amplification and sequencing, can include a payload 144 that includes the subunits of the data unit, a redundancy portion 146 of error correction code (ECC) data for the subunits, and an address portion 150 for determining the sequence of the payloads for recombining the data block. In some configurations, a Reed-Solomon error correction code can be used to determine the redundancy portion 146 of the payload 144.

[0033] At block 160, the DNA base data signal can be read from the sequenced DNA. For example, the analog signal from the nanopore reader can be conditioned (equalized, filtered, etc.) and converted to a digital data signal for each oligonucleotide.

[0034] At block 162, a number of copies of the oligonucleotides can be determined. Through the amplification process, a number of copies of each oligonucleotide can be produced, and the decoding system can determine groups of the same oligonucleotides to be processed together.

[0035] ​At block 164, each set of identical oligonucleotides can be aligned, and a consensus sequence between the multiple copies can be determined. For example, a set of four copies can be aligned based on their primers, and a consensus sequence algorithm can be applied along each base position of the set of base values to determine the most likely version of the oligonucleotide for further processing, such as using the value in the case of 3 out of 4 agreement.

[0036] At block 166, the primers can be dropped. For example, primers 142 and primers 148 can be removed from the data sets corresponding to payload data 144, redundancy data 146, and addresses 150.

[0037] At block 168, an error check can be performed on the resulting data sets. For example, ECC processing of the payload 144 based on redundancy data 146 can allow correction of errors in the resulting consensus sequence data set for the oligonucleotides. The number of errors that can be corrected can depend on the ECC code used. ECC codes can have difficulty correcting errors resulting from insertions or deletions that cause a shift of all subsequent base values. The size of the oligonucleotide payload 144 and the portion allocated to redundancy data 146 can determine and limit the efficiency of correctable errors and data format.

[0038] At block 170, the base or base symbols can be reverse mapped back to the original bit data. For example, the symbol encoding scheme used to generate the DNA code can be reversed to determine the corresponding bit data sequence.

[0039] At block 172, a file or similar data unit can be reassembled from the bit data corresponding to the set of oligonucleotides. For example, the addresses 150 from each oligonucleotide payload can be used to sort the decoded bit data and reassemble the original file.

[0040] Figure 2 An improved DNA storage system 200 is shown, and more specifically, an improved encoding system 210 and decoding system 240 for using a multi-layer error correction code distributed across oligonucleotides in a pool of oligonucleotides to improve data retrieval and efficiency. In some configurations, the encoding system 210 can be a first computer system for determining target binary data, such as a regular binary data unit, and converting it to a DNA base sequence to synthesize into DNA for storage, and the decoding system 240 can be a second computer system for receiving a data signal corresponding to a base sequence read from the DNA.

[0041] In some configurations, DNA decoding can be divided into two stages. First, a correlation analysis is performed on individual oligonucleotides to eliminate deletion and insertion errors. Second, a final decoding is performed using a more conventional scheme, such as Reed-Solomon, and the block size is larger than an individual oligonucleotide. The use of multi-layer error correction codes can be implemented in the second part of the decoding process. The error correction process can be configured to handle only traditional errors or erasures. Insertion and deletion errors can be assumed to be resolved in the first stage of decoding. A relatively simple decoding scheme that utilizes short code words can be selected and combined with a larger multi-layer scheme that covers larger block sizes. This scheme can take advantage of the fact that DNA read operations are performed on larger pools of oligonucleotides at the same time, unlike typical stream reads used in other storage devices (e.g., magnetic tape, disk drives, etc.).

[0042] Encoding system 210 can include a processor 212, a memory 214, and a synthesis system interface 216. For example, encoding system 210 can be a computer system configured to receive or access conventional computer data (such as data stored as binary files, blocks, data objects, databases, etc.) and map that data to DNA base sequences for synthesis into DNA storage units (such as a set of DNA oligonucleotides). Processor 212 can include any type of conventional processor or microprocessor that interprets and executes instructions. Memory 214 can include a random access memory (RAM) or another type of dynamic storage device that stores information and instructions for execution by processor 212, and / or a read only memory (ROM) or another type of static storage device that stores static information and instructions for use by processor 212. Encoding system 210 can also include any number of input / output devices and / or interfaces. Input devices can include one or more conventional mechanisms that permit an operator to input information to encoding system 210, such as a keyboard, a mouse, a pen, voice recognition and / or biometric mechanisms, etc. Output devices can include one or more conventional mechanisms that output information to the operator, such as a display, a printer, a speaker, etc. Interfaces can include any transceiver-like mechanism that enables encoding system 210 to communicate with other devices and / or systems. For example, synthesis system interface 216 can include a connection to a bus interface (e.g., a Peripheral Component Interface Express (PCIe) bus, a Universal Serial Bus (USB), etc.) or a network for communicating DNA base sequences for storing data to a DNA synthesis system. In some configurations, synthesis system interface 216 can include a network connection using the Internet or a similar communication protocol to send a conventional data file listing DNA base sequences for synthesis, such as the required base sequences for each oligonucleotide to be synthesized, to a DNA synthesis system. In some configurations, the DNA base sequences can be stored to a conventional removable medium, such as a USB drive or a flash memory card, and transferred from encoding system 210 to the DNA synthesis system using the removable medium.

[0043] In some configurations, a series of processing components 218 can be used to process target binary data, such as a target data file or other unit of data, into a DNA base sequence to output to a synthesis system. For example, the processing components 218 can be embodied in software and / or hardware circuitry, and collectively referred to as an encoder for encoding binary database DNA base pair data. In some configurations, the processing components 218 can be embodied in one or more software modules stored in memory 214 for execution by processor 212. Note that the series of processing components 218 is an example and different configurations and ordering of components are possible without materially changing the operation of processing components 218. For example, in an alternative configuration, additional data processing such as a data randomizer for whitening an input data sequence can be used to pre-process data prior to encoding. In another configuration, data can be divided according to the codeword size of one or more ECC layers and / or CRC, ECC, and layering can be applied in a different order. Other variations are possible.

[0044] In some configurations, processing target data can begin with a run length limited (RLL) encoder 220. RLL encoder 220 can modulate the length of runs in input data. RLL encoder 220 can employ line encoding techniques that process arbitrary data with bandwidth limitations. In particular, RLL encoder 220 can limit the length of extensions of repeated bits or specific repeated bit patterns such that extensions are not too long or too short. By modulating the data, RLL encoder 220 can reduce problematic data sequences that can create additional errors in subsequent encoding and / or DNA synthesis or sequencing.

[0045] In some configurations, the symbol encoder 222 can include logic to convert binary data to symbols based on four DNA bases (adenine (A), cytosine (C), guanine (G), and thymine (T)). In some configurations, the symbol encoder 222 can encode each bit as a single base pair, such as 1 mapping to A or T, and 0 mapping to G or C. In some configurations, the symbol encoder 222 can encode two-bit symbols into a single base, such as 11 mapping to A, 00 mapping to T, 01 mapping to G, and 10 mapping to C. More complex symbol mappings can be implemented based on multi-base symbol mappings to corresponding longer sequences of bit data. For example, two-base symbols can correspond to 16 states for mapping four-bit symbols, or four-base symbols can map 256 states of byte symbols. From an oligonucleotide synthesis perspective, multi-base pair symbols can be preferable. For example, synthesis can not be performed on base pairs, but on larger blocks, like “bytes” related to the symbol size, which are prepared and cleaned earlier in the synthesis process (e.g., pre-synthesis). This can reduce the amount of synthesis errors. From an encoder / decoder perspective, these physically larger blocks can be treated as symbols or smaller sets of symbols.

[0046] In some configurations, the oligonucleotide formatter 224 can include logic to allocate portions of the target data unit to the set of oligonucleotides. For example, the oligonucleotide formatter 224 can be configured for a predetermined payload size of each oligonucleotide, and determine a number of symbols corresponding to the payload size of each oligonucleotide in the set. In some configurations, the payload size can be determined based on the oligonucleotide size used by the synthesis system and any portion of the total length of the oligonucleotide allocated to redundant data, address data, synchronization marker data, or other data formatting constraints. For example, for a 150 base pair oligonucleotide using two-base symbols, an eight-base addressing scheme and six four-base synchronization markers can be included, resulting in 118 base pairs of target data allocated to each oligonucleotide. In some configurations, the oligonucleotide formatter 224 can insert a unique address for each oligonucleotide in the set, such as at the beginning or end of the data payload. In some configurations, the set of symbols to be written for each oligonucleotide can be determined by the pool formatter 324 based on the configuration and distribution of codewords across the oligonucleotides in the pool as a whole.

[0047] In some configurations, the synchronization marker formatter 226 can include logic to insert synchronization markers at predetermined intervals within the symbols. For example, a synchronization marker can be inserted every 20 base pairs to divide the data in the oligonucleotide into a predetermined number of shorter segments. The synchronization marker can be any sequence of base pairs that has a good signal-to-noise ratio (SNR). To avoid false synchronization marker detection, the sequence can be excluded from the user data by a specific modulation code. The use of synchronization markers to determine insertions and deletions is further described below with respect to the decoding system 240.

[0048] In some configurations, the encoding system 210 can use one or more layers of CRC and / or ECC encoding based on treating the pool of oligonucleotides as a storage pool and distributing codewords, CRC values, and redundancy data across the oligonucleotides to minimize known error correlations that can affect the performance and / or reliability of ECC decoding. For example, the encoding system 210 can use relatively small codewords that can be efficiently encoded and decoded using a Reed-Solomon error correction code, but distribute the symbols of the codewords across the oligonucleotides such that any error correlations in the oligonucleotides do not transfer to the codewords. If each oligonucleotide includes only one symbol of each codeword it participates in, error conditions specific to the oligonucleotide do not propagate into the codewords (i.e., they produce at most a single symbol error in each codeword). While the example configurations described herein focus on error correlations based on the oligonucleotides, any known error correlations can be used to produce similar codeword construction rules if the data exhibits error correlations other than within the oligonucleotides.

[0049] A multi-layer encoding and decoding scheme can be used that provides multiple protection of the user data symbols when employing standard codeword sizes and ECC encoding / decoding schemes. The effective size of the encoding block can be equal to the pool size, with multi-layer protection encoded within the set of oligonucleotides, providing a known format efficiency equal to the symbols in the input user data relative to the aggregate symbols encoded across all of the oligonucleotides in the pool. For example, the difference between the input user data and the aggregate symbols in the pool can be composed of CRC values, ECC redundancy data, and their permutations that provide multi-layer ECC protection. Thus, a more robust overall data protection configuration can be achieved while keeping the base codeword and symbol size small. The layered nature of the ECC protection also supports selective and iterative decoding of multiple passes with the same ECC decoding scheme (and in some cases, decoding hardware or software), making the decoding system 240 more efficient. In some configurations, each user symbol in the DNA oligonucleotide pool will be encoded in multiple ECC codewords and CRCs to have a very low error floor for successful detection and decoding of the user data. The error floor can be adjusted by the ECC redundancy, number of layers, and CRC size relative to the user data size and aggregate oligonucleotide pool size.

[0050] The CRC encoder 228 can include logic to determine a CRC value for a particular set of user data symbols, such as for a particular codeword. For example, each first layer codeword can include a predetermined number of user data symbols, a CRC value (written as one or more symbols), and ECC redundancy symbols. In some configurations, the CRC encoder 228 can be configured to aggregate the size of the user data symbols and a predetermined CRC divisor to generate a CRC value, one or more symbols corresponding to a remainder or checksum value. The CRC value can support a CRC check by the decoding system 240 to detect errors in the user data symbols. Note that a CRC is an error detection code, not an error correction code, as it can not itself correct user data symbols in error. The decoding system 240 can use the CRC value and CRC check to determine whether the set of user data symbols requires additional iterations of ECC decoding.

[0051] The ECC encoder 230 can append one or more parity bits to the set of user data symbols to generate a codeword for later detection and correction of errors that occur during a data read process. For example, the added redundancy data can include one or more symbols of additional binary bits (parity bits) that can be associated with a string of binary bits (symbols) in the user data symbols (or other unit of data being encoded) and enable a corresponding decoder to locate and correct errors in the user data symbols (up to a particular error correction capability defined by the ECC used). In some configurations, a Reed-Solomon ECC or similar erasure code (e.g., Bose-Chaudhuri-Hocquenghem (BCH) ECC) can be used to encode the set of user data symbols by adding one or more check symbols as redundancy data. In some configurations, the ECC encoder 230 can be used to encode codewords for each layer of ECC configuration in the pool. For example, the ECC encoder 230 can operate on an initial set of user data symbols for a target data block to be encoded, then permute the encoding of the first layer data (user data, CRC value, and redundancy data) in the codeword for additional layers. In some configurations, the redundancy data from additional layers can also be encoded in the codeword and added to the first layer.

[0052] The permutor 232 can include logic to generate permutations of codewords that re-group them for an additional layer of ECC encoding. For example, the permutor 232 can operate on the codewords in the first layer (including user data symbols, CRC values, and redundancy data) to re-group new codewords in an additional layer to provide tiered data protection. In some configurations, the permutor 232 can operate on the first layer data multiple times to generate multiple different permutations of the data that can be distributed among the oligos to provide for each ECC layer. More will be described regarding Figure 3 Further examples of possible permutations are described. In some configurations, the permuted data based on the first layer codewords can be fed back to the ECC encoder 230 or a similar ECC encoder to generate new codewords with additional sets of redundancy data. In some configurations, the redundancy data from one or more of the permuted layers can be encoded in additional codewords by the CRC encoder 228 and / or the ECC encoder 230. For example, new codewords including redundancy data from the permuted data, corresponding CRC values, and corresponding ECC redundancy data can be appended to the first layer for distribution across the oligos.

[0053] The pool formatter 234 can include logic to distribute symbols from the codewords calculated by the CRC encoder 228, the ECC encoder 230, and the permutor 232 among the oligo pool. For example, the pool formatter 234 can include logic to ensure that two or more symbols from the same codeword are not distributed to the same oligo. In some configurations, the pool formatter 234 can define a symbol matrix for distributing codeword symbols among a set of oligos in a pool. For example, the oligos can be arranged in a grid with each row having a number of oligos equal to the number of symbols N in each codeword. The pool and grid can include NxK oligos, meaning that K codewords can be distributed in each layer of the oligo pool, where a layer is defined as a set of symbols having the same sequential position along the length of the oligos. The oligo pool and matrix can have M layers, where M is the number of symbols in the payload of each oligo. The pool formatter 234 can distribute all of the codewords from a combination of layers in the oligo pool in any order and using a set of configuration rules determined to prevent or reduce correlation. For example, the pool formatter 234 can only distribute one symbol from any given codeword to any given oligo. In some configurations, the codewords can be distributed in layer order and fill sequential layers of the symbol matrix.

[0054] Base pair converter 236 can include logic to convert binary symbol data and any associated addresses, synchronization markers, or similar formatted data into base pair sequences for each oligonucleotide. For example, base pair converter 236 can receive an array of data or similar data structure from pool formatter 234 in which each oligonucleotide is represented by a string of binary data, and base pair converter 236 can convert the binary data into a series of base pair indicators based on the DNA encoding scheme used. The resulting DNA base sequence corresponding to the encoded target data unit can be output as DNA data 238 from processing component 218. For example, the base pair sequence for each oligonucleotide in the set of oligonucleotides corresponding to the target data unit can be stored as a sequence list for transfer to a synthesis system via synthesis system interface 216.

[0055] Decoding system 240 can include a processor 242, a memory 244, and a sequencing system interface 246. For example, decoding system 240 can be a computer system configured to receive or access analog and / or digital signal data from a read sequenced DNA, such as data signals associated with a set of oligonucleotides that have been amplified, sequenced, and read from a stored DNA medium. Processor 242 can include any type of conventional processor or microprocessor that interprets and executes instructions. Memory 244 can include a random access memory (RAM) or another type of dynamic storage device that stores information and instructions for execution by processor 242 and / or a read only memory (ROM) or another type of static storage device that stores static information and instructions for use by processor 242. Decoding system 240 can also include any number of input / output devices and / or interfaces. Input devices can include one or more conventional mechanisms that permit an operator to input information to decoding system 240, such as a keyboard, mouse, pen, voice recognition and / or biometric mechanisms, etc. Output devices can include one or more conventional mechanisms that output information to the operator, such as a display, printer, speaker, etc. Interfaces can include any transceiver-like mechanism that enables decoding system 240 to communicate with other devices and / or systems. For example, sequencing system interface 246 can include a connection to an interface bus (e.g., a PCIe bus, USB, etc.) or a network for receiving analog or digital representations of DNA sequences from a DNA sequencing system. In some configurations, sequencing system interface 246 can include a network connection using the Internet or a similar communication protocol to receive a conventional data file that lists DNA base sequences and / or corresponding digital sample values generated by analog-to-digital sampling of sequencing read signals from a DNA sequencing system. In some configurations, the DNA base sequence list can be stored to a conventional removable medium, such as a USB drive or flash memory card, and transferred from the DNA sequencing system to encoding system 240 using the removable medium.

[0056] The decoding system 240 can include a set of processing components 248 for decoding the oligonucleotide read data sets into the user data units that were originally encoded by the encoding system 210. In some configurations, the processing components 248 can be collectively embodied in a software or hardware decoder that operates in the decoding system 240 to receive the read data and output the user data. The decoding system 240 can use a first stage of error correction that targets elimination of insertion and deletion errors (which create shifts in all subsequent base pairs in the sequence), followed by ECC error correction to address mutation or erasure errors. DNA media and sequencing face three main types of errors: deletions, insertions, and mutations. Mutation errors are most similar to traditional errors in data storage devices, and can be effectively handled using ECC correction. Insertion and deletion errors affect the position of all subsequent bits or symbols, and can not be effectively handled by ECC. Thus, pre-processing the oligonucleotide sequences to obtain sequence position shifts, and correcting those position shifts when possible, can help more effective and reliable data reading. The pre-processing stage can significantly reduce the error rate in the oligonucleotide sequences prior to applying ECC correction, enabling more effective ECC codes and more reliable data retrieval using the first level ECC encoding in a nested ECC configuration. In some configurations, the pre-processing stage can include an oligonucleotide set sorter 250, a cross-correlation analyzer 252, a synchronization marker detector 254, and an insertion / deletion correction 256.

[0057] In some configurations, the oligonucleotide set sorter 250 can sort the received oligonucleotide data sequences into sets of copies. For example, the DNA amplification process can produce multiple copies of some or all of the oligonucleotides, and the oligonucleotide set sorter 250 can sort the oligonucleotide data sequences into similar sequences. The sorting can be based on markers during the sequencing process, address data, and / or statistical analysis of the sequences (or samples thereof) to determine duplicate copies of each oligonucleotide. Note that different copies can include different errors, and at this stage, exact matching of all bases in the sequence can not be a sorting criterion.

[0058] In some configurations, the cross-correlation analyzer 252 can include logic to compare two or more copies of an oligonucleotide to determine insertions and deletions. For example, after synthesizing and identifying multiple copies of an oligonucleotide, insertion and deletion errors will have different positions for different copies, and those insertions / deletions can be located by a correlation analysis. Cross-correlation analysis can enable elimination of insertion and deletion base errors. For example, in the case that a correlation analysis identifies an insertion, the particular insertion can be identified and removed to realign the following bases in the oligonucleotide. In the case that a correlation analysis identifies a deletion, a placeholder value can be added to realign the following bases in the oligonucleotide, and the location of the placeholder can be identified as an erasure for correction in the ECC stage. After analysis, regions where there is insufficient SNR across copies, base fragments, and / or regions of uncertainty of the consensus sequence can be identified and flagged as erasures. For example, short shift regions can have small correlation signals, and if the SNR is insufficient, can be treated as erasures. In some configurations, averaging the correlation fragments of a base across multiple copies can provide soft information for subsequent ECC processing. For example, the signals from multiple aligned copies can be averaged and statistically processed, providing soft information for each symbol.

[0059] In some configurations, the synchronization marker detector 254 can include logic to identify periodic synchronization markers along the length of the oligonucleotide to help identify insertion and deletion errors. For example, the synchronization markers inserted by the synchronization marker formatter 226 can have a defined synchronization marker spacing that corresponds to the number of base pairs that should occur between synchronization markers. The synchronization marker detector 254 can detect the synchronization markers in the oligonucleotide and determine the number of base pairs between the synchronization markers. If the number of synchronization markers is greater than the synchronization marker spacing, an insertion error has occurred. If the number of synchronization markers is less than the synchronization marker spacing, a deletion error has occurred. Thus, it can be determined whether an insertion or deletion error has occurred in any piece of data between two synchronization markers without affecting the data between other synchronization markers in the oligonucleotide. In some configurations, the synchronization marker detection can be followed by cross-correlation analysis of the synchronization markers to determine insertions and deletions. After synchronization marker analysis, there should be no insertion or deletion errors in the regions of the oligonucleotide that were aligned to the synchronization markers. An insertion or deletion within a region of data between synchronization markers can cause the entire region to be identified as an erasure (as it can not be possible to identify where in the piece the insertion or deletion occurred based on the synchronization markers alone). An insertion or deletion within a synchronization marker itself can determine that the pieces on either side of the synchronization marker are to be treated as erasures. In cases where multiple copies are not available, the synchronization marker detection and analysis can be performed on a single copy of the oligonucleotide. In some configurations, the operations of the synchronization marker detector 254 can be combined with the operations of the cross-correlation analyzer 252 where multiple copies are available. For example, the synchronization marker detector 254 can be used to identify pieces where there are insertions and / or deletions, and the cross-correlation analyzer 252 can target those regions for cross-correlation analysis to determine the specific location of the insertions or deletions and related soft information. This can greatly reduce the amount of cross-correlation analysis to be performed.

[0060] In some configurations, the insertion / deletion correction 256 includes logic for selectively correcting insertion and deletion errors in the oligonucleotides where possible. For example, the insertion / deletion correction 256 can use output from the cross-correlation analyzer 252 and / or the sync marker detector 254 to determine correctable insertion / deletion errors. For example, when an insertion or deletion error occurs between sync markers, the position of a subsequent fragment can be corrected for a previous shift in base pair position. In some configurations, the cross-correlation analysis can enable the insertion / deletion correction 256 to specifically identify likely locations of insertion and / or deletion errors at a finer granularity level, such as a symbol, base pair, or other step. For example, a correlation analysis of two or more copies of an oligonucleotide can allow statistical methods and soft information to be compared to a correction threshold to cause an inserted base pair and / or an insertion padding or placeholder base pair, which can be identified as an erasure, to be deleted to at least partially correct the insertion and / or deletion. In some configurations, the insertion / deletion correction 256 can mark a fragment of base pairs in the oligonucleotide as an erasure that requires ECC error correction. For example, the sync marker detector 254 and correlation analysis can determine that one or more fragments between sync markers are identified as erasures for error correction, and / or the cross-correlation analyzer 252 can determine that a deletion location of an unknown base or symbol deletion is identified as an erasure. In some configurations, a signal quality, statistical uncertainty, and / or a particular threshold for a consensus sequence can identify one or more fragments as erasures because the processing of the sync marker detector and / or the cross-correlation analyzer is uncertain. For example, the insertion / deletion correction 256 can be configured to output an oligonucleotide with as many base pairs or symbols as possible that are positively identified as not containing an insertion or deletion error, and to identify (and locate) any fragments that cannot be determined to be free of an insertion or deletion error as erasures prior to performing ECC error correction.

[0061] Using multiple layers of ECC can provide improved and efficient error detection and correction of erasures using relatively short codewords and corresponding ECC codes. By using multiple layers and distributing the code symbols across the oligonucleotide, a user data unit can be reliably decoded from the oligonucleotide pool, and the redundant layers of permutation data can be selectively decoded to provide additional redundancy. For example, the decoding system 240 can use the ECC decoder 260 and the CRC check 262 to process the first layer codewords, identify any codewords that cannot be successfully decoded, and selectively continue decoding using one or more additional layers. In some configurations, the first layer codewords can be reprocessed after decoding each layer of permutation data to determine whether additional redundant data can successfully decode a previously failed codeword.

[0062] Code word logic 258 can include logic to map the set of oligonucleotides back to the configuration of code words in which they were written. For example, code word logic 258 can include inverse logic of pool formatter 234 to parse the location of each code word (and code word layer) from the payload of the set of oligonucleotides. In some configurations, code word logic 258 can organize the payload of the oligonucleotides and the sequence of symbols they contain into a matrix of symbols that can be divided into different code words and layers in which the code words are encoded according to their location in the matrix. For example, code word logic 258 can use the address values from each oligonucleotide to position its data in the same grid location used to encode it. Code word logic 258 can output the set of code words and / or portions thereof for processing by subsequent components in processing component 248.

[0063] ECC decoder 260 can include logic to receive code words and decode the input data (e.g., user data symbols) using the redundant data in the code words and the ECC encoding scheme. For example, ECC decoder 260 can operate on code words in a first layer that include user data, CRC values, and a first set of redundant data computed by ECC encoder 230. In some cases, ECC decoder 260 can successfully decode the code words, determining the corresponding user data symbols. In some cases, ECC decoder 260 can fail to decode and determine that there are more errors than allowed by the correction capability of the ECC. In some cases, ECC decoder 260 can return a successful decode of a data unit, but the data unit can be incorrect (as determined by CRC check 262). In some configurations, ECC decoder 260 can be invoked through multiple processing layers. For example, in first layer processing, ECC decoder 260 can decode all code words in a first layer, which can include a complete set of user data code words, as well as code words containing redundant data for any permutation layers.

[0064] CRC check 262 can include logic to perform a CRC check on code words to verify those symbols and determine if additional processing layers are needed. For example, CRC check 262 can evaluate code words based on CRC values to determine if the code words are valid or if errors have been detected. In some configurations, CRC check 262 can perform a CRC check on each code word containing a data block of user data from a first layer in the code word. For example, CRC check 262 can determine CRC validity for each code word in a set of code words for a data unit and use the ECC results to determine validity of each code word for triggering additional decoding using one or more permutation layers.

[0065] The CRC mask 264 can include logic for controlling selective decoding through additional layers of oligo codewords. The CRC mask 264 can generate a mask data structure that indicates whether particular symbols in a codeword are verified and how those symbols map to a codeword in the next layer of permuted data. The CRC mask 264 can improve processing efficiency by determining additional processing that is performed using the permuted data and corresponding redundancy data to decode each codeword.

[0066] The permuted data logic 266.1-266.n can include logic for managing the processing of each iterative decoding process through each layer of permuted data by the ECC decoder 260 using the CRC mask 264 and the return of data to update the first layer decoding between layers of permuted data. For example, a codeword can be selectively passed through the CRC mask 264 for ECC decoding based on the number of invalid symbols in each permuted layer codeword and determining decoder processing for that permuted layer codeword. For example, a permuted layer codeword can include a number of unverified symbols that is less than or equal to the error correction capability of that codeword. In this case, the codeword can be successfully decoded in the permuted layer and the correct values of those symbols are returned. The permuted layer codeword can not include invalid symbols and it has already been verified. In this case, no additional decoding in the permuted layer is needed. In some configurations, the CRC can still be used to verify the permuted layer codeword upon return to the first layer to provide a second level of verification of the CRC result. The number of unverified symbols from a non-recovered codeword in the first layer can be higher than the correction capability of the permuted layer codeword. In this case, decoding can still be performed on the permuted layer codeword to see if one or more symbols can be recovered even if a full recovery can not be possible. In some configurations, the corrected symbols from a permitted layer can be mapped back to their corresponding codeword in the first layer for additional decoding attempts between each permuted layer (and selective sequential operation of the permuted data logic 266.1-266.n) until all user data codewords in the first layer are recovered. For example, in a configuration with a permuted data second layer and a permuted data third layer, the decoding operations can include: ECC decoding and CRC check of the first layer, selective ECC decoding of the second layer (based on the CRC mask), return to the first layer ECC decoding and CRC check, selective ECC decoding of the third layer (based on the CRC mask), and return to the first layer ECC decoding and CRC check. This pattern can be repeated for any number of layers of permuted data. In some configurations, the CRC mask 264 and the permuted data logic 266 can include logic (based on the permuter 232) and / or mapping data structures for determining the correlation among symbols in the user data layer (first layer) and each permuted data layer. For example, the CRC mask 264 can map symbols from a codeword in the first layer to symbols in a codeword of the second layer and the permuted data logic 266 can map symbols from a codeword in the second layer back to symbols in the first layer.

[0067] In some configurations, the symbol decoder 270 can be configured to convert the symbols used to encode the bit data back to its bit data representation. For example, the symbol decoder 270 can invert the symbols generated by the symbol encoder 222. In some configurations, the symbol decoder 270 can receive the error corrected sequence from the iterative decoding through the ECC decoder 260, CRC check 262, CRC masking 264, and permutation data logic 266, and output a digital bit stream or bit data representation. For example, the symbol decoder 270 can receive an array of symbols from the corrected codeword corresponding to the user data and convert them to a bit representation of the data units, subject to further post-processing decoding (to invert the pre-processing encoding from the encoding system 210).

[0068] In some configurations, the RLL decoder 272 can decode the run length limited code encoded by the RLL encoder 220 during the data encoding process. In some configurations, the data can be subject to additional post-processing or formatting to place the digital data in a regular binary data format. The output data 274 can then be output to a regular binary data storage medium, network, or device, such as a host computer, network node, etc., for storage, display, and / or use as a regular binary data file or other data unit.

[0069] Figure 3An example matrix 300 is shown for constructing codewords in an oligonucleotide pool 310 of a multi-layer error correction code. The matrix 300 can be used to organize and allocate symbols 352 in a set of oligonucleotides in the oligonucleotide pool 310. In an example oligonucleotide format 340, a plurality of symbols 352 can be arranged in order from symbol 352.1 to 352.M in the payload of each oligonucleotide 312. Each oligonucleotide 312 can have a symbol length 314 equal to the number of sequential symbols it holds. Each oligonucleotide 312 can be assigned to a two-dimensional matrix or grid having a width 316 and a depth 318, where the vertical orientation of the oligonucleotides defines the symbol positions in the three-dimensional matrix 300. That is, the oligonucleotides can be viewed as vertical stacks of symbols arranged in a grid. For example, the symbol length 314 can be M symbols, the grid width 316 can be N symbols, and the grid depth 318 can be K symbols. The total number of symbols that can be stored to the oligonucleotide pool 310 can equal M x N x K. In some configurations, the encoding scheme can use multi-layer encoding, where the total number of symbols in the oligonucleotide pool 310 substantially equals the total number of symbols used across the codewords in all layers in the encoding scheme for the block size, such that the oligonucleotide length and / or payload is substantially utilized. In this context, substantially can refer to less than 1% loss of format efficiency from unused symbols. Similarly, the position of any symbol 352 can be referenced by a three-number indication of (1 to M), (1 to N), (1 to K), and any oligonucleotide 312 can be referenced by a two-number indication of (1 to N), (1 to K). While the (1 to M) value can refer to a particular symbol in the context of a particular oligonucleotide, in the context of the matrix 300, it can refer to the layer of all symbols having the same sequential position included in the oligonucleotide pool. Note that the matrix 300 is a data construct for mapping symbols 352 to oligonucleotide positions based on oligonucleotide addresses 350 and sequential symbol positions in those oligonucleotides. The actual oligonucleotides can not be physically arranged in a grid, and the symbols written to and read from those oligonucleotides can be managed based on the matrix position assignment for encoding and decoding data in the oligonucleotides.

[0070] In the example shown, each oligo 312 includes four symbols arranged in an N x K matrix. The front left oligo 312.1.1, the back left oligo 312.1.K, the front right oligo 312.N.1, and the back right oligo 312.N.K provide an example of how oligo positions can be specified. In the example shown, N is chosen to equal the codeword size used to encode the oligo pool. By standardizing the codeword size and distributing the codewords 320 across the grid, it can be guaranteed that no codeword includes multiple symbols on the same oligo, and the efficiency of the oligo pool format can be more easily managed. This configuration prevents correlation between oligos and codewords, as no codeword can contain two or more symbols from the same oligo, where errors are strongly correlated. The codewords 320 can be grouped only in the horizontal direction, and a given ECC layer combines all codewords (M x K of which) from different planes. After the first layer, other layers re-group the same symbols in different codewords, but the rule that "no two symbols in a codeword come from the same oligo" makes it natural to permute the symbols within a single horizontal plane of the matrix 300. Layer 2 can be a set of the same symbols that form a new codeword. Additional redundancy can be needed for these additional codewords. This point will be further explained in Figure 4 and Figure 5 In some configurations, the permutation rules used to form layers 2, 3, etc. can be the same, where the symbol permutation for forming new codewords is done within the horizontal planes of the matrix 300. If the codewords are still formed on the horizontal planes, then it is allowed to make vertical direction permutations within one oligo.

[0071] Note that this configuration rule set is intended to reduce oligo correlation, but other correlations can be identified that would prompt additional or alternative rules for assigning symbols to codewords and assigning codewords to positions in the matrix 300.

[0072] An example oligonucleotide format 340 illustrates how data locations from matrix 300 can be mapped to oligonucleotide storage devices, and vice versa. Each oligonucleotide 312 can include a payload 344 that corresponds to the amount of data that can be encoded in the base pairs of the oligonucleotide. Payload 344 can be in addition to any primers 342, 346 or other elements that can be attached to the payload to aid in DNA synthesis or sequencing. Each oligonucleotide 312 can include an address 350 that includes a unique identifier for the oligonucleotide to aid in organizing the oligonucleotide pool 310 to reconstruct the data stored therein. For example, address 350 can include a two-dimensional matrix location for the oligonucleotide as well as a unique identifier associated with the data unit being stored. Payload 344 can then include a sequence of symbols 352.1-352.M that can be indexed by their position along the symbol sequence given a known symbol size. For example, symbols 352.1-352.M can be encoded in sequential positions along the length of oligonucleotide 312 and decoded based on those symbol positions that correspond to a codeword written across multiple oligonucleotides. In some configurations, additional header information (e.g., symbol size, format standard, etc.) can be included in payload 344, as can synchronization markers or other formatting to aid in encoding and decoding the data therein.

[0073] Figure 4 An example configuration 400 for oligonucleotide pool encoding is shown. Configuration 400 includes a first layer 402, a second layer 404, and a third layer 406. First layer 402 can include a set of symbols that include a complete user data unit (shown as user data 410). Second layer 404 and third layer 406 can be permutation data layers that encode permutations of the data in first layer 402 to provide additional redundancy for decoding the user data when using the same codeword size 408 and ECC encoder / decoder. The permutation codewords in second layer 404 and third layer 406 use different codewords to additionally protect the same user data symbols.

[0074] In the first layer 402, the user data 410 includes a set of symbols corresponding to the user data units stored in the pool of oligonucleotides. The user data 410 can be used to generate a set of codewords using the construction rules described above with respect to the system 200 and the matrix 300. Each codeword in the first portion 420 of the first layer 402 can include symbols from the user data 410, a CRC value 412.1 (CRC bits) for the codeword, and a first set of redundancy data 414.1. In some configurations, the first portion 420 can be a primary set of codewords for the user data 410, and additional permuted data layers and their additional redundancy data can be selectively used to provide additional protection to recover codewords that are not recovered from the first portion 420 by iterative decoding through the first layer 402 as well as each of the permuted second layer 404 and third layer 406.

[0075] In the second layer 404, the permuted data 416.1 can include permuted symbols of the user data 410, CRC values 412.1, and redundancy data 414.1. The permutation can generate new codewords based on a set of permutation rules, the new codewords consisting of the permuted data 416.1 and corresponding redundancy data 414.2 encoded to the same codeword size 408. In some configurations, the redundancy data 414.2 can be less than the redundancy data 414.1 and use an ECC with the same codeword size but a lower number of correctable errors, such as 2 redundancy data symbols. Additionally, the second layer 404 can not include CRC values. In some configurations, the redundancy data 414.2 can be permuted and added to the first layer 402 in the second portion 422. The second portion 422 can include redundancy data 414.2 protected by corresponding CRC values 412.2 and new redundancy data 414.3 computed for those codewords.

[0076] On the third layer 406, the permuted data 416.2 can include another set of permuted symbols of the user data 410, the CRC value 412.1, and the redundancy data 414.1. The permutation can generate additional new codewords based on the permutation rule set, which consist of the permuted data 416.1 and the corresponding redundancy data 414.4 encoded to the same codeword size 408. In some configurations, the redundancy data 414.4 can also be smaller than the redundancy data 414.1 and use an ECC with the same codeword size but a lower number of correctable errors, such as 2 redundancy data symbols. Additionally, the third layer 406 can not include a CRC value. In some configurations, the redundancy data 414.4 can be permuted and added to the first layer 402 in the third portion 424. The third portion 424 can include the redundancy data 414.4 protected by the corresponding CRC value 412.3 and new redundancy data 414.5 computed for those codewords. While two layers of permuted data are shown for the example configuration 400, any number of layers of permuted data can be included, as appropriate for the size of the user data unit, the oligonucleotide pool, and the desired threshold of recoverability. The multiple layers of ECC based on permuted data provide multiple protection of the user data. The codewords in different layers are interconnected and effectively make the ECC block size equal to the aggregate symbol size of the entire oligonucleotide pool. Using a common codeword size 408 allows any codeword in any layer to be encoded and decoded using similar ECC encoders / decoders to simplify the implementation.

[0077] Figure 5 An example configuration 500 for oligonucleotide pool decoding is shown. The configuration 500 includes a first layer 502 and a second layer CRC mask 504 or similar data validation mask. The first layer 502 can include a set of symbols containing a complete user data unit, shown as user data 510. The second layer, corresponding to the second layer CRC mask 504, can be a permuted data layer that encodes a permutation of the data in the first layer 502 to provide additional redundancy for decoding the user data when using the same codeword size 508 and ECC encoder / decoder. Note that the first layer 502 is configured to support additional layers of permuted data (not shown), similar to the configuration 400 in Figure 4

[0078] ​Decoding can begin with decoding the codewords in portion 530 using an ECC decoder for the ECC configuration. Each codeword can be successfully decoded to recover the corresponding symbols from user data 510, or can have a number of errors that exceed the decoding capability of the ECC and result in a decoding failure (in that iteration). CRC values 512.1 can then be used for CRC checking to verify the decoding by the ECC decoder. A second layer CRC mask 504 can be generated from the output of the ECC decoding and CRC checking of portion 530. The second layer CRC mask 504 can use the mapping of invalid symbols from portion 530 to permuted data 516.1 in the second layer. In some configurations, first layer 502 includes redundancy data 514.2 for encoding of the second layer, and this redundancy data can also be decoded using redundancy data 514.3, and CRC checked using CRC values 512.2, for second layer processing. Second layer CRC mask 504 can be used to determine whether additional decoding should be processed using the second layer data and what additional decoding should be processed using the second layer data. For example, a codeword can be selectively passed through CRC mask 504 based on the number of invalid symbols in each permuted layer codeword for ECC decoding, and a determination of decoder processing to selectively decode that permuted layer codeword.

[0079] As shown in codeword 520, the permuted layer codeword can include no invalid symbols and be verified from an ECC perspective. In this case, no additional decoding in the second layer is needed. In some configurations, upon returning to the first layer, CRC can still be used to verify codeword 520 to provide a second level of verification of the CRC results. As shown for codeword 522, the permuted layer codeword can include a number of unverified symbols that is less than or equal to the correction capability of the codeword. In this case, the codeword can be successfully decoded in the permuted layer and return the correct values for those symbols. As shown for codeword 524, the number of unverified symbols from the unrecovered codeword in the first layer can be higher than the correction capability of the permuted layer codeword. The symbols that were not verified by the CRC of the layer 1 symbols in 524 can still be uncorrupted, and codeword 524 can be valid. In any case, decoding can still be performed on the permuted layer codeword to see if one or more symbols can be recovered, even if a full recovery can not be possible.

[0080] In some configurations, corrected symbols from an authorized layer can be mapped back to their respective codewords in the first layer for additional decoding attempts between each permutation layer. For example, in a configuration with a second and third permutation data layer, the decoding operation may include: ECC decoding and CRC check of the first layer, selective ECC decoding (based on CRC masking) of the second layer, a return to ECC decoding and CRC check of the first layer, selective ECC decoding (based on CRC masking) of the third layer, and a return to ECC decoding and CRC check of the first layer. After decoding the second-layer codeword, the verified symbol can be returned to the first layer to correct the corresponding symbol in the first layer. The first-layer codeword can be decoded again with the corrected symbol, allowing additional codewords to be successfully decoded and verified using CRC check. The same process of CRC masking and decoding can then be attempted with the next permutation data layer. This pattern can be repeated for any number of permutation data layers.

[0081] like Figure 6 As shown, the processing component 218 (or encoder) can operate according to an example method of encoding data units in an oligonucleotide pool using multilayer ECC (i.e., according to method 600 illustrated by boxes 610-642).

[0082] At box 610, user data for data blocks to be stored in DNA can be received. For example, the encoder can receive data units, such as data files or data objects, from a conventional binary data system.

[0083] At box 612, the number of oligonucleotides in the data block can be determined. For example, the encoder can be configured to include an oligonucleotide pool size that has a predetermined number of oligonucleotides.

[0084] At box 614, the error correction code can be determined. For example, the encoder can be configured with ECC settings for codeword size and ECC algorithms (such as the Reed-Solomon ECC algorithm).

[0085] At box 616, a number of substitution layers can be defined. For example, based on the aggregation symbols in the oligonucleotide pool, the size of the data units, and the desired redundancy, the encoder can be configured to use one or more substitution ECC layers for additional data protection.

[0086] At box 618, the primary codewords for the first layer can be determined. For example, the encoder can generate a codeword set for the data unit based on the ECC configuration, CRC, and data symbols from the user data unit.

[0087] In box 620, the user data symbols for the codeword can be determined. For example, the encoder can divide the data units into symbols and assign them to the main codeword.

[0088] At block 622, a CRC value for the codeword can be determined. For example, the encoder can compute CRC bits for the user data symbols and add them to the codeword.

[0089] At block 624, ECC redundancy data for the codeword can be computed. For example, the encoder can compute parity bits based on the ECC configuration and the user data symbols.

[0090] At block 626, a permutation data set for one or more permutation layers can be determined. For example, the encoder can rearrange symbols from the user data, CRC value, and redundancy data of the primary codeword into permutation data for a new codeword using a permutation rule.

[0091] At block 628, a permitted layer codeword can be determined. For example, the encoder can compute a new codeword using the ECC configuration and symbols from the permutation data.

[0092] At block 630, permutation data symbols can be determined. For example, the encoder can select a set of symbols for the new codeword from the permutation data.

[0093] At block 632, ECC redundancy data for the new codeword can be computed. For example, the encoder can compute redundancy data for the set of symbols from the permutation data to generate the new codeword.

[0094] At block 634, redundancy data symbols for a permutation layer can be determined. For example, for each permutation layer, the encoder can select redundancy data symbols for that permutation layer and permute them into a set of redundancy data symbols for the new codeword.

[0095] At block 636, ECC redundancy data can be computed for the permutation layer redundancy data. For example, the encoder can partition the redundancy data symbols and compute redundancy data that protects the new codeword for the permutation layer redundancy data. In some configurations, the encoder can also compute a CRC value for the new codeword.

[0096] At block 638, the permutation layer redundancy codeword can be added to the first layer. For example, the encoder can add the new codeword for the permutation layer redundancy data to the first layer codeword for each permutation layer.

[0097] At block 640, the codewords can be distributed among the oligonucleotides in the oligonucleotide pool. For example, the encoder can map the codewords to the oligonucleotides using the rule set and the symbol position matrix in the oligonucleotides such that the symbols in the same codeword are distributed across the oligonucleotides with no more than one symbol per codeword on any given oligonucleotide.

[0098] At block 642, write data for the pool of oligonucleotides can be output. For example, the encoder can output write data for the payload of each oligonucleotide in the set of oligonucleotides in the pool.

[0099] As shown, the processing component 248 (or decoder) can operate according to an example method of decoding data units from a pool of oligonucleotides using multi-layer ECC (i.e., according to the method 700 illustrated by blocks 710-728). Figure 7

[0100] At block 710, read data can be received from sequencing a set of oligonucleotides. For example, the decoder can receive read data from sequencing a pool of oligonucleotides corresponding to user data units, such as an address and a sequence of symbols for each oligonucleotide.

[0101] At block 712, first layer primary codewords can be determined. For example, the decoder can map symbols from the oligonucleotides into a matrix of symbol positions and identify a set of symbols for first layer decoding, such as symbols from one or more layers, and group them into codewords based on the matrix format.

[0102] At block 714, the first layer primary codewords can be decoded using ECC. For example, the decoder can decode each codeword in the first layer using an ECC decoder configuration, which can include primary codewords corresponding to user data symbols and permutation redundancy data codewords for one or more permutation layers.

[0103] At block 716, a CRC check can be performed. For example, the decoder can compute a CRC check for each codeword in the first layer based on CRC values in those codewords. If all primary codewords are successfully decoded and verified by the CRC check, the method 700 can proceed to block 728. Otherwise, the method 700 can proceed to block 718.

[0104] At block 718, a next permutation layer can be determined. For example, the encoder can determine, based on a format and / or configuration data for the data units and / or data pool, whether there are permutation layers stored in the pool of oligonucleotides in addition to the first layer and how many permutation layers are stored in the pool of oligonucleotides.

[0105] At block 720, permutation layer codewords can be determined. For example, the decoder can select a next set of codewords for the permutation layer corresponding to new codewords permuted from the first layer primary codewords, such as from another layer in the matrix.

[0106] At block 722, a CRC mask can be determined. For example, the decoder can map symbols from the primary codewords to permutation layer symbols and apply mask logic based on the primary codeword decoding output and CRC check to identify symbols that have not been successfully decoded.​

[0107] At box 724, a CRC mask can be applied to the permutation layer codeword. For example, the decoder can use symbols that have not yet been decoded to determine the permutation codeword to be decoded.

[0108] At box 726, the permutation layer codewords can be selectively decoded. For example, the decoder can perform ECC decoding on permutation layer codewords that include at least one symbol that has not yet been decoded. Method 700 can return to box 714 to iterate through additional decoding of the first layer, and, if necessary, through additional decoding of additional permutation layers.

[0109] like Figure 8 As shown, the DNA storage system 200 can operate according to an example method of storing data units distributed in an oligonucleotide pool of DNA data storage, namely method 800 illustrated in boxes 810-836. In some configurations, boxes 810-816 can be performed by an encoding system to encode the data units 802, boxes 818-826 can be performed using DNA synthesis, storage, and sequencing hardware to write and read the data units 804, and boxes 828-836 can be performed by a decoding system to decode the data units 806.

[0110] At box 810, the set of oligonucleotides for the data unit can be determined. For example, the storage system can generate or receive conventional binary data units for storage in DNA.

[0111] At box 812, ECC can be used to determine codewords. For example, the storage system can divide the codeword into symbols from data units and generate corresponding ECC redundancy data for those codewords based on the data unit symbols and ECC configuration.

[0112] At box 814, symbols can be assigned among the oligonucleotides in the oligonucleotide pool. For example, the storage system can distribute symbols from each codeword among the oligonucleotides in the oligonucleotide pool such that no oligonucleotide has more than one symbol from the same codeword.

[0113] At box 816, write data can be output. For example, the storage system can output write data corresponding to the symbols in the payload of the oligonucleotide set to be written.

[0114] At box 818, binary data can be converted into base pairs. For example, a storage system can provide instructions to convert the binary data sequence of each oligonucleotide into the corresponding sequence of base pairs that correspond to the binary data sequence.

[0115] At box 820, oligonucleotides can be synthesized. For example, the storage system may include a DNA synthesizer to synthesize a set of oligonucleotides encoding base pair sequences.

[0116] At box 822, oligonucleotides can be stored. For example, a storage system can place a synthetic set of oligonucleotides in a medium for storing the oligonucleotides for any period of time.

[0117] At box 824, oligonucleotides can be sequenced. For example, a storage system (or another storage system) can be processed by a sequencer to generate sequence data corresponding to the stored base pairs.

[0118] At box 826, sequence data can be converted into read data. For example, a storage system can sample sequencer signals as digital signal data.

[0119] At box 828, the read data can be received by the decoding system. For example, the storage system can provide digital read data from the sequencing of oligonucleotides to the decoding system.

[0120] At box 830, the symbol of the codeword can be determined from the oligonucleotide. For example, a storage system can organize the read data into a matrix based on the oligonucleotide address and the sequential set of symbols read from each oligonucleotide.

[0121] At box 832, codewords can be assembled from symbols derived from oligonucleotides. For example, the storage system can select a symbol set from the set of corresponding oligonucleotides that correspond to the codewords encoded during encoding 802.

[0122] At box 834, codewords can be decoded. For example, a storage system can use an ECC configuration where the codeword is encoded to decode it. In some configurations, this can be done as follows: Figure 6 and Figure 7 It uses multiple layers of encoding and decoding as described.

[0123] At box 836, a data unit can be output. For example, a storage system can reassemble a data unit from decoded symbols in one or more codewords and output that data unit for use by a conventional binary computing system.

[0124] The foregoing describes a technique for improving the encoding and decoding of DNA data storage using multilayer ECC distributed across oligonucleotide pools. In the above description, numerous specific details have been set forth for illustrative purposes. However, it will be apparent that the disclosed techniques can be practiced without any given subset of these specific details. In other instances, structures and devices are shown in block diagram form. For example, the disclosed techniques are described with reference to specific hardware in some of the specific embodiments described above.

[0125] References in the specification to "one embodiment" or "an embodiment” mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment or specific implementation of the technology disclosed. The appearances of the phrase "in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment or specific implementation.

[0126] Some portions of the detailed description included herein can be presented in terms of processes and symbolic representations of operations on data stored as bits in computer memory. These processes and symbolic representations can be the techniques used by those in the art to convey the substance of their work to others. A process is generally considered to be a self-consistent sequence of operations leading to a desired result. Operations can involve physical manipulations to physical quantities. These quantities can take the form of electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.

[0127] These and similar terms can be associated with the appropriate physical quantities and can be considered labels applied to these quantities. Unless specifically stated otherwise as apparent from the prior discussion, it is appreciated that throughout the description, discussions utilizing terms such as "processing" or "computing" or "calculating" or "determining" or "displaying" or the like can refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system memories registers and other such information storage, transmission or display devices.

[0128] The disclosed technology can also relate to an apparatus for performing the operations herein. This apparatus can be specially constructed for the required purposes, or it can comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a computer readable storage medium, such as, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs and magnetic disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memories including USB keys having non-volatile memory, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.

[0129] The disclosed technology can take form in a completely hardware embodiment, a completely software embodiment or an embodiment containing both hardware and software elements. In some embodiments, the technology is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0130] Moreover, the disclosed technology can take the form of a computer program product accessible from a non-transitory computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0131] A computing system or data processing system suitable for storing and / or executing program code will include at least one processor (e.g., hardware processor) coupled, either directly or through intervening circuitry, to memory elements. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memory providing temporary storage of at least some program code to reduce the number of times code must be retrieved from bulk storage during execution.

[0132] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system either directly or through intervening I / O controllers.

[0133] Network adapters can also be coupled to the system to enable the data processing system to become coupled to other data processing systems, remote printer / storage devices, through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters.

[0134] The terms storage medium, storage device, and data block are used interchangeably throughout this disclosure to refer to the physical medium that stores data.

[0135] Finally, the processes and displays presented herein can not inherently be related to any particular computer or other apparatus. Various general purpose systems can be used with programs in accordance with the teachings herein, or it can prove convenient to construct a more specialized apparatus to perform the required method operations. The required structure for a variety of these systems will be apparent from the description above. In addition, the disclosed technology is not described with reference to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the technology as described herein.

[0136] The foregoing description of implementations of the technology and science of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the technology and science of the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the technology and science of the application be limited not by this detailed description, but rather by the claims appended hereto. The technology and science of the application can be implemented in other specific forms without departing from the spirit or essential characteristics thereof. Likewise, the specific naming and division of modules, routines, features, attributes, methodologies and other aspects are not mandatory or significant, and the mechanisms that implement the technology and science of the application or its features can have different names, divisions and / or formats. Additionally, the modules, routines, features, attributes, methodologies and other aspects of the technology and science of the application can be implemented as software, hardware, firmware or any combination of the three. Further, where implemented as software, any component, including modules, can be implemented as a standalone program, as part of a larger program, as a plurality of separate programs, as a static or dynamic library, as a kernel loadable module, as a device driver, and / or in any other way now known or later developed. Also, the technology and science of the application are under no requirement to be implemented in any particular programming language or for any particular operating system or environment. Thus, the disclosure of the technology and science of the application is intended to be illustrative, but not limiting.

Claims

1. A system comprising: an encoder configured to: determine a set of oligonucleotides for encoding a data unit, wherein each oligonucleotide in the set of oligonucleotides encodes a plurality of symbols; determine a first codeword of an error correcting code, wherein the first codeword comprises a first set of symbols for encoding the data unit; allocate a symbol from the first set of symbols among a plurality of oligonucleotides from the set of oligonucleotides, wherein each oligonucleotide in the plurality of oligonucleotides receives one symbol from the first set of symbols; and output write data for the set of oligonucleotides to a synthesis interface for synthesizing the set of oligonucleotides.

2. The system of claim 1, wherein: the plurality of symbols are encoded in sequential positions along a length of each oligonucleotide; and the first set of symbols occupies the same sequential positions in the plurality of oligonucleotides.

3. The system of claim 1, wherein the first codeword comprises: a plurality of symbols corresponding to user data in the data unit; and at least one symbol corresponding to redundancy data of the error correcting code.

4. The system of claim 3, wherein the first codeword further comprises at least one symbol corresponding to a cyclic redundancy check value.

5. The system of claim 1, wherein the encoder is further configured to: determine a first set of codewords comprising a plurality of codewords corresponding to the data unit and a first set of redundancy data for the data unit, wherein the first set of codewords comprises the first codeword; determine at least one set of permutation data based on the data unit and the first set of redundancy data; determine a second set of codewords comprising: the at least one set of permutation data; and a second set of redundancy data for the at least one set of permutation data; and allocate symbols for the first set of codewords and the second set of codewords to the set of oligonucleotides.

6. The system of claim 5, wherein the encoder is further configured to: in response to determining the second set of redundancy data, add a codeword to the first set of codewords, the codeword comprising: the second set of redundancy data; and a third set of redundancy data for the second set of redundancy data.

7. The system of claim 6, wherein: the set of oligonucleotides stores an aggregate number of symbols; the at least one set of permutation data comprises a plurality of sets of permutation data based on the data unit and the first set of redundancy data; and the aggregate number of symbols is substantially equal to a number of symbols in a combination of the first set of codewords and the second set of codewords.

8. The system of claim 1, further comprising: a decoder configured to: receive read data determined from sequencing the set of oligonucleotides; determine the first set of symbols of the first codeword from the read data; assemble the first codeword; decode the first codeword using the error correcting code; and output the data unit based on the decoded first codeword.

9. The system of claim 8, wherein the decoder is further configured to: determining a first set of code words from the read data, the first set of code words comprising a plurality of code words corresponding to the data unit and a first set of redundancy data for the data unit, wherein the first set of code words comprises the first code word; determining a second set of code words from the read data, the second set of code words comprising: at least one permuted set of data based on the data unit and the first set of redundancy data; and a second set of redundancy data for the at least one permuted set of data; decoding the data unit using the first set of code words; and selectively decoding the second set of code words in response to failing to decode at least one code word in the first set of code words.

10. The system of claim 9, wherein the decoder is further configured to: determine a cyclic redundancy check value for each code word in the first set of code words; determine a verification mask by evaluating the cyclic redundancy check value for each code word in the first set of code words; and use the verification mask to determine a target code word for the selective decoding of the second set of code words.

11. A method, the method comprising: receiving read data determined from sequencing a set of oligonucleotides, wherein each oligonucleotide in the set of oligonucleotides encodes a plurality of symbols; determining a first set of symbols of a first code word from the read data, wherein the first set of symbols encodes a portion of a data unit using an error correcting code; assembling the first code word from the first set of symbols, wherein: the first set of symbols is distributed among a plurality of oligonucleotides from the set of oligonucleotides; and each oligonucleotide in the plurality of oligonucleotides receives one symbol from the first set of symbols; decoding the first code word using the error correcting code; and outputting the data unit based on the decoded first code word.

12. The method of claim 11, wherein: the plurality of symbols are encoded in sequential positions along a length of each oligonucleotide; and the first set of symbols occupies the same sequential positions in the plurality of oligonucleotides.

13. The method of claim 11, wherein the first code word comprises: a plurality of symbols corresponding to user data in the data unit; and at least one symbol corresponding to redundancy data for the error correcting code.

14. The method of claim 13, wherein the first code word further comprises at least one symbol corresponding to a cyclic redundancy check value.

15. The method of claim 11, further comprising: determining a first set of code words from the read data, the first set of code words comprising a plurality of code words corresponding to the data unit and a first set of redundancy data for the data unit, wherein the first set of code words comprises the first code word; determining a second set of code words from the read data, the second set of code words comprising: at least one permuted set of data based on the data unit and the first set of redundancy data; and a second set of redundancy data for the at least one permuted set of data; decoding the data unit using the first set of code words; and selectively decoding the second set of code words in response to failing to decode at least one code word in the first set of code words.

16. The method of claim 15, further comprising: determining a cyclic redundancy check value for each codeword in the first set of codewords; determining a validation mask by evaluating the cyclic redundancy check value for each codeword in the first set of codewords; and using the validation mask to determine a target codeword for the selective decoding of the second set of codewords.

17. The method of claim 15, further comprising: iteratively decoding the data unit by alternating between decoding using the first set of codewords and decoding using the second set of codewords, wherein: the first set of codewords corresponds to a first codeword layer that encoded the data unit; the at least one permuted data set comprises a plurality of permuted data sets based on the data unit and the first redundancy data set; and the second set of codewords corresponds to a plurality of additional codeword layers that encoded the plurality of permuted data sets.

18. The method of claim 11, further comprising: receiving the data unit; determining a first set of codewords comprising a plurality of codewords corresponding to the data unit and a first redundancy data set for the data unit, wherein the first set of codewords comprises the first codeword; determining at least one permuted data set based on the data unit and the first redundancy data set; determining a second set of codewords comprising: the at least one permuted data set; and a second redundancy data set for the at least one permuted data set; allocating symbols for the first set of codewords and the second set of codewords to the set of oligonucleotides; and outputting write data for the set of oligonucleotides to a synthesis interface for synthesizing the set of oligonucleotides.

19. The method of claim 18, further comprising: in response to determining the second redundancy data set, adding a codeword to the first set of codewords, the codeword comprising: the second redundancy data set; and a third redundancy data set for the second redundancy data set.

20. A system, comprising: means for receiving read data determined from sequencing a set of oligonucleotides, wherein each oligonucleotide in the set of oligonucleotides encodes a plurality of symbols; means for determining a first set of symbols for a first codeword from the read data, wherein the first set of symbols encodes a portion of a data unit using an error correcting code; means for assembling the first codeword from the first set of symbols, wherein: the first set of symbols is distributed among a plurality of oligonucleotides from the set of oligonucleotides; and each oligonucleotide in the plurality of oligonucleotides receives one symbol from the first set of symbols; means for decoding the first codeword using the error correcting code; and means for outputting the data unit based on the decoded first codeword.