Quality control of DNA data storage
QC methods for DNA data storage using QC polynucleotides with primer sequences and internal codecs address errors in polynucleotide pools, enhancing data integrity and reducing sequencing costs.
Patent Information
- Application Number
- JP2025543913
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-26
- Filing Date
- 2024-01-26
- Publication Date
- 2026-02-25
AI Technical Summary
DNA data storage is prone to errors and ambiguities during storage and sequencing, necessitating efficient quality control methods for verifying data integrity and correcting polynucleotide pools, which are often large and expensive to sequence fully.
Implementing quality control (QC) methods involving QC polynucleotides with primer sequences, amplification, sequencing, and alignment to estimate error rates and uniformity, as well as using internal codecs for probabilistic decoding to correct errors in polynucleotide pools.
Provides cost-effective and scalable quality control of polynucleotide pools, ensuring data integrity and reducing sequencing costs by using QC polynucleotides as surrogates for error estimation and correction.
Smart Images

Figure 2026506505000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 481,747, filed January 26, 2023, which is incorporated herein by reference in its entirety. All published patents, patents, and patent applications cited herein are incorporated herein by reference to the same extent as if each individual published patent, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Background technology]
[0002] background
[0002] DNA is a compelling data storage medium, given its superior density, stability, energy efficiency, and long lifespan compared to currently used electronic media. However, errors and ambiguities can be introduced or otherwise generated at or during various stages of storage, sequencing, and sequencing-related operations and processes. Therefore, there is a need to develop methods for efficiently performing DNA quality control. Summary of the Invention [Means for solving the problem]
[0003] overview
[0003] A goal of DNA data storage can be to provide a long-lasting backup, especially if other backups fail. In such cases, verifying proper synthesis and / or storage of polynucleotide pools can be crucial. However, these pools can be extremely large, and a typical full sequencing run can be very expensive and not scalable. In such cases, it may also be necessary to develop a method for quality control of polynucleotide pools over time. Quality control of polynucleotide pools can include verifying data integrity and / or quantifying DNA degradation. In some cases, quality control further includes correcting the polynucleotide pool or a subset thereof as needed. In this case, the original data in the polynucleotide pool is unavailable until the polynucleotide pool is fully decoded and / or sequenced.
[0004]
[0004] Provided herein are designs and embodiments for quality control of a plurality of polynucleotides. The plurality of polynucleotides can store digital information (e.g., binary data). The quality control of the plurality of polynucleotides can be performed before, during, and / or after synthesis or storage of the plurality of polynucleotides. The quality control of the plurality of polynucleotides can be performed in any suitable synthesis or storage device, such as those described herein. The quality control can further include quality control polynucleotides and / or one or more codecs, as further described herein. The provided designs and embodiments can provide cost-effective and / or scalable quality control of a plurality of polynucleotides.
[0005] In one aspect, provided herein is a quality control (QC) method for data polynucleotides, the method including: (i) providing a plurality of QC polynucleotides on a surface, wherein the plurality of QC polynucleotides comprises a first primer sequence; (ii) amplifying the plurality of QC polynucleotides based on the first primer sequence; (iii) sequencing the plurality of QC polynucleotides; and (iv) aligning the plurality of QC polynucleotides with a reference to estimate an error rate in the data polynucleotides, a composition uniformity in the data polynucleotides, or a combination thereof, for the data polynucleotide QC. In some cases, the error rate, composition uniformity, or a combination thereof is based at least in part on the relative read counts of the plurality of QC polynucleotides. In some cases, the plurality of QC polynucleotides is about 1% or less than 1% of the polynucleotides on the surface. In some cases, the plurality of QC polynucleotides is provided on a portion of the surface. In some cases, the plurality of QC polynucleotides is provided uniformly on the surface. In some cases, the data polynucleotides comprise a second primer sequence. In some cases, the first primer sequence is different from the second primer sequence. In some cases, the first primer sequence and the second primer sequence are different lengths. In some cases, quality control is performed after synthesis of the data polynucleotides. In some cases, QC is performed before cleavage of the data polynucleotides from the surface. In some cases, each of the QC polynucleotides is about 50 to 200 nucleobases in length. In some cases, each of the data polynucleotides is about 100 to about 300 nucleobases in length.
[0006]
[0006] Also provided herein is a quality control (QC) method for data polynucleotides, the method including: (i) selecting a subset of a plurality of data polynucleotides; (ii) applying an internal codec to the subset of the plurality of data polynucleotides, the internal codec including probabilistic decoding; and (iii) estimating an error rate in the plurality of polynucleotides based, at least in part, on a likelihood associated with each decoded sequence in the subset of the plurality of data polynucleotides. In some cases, a lower error rate is associated with a higher likelihood. In some cases, a higher error rate is associated with a lower likelihood. In some cases, the method further includes decoding an index for the subset of data polynucleotides. In some cases, the index is decoded using an internal codec, an external codec, or a combination thereof. In some cases, the index is used to estimate the relative distribution of the subset of the plurality of data polynucleotides. In some cases, QC is performed during synthesis of the polynucleotides, during QC of stored polynucleotides, or a combination thereof. In some cases, the subset of the plurality of data polynucleotides is randomly selected. In some cases, the subset of the plurality of data polynucleotides is selected based at least in part on their position on the surface. In some cases, the plurality of data polynucleotides comprises about 100,000 polynucleotides. In some cases, the subset of the plurality of data polynucleotides is about 0.1% of the plurality of data polynucleotides. In some cases, the method is used in combination with current sensing, optical imaging, flow sensing, size estimation, quality estimation, mass estimation, or any combination thereof. In some cases, the current sensing includes measuring a current through the tip or a section of the tip. In some cases, the current is compared to a reference value. In some cases, a difference between the current and the reference value indicates a tip failure, a deblocking failure, or a combination thereof. In some cases, the current sensing is performed prior to synthesis of the plurality of data polynucleotides.In some cases, current sensing is used to detect chip defects, adjust polynucleotide synthesis locations on the chip, or a combination thereof. In some cases, mass estimation is performed using fluorescence. In some cases, fluorescence is used to detect yields of multiple polynucleotides. In some cases, optical imaging includes detecting chip defects, non-uniformities, or a combination thereof.
[0007] Also provided herein is a method for performing QC of a plurality of cells on a surface, the method including: (i) measuring the current of each cell in the plurality of cells on the surface; (ii) determining whether one or more cells in the plurality of cells have a defect based at least in part on the current; and (iii) synthesizing and / or storing polynucleotides in a second one or more cells in the plurality of cells, wherein the second one or more cells do not have the defect. The defect includes a physical defect. In some cases, the surface is a synthesis surface, a storage surface, or a combination thereof. In some cases, the method further includes blocking one or more cells having the defect. In some cases, the blocking is performed by a protecting group on the surface. In some cases, the blocking is performed by a photolabile protecting group on the surface. In some cases, the blocking is performed by selectively applying energy to one or more cells. In some cases, the blocking is performed by a masking material. In some cases, the blocking is performed by addressable control of each cell in the plurality of cells.
[0008] BRIEF DESCRIPTION OF THE DRAWINGS
[0008] A better understanding of the features and advantages of the present subject matter will be obtained by reference to the following detailed description setting forth exemplary embodiments and the accompanying drawings. [Brief explanation of the drawings]
[0009] [Figure 1]
[0009] Figure 1 shows a non-limiting example of post-polynucleotide synthesis quality control according to some embodiments. [Figure 2]
[0010] 1 shows a non-limiting example of periodic quality control of stored polynucleotides, according to some embodiments. [Figure 3]
[0011] 1 illustrates one non-limiting example of digital information storage according to some embodiments. [Figure 4]
[0012] 1 illustrates a non-limiting example of generating a hash according to some embodiments. [Figure 5]
[0013] 1 illustrates a non-limiting example of an encoding scheme including an outer codec according to some embodiments. [Figure 6]
[0014] 1 illustrates one non-limiting example of an encoding scheme including shuffle lanes of binary data, according to some embodiments. [Figure 7]
[0015] 1 illustrates a non-limiting example of an encoding scheme including an inner codec, according to some embodiments. [Figure 8]
[0016] 1 illustrates a non-limiting example of a f-in-one scheme including alternative inner codecs, according to some embodiments. [Figure 9]
[0017] 1 illustrates a non-limiting example of a decoding scheme including an inner codec and an outer codec, according to some embodiments. [Figure 10]
[0018] 1 illustrates a non-limiting example of a greedy algorithm for decoding according to some embodiments. [Figure 11]
[0019] 1 illustrates a non-limiting example of a maximum likelihood (ML) algorithm for decoding according to some embodiments. [Figure 12]
[0020] A non-limiting example of a computing device is shown, where the device has one or more processors, memory, storage, and a network interface. DETAILED DESCRIPTION OF THE INVENTION
[0010] Detailed Description
[0021] Provided herein are methods and systems for quality control of digital information stored in nucleic acids. To make DNA data storage a viable option for long-term storage, scalable, efficient, and cost-effective methods for verifying proper synthesis and / or storage of polynucleotide pools can be crucial. Accordingly, provided herein are quality control methods for verifying digital information encoded in nucleic acids, referred to herein as data polynucleotides. In some cases, the quality control method involves using designated quality control (QC) polynucleotides that are synthesized and / or stored along with the data polynucleotides. In some cases, the quality control method involves verifying a subset of the data polynucleotides. The methods provided herein use the QC polynucleotides or the subset of data polynucleotides as a surrogate for estimating the error rate, uniformity, or a combination thereof, in the data polynucleotides.
[0011]
[0022] In some cases, the method provides quality control (QC) of the data polynucleotides. In some cases, the method includes providing a plurality of QC polynucleotides on a surface. In some cases, the plurality of QC polynucleotides includes a first primer sequence. In some cases, the method includes amplifying the plurality of QC polynucleotides based on the first primer sequence. In some cases, the method includes sequencing the plurality of QC polynucleotides. In some cases, the method includes aligning the plurality of QC polynucleotides with a standard. In some cases, the aligning is performed to estimate, for data polynucleotide QC, an error rate in the data polynucleotides, a synthetic uniformity in the data polynucleotides, or a combination thereof.
[0012]
[0023] In some cases, the method provides quality control (QC) of the data polynucleotides. In some cases, the method includes selecting a subset of the plurality of data polynucleotides. In some cases, the method includes applying an internal codec to the plurality of data polynucleotides. In some cases, the internal codec includes stochastic decoding. In some cases, the method includes estimating an error rate in the plurality of polynucleotides. In some cases, the error rate in the plurality of polynucleotides is based, at least in part, on a likelihood associated with each decoded sequence in the subset of the plurality of data polynucleotides.
[0013]
[0024] In some cases, the method is for performing quality control (QC) of a plurality of cells on a synthetic surface. In some cases, the method includes measuring the current of each cell in the plurality of cells on the surface. In some cases, the method includes determining whether one or more cells in the plurality of cells have a defect. In some cases, determining the defect is based at least in part on the current. In some cases, the method includes synthesizing and / or storing a polynucleotide in a second one or more cells in the plurality of cells. In some cases, the second one or more cells do not have a defect.
[0014]
[0025] Nucleic acid-based information storage
[0026] Provided herein are devices, compositions, systems, and methods for nucleic acid-based information (data) storage. Biomolecules, such as DNA molecules, provide suitable hosts for information storage, due in part to their stability over time and their ability to enhance information encoding, in contrast to traditional binary information encoding. In a first step, an information item (i.e., digital information in binary code for processing by a computer) is received. A cryptographic scheme is applied to convert the digital sequence from binary code to a nucleic acid sequence. A surface material for nucleic acid extension, a design of a locus for nucleic acid extension (also known as an arrangement spot), and reagents for nucleic acid synthesis are selected. The surface of the structure is prepared for nucleic acid synthesis. De novo polynucleotide synthesis is then performed. The synthesized polynucleotides are stored, in whole or in part, and made available for subsequent release. Once released, the polynucleotides are sequenced, in whole or in part, and subjected to decoding to convert the nucleic acid sequence back into the original digital sequence. The digital sequences are then assembled to obtain an aligned encoding of the original information item.
[0015]
[0027] Information item
[0028] Optionally, an initial step in the data storage process disclosed herein includes acquiring or receiving one or more information items in the form of initial code. Information items (e.g., digital information) include, but are not limited to, textual information, audio information, and visual information. Exemplary sources of information items include, but are not limited to, books, periodicals, electronic databases, medical records, letters, forms, audio recordings, animal records, biological profiles, broadcasts, movies, short videos, emails, bookkeeping phone logs, internet activity logs, drawings, paintings, printouts, photographs, pixelated graphics, and software code. Exemplary biological profile sources of information items include, but are not limited to, gene libraries, genomes, gene expression data, and protein activity data. Exemplary formats of information items include, but are not limited to, .txt, .PDF, .doc, .docx, .ppt, .pptx, .xls, .xlsx, .rtf, .jpg, .gif, .psd, .bmp, .tiff, .png, and .mpeg. The size of an individual file encoding an information item or the size of multiple files encoding information items, in digital format, may be, without limitation, up to 1024 bytes (1 KB equivalent), 1024 KB (1 MB equivalent), 1024 MB (1 GB equivalent), 1024 GB (1 TB equivalent), 1024 TB (1 PB equivalent), 1 exabyte, 1 zettabyte, 1 yottabyte, 1 xenotabyte, or more. In some cases, the amount of digital information is at least 1 gigabyte (GB). In some cases, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more gigabytes. In some cases, the amount of digital information is at least 1 terabyte (TB). In some cases, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more terabytes.In some cases, the amount of digital information is at least 1 petabyte (PB). In some cases, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more petabytes. In some cases, the digital information does not include genomic data obtained from an organism. In some cases, the information items are coded. Non-limiting example coding methods include 1 bit / base, 2 bits / base, 4 bits / base, or other coding methods.
[0016]
[0029] Systems and methods for quality control of polynucleotide pools
[0030] Provided herein are systems and methods for QC of one or more polynucleotide pools. In some cases, the one or more polynucleotide pools comprise data polynucleotides. In some cases, the data polynucleotides comprise digital information, such as binary data. In some cases, the digital information comprises an item of information, such as, but not limited to, those described herein. In some cases, the one or more polynucleotide pools comprise one or more items of information. In some cases, the one or more items of information are encoded by data polynucleotides in the one or more polynucleotide pools.
[0017]
[0031] Provided herein are systems and methods for QC of data polynucleotides. The data polynucleotides can encode digital information as described herein. In some cases, the digital information is encoded as a data polynucleotide using the systems and methods described herein. However, in some cases, the QC methods described herein are independent of the systems and methods for encoding digital information into polynucleotides. In some cases, the QC methods described herein are independent of the size or type of digital information encoded in the polynucleotides. In some cases, the QC methods described herein are independent of the size of the polynucleotide pools described herein.
[0018]
[0032] In some cases, the QC of a data polynucleotide comprises multiple QC polynucleotides. The QC polynucleotides can be provided on a surface, e.g., for synthesis and / or storage of data polynucleotides, such as those described herein. In some cases, the QC polynucleotides are synthesized simultaneously with the data polynucleotides. In some cases, the QC polynucleotides are provided on a surface or a portion of a surface. In some cases, the QC polynucleotides are provided heterogeneously on a surface or a portion of a surface. In some cases, the portion of a surface comprises one discrete location (e.g., a gene locus, a cell, a mechanism, etc.) or multiple discrete locations on the surface. In some cases, the QC polynucleotides are about 1% to about 10% of the polynucleotides on the surface. In some cases, the polynucleotides on the surface comprise QC polynucleotides and data polynucleotides encoding digital information. In some cases, QC polynucleotides comprise about 1% to about 2%, about 1% to about 3%, about 1% to about 4%, about 1% to about 5%, about 1% to about 6%, about 1% to about 7%, about 1% to about 8%, about 1% to about 9%, about 1% to about 10%, about 2% to about 3%, about 2% to about 4%, about 2% to about 5%, about 2% to about 6%, about 2% to about 7%, about 2% to about 8%, about 2% to about 9%, about 2% to about 10%, about 3% to about 4%, about 3% to about 5%, about 3% to about 6%, about 3% to about 7%, about 3% to about 8%, about 3% to about 9%, about 3% to about 10%, about 4% to about 5%, about 4% to about 6%, about 4% to about 7%, about 4% to about 8%, about 4% to about 9%, about 4% to about 10%, about 5% to about 6%, about 5% to about 7%, about 5% to about 8%, about 5% to about 9%, about 5% to about 10%, about 6% to about 7%, about 6% to about 8%, about 6% to about 9%, about 6% to about 10%, about 7% to about 8%, about 7% to about 9%, about 7% to about 10%, about 8% to about 9%, about 8% to about 10%, and about 9% to about 10%. In some cases, the QC polynucleotides are about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the polynucleotides on the surface.In some cases, QC polynucleotides are at least about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, or about 9% of the polynucleotides on the surface, and in some cases, QC polynucleotides are at most about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the polynucleotides on the surface.
[0019]
[0033] In some cases, the length of each of the plurality of QC polynucleotides is about 20 to about 500 bases. In some cases, the length of each of the plurality of QC polynucleotides is about 20 to about 50 bases, about 20 to about 100 bases, about 20 to about 200 bases, about 20 to about 300 bases, about 20 to about 400 bases, about 20 to about 500 bases, about 50 to about 100 bases, about 50 to about 200 bases, about 50 to about 300 bases, about 50 to about 400 bases, The length of each of the plurality of QC polynucleotides is about 50 bases to about 500 bases, about 100 bases to about 200 bases, about 100 bases to about 300 bases, about 100 bases to about 400 bases, about 100 bases to about 500 bases, about 200 bases to about 300 bases, about 200 bases to about 400 bases, about 200 bases to about 500 bases, about 300 bases to about 400 bases, about 300 bases to about 500 bases, or about 400 bases to about 500 bases. In some cases, the length of each of the plurality of QC polynucleotides is about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases. In some cases, the length of each of the plurality of QC polynucleotides is at least about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, or about 400 bases. In some cases, the length of each of the plurality of QC polynucleotides is at most about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases.
[0020]
[0034] In some cases, each of the multiple QC polynucleotides comprises a first primer sequence. In some cases, the first primer sequence of each of the multiple QC polynucleotides is different from the second primer sequence of each of the data polynucleotides. In some cases, the first primer sequence of each of the multiple QC polynucleotides is different in length from the second primer sequence of each of the data polynucleotides. In some cases, the first primer sequence is unique to the polynucleotide pool. As an example, a first plurality of QC polynucleotides for QC of a first polynucleotide pool encoding a first file can have a different primer sequence from a second plurality of QC polynucleotides for QC of a second polynucleotide pool encoding a second file. Alternatively, the polynucleotide pool can include two or more files, and each of the two or more QC polynucleotides can have a primer sequence unique to the QC of each file. Alternatively, multiple QC polynucleotides with a single primer sequence can be used for QC of multiple polynucleotide pools. In some cases, the length of the first primer sequence is about 10 bases to about 50 bases.In some cases, the length of the first primer sequence is from about 10 bases to about 15 bases, from about 10 bases to about 18 bases, from about 10 bases to about 20 bases, from about 10 bases to about 22 bases, from about 10 bases to about 25 bases, from about 10 bases to about 28 bases, from about 10 bases to about 30 bases, from about 10 bases to about 35 bases, from about 10 bases to about 40 bases, from about 10 bases to about 45 bases, from about 10 bases to about 50 bases, from about 15 bases to about 18 bases, from about 15 bases to about 20 bases, from about 15 bases to about 22 bases, from about 15 bases to about 25 bases, from about 15 bases to about 30 bases, from about 15 ...5 bases to about 40 bases, from about 15 bases to about 45 bases, from about 15 bases to about 50 bases, from about 15 bases to about 18 bases, from about 15 bases to about 20 bases, from about 15 bases to about 22 bases, from about 15 bases to about 25 bases, from about 1 bases to about 28 bases, about 15 bases to about 30 bases, about 15 bases to about 35 bases, about 15 bases to about 40 bases, about 15 bases to about 45 bases, about 15 bases to about 50 bases, about 18 bases to about 20 bases, about 18 bases to about 22 bases, about 18 bases to about 25 bases, about 18 bases to about 28 bases, about 18 bases to about 30 bases, about 18 bases to about 35 bases, about 18 bases to about 40 bases, about 18 bases to about 45 bases, about 18 bases to about 50 bases, about 20 bases to about 22 bases, about 20 bases to about 25 bases, about 20 bases bases to about 28 bases, about 20 bases to about 30 bases, about 20 bases to about 35 bases, about 20 bases to about 40 bases, about 20 bases to about 45 bases, about 20 bases to about 50 bases, about 22 bases to about 25 bases, about 22 bases to about 28 bases, about 22 bases to about 30 bases, about 22 bases to about 35 bases, about 22 bases to about 40 bases, about 22 bases to about 45 bases, about 22 bases to about 50 bases, about 25 bases to about 28 bases, about 25 bases to about 30 bases, about 25 bases to about 35 bases, about 25 bases to about 4 ... The length of the first primer sequence may be about 45 bases, about 25 bases to about 50 bases, about 28 bases to about 30 bases, about 28 bases to about 35 bases, about 28 bases to about 40 bases, about 28 bases to about 45 bases, about 28 bases to about 50 bases, about 30 bases to about 35 bases, about 30 bases to about 40 bases, about 30 bases to about 45 bases, about 30 bases to about 50 bases, about 35 bases to about 40 bases, about 35 bases to about 45 bases, about 35 bases to about 50 bases, about 40 bases to about 45 bases, about 40 bases to about 50 bases, or about 45 bases to about 50 bases. In some cases, the length of the first primer sequence is about 10 bases, about 15 bases, about 18 bases, about 20 bases, about 22 bases, about 25 bases, about 28 bases, about 30 bases, about 35 bases, about 40 bases, about 45 bases, or about 50 bases.In some cases, the length of the first primer sequence is at least about 10 bases, about 15 bases, about 18 bases, about 20 bases, about 22 bases, about 25 bases, about 28 bases, about 30 bases, about 35 bases, about 40 bases, or about 45 bases, In some cases, the length of the first primer sequence is at most about 15 bases, about 18 bases, about 20 bases, about 22 bases, about 25 bases, about 28 bases, about 30 bases, about 35 bases, about 40 bases, about 45 bases, or about 50 bases.
[0021]
[0035] The QC polynucleotides described herein can be extracted and / or amplified. In some cases, multiple QC polynucleotides are amplified. In some cases, multiple QC polynucleotides are amplified based on their primer sequences (e.g., first primer sequences). In some cases, multiple QC polynucleotides are extracted and / or amplified from the surface on which they are synthesized or stored. After extraction and / or amplification of the QC polynucleotides from the surface of the structure, the polynucleotides can be sequenced using a suitable sequencing technique, as further described herein. In some cases, the DNA sequence is read on the substrate or within the structure.
[0022]
[0036] In some cases, the plurality of QC polynucleotides are aligned. In some cases, the plurality of QC polynucleotides are aligned with a reference. In some cases, the reference is a known sequence or a preselected sequence. In some cases, the known sequence or preselected sequence is the source sequence of the plurality of QC polynucleotides. In some cases, the plurality of QC polynucleotides are aligned with a reference to estimate the error rate in the data polynucleotides, the composite uniformity in the data polynucleotides, or a combination thereof.
[0023]
[0037] The error rate in data polynucleotides can be estimated by aligning multiple QC polynucleotides with a reference. In some cases, as described herein above, the reference is a known sequence. In some cases, aligning multiple QC polynucleotides with a reference generates a relative read count. In some cases, the relative read count includes the number of sequence QC polynucleotides that have the same sequence as the reference sequence. In some cases, the relative read count is used to estimate the error rate in multiple QC polynucleotides by identifying the number of sequenced QC polynucleotides that have the same sequence as the reference sequence from all sequenced QC polynucleotides. In some cases, the error rate in multiple QC polynucleotides is used to estimate the error rate in data polynucleotides. In some cases, the error rate in data polynucleotides is based at least in part on the relative read count.
[0024]
[0038] The synthetic uniformity in data polynucleotides can be estimated by aligning multiple QC polynucleotides with a reference. In some cases, the reference is a known sequence, as described herein above. In some cases, aligning multiple QC polynucleotides with a reference generates relative read counts, as described herein. In some cases, the relative read counts are used to estimate the synthetic uniformity in multiple QC polynucleotides by identifying the number of sequenced QC polynucleotides that have the same sequence as the reference from all sequenced QC polynucleotides. In some cases, the relative read counts are used to estimate the synthetic uniformity in one or more discrete locations (e.g., loci, cells, mechanisms, etc.) where the sequenced QC polynucleotides have the same sequence as the reference sequence. For example, a particular cell from a plurality of cells may be determined to contain fewer QC polynucleotides or contain QC polynucleotides that have fewer alignments with the reference compared to QC polynucleotides in other cells. In some cases, the synthetic uniformity in multiple QC polynucleotides is used to estimate the synthetic uniformity in data polynucleotides. In some cases, the synthetic uniformity in data polynucleotides is based at least in part on the relative read counts.
[0025]
[0039] In some cases, polynucleotide QC methods comprising QC polynucleotides described herein are performed after synthesis of data polynucleotides. In some cases, polynucleotide QC methods comprising QC polynucleotides described herein are performed after initial synthesis of data polynucleotides, as exemplarily shown in FIG. 1. In some cases, polynucleotide QC methods comprising QC polynucleotides described herein are performed after resynthesis of data polynucleotides in the event of a synthesis or storage error. In some cases, polynucleotide QC methods comprising QC polynucleotides described herein are performed on stored data polynucleotides.
[0026]
[0040] In some cases, QC of the data polynucleotides includes selecting a subset of the plurality of data polynucleotides. In some cases, the subset of the plurality of data polynucleotides is randomly selected. In some cases, the subset of the plurality of data polynucleotides is pseudo-randomly selected. In some cases, the subset of the plurality of data polynucleotides is selected based, at least in part, on their location on a synthesis or storage surface, such as described herein. In some cases, the subset of the plurality of data polynucleotides is selected based on one or more physical or chemical properties, such as, but not limited to, as measured by current sensing, optical imaging, flow sensing, etc. In some cases, the subset of the plurality of data polynucleotides comprises about 0.01% to about 5% of the plurality of data polynucleotides.In some cases, the subset of the plurality of data polynucleotides is from about 0.01% to about 0.02%, about 0.01% to about 0.05%, about 0.01% to about 0.08%, about 0.01% to about 0.1%, about 0.01% to about 0.2%, about 0.01% to about 0.5%, about 0.01% to about 1%, about 0.01% to about 2%, about 0.01% to about 3%, about 0.01% to about 4%, about 0.01% to about 5%, about 0.02% to about 0.05, about 0.01% to about 1. 0.02% to about 0.08%, about 0.02% to about 0.1%, about 0.02% to about 0.2%, about 0.02% to about 0.5%, about 0.02% to about 1%, about 0.02% to about 2%, about 0.02% to about 3%, about 0.02% to about 4%, about 0.02% to about 5%, about 0.05% to about 0.08%, about 0.05% to about 0.1%, about 0.05% to about 0.2%, about 0.05% to about 0.5%, about 0.05% to about 1%, about 0.05% to about 2%, about 0.05% to about 3%, about 0.0 5% to about 4%, about 0.05% to about 5%, about 0.08% to about 0.1%, about 0.08% to about 0.2%, about 0.08% to about 0.5%, about 0.08% to about 1%, about 0.08% to about 2%, about 0.08% to about 3%, about 0.08% to about 4%, about 0.08% to about 5%, about 0.1% to about 0.2%, about 0.1% to about 0.5%, about 0.1% to about 1%, about 0.1% to about 2%, about 0.1% to about 3%, about 0.1% to about 4%, about 0.1 to about 5%, about 0.2 to Including about 0.5%, about 0.2% to about 1%, about 0.2% to about 2%, about 0.2% to about 3%, about 0.2% to about 4%, about 0.2% to about 5%, about 0.5% to about 1%, about 0.5% to about 2%, about 0.5% to about 3%, about 0.5% to about 4%, about 0.5% to about 5%, about 1% to about 2%, about 1% to about 3%, about 1% to about 4%, about 1% to about 5%, about 2% to about 3%, about 2% to about 4%, about 2% to about 5%, about 3% to about 4%, about 3% to about 5%, or about 4% to about 5%. In some cases, the subset of the plurality of data polynucleotides comprises about 0.01%, about 0.02%, about 0.05%, about 0.08%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, or about 5% of the plurality of data polynucleotides.In some cases, the subset of the plurality of data polynucleotides comprises at least about 0.01%, about 0.02%, about 0.05%, about 0.08%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, or about 4% of the plurality of data polynucleotides. In some cases, the subset of the plurality of data polynucleotides comprises at most about 0.02%, about 0.05%, about 0.08%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, or about 5% of the plurality of data polynucleotides.
[0027]
[0041] In some cases, the polynucleotide pool comprises a plurality of data polynucleotides. In some cases, the plurality of data polynucleotides comprises about 100 to 500,000 polynucleotides. In some cases, the plurality of data polynucleotides is about 100 to about 500, about 100 to about 1,000, about 100 to about 5,000, about 100 to about 10,000, about 100 to about 50,000, about 100 to about 100,000, about 100 to about 200,000, about 100 to about 300,000, about 100 to about 400,000, about 100 to about 500,000, about 500 to about 1,000, about 500 to about 5,000, about 500 to about 10,000 0, about 500 to about 50,000, about 500 to about 100,000, about 500 to about 200,000, about 500 to about 300,000, about 500 to about 400,000, about 500 to about 500,000, about 1,000 to about 5,000, about 1,000 to about 10,000, about 1,000 to about 50,000, about 1,000 to about 100,000, about 1,000 to about 200,000, about 1,000 to about 300,000, about 1,000 to about 400 ,000, about 1,000 to about 500,000, about 5,000 to about 10,000, about 5,000 to about 50,000, about 5,000 to about 100,000, about 5,000 to about 200,000, about 5,000 to about 300,000, about 5,000 to about 400,000, about 5,000 to about 500,000, about 10,000 to about 50,000, about 10,000 to about 100,000, about 10,000 to about 200,000, about 10,000 to about 30 0,000 approximately, about 10,000 to about 400,000 approximately, about 10,000 to about 500,000 approximately, about 50,000 to about 100,000 approximately, about 50,000 to about 200,000 approximately, about 50,000 to about 300,000 approximately, about 50,000 to about 400,000 approximately, about 50,000 to about 500,000 approximately, about 100,000 to about 200,000 approximately, about 100,000 to about 300,000 approximately, about 100,000 to about 400,000 approximately, about 100,000 to about 500,In some cases, the plurality of data polynucleotides comprises about 100, about 500, about 1,000, about 5,000, about 10,000, about 50,000, about 100,000, about 200,000, about 300,000, about 400,000, about 500,000, or about 500,000 polynucleotides. In some cases, the plurality of data polynucleotides comprises at least about 100, about 500, about 1,000, about 5,000 to 10,000, about 50,000, about 100,000, about 200,000, about 300,000, or about 400,000 polynucleotides. In some cases, the plurality of data polynucleotides comprises at most about 500, about 1,000, about 5,000 to 10,000, about 50,000, about 100,000, about 200,000, about 300,000, about 400,000, or about 500,000 polynucleotides.
[0028]
[0042] In some cases, the internal codec is applied to a subset of the plurality of data polynucleotides. In some cases, the internal codec includes probabilistic decoding. The internal codec generally includes decoding polynucleotides into digital information. In some cases, the internal codec includes converting or transforming each of the subset of the plurality of data polynucleotides into binary data. In some cases, the entire length of the subset of the plurality of data polynucleotides is transformed or converted into binary data (e.g., full decoding). In some cases, a partial length of the subset of the plurality of data polynucleotides is transformed or converted into binary data (e.g., partial decoding). In some examples, the partial length includes an index (e.g., a lane index, a frame index, a UUID index, a content ID, etc.) such as described herein. In some cases, the internal codec is applied to a subset of the plurality of data polynucleotides that has or has not been ordered, aligned, clustered, or any combination thereof.
[0029]
[0043] In some cases, the plurality of data polynucleotides and / or a subset of the data polynucleotides are encoded using the methods described herein. In some cases, the plurality of data polynucleotides and / or a subset of the plurality of data polynucleotides are decoded using the methods described herein. In some cases, the inner codec comprises a greedy algorithm. In some cases, the inner codec comprises a maximum likelihood (ML) algorithm. In some cases, the inner codec comprises a greedy ML hybrid algorithm.
[0030]
[0044] In some cases, the probabilistic decoding of the inner codec provides a likelihood of the overall decoded sequence. In some cases, redundancy within each polynucleotide sequence helps estimate the error rate without knowing the reference polynucleotide. As an example, if the inner codec decodes a sequence with high probability and / or takes very few steps, the error rate is likely to be low. As a further example, if the inner codec decodes a sequence with low probability and / or takes more steps, the error rate is likely to be high.
[0031]
[0045] In some cases, the data polynucleotides include an index. In some cases, the index of a subset of the plurality of data polynucleotides is decoded. In some cases, the index is decoded using an internal codec, an external codec, or a combination thereof, such as, but not limited to, those described herein. In some cases, the index is used to estimate the relative distribution of the subset of the plurality of polynucleotides. In some examples, the relative distribution is used to estimate the uniformity of the data polynucleotides. For example, if the plurality of data polynucleotides includes approximately 100,000 polynucleotide sequences and the selected subset is 0.1% of the data polynucleotides, a distribution centered around a decoded index of around 100 can be expected. In some examples, the relative distribution varies between the subsets of data polynucleotides, indicating a loss of uniformity across the data polynucleotides.
[0032]
[0046] In some cases, a data polynucleotide QC method comprising selecting a subset as described herein is performed after synthesis of the data polynucleotides. In some cases, a data polynucleotide QC method comprising selecting a subset as described herein is performed after initial synthesis of the data polynucleotides, as exemplarily shown in Figure 1. In some cases, a data polynucleotide QC method comprising selecting a subset as described herein is performed after re-synthesis of the data polynucleotides in the face of a synthesis or storage error. In some cases, a data polynucleotide QC method comprising selecting a subset as described herein is performed on stored data polynucleotides, as exemplarily shown in Figure 2.
[0033]
[0047] In some cases, the polynucleotide QC methods and systems are used in conjunction with one or more additional QC methods. In some cases, the one or more additional methods include any one of current sensing, resistance sensing, optical imaging, flow sensing, size estimation, quality estimation, mass estimation, or any combination thereof. In some cases, sensing the current or resistance includes measuring the current or resistance of a chip or a section of a chip, respectively. In some cases, the current or resistance is compared to a reference value. In some cases, a difference between the current or resistance and the reference value indicates a chip failure, a synthesis error, or a combination thereof. In some examples, the chip failure includes a defect in the chip. In some cases, the defect in the chip causes polynucleotide synthesis and / or storage problems. In some examples, the synthesis error includes a deblocking failure. In some cases, mass estimation includes measuring absorbance to estimate the mass of a polynucleotide sequence. In some cases, mass estimation includes using fluorescence to measure the mass of a polynucleotide sequence. In some cases, fluorescence is used to detect the yield of a plurality of polynucleotides. In some cases, optical imaging includes detecting chip defects, non-uniformities, or a combination thereof. In some cases, flow sensing is used to detect the flow of liquids, gases, or combinations thereof across the synthesis and / or storage chip.
[0034]
[0048] An exemplary flow diagram for QC of a polynucleotide pool is provided in FIG. 1. As shown, amperometric sensing can be employed prior to synthesis to provide quality control for the synthesis chip. Pre-synthesis amperometric sensing can be performed to detect chip defects, adjust polynucleotide synthesis locations on the chip, or a combination thereof. This can be followed by synthesis geometry optimization, followed by synthesis of data polynucleotides. In some cases, data polynucleotides are synthesized along with QC polynucleotides, as described herein above. In some cases, synthesis of multiple data polynucleotides is performed along with sequential QC using one or more additional methods described herein. In some cases, sequential QC includes amperometric sensing, resistive sensing, optical imaging, flow sensing, or a combination thereof. In some cases, post-synthesis QC includes determining the oligo length distribution and / or mass estimate of the multiple data polynucleotides, for example, using the techniques described herein.
[0035]
[0049] In some cases, the QC polynucleotides are amplified and sequenced for QC of the data polynucleotides, as described herein above. In some cases, the QC polynucleotides are fully sequenced. In some cases, the QC polynucleotides are partially sequenced. In some cases, the QC polynucleotides are aligned and the error rate and / or uniformity are estimated, as described herein above.
[0036]
[0050] Alternatively or in combination, in some cases, the data polynucleotides are amplified and a subset (or subsample) is sequenced. In some cases, the subset is fully decoded. In some cases, the subset is partially decoded. In some cases, the partial decoding includes applying an inner codec, an outer codec, or a combination thereof. In some examples, an inner codec is applied to estimate an error rate, as described above, in some cases. In some cases, the index of the subset is partially decoded. In some examples, an outer codec is applied to estimate uniformity, as described above, in some cases.
[0037]
[0051] In some cases, QC of the PC polynucleotides, data polynucleotides, or a combination thereof is used to determine a final QC decision. In some cases, the final QC decision is based on error rate, uniformity, or both. In some cases, the final QC decision includes passing or failing the synthesized data polynucleotides. In some cases, if the final QC decision is passing, the data polynucleotides are stored. In some cases, if the final QC decision is failing, the data polynucleotides are resynthesized. In some cases, the final QC decision includes passing some sections of the synthesis surface and failing some sections of the synthesis surface. In some cases, only data polynucleotides from passing sections of the synthesis surface are stored. In some cases, data polynucleotides from failing sections of the synthesis surface are resynthesized.
[0038]
[0052] In some cases, the final QC decision includes a threshold. In some cases, the threshold includes a static value or a dynamic value. In some cases, the threshold includes a static range or a dynamic range. In some cases, the threshold is based on one or more combined values or ranges, such as error rate, uniformity, or both. In some cases, if the error rate is low and the uniformity is high, the final QC decision is pass. For example, the error rate can be less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.0005%, or less than 0.0001%. In further examples, the uniformity can be greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9%, greater than 99.95%, or greater than 99.99%. In some cases, when the error rate is high and the uniformity is low, the final QC decision is pass. For example, the error rate may be greater than 10%, greater than 15%, greater than 20%, greater than 25%, greater than 30%, greater than 35%, greater than 40%, greater than 45%, or greater than 50%. In a further example, the uniformity may be less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%.
[0039]
[0053] A further exemplary flow diagram for QC of a polynucleotide pool is provided in FIG. 2 , where a plurality of data polynucleotides in the pool have already been synthesized and stored. In some cases, QC of the data polynucleotides in the pool is performed periodically. In some cases, QC of the data polynucleotides is performed over days, weeks, months, or years. In some cases, a plurality of data polynucleotides are retrieved from storage. In some cases, the data polynucleotides are amplified and a subset (or subsample) is sequenced. In some cases, the subset is fully decoded. In some cases, the subset is partially decoded. In some cases, the partial decoding includes applying an internal codec, an external codec, or a combination thereof. In some examples, an internal codec is applied to estimate an error rate, as described above, in this specification. In some cases, an index of the subset is partially decoded. In some examples, an external codec is applied to estimate uniformity, as described above, in this specification.
[0040]
[0054] In some cases, QC of the inspected data polynucleotides further includes pool-level QC, which includes determining an oligo length distribution and / or mass estimate of a plurality of data polynucleotides, for example, using the techniques described herein. In some cases, as shown in FIG. 2, pool-level QC is combined with subset QC to determine a final QC judgment. In some cases, the final QC judgment includes passing or failing the synthesized data polynucleotides. In some cases, if the final QC judgment is passing, the data polynucleotides are returned to storage. In some cases, if the final QC judgment is failing, the data polynucleotides are fully sequenced and / or decoded. In some cases, fully sequencing or decoding the data polynucleotides includes sequencing and / or decoding duplicate data polynucleotides if a sample of the data polynucleotide cannot be decoded. In some cases, the data polynucleotides are resynthesized. In some cases, the final QC judgment includes passing some sections of the storage surface and failing some sections of the storage surface. In some cases, only data polynucleotides from the passing sections of the storage surface are returned to storage. In some cases, data polynucleotides from the rejected sections of the storage surface are fully sequenced, decoded, and / or resynthesized.
[0041]
[0055] In some cases, the final QC decision includes a threshold value. In some cases, the threshold value includes a static value or a dynamic value. In some cases, the threshold value includes a static range or a dynamic range. In some cases, the threshold value is based on one or more combined values or ranges, such as error rate, uniformity, or both. In some cases, if the error rate is low and the uniformity is high, the final QC decision is pass. For example, the error rate can be less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.0005%, or less than 0.0001%. In further examples, the uniformity can be greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9%, greater than 99.95%, or greater than 99.99%. In some cases, when the error rate is high and the uniformity is low, the final QC decision is pass. For example, the error rate may be greater than 10%, greater than 15%, greater than 20%, greater than 25%, greater than 30%, greater than 35%, greater than 40%, greater than 45%, or greater than 50%. In a further example, the uniformity may be less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%.
[0042]
[0056] Also provided herein are systems and methods for performing quality control (QC) of a plurality of cells on a surface, such as a synthesis surface or a storage surface. In some cases, the cells comprise active areas on the surface, such as compartments, locations, loci, features, spots, or any combination thereof, suitable for polynucleotide synthesis and / or storage. In some cases, the surface comprises, but is not limited to, a synthesis and / or storage surface of a device, such as those described herein.
[0043]
[0057] In some cases, the method for performing QC of a plurality of cells includes one or more steps. In some cases, the step includes measuring a physical property and / or a chemical property of each of the plurality of cells on the surface. In some cases, the step includes amperometric sensing, resistive sensing, optical imaging, or a combination thereof. In some cases, a voltage is applied to the surface. In some cases, the voltage is about 0.1 V to about 3 V. In some cases, the voltage is about 0.1 V to about 0.25 V, about 0.1 V to about 0.5 V, about 0.1 V to about 0.75 V, about 0.1 V to about 1 V, about 0.1 V to about 1.25 V, about 0.1 V to about 1.5 V, about 0.1 V to about 1.75 V, about 0.1 V to about 2 V, about 0.1 V to about 2.5 V, about 0.1 V to about 3 V, about 0.25 V to about 0.5 V, about 0.25 V to about 0.75 V, about 0.25 V to about 1V, about 0.25V to about 1.25V, about 0.25V to about 1.5V, about 0.25V to about 1.75V, about 0.25V to about 2V, about 0.25V to about 2.5V, about 0.25V to about 3V, about 0.5V to about 0.75V, about 0.5V to about 1V, about 0.5V to about 1.25V, about 0.5V to about 1.5V, about 0.5V to about 1.75V, about 0.5V to about 2V, about 0.5V to about 2.5V, about 0. 5V to about 3V, about 0.75V to about 1V, about 0.75V to about 1.25V, about 0.75V to about 1.5V, about 0.75V to about 1.75V, about 0.75V to about 2V, about 0.75V to about 2.5V, about 0.75V to about 3V, about 1V to about 1.25V, about 1V to about 1.5V, about 1V to about 1.75V, about 1V to about 2V, about 1V to about 2.5V, about 1V to about 3V, about 1.25V to about 1.5 V, about 1.25 V to about 1.75 V, about 1.25 V to about 2 V, about 1.25 V to about 2.5 V, about 1.25 V to about 3 V, about 1.5 V to about 1.75 V, about 1.5 V to about 2 V, about 1.5 V to about 2.5 V, about 1.5 V to about 3 V, about 1.75 V to about 2 V, about 1.75 V to about 2.5 V, about 1.75 V to about 3 V, about 2 V to about 2.5 V, about 2 V to about 3 V, or about 2.5 V to about 3 V. In some cases, the voltage is about 0.1 V, about 0.25 V, about 0.5 V, about 0.75 V, about 1 V, about 1.25 V, about 1.5 V, about 1.75 V, about 2 V, about 2.5 V, or about 3 V.In some cases, the voltage is at least about 0.1 V, about 0.25 V, about 0.5 V, about 0.75 V, about 1 V, about 1.25 V, about 1.5 V, about 1.75 V, about 2 V, or about 2.5 V. In some cases, the voltage is at most about 0.25 V, about 0.5 V, about 0.75 V, about 1 V, about 1.25 V, about 1.5 V, about 1.75 V, about 2 V, about 2.5 V, or about 3 V. In some cases, the process includes measuring the current, resistance, or a combination thereof, of each of the plurality of cells on the surface when the voltage is applied.
[0044]
[0058] In some cases, current sensing, resistance sensing, optical imaging, or a combination thereof is used to determine whether one or more cells in the plurality of cells have a defect. In some cases, the defect includes a physical defect in the surface. In some cases, the defect is determined at least in part based on current, resistance, or a combination thereof. In some cases, the current, resistance, or a combination thereof measured in a cell having a defect differs from a corresponding standard value measured in a cell without the defect. In some cases, the current, resistance, or a combination thereof measured in a cell having a defect differs from a corresponding standard value by about 1% to about 40%. In some cases, the current, resistance, or a combination thereof may be between about 1% and about 5%, about 1% and about 10%, about 1% and about 15%, about 1% and about 20%, about 1% and about 25%, about 1% and about 30%, about 1% and about 35%, about 1% and about 40%, about 5% and about 10%, about 5% and about 15%, about 5% and about 20%, about 5% and about 25%, about 5% and about 30%, about 5% and about 35%, about 5% and about 40%, about 10% and about 15%, about 10% and about 20%, about 10% and about 25%, %, about 10% to about 30%, about 10% to about 35%, about 10% to about 40%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 35%, about 15% to about 40%, about 20% to about 25%, about 20% to about 30%, about 20% to about 35%, about 20% to about 40%, about 25% to about 30%, about 25% to about 35%, about 25% to about 40%, about 30% to about 35%, about 30% to about 40%, or about 35% to about 40%. In some cases, the current, resistance, or a combination thereof differs by about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, or about 40%. In some cases, the current, resistance, or a combination thereof differs by at least about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, or about 35%. In some cases, the current, resistance, or a combination thereof differs by at most about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, or about 40%.
[0045]
[0059] In some cases, if a defect is detected, one or more cells having the defect are blocked. In some cases, blocking one or more cells includes leaving a protecting group, such as DMT, on the surface. In some cases, blocking one or more cells includes avoiding deblocking of the protecting group. In some cases, blocking includes one or more photolabile protecting groups, and hydroxyl groups generated on the surface are blocked by the photolabile protecting groups. For example, when the surface is exposed to UV light, such as through a photolithography mask, a pattern of free hydroxyl groups can be generated on the surface. These hydroxyl groups can react with a photoprotective nucleoside phosphoramidite according to phosphoramidite chemistry. In some examples, a second photolithography mask can be applied, and the surface can be exposed to UV light to generate a second pattern of hydroxyl groups, followed by coupling with a 5'-photoprotective nucleoside phosphoramidite. Similarly, patterns can be generated and oligomer chains can be extended. Without being bound by theory, the lability of the photocleavable group depends on the wavelength and polarity of the solvent employed, and the rate of photocleavage can be affected by the duration of exposure and the intensity of the light. This method can be leveraged by many factors, such as the accuracy of mask alignment, the efficiency of photoprotective group removal, and the yield of the phosphoramidite coupling step.
[0046]
[0060] In some cases, blocking further comprises selectively supplying energy to one or more cells. In some cases, a mask is created on a surface through a heating element on or near the surface. In some cases, a layer of masking material is applied to the surface, and a heating element is employed to apply energy to selected portions of the masking material, whereby the applied energy causes a phase transition in the selected portions of the masking material such that the masking material adheres to the surface or can displace the masking material from the surface to mask or unmask the selected portions, respectively. In some cases, the masking material is a solid, gas, liquid, or a combination thereof. In some cases, the masking material is, for example, a C 15~C 30 n-alkanes (e.g., tetracosane (C 24 ), Icosan (C 20 ) etc. In some cases, the masking material may be, for example, C 16 ~C 30 n-alkane or C 18 ~C 28 It includes a mixture of two or more straight chain higher alkanes, such as n-alkanes. In some cases, the masking material, e.g., nanospheres, can be deposited on the surface in the form of a dispersion, e.g., in acetonitrile.
[0047]
[0061] In some cases, the blocking includes addressable locations on the surface. In some cases, the locations are addressable through one or more electrodes near the locations. In some cases, one or more electrodes are independently addressable. In some cases, one or more electrodes at each location on the surface are independently addressable. In some cases, each electrode controls nucleoside (nucleoside phosphoramidite) binding through electrochemistry at a specific location on the surface. In some cases, the reagent includes a protein or other acidic molecule. In some cases, the electrodes are positioned at locations around the edge of the well surface. In some cases, the electrodes control chemical reactions that occur near the synthesis surface. For example, if acidic or other reagents are generated near the synthesis surface, portions of the polynucleotide bound to this surface will come into contact with a higher concentration of acid than portions of the polynucleotide farther from the acid generation site. This can lead to degradation of portions of the polynucleotide exposed to higher concentrations of acid. In some cases, electrodes, such as those positioned near the surface of the well, generate or control a proton gradient that leads to uniform or targeted exposure of portions of the polynucleotide to acid. Sites near the uncharged electrodes do not bind to nucleosides deposited on the synthesis surface, and the pattern of charged electrodes is altered before the addition of the next nucleoside. By applying a series of electrode-controlled masks to the surface, desired polynucleotides are synthesized at precise locations on the surface.
[0048]
[0062] In some cases, the polynucleotide is synthesized and / or stored in a second one or more cells of the plurality of cells. In some cases, the second one or more cells do not have a defect. In some cases, the polynucleotide is synthesized and / or stored using the systems and methods described herein. In some cases, the polynucleotide is encoded according to the systems and methods described herein.
[0049]
[0063] Digital information storage system and method
[0064] Provided herein are methods and systems for storing digital information. In some cases, the digital information includes one or more objects. In some cases, the one or more objects include items of information, such as, but not limited to, those described herein. In some cases, the one or more objects include files or metadata associated with files. In some cases, the digital information includes binary data. In some cases, the binary data is a byte stream or byte array. In some cases, each of the one or more objects is between about 1 GB and about 1 TB. In some cases, each of the one or more objects is between about 1 GB and about 1 TB. In some cases, each of the one or more objects is about 1 GB to about 10 GB, about 1 GB to about 50 GB, about 1 GB to about 100 GB, about 1 GB to about 500 GB, about 1 GB to about 1 TB, about 10 GB to about 50 GB, about 10 GB to about 100 GB, about 10 GB to about 500 GB, about 10 GB to about 1 TB, about 50 GB to about 100 GB, about 50 GB to about 500 GB, about 50 GB to about 1 TB, about 100 GB to about 500 GB, about 100 GB to about 1 TB, or about 500 GB to about 1 TB. In some cases, each of the one or more objects is about 1 GB, about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB. In some cases, each of the one or more objects is at least about 1 GB, about 10 GB, about 50 GB, about 100 GB, or about 500 GB. In some cases, each of the one or more objects is at most about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB.
[0050]
[0065] A system for storing digital information may include one or more processing units, a memory in communication with the one or more processing units, where instructions are stored in the memory and executed by the one or more processing units, or any combination thereof. In some cases, the one or more processing units and the memory are distributed across one or more physical or logical locations. In some cases, the one or more processing units include a central processing unit (CPU), a graphical processing unit (GPU), a single-core processor, a multi-core processor, a processor cluster, an application-specific integrated circuit (ASIC), a programmable circuit such as a field-programmable gate array (FPGA), an AI accelerator, and any combination thereof. In some cases, one or more of the processing units include a single-instruction, multiple-data (SIMD) or single-program, multiple-data (SPMD) parallel architecture. As an example, the one or more processing units include one or more GPUs or CPUs implementing SIMD or SPMD. In some cases, the AI accelerator includes Google-TPU, Graphcore, Cerebras, SambaNova, or a combination thereof. In some embodiments, one or more of the processing units are implemented in software and / or firmware in addition to hardware implementations. A software or firmware implementation of a processing unit may include computer- or machine-executable instructions written in any suitable programming language to perform various functions described herein. A software implementation of one or more processing units may be stored, in whole or in part, in memory. Alternatively or additionally, the system may include one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard packages (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc.In some cases, the memory includes removable, non-removable, local, and / or remote storage to provide instructions, data structures, program modules (e.g., hashing modules), and any other data described herein. In some cases, the memory is used to store information (e.g., software code, parameters, executable instructions, etc.) related to the algorithms described herein.
[0051]
[0066] The instructions stored in the memory can include one or more steps of storing the digital information. In some cases, the one or more steps include dividing the digital information of one or more objects into multiple pools. In some cases, each of the multiple pools is about 1 GB to about 1 TB. In some cases, each of the multiple pools is about 1 GB to about 10 GB, about 1 GB to about 50 GB, about 1 GB to about 100 GB, about 1 GB to about 500 GB, about 1 GB to about 1 TB, about 10 GB to about 50 GB, about 10 GB to about 100 GB, about 10 GB to about 500 GB, about 10 GB to about 1 TB, about 50 GB to about 100 GB, about 50 GB to about 500 GB, about 50 GB to about 1 TB, about 100 GB to about 500 GB, about 100 GB to about 1 TB, or about 500 GB to about 1 TB. In some cases, each of the plurality of pools is about 1 GB, about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB. In some cases, each of the plurality of pools is at least about 1 GB, about 10 GB, about 50 GB, about 100 GB, or about 500 GB. In some cases, each of the plurality of pools is at most about 10 GB, about 50 GB, about 100 GB, about 500 GB, or about 1 TB.
[0052]
[0067] In some cases, one or more objects include an item of information, such as a file, as described herein above. In some cases, one or more objects include metadata associated with the item of information (e.g., metadata associated with a file). Non-limiting examples of metadata associated with an object include a list of keywords attached to the object, object size, a thumbnail image, a text summary, an ID range in a sorted key-value database, a timestamp of the object, a version, or any other data that provides information about one or more aspects of the object, or any combination thereof. In some examples, the metadata is customizable. In some examples, the metadata is used to search for objects within multiple pools.
[0053]
[0068] An exemplary diagram of a digital information store is shown in Figure 3. As shown, one or more objects 305 can be divided into multiple pools 310. In some cases, a single object is divided into multiple pools. In some cases, a single object is divided into 2, 3, 4, 5, 6, 7, 8, 9, or 10 pools. In some cases, two or more objects are divided into multiple pools. In some cases, one or more objects are in a pool. In some cases, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 objects are in a pool. In some cases, multiple pools are replicated. In some cases, multiple pools comprise information pools, where two or more pools contain the same one or more objects. In some cases, 2, 3, 4, 5, 6, 7, 8, 9, or 10 pools contain the same one or more objects.
[0054]
[0069] Each pool in the plurality of pools can include any one or combination of pool descriptors, pool items, or end descriptors. In some cases, a pool includes at least one pool item. In some cases, a pool includes two or more pool items. In some cases, a pool includes at least one pool descriptor. In some cases, a pool includes two or more pool descriptors. In some cases, a pool includes at least one end descriptor. In some cases, a pool includes two or more end descriptors. As an example, each pool includes a pool descriptor 315, one or more pool items 320, and an end descriptor 325. In some cases, a pool includes redundant pool items, pool descriptors, pool end descriptors, or combinations thereof. In such cases, two or more pool items, pool descriptors, pool end descriptors, or combinations thereof are identical. In some cases, 2, 3, 4, 5, 6, 7, 8, 9, or 10 pool descriptors, pool end descriptors, or combinations thereof are identical.
[0055]
[0070] In some cases, the one or more steps in the instructions include generating a pool descriptor, a pool item, an end descriptor, or any combination thereof for each pool of the plurality of pools. In some cases, the pool descriptor includes a version, a pool ID, a list of pool item descriptors, or any combination thereof. In some cases, the version includes a version of the information (e.g., if the information has been updated). In some cases, the pool ID includes a unique ID for the pool. In some examples, the unique ID includes a universally unique identifier (UUID). In some examples, the unique ID includes a content ID. In some examples, the content ID includes a digital fingerprinting system, which can be used to identify and / or manage copyright or ownership of content. In some cases, the list of pool item descriptors includes a path to the object, a size of the object (e.g., a total size of the object), a range of pool items within the object, an offset of the pool item within the pool, or any combination thereof. In some examples, the range of pool items within the object includes one or more locations of payloads in pool items within the object. In some examples, the one or more locations include a start and / or end range of a payload in a pool item (e.g., lines 1-6 of pool item 1 in a pool, lines 7-13 of pool item 2, etc.). In some examples, a pool item offset includes one or more payload locations in a pool. In some cases, a pool item includes a data payload and / or a hash of the pool item. In some cases, the data payload includes a stored object or portion of an object. In some cases, a pool item hash includes a hash value of a stored object or portion of an object. In some cases, a pool end descriptor includes a list of object descriptors. In some cases, a list of object descriptors includes an object path and / or an object hash. In some examples, an object path includes a unique path. In some examples, an object path includes a hierarchy (e.g., a directory hierarchy). In some examples, an object path does not include a hierarchy.
[0056]
[0071] Systems and methods for storing digital information can include one or more hashes. In some cases, the one or more hashes are determined using a hashing module. In some cases, the hashing module executes on one or more processing units, such as described herein. In some cases, the hashing module includes instructions (e.g., a hash function) for determining the one or more hashes. In some cases, the instructions (e.g., the hash function) are stored in a memory, such as described herein. In some cases, information including objects, portions of objects, or pool items is stored using hashes. In some cases, one or more first hashes of each of one or more objects are determined and / or one or more second hashes of each of one or more pool items are determined. In some cases, hashes of pool items are attached to a data payload. In some cases, hashes of objects are attached to a pool end descriptor.
[0057]
[0072] A hash may be specified as a hash function (FIG. 4). A hash function generally includes a function that converts an input of any length into an output of a fixed length (e.g., 224, 256, 384, 512 bits or characters). In some cases, the hash function includes a cryptographic hash function. In some cases, the hash function includes MD-5, SHA-1, SHA-2, SHA-3, RIPEMD-160, Whirlpool, BLAKE, BLAKE2, BLAKE3, or variations thereof. In some cases, the hash function includes SHA-2. In some examples, SHA-2 includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256. The output of a hash function can be deterministic and impossible to reverse engineer. Furthermore, generating a fixed-length output can increase security because any party involved in cracking the hash cannot know the length of the input. In some examples, a hash is generated upon input of an identification code, a cryptographic key, a password, or any variation thereof. In some examples, the hash allows for verification of the contents (e.g., an information item or digital information stored in a pool) during decryption.
[0058]
[0073] In some cases, input 405 includes an object. In some examples, hash function 410 is used to determine hash output (or hash) 415. In some examples, input 420 includes an object. In some examples, hash function 425 is used to determine hash output (or hash) 430. In some examples, hash function 410 and hash function 425 are the same hash function. In some examples, hash function 410 and hash function 425 are both SHA-256. In some examples, hash function 410 and hash function 425 are different hash functions. In some examples, output 415 and output 430 are the same length. In some examples, output 415 and output 430 are both 256 bits. In some examples, output 415 and output 430 are different lengths.
[0059]
[0074] A hash function can include one or more steps to generate a hash. In some cases, one or more steps in a hash function include padding bits. In some cases, additional bits are added to the digital information (or message) during the hash calculation. In some examples, the additional bits are added to the message such that the length of the digital message is the modulus value less than the total number of bits. In some examples, the modulus value is 64 bits. In some examples, the number of bits is 512 bits and the length of the digital information is 448 bits (e.g., for SHA-256). In some examples, the first additional bit has a binary value of 1. In some examples, subsequent additional bits have a binary value of 0.
[0060]
[0075] In some cases, one or more steps in the hash function include length padding. In some cases, length padding includes adding a modulus value to the digital information (e.g., also called a bi-endian (BE) integer). The modulus value or BE integer generally represents the length of the original input, which contains the original digital information in binary. In some examples, the modulus value is 64 bits. In some examples, 64 bits are added to a 448-bit digital message, for a total number of bits of 512 (e.g., in the case of SHA-256). In some cases, the modulus value is calculated by applying the modulus to the original digital information. In one example, if the original digital information is the binary value "hello world," the length of the original input is 88 bits, and 88 is the binary value "1011000." Therefore, a 0 followed by "1011000" is added to the end of the 448-bit digital information so that the total number of bits is 512.
[0061]
[0076] In some cases, one or more steps in the hash function include initializing one or more hash values or buffers. In some cases, eight hash values or buffers are initialized. In some cases, the initialized hash values are hard-coded (e.g., constants). In some cases, the initialized hash values represent the first 32 bits of the fractional part of the square root of the first eight prime numbers (e.g., 2, 3, 5, 7, 11, 13, 17, 19). In some cases, one or more steps in the hash function further include initializing round constants (or keys). In some cases, 64 round constants are initialized. In some examples, each of the 64 round constants represents the first 32 bits of the fractional part of the cube root of the first 64 prime numbers (e.g., 2 to 311). In some cases, the 64 different round constants are stored in an array.
[0062]
[0077] In some cases, one or more steps in the hash function include compression. In some cases, each information block (e.g., every 512 bits) undergoes compression. During compression, each information block goes through a fixed number of rounds. In some cases, the number of rounds is 64 or less. In some cases, compression is performed by a one-way compression function. In some cases, the one-way compression function is a single block length compression function. In some examples, the compression function is a Davies-Meyer, Matyas-Meyer-Oseas, or Miyaguchi-Preneel compression function. In some cases, the one-way compression function is a double block length compression function. In some examples, the compression function is an MDC-2 / Meyer-Schilling, MDC-4, or Hirose compression function. In some cases, the output from the compression function is smaller than the information block. In some examples, the output is 256 bits long.
[0063]
[0078] In some cases, one or more of the hashes (e.g., pool item hashes, object hashes) are calculated while the information is being stored. In some cases, all of the hashes (e.g., pool item hashes, object hashes) are calculated while the information is being stored. In some examples, this allows for consistently low memory usage regardless of the size of the object. In some cases, the first one or more hashes of each of one or more objects require less memory than the one or more objects. In some cases, the second one or more hashes of each of one or more pool items require less memory than the one or more pool items. In some cases, the source data (e.g., information item 9) is read only once. In some cases, each pool is written once with no seeks. In some examples, this minimizes data transfer and latency.
[0064]
[0079] In some cases, the hashes described herein may serve one or more purposes, including, by way of non-limiting example, one or more of the following: integrity of one or more information items (e.g., objects), signature generation and verification (e.g., in the case of digital signatures), password verification, proof-of-work, or an identifier for an information item.
[0065]
[0080] In some cases, further encryption and / or compression can be added. In some examples, encryption and / or compression is performed using a streaming application programmable interface (API). In some examples, this avoids the need to store intermediate results. In some cases, the digital information to be stored is already compressed, e.g., to reduce data transfer costs. In some cases, the digital information to be stored is already encrypted, e.g., for security reasons.
[0066]
[0081] One or more steps in the instructions stored in the memory can further include creating a plurality of index pools. In some cases, the plurality of index pools includes only indexes. In some cases, the index pools are used to retrieve objects stored in the plurality of pools encoded by the plurality of polynucleotides. In some cases, the index pools are sequenced and temporarily stored in a digital storage system (e.g., a flash drive) for searching for objects. In some examples, once the pools are identified, the plurality of polynucleotides encoding the pools are sequenced.
[0067]
[0082] In some cases, one or more index pools include an index pool descriptor and / or a list of object indexes. In some cases, the index pool descriptor includes a version, a pool ID, a pool size, a timestamp, or a combination thereof. In some examples, the pool ID includes a unique ID for the pool. In some examples, the unique ID includes a universally unique identifier (UUID). In some examples, the unique ID includes a content ID. In some examples, the content ID includes a digital fingerprinting system, which can be used to identify and / or manage copyright or ownership of content. In some examples, the size of each of the multiple index pools is about 1 GB to about 1 TB. In some cases, the list of object indexes includes patches of the object, hashes of the object, a list of object fragments, a list of object metadata, or any combination thereof. In some examples, the object path includes a unique path. In some examples, the object path includes a hierarchy (e.g., a directory hierarchy). In some examples, the object path does not include a hierarchy. In some examples, the object hash is a hash (e.g., SHA-256) as described herein. In some examples, the list of object fragments includes the pool ID of the pool that includes the fragments, ranges of fragments, or a combination thereof. In some examples, the list of object metadata includes a metadata type, a metadata payload, or a combination thereof. In some examples, the metadata type includes a list of keywords attached to the object, a thumbnail image, a text summary, a sorted key-value database ID range, a timestamp, a version, or any combination thereof. In some examples, the metadata is customizable. In some examples, the metadata is used to search for objects in multiple pools.
[0068]
[0083] In some cases, the index pool can store information from about 1 to about 1 million pools. In some cases, the index pool can store information from about 1 to about 10 pools, from about 1 to about 100 pools, from about 1 to about 1,000 pools, from about 1 to about 5,000 pools, from about 1 to about 10,000 pools, from about 1 to about 50,000 pools, from about 1 to about 100,000 pools, from about 1 to about 500,000 pools, from about 1 to about 1 million pools, from about 10 to about 100 pools, from about 10 to about 100 pools, from about 1,000 pools, from about 10 pools to about 5,000 pools, from about 10 pools to about 10,000 pools, from about 10 pools to about 50,000 pools, from about 10 pools to about 100,000 pools, from about 10 pools to about 500,000 pools, from about 10 pools to about 1 million pools, from about 100 pools to about 1,000 pools, from about 100 pools to about 5,000 pools, from about 100 pools to about 10,000 pools, from about 100 pools to about 50,000 pools pools, from about 100 pools to about 100,000 pools, from about 100 pools to about 500,000 pools, from about 100 pools to about 1 million pools, from about 1,000 pools to about 5,000 pools, from about 1,000 pools to about 10,000 pools, from about 1,000 pools to about 50,000 pools, from about 1,000 pools to about 100,000 pools, from about 1,000 pools to about 500,000 pools, from about 1,000 pools to about 1 million pools, about 5,0 00 pools to about 10,000 pools, about 5,000 pools to about 50,000 pools, about 5,000 pools to about 100,000 pools, about 5,000 pools to about 500,000 pools, about 5,000 pools to about 1 million pools, about 10,000 pools to about 50,000 pools, about 10,000 pools to about 100,000 pools, about 10,000 pools to about 500,000 pools, about 10,000 pools to about 1 million pools, about 50,10,000 to about 100,000 pools, about 50,000 to about 500,000 pools, about 50,000 to about 1 million pools, about 100,000 to about 500,000 pools, about 100,000 to about 1 million pools, or about 500,000 to about 1 million pools. In some cases, the index pools may store information for about 1 pool, about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, about 500,000 pools, or about 1 million pools. In some cases, the index pool can store information for at least about 1 pool, about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, or about 500,000 pools. In some cases, the index pool can store information for at most about 10 pools, about 100 pools, about 1,000 pools, about 5,000 pools, about 10,000 pools, about 50,000 pools, about 100,000 pools, about 500,000 pools, or about 1 million pools.
[0069]
[0084] In some cases, each of the one or more index pools is about 1 GB to about 1 TB. In some cases, each of the multiple pools is about 1 GB to about 1 TB. In some cases, each of the one or more index pools is about 1 GB to about 10 GB, about 1 GB to about 50 GB, about 1 GB to about 100 GB, about 1 GB to about 500 GB, about 1 GB to about 1 TB, about 10 GB to about 50 GB, about 10 GB to about 100 GB, about 10 GB to about 500 GB, about 10 GB to about 1 TB, about 50 GB to about 100 GB, about 50 GB to about 500 GB, about 50 GB to about 1 TB, about 100 GB to about 500 GB, about 100 GB to about 1 TB, or about 500 GB to about 1 TB. In some cases, each of the one or more index pools is about 1 GB, 10 GB, 50 GB, 100 GB, 500 GB, or about 1 TB. In some cases, each of the one or more index pools is at least about 1 GB, 10 GB, 50 GB, 100 GB, or 500 GB. In some cases, each of the one or more index pools is at most about 10 GB, 50 GB, 100 GB, 500 GB, or about 1 TB.
[0070]
[0085] Encoding method
[0086] An encoding scheme can be applied to each of the multiple pools and / or index pools. In some cases, the encoding scheme encodes the digital information into the multiple pools as multiple polynucleotides. In some cases, the encoding scheme encodes the digital information into the index pool as multiple polynucleotides. In some cases, the encoding scheme includes a codec (e.g., an internal codec) that encodes binary data as nucleic acid sequences. In some cases, the encoding scheme includes an error correction code (ECC) (e.g., an external codec). In some cases, the encoding scheme (e.g., an internal codec or a low-level codec) is also designed and implemented to enable streaming read / write API access. In some cases, the encoding scheme (e.g., an internal codec or a low-level codec) is also designed and implemented to be compatible with streaming digital storage systems and / or methods (e.g., an external codec or a high-level codec) described herein.
[0071]
[0087] An encoding scheme generally includes one or more operations. The one or more operations may include one or more operations that manipulate or transform data (e.g., digital information). The one or more operations may include, by way of non-limiting example, splitting, shuffling, concatenating, transposing, translating, duplicating, labeling (e.g., using an index), or combinations thereof, of data or portions of data.
[0072]
[0088] As an example, a method for encoding digital or binary data into a plurality of nucleotide sequences can include dividing the binary data into a plurality of frames. In some cases, the plurality of frames can include about 100 to about 10,000 frames. In some cases, the plurality of frames can include about 100 to about 250 frames, about 100 to about 500 frames, about 100 to about 750 frames, about 100 to about 1,000 frames, about 100 to about 2,500 frames, about 100 to about 5,000 frames, about 100 to about 7,500 frames, about 100 to about 10,000 frames, or about 250 to about 500 frames. , approximately 250 frames to approximately 750 frames, approximately 250 frames to approximately 1,000 frames, approximately 250 frames to approximately 2,500 frames, approximately 250 frames to approximately 5,000 frames, approximately 250 frames to approximately 7,500 frames, approximately 250 frames to approximately 10,000 frames, approximately 500 frames to approximately 750 frames, approximately 500 frames to approximately 1,000 frames, approximately 500 frames to approximately 2,500 frames, approximately 500 frames to approximately 5 ,000 frames, approximately 500 frames to approximately 7,500 frames, approximately 500 frames to approximately 10,000 frames, approximately 750 frames to approximately 1,000 frames, approximately 750 frames to approximately 2,500 frames, approximately 750 frames to approximately 5,000 frames, approximately 750 frames to approximately 7,500 frames, approximately 750 frames to approximately 10,000 frames, approximately 1,000 frames to approximately 2,500 frames, approximately 1,000 frames to approximately 5,00 The range may include 0 frames, about 1,000 frames to about 7,500 frames, about 1,000 frames to about 10,000 frames, about 2,500 frames to about 5,000 frames, about 2,500 frames to about 7,500 frames, about 2,500 frames to about 10,000 frames, about 5,000 frames to about 7,500 frames, about 5,000 frames to about 10,000 frames, or about 7,500 frames to about 10,000 frames.In some cases, the plurality of frames comprises about 100 frames, about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, about 7,500 frames, or about 10,000 frames. In some cases, the plurality of frames comprises at least about 100 frames, about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, or about 7,500 frames. In some cases, the plurality of frames comprises at most about 250 frames, about 500 frames, about 750 frames, about 1,000 frames, about 2,500 frames, about 5,000 frames, about 7,500 frames, or about 10,000 frames. In some cases, the frames each comprise the same amount of data. Alternatively, the frames each comprise different amounts of data. In some cases, each frame is assigned a frame index. In some examples, the frame index increases with each frame index (e.g., 0, 1, 2, 3, 4, 5, ..., etc.). In some examples, the frame index increases monotonically with each frame index.
[0073]
[0089] The method for encoding digital or binary data includes an external codec. In some cases, the method for encoding digital or binary data into a plurality of nucleotide sequences includes an external codec. In some cases, the external codec is applied to the binary data. In some cases, the external codec is applied to the binary data when the binary data is divided into a plurality of frames. In such cases, the external codec is applied to each of the plurality of frames. An exemplary diagram of dividing a data stream into frames and applying an external codec is illustratively shown in FIG. 5.
[0074]
[0090] In some cases, the outer codec includes an error correction code or scheme, such as a Reed-Solomon (RS) code. This outer codec is used to spread the digital or binary data to be stored across many oligonucleotides. In some cases, the spreading of the data builds in redundancy to correct erasures (e.g., oligo losses). In some further embodiments, the spreading of the data also builds in redundancy to correct errors from the inner codec.
[0075]
[0091] In some cases, the error correction scheme includes a Reed-Solomon (RS) code. In such cases, an RS encoder is used to encode binary data or multiple frames containing binary data. Generally, RS codes operate on blocks of data that are treated as a set of finite field elements. In some cases, RS codes operate on blocks of data, e.g., x=(x1,...,x k )∈F k Let p be the polynomial x where:
number
[0076]
[0092] In some further embodiments, the RS code includes an encoding scheme in which each codeword contains a message as a prefix and error correction symbols are appended as a suffix. In some cases, the RS code is designated as RS(n,k) using m-bit symbols. In such cases, the encoder takes k data symbols of m bits each and appends parity symbols (error correction symbols or check symbols) to create an n-symbol codeword, where there are nk parity symbols (or check symbols, t) of m bits each. In some cases, the RS decoder corrects up to t symbols in the codeword that contain errors, where 2t = nk. The codeword C(x) contains parity check information CK(x) that is systematically appended to the message information M(x). The codeword C(x) is expressed as C(x) = x n-k M(x)+CK(x)=x n-k M(x)+x n-k It can be calculated as M(x) mod g(x), where k is the message length (e.g., symbols), t is the number of errors to be corrected, n is the block length (e.g., message length n plus correction length t), and m is the symbol width. Given the symbol size m, the maximum codeword length n in an RS code is n=2. m -1. Furthermore, x n-k refers to the displacement shift in the message, and g(X) refers to the generator polynomial, defined as a polynomial whose roots are successive powers of a primitive element α of the Galois Field (GF) (e.g., g(x)-(xα i )(xa i+1 )…(xa i+n-k-1 )-g0+g1x+…+g n-k-1 x n-1-1 +x n-k ).
[0077]
[0093] For example, in RS(255,223) using 8-bit symbols, the block length n is 255 codeword bytes, the message length k is 223 bytes, and the parity 2t is 32 bytes. In such an example, the RS decoder corrects up to 16 symbol errors within the codeword, meaning that up to 16 byte errors can be corrected by the decoder. RS codes are based on a Galois field in GF(2 m ) can also be shown as RS GF(2 12 ) encoding method, n is 4095 (e.g., n=2 12 -1=4096-1=4095). If k is, for example, 2499, then 2t=4095-2499=1596, and therefore t is 798.
[0078]
[0094] In some cases, the error correction scheme includes a linear error correcting code (or linear block code), such as a low-density parity check (LDPC) code. In some cases, the error correction scheme includes a linear block error correcting code, such as a polar code. In some further embodiments, the error correction scheme includes a high-performance forward error correction (FEC), such as a turbo code. In some cases, the error correction scheme includes an RS code, an LDPCv, a turbo code, a polar code, or any combination thereof (e.g., an RS-based LDPC code).
[0079]
[0095] In some cases, the error correction scheme includes a low-density parity-check (LDPC) code. In such cases, the LDPC code is used to encode binary data or multiple frames containing binary data. Generally, the structure of an LDPC code is defined by a parity-check matrix with most entries containing zeros and others containing ones. For example, an (N,K) LDPC code for K information bits is a linear block code of block size N defined by a sparse (NK) × N parity-check matrix where all elements other than 1 are zeros. The number of ones in a row or column is called the degree of the row or column. In some cases, a codeword of length N is represented as a vector C, and in the case of length K information bits, an (N,K) code with 2K codewords is used. In some cases, an (N,K) LDPC code is defined by an (NK) × N parity-check matrix H, where H is the number of bits that satisfy the condition: T =0 is satisfied.
[0080]
[0096] In some cases, an LDPC code is regular if each row and column of the parity check matrix has a constant degree, and in other cases, it is irregular. In some cases, irregular LDPC codes outperform regular LDPC codes. In some cases, due to the different degrees between rows and columns, irregular LDPC codes promise improved performance only if the row degrees and column degrees are adjusted appropriately.
[0081]
[0097] In some cases, the error correction scheme includes a polar code. In some cases, the polar code can achieve Shannon capacity by theoretical proof. In some cases, the polar code has low encoding and decoding complexity. Polar codes are generally based on a generator matrix G N The information includes
number
number
number
number
number
number
number
[0082]
[0098] In some cases, polar codes are expressed using cosec codes.
number
number
number
number
number
[0083]
[0099] In some cases, the error correction scheme includes a turbo code. A turbo code typically includes a parallel concatenation of two or more component codes applied to different interleaved versions of the same information sequence. Typically, recursive systematic convolutional (RSC) codes are used as the component codes. A turbo code structure, for example, includes two parallel-concatenated RSC encoders (e.g., M=2), with a code rate R=1 / (M+1) (approximately), so that R=1 / 3. The input to the first RSC encoder is the original information sequence. The original information sequence d is also applied to an interleaver to generate an interleaved version d'. The interleaved version d' of the information sequence is the input to the second RSC encoder. The output from the turbo encoder includes a systematic sequence of u and redundancies x(1) (output from the first RSC encoder) and x(2) (output from the second encoder). Thus, the output from the encoder is u1, x1(1), x1(2), u2, x 2(1) , x 2(2) where u k is the k-th systematic bit (i.e., data bit), and x k(1) is the k-th systematic bit u k is the parity output from the first RSC encoder associated with x k(2) is the k-th systematic bit u kand a parity output from a second RSC encoder associated with the first RSC encoder. A turbo code decoding procedure includes iterative decoding. The turbo code decoding procedure can include two component decoders (corresponding to the two RSC encoders), an interleaver, and a deinterleaver. In some cases, the two component decoders are soft-input soft-output (SISO) decoders. In some cases, the outputs of the two component decoders include likelihood information about the coded data sequence.
[0084]
[0100] In some cases, the size of the binary data is increased when an external codec (e.g., ECC) is applied. In some cases, the frame size is increased when ECC is applied to each frame containing the binary data. In some cases, the frame is divided into multiple lanes. In some cases, each lane includes a lane index. In some cases, each frame includes about 1,000 to about 10,000 lanes. In some cases, each frame includes about 5,000 lanes. In some cases, each frame includes about 1,000 to about 2,500 lanes, about 1,000 to about 5,000 lanes, about 1,000 to about 7,500 lanes, about 1,000 to about 10,000 lanes, about 2,500 to about 5,000 lanes, about 2,500 to about 7,500 lanes, about 2,500 to about 10,000 lanes, about 5,000 to about 7,500 lanes, about 5,000 to about 10,000 lanes, or about 7,500 to about 10,000 lanes. In some cases, each frame includes about 1,000 lanes, about 2,500 lanes, about 5,000 lanes, about 7,500 lanes, or about 10,000 lanes. In some cases, each frame includes at least about 1,000 lanes, about 2,500 lanes, about 5,000 lanes, or about 7,500 lanes. In some cases, each frame includes at most about 2,500 lanes, about 5,000 lanes, about 7,500 lanes, or about 10,000 lanes. Each lane can further include about 100 bits to about 300 bits. In some cases, each lane includes about 100 bits to about 150 bits, about 100 bits to about 200 bits, about 100 bits to about 250 bits, about 100 bits to about 300 bits, about 150 bits to about 200 bits, about 150 bits to about 250 bits, about 150 bits to about 300 bits, about 200 bits to about 250 bits, about 200 bits to about 300 bits, or about 250 bits to about 300 bits. In some cases, each lane includes approximately 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bits.In some cases, each lane includes at least about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bits. In some cases, each lane includes at most about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300 bits.
[0085]
[0101] In some cases, the method of encoding digital or binary data into multiple nucleotide sequences includes shuffling the binary data. In some cases, each lane is shuffled at least in part based on the lane index. In some cases, each lane is shuffled after applying an external codec (e.g., ECC) to the binary data. In some cases, shuffling each lane allows for tolerance to errors that may occur during synthesis or sequencing, such as those that affect an entire oligonucleotide library. Errors can include insertions, deletions, substitutions, or combinations thereof. In some cases, shuffling includes a rotation scheme within each lane based at least in part on the lane index. For example, each bit within a lane can be shifted by the lane index (e.g., no shuffling for lane 0, a one-bit shift for lane 1, a two-bit shift for lane 2, etc.).
[0086]
[0102] In some further cases, the shuffling includes a pseudo-random process within each lane. In this pseudo-random shuffling process, a random seed is used to initialize a pseudo-random number generator. In some cases, the numbers generated by the pseudo-random number generator depend on the random seed. Thus, the same sequence of numbers is generated by the pseudo-random number generator using the same seed. As an example, the shuffling includes a pseudo-random process in which each bit within a lane is shifted according to the number generated by the pseudo-random number generator.
[0087]
[0103] In some further cases, the lane index is used as a seed to generate permutations for some or all bits in that lane. In some cases, permutations for some or all bits are created by sampling from a random number generator. In some cases, the permutations are stored in precompiled form. In some cases, the use of pseudo-random number generation can result in smaller implementation source code.
[0088]
[0104] In some cases, a frame index and a lane index are prepended. In some cases, a frame index and a lane index are prepended to each lane when the lane is shuffled. An exemplary diagram of shuffling lanes and prepending a frame index and a lane index is shown in FIG. 6. In some cases, the frame index includes about 12 bits to about 20 bits. In some cases, the frame index includes about 12 bits to about 14 bits, about 12 bits to about 16 bits, about 12 bits to about 18 bits, about 12 bits to about 20 bits, about 14 bits to about 16 bits, about 14 bits to about 18 bits, about 14 bits to about 20 bits, about 16 bits to about 18 bits, about 16 bits to about 20 bits, or about 18 bits to about 20 bits. In some cases, the frame index includes about 12 bits, about 14 bits, about 16 bits, about 18 bits, or about 20 bits. In some cases, the frame index includes at least about 12 bits, about 14 bits, about 16 bits, or about 18 bits. In some cases, the frame index includes at most about 14 bits, about 16 bits, about 18 bits, or about 20 bits. In some cases, the lane index includes about 12 bits to about 16 bits. In some cases, the lane index includes about 12 bits to about 14 bits, about 12 bits to about 16 bits, or about 14 bits to about 16 bits. In some cases, the lane index includes about 12 bits, about 14 bits, or about 16 bits. In some cases, the lane index includes at least about 12 bits or about 14 bits. In some cases, the lane index includes at most about 14 bits or about 16 bits. As shown in FIG. 6, in some cases, the lane index is 12 bits and the frame index is 20 bits. In some cases, the lane index is the symbol width m from the RS code.
[0089]
[0105] In some cases, the method of encoding digital or binary data into a plurality of nucleotide sequences includes an internal codec. In some cases, the internal codec is applied to the binary data. In some cases, the internal codec is applied to the binary data from the ECC. In some cases, the internal codec is applied to lanes of binary data. In some cases, the internal codec is applied to lanes of binary data when the lanes are shuffled.
[0090]
[0106] In some cases, the encoding scheme includes an inner codec. In some cases, the inner codec is applied to each lane to encode binary data as a nucleotide sequence. The inner codec is used to convert digital or binary data into nucleotide bases. In some cases, the inner codec is capable of correcting deletion, substitution, or insertion errors. In some further embodiments, the inner codec is used to validate oligos and discard any suspect oligos, thereby avoiding contamination of the outer decoding. The inner codec may further encode indexes (frame index and lane index), thereby enabling efficient clustering during decoding.
[0091]
[0107] In some cases, the encoding scheme adds redundancy across multiple oligonucleotide sequences. In some cases, the redundancy is about 5% to about 10%. In some cases, the redundancy is about 5% to about 6%, about 5% to about 7%, about 5% to about 8%, about 5% to about 9%, about 5% to about 10%, about 6% to about 7%, about 6% to about 8%, about 6% to about 9%, about 6% to about 10%, about 7% to about 8%, about 7% to about 9%, about 7% to about 10%, about 8% to about 9%, about 8% to about 10%, or about 9% to about 10%. In some cases, the redundancy is about 5%, about 6%, about 7%, about 8%, about 9%, or about 10%. In some cases, the redundancy is at least about 5%, about 6%, about 7%, about 8%, or about 9%. In some cases, the redundancy is at most about 61 GB%, about 7%, about 8%, about 9%, or about 10%. In some cases, the redundancy allows the library of oligos to be decoded in the presence of errors in individual oligos, such as insertions, deletions, substitutions, or any combination thereof.
[0092]
[0108] An exemplary diagram of an encoding scheme is shown in FIG. 7. In this exemplary diagram, the encoding scheme of the inner codec combines two or more of the bits from each lane, the bit history, and the bit position. In some cases, a model (e.g., an adaptive model) is used to separate the known bits into contexts, and each context is mapped to a bit history. In some cases, the bit history is represented by an 8-bit state. In some cases, the bit history is updated each time a context is encountered, for example, through the use of a lookup table. The bit position includes the least significant bit (LSB) from the bit index of the bit to be encoded. For example, if 100 bits encode a 100-mer oligonucleotide, the "bit index" refers to an index from 0 to 99 in the bit to be encoded. The LSB includes the bit position in a binary integer that represents the location of a binary 1 in the integer. In some cases, the LSB index is of any length. In some cases, the LSB index is represented by a 4-bit state.
[0093]
[0109] In some cases, the internal codec includes generating base candidates for bits of the binary data. The base candidates are generated for the binary data using a lookup table, a hash, or a combination thereof. In some cases, the hash is determined using the methods described above. In some cases, the binary data includes two or more of bits from each lane, bit history, and bit position. In some cases, the encoding bit rate is from about 1 bit / base to about 2 bits / base. In some cases, the encoding bit rate is from about 1 bit / base to about 1.1 bits / base, from about 1 bit / base to about 1.2 bits / base, from about 1 bit / base to about 1.3 bits / base, from about 1 bit / base to about 1.4 bits / base, from about 1 bit / base to about 1.5 bits / base, from about 1 bit / base to about 1.6 bits / base, from about 1 bit / base to about 1.7 bits / base, from about 1 bit / base to about 1.8 bits / base, or 1 bit / base to about 1.9 bits / base, about 1 bit / base to about 2 bits / base, about 1.1 bits / base to about 1.2 bits / base, about 1.1 bits / base to about 1.3 bits / base, about 1.1 bits / base to about 1.4 bits / base, about 1.1 bits / base to about 1.5 bits / base, about 1.1 bits / base to about 1.6 bits / base, about 1.1 bits / base to about 1.7 bits / base, about 1.1 bits / base to about 1.8 bits / base About 1.8 bits / base, about 1.1 bits / base to about 1.9 bits / base, about 1.1 bits / base to about 2 bits / base, about 1.2 bits / base to about 1.3 bits / base, about 1.2 bits / base to about 1.4 bits / base, about 1.2 bits / base to about 1.5 bits / base, about 1.2 bits / base to about 1.6 bits / base, about 1.2 bits / base to about 1.7 bits / base, about 1.2 bits / base to about 1.8 bits / base, about 1.2 bits / base to about 1.9 bits / base, about 1.2 bits / base to about 2 bits / base, about 1.3 bits / base to about 1.4 bits / base, about 1.3 bits / base to about 1.5 bits / base, about 1.3 bits / base to about 1.6 bits / base, about 1.3 bits / base to about 1.7 bits / base, about 1.3 bits / base to about 1.8 bits / base, about 1.3 bits / base to about 1.9 bits / base, about 1.3 bits / base to about 2 bits / base, about 1.4 bits / base to about 1.5 bits / base, about 1.4 bits / base to about 1.6 bits / base, about 1.4 bits / base to about 1.7 bits / base, about 1.4 bits / base to about 1.8 bits / base, about 1.4 bits / base to about 1.9 bits / base, about 1.4 bits / base to about 2 bits / base, about 1.5 bits / base to about 1.6 bits / base, about 1.5 bits / base to about 1.7 bits / base, about 1.5 bits / base to about 1.8 bits / base, about 1.5 bits / base to about 1.9 bits / base, about 1.5 bits / base to about 2 bits / base, about 1.6 bits / base to about 1.7 bits / base, about 1.6 bits / base to about 1.8 bits / base, about 1.6 bits / base to about 1.9 bits / base, about 1.6 bits / base to about 2 bits / base, about 1.7 bits / base to about 1.8 bits / base, about 1.7 bits / base to about 1.9 bits / base, about 1.7 bits / base to about 2 bits / base, about 1.8 bits / base to about 1.9 bits / base, about 1.8 bits / base to about 2 bits / base, or about 1.9 bits / base to about 2 bits / base. In some cases, the encoding bit rate is about 1 bit / base, about 1.1 bits / base, about 1.2 bits / base, about 1.3 bits / base, about 1.4 bits / base, about 1.5 bits / base, about 1.6 bits / base, about 1.7 bits / base, about 1.8 bits / base, about 1.9 bits / base, or about 2 bits / base. In some cases, the encoding bit rate is at least about 1 bit / base, about 1.1 bits / base, about 1.2 bits / base, about 1.3 bits / base, about 1.4 bits / base, about 1.5 bits / base, about 1.6 bits / base, about 1.7 bits / base, about 1.8 bits / base, or about 1.9 bits / base. In some cases, the encoding bit rate is at most about 1.1 bits / base, about 1.2 bits / base, about 1.3 bits / base, about 1.4 bits / base, about 1.5 bits / base, about 1.6 bits / base, about 1.7 bits / base, about 1.8 bits / base, about 1.9 bits / base, or approximately 2 bits / base. In some cases, a lookup table is used to map bits to nucleotides (e.g., A=00, T=10, C=01, G=11). In some cases, the hash involves a function that can map data of any size (e.g., any number of bits) to a fixed-size value (e.g., a hash value). In some examples, the hash value is mapped to a nucleotide sequence.
[0094]
[0110] In some cases, the internal codec includes a base repeat check. In some cases, the base repeat check is performed once a base candidate is selected. In some cases, the base repeat check checks for repeats in two or more sequential bases. In some cases, the base repeat check substitutes a base if there is a repeat in two or more sequential bases. In some cases, a lookup table or hash is updated based on the bases updated during the base repeat check. Additionally, after the base repeat check, the bit history is updated. In some cases, the frame index and / or lane index are incremented. In some cases, this process is repeated until the entire sequence of the plurality of nucleotide sequences is determined.
[0095]
[0111] In some cases, the internal codec further comprises performing GC filtering before synthesizing the plurality of nucleotide sequences. In some cases, the GC filtering removes about 1% to about 10% of the lanes within the plurality of lanes. In some cases, the GC filtering removes about 5% to about 10% of the lanes within the plurality of lanes. In some cases, the GC filtering does not remove any lanes within the plurality of lanes. In some cases, the GC filtering removes about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10%. In some cases, the GC filtering removes at least about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, or about 9%. In some cases, the GC filtering removes at most about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10%. In some cases, the plurality of nucleotide sequences comprises a GC content of about 40% to about 60%. In some cases, the plurality of nucleotide sequences comprises a GC content of about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 50% to about 55%, about 50% to about 60%, or about 55% to about 60%. In some cases, the plurality of nucleotide sequences comprises a GC content of about 40%, about 45%, about 50%, about 55%, or about 60%. In some cases, the plurality of nucleotide sequences comprises a GC content of at least about 40%, about 45%, about 50%, or about 55%. In some cases, the plurality of nucleotide sequences comprises at most about 45%, about 50%, about 55%, or about 60% GC content. In some cases, at least 90% of the plurality of nucleotide sequences comprise a GC content of about 40% to about 60%. In some cases, at least 90% of the plurality of nucleotide sequences comprise a GC content of about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 50% to about 55%, about 50% to about 60%, or about 55% to about 60%.In some cases, at least 90% of the plurality of nucleotide sequences comprise a GC content of about 40%, about 45%, about 50%, about 55%, or about 60%. In some cases, at least 90% of the plurality of nucleotide sequences comprise a GC content of at least about 40%, about 45%, about 50%, or about 55%. In some cases, at least 90% of the plurality of nucleotide sequences comprise a GC content of at most about 45%, about 50%, about 55%, or about 60%. Output from the internal codec comprises the final oligonucleotide library.
[0096]
[0112] An exemplary diagram of an alternative encoding scheme is shown in Figure 8. In some cases, the encoding scheme in the internal codec includes starting with a default lookup table. The default lookup table is used to select a word to encode in each lane. In some cases, the word is an 8-bit word or a byte. A lookup table is applied to generate base candidates for each word or byte in each lane. The next lookup table is selected based on the previously coded word or byte. In some cases, the encoding scheme further includes performing base repetition checking, GC filtering, or a combination thereof, as described above herein. In some cases, this process is repeated until the entire sequence of the plurality of nucleotide sequences can be determined. The output from the internal codec comprises the final oligonucleotide library.
[0097]
[0113] In some cases, the length of each oligonucleotide (or polynucleotide) in the library is from about 20 to about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is from about 20 bases to about 50 bases, from about 20 bases to about 100 bases, from about 20 bases to about 200 bases, from about 20 bases to about 300 bases, from about 20 bases to about 400 bases, from about 20 bases to about 500 bases, from about 50 bases to about 100 bases, from about 50 bases to about 200 bases, from about 50 bases to about 300 bases, from about 50 bases to about 400 bases, about 50 bases to about 500 bases, about 100 bases to about 200 bases, about 100 bases to about 300 bases, about 100 bases to about 400 bases, about 100 bases to about 500 bases, about 200 bases to about 300 bases, about 200 bases to about 400 bases, about 200 bases to about 500 bases, about 300 bases to about 400 bases, about 300 bases to about 500 bases, or about 400 bases to about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is at least about 20 bases, about 50 bases, about 100 bases, about 200 bases, about 300 bases, or about 400 bases. In some cases, the length of each oligonucleotide (or polynucleotide) in the library is at most about 50 bases, about 100 bases, about 200 bases, about 300 bases, about 400 bases, or about 500 bases.
[0098]
[0114] De novo polynucleotide synthesis
[0115] Provided herein are systems and methods for synthesizing a library of polynucleotides on a substrate. In some cases, a library comprising a plurality of polynucleotides from an encoding scheme is synthesized. In some examples, the library comprising a plurality of polynucleotides from an encoding scheme encodes a pool of a plurality of pools. In some examples, the library comprising a plurality of polynucleotides from an encoding scheme encodes an index pool. In some cases, the method includes the use of electrochemical deprotection. In some cases, the substrate is a flexible substrate. In some cases, at least 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , or 10 15 bases are synthesized per day. In some cases, at least 10 x 10 8 , 10×10 9 , 10×10 10 , 10×10 11 , or 10 x 10 12polynucleotides are synthesized per day. In some cases, each polynucleotide synthesized contains at least 20, 50, 100, 200, 300, 400, or 500 nucleic acid bases. In some cases, these bases are synthesized with an average total error rate of less than about 1 error per 100, 200, 300, 400, 500, 1000, 2000, 5000, 10,000, 15,000, or 20,000 bases. In some cases, these error rates are for at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or more of the synthesized polynucleotides. In some cases, at least 90%, 95%, 98%, 99%, 99.5%, or more of the synthesized polynucleotides do not differ from the predetermined encoded sequence. In some cases, the error rate of polynucleotides synthesized on a substrate using the methods and / or systems described herein is less than about 1 in 200, less than about 1 in 1,000, less than about 1 in 2,000, less than about 1 in 3,000, or less than about 1 in 5,000. Each type of error rate includes mismatches, deletions, insertions, and / or substitutions in the polynucleotides synthesized on the substrate. The term "error rate" refers to the comparison of the total amount of polynucleotides synthesized to the total amount of a predetermined polynucleotide sequence. In some cases, the synthesized polynucleotides disclosed herein include a tether of 12 to 25 bases. In some cases, the tether comprises 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more bases.
[0099]
[0116] Described herein are methods, systems, devices, and compositions in which chemical reactions used in polynucleotide synthesis are controlled using electrochemistry. In some cases, electrochemical reactions are controlled by any energy source, such as light, heat, radiation, or electricity. For example, electrodes are used to control chemical reactions at all or some of the discrete loci on a substrate. In some cases, electrodes are charged by applying an electrical potential to the electrodes to control one or more chemical steps in polynucleotide synthesis. In some cases, these electrodes are addressable. Any number of chemical steps described herein are controlled using one or more electrodes in some cases. Electrochemical reactions can include oxidation, reduction, acid / base chemistry, or other reactions controlled by electrodes. In some cases, electrodes generate electrons or protons, which are used as reagents for chemical transformations. In some cases, electrodes directly generate reagents, such as acids. In some cases, acids are protons. In some cases, electrodes directly generate reagents, such as bases. Acids or bases are often used to cleave protecting groups or affect the kinetics of various polynucleotide synthesis reactions, for example, by adjusting the pH of the reaction solution. In some cases, electrochemically controlled polynucleotide synthesis reactions involve redox-active metals or other redox-active organic materials. In some cases, metal or organic catalysts are employed in conjunction with these electrochemical reactions. In some cases, acids are generated from the oxidation of quinones.
[0100]
[0117] Control of chemical reactions using electrochemical generation of reagents is not limited; chemical reactivity may be indirectly influenced through biophysical changes to substrates or reagents through electric fields (or gradients) generated by electrodes. In some cases, substrates include, but are not limited to, nucleic acids. In some cases, electric fields are generated that attract or repel specific reagents or substrates from electrodes or surfaces. In some cases, such fields are generated by applying a potential to one or more electrodes. For example, negatively charged nucleic acids are repelled from negatively charged electrode surfaces. In some cases, such repulsion or attraction of polynucleotides or other reagents caused by local electric fields provides for the movement of polynucleotides or other reagents into or out of regions of a synthesis device or structure. In some cases, electrodes generate electric fields that repel polynucleotides from a synthesis surface, structure, or device. In some cases, electrodes generate electric fields that attract polynucleotides toward a synthesis surface, structure, or device. In some cases, protons are repelled from a positively charged surface, limiting their contact with a substrate or portion thereof. In some cases, repulsive or attractive forces are used to allow or block the entry of reagents or substrates into specific areas of the synthesis surface. In some cases, nucleoside monomers are prevented from contacting polynucleotide chains by applying an electric field near one or both components. Such an arrangement allows for gating of specific reagents and may obviate the need for protecting groups if the concentration or rate of contact between reagents and / or substrates is controlled. In some cases, unprotected nucleoside monomers are used in polynucleotide synthesis. Alternatively, applying a field near one or both components promotes contact of nucleoside monomers with polynucleotide chains. Furthermore, applying an electric field to a substrate can alter the reactivity or conformation of the substrate. In one exemplary application, an electric field generated by electrodes is used to prevent polynucleotides at adjacent loci from interacting. In some cases, the substrate is a polynucleotide, optionally attached to a surface. In some cases, applying an electric field alters the three-dimensional structure of the polynucleotide.Such modifications include folding or unfolding various structures, such as helices, hairpins, loops, or other three-dimensional nucleic acid structures. Such modifications are useful for manipulating nucleic acids within wells, channels, or other structures. In some cases, an electric field is applied to the nucleic acid substrate to prevent secondary structure. In some cases, the electric field eliminates the need for linkers or attachment to a solid support during polynucleotide synthesis.
[0101]
[0118] A suitable method for synthesizing polynucleotides on substrates of the present disclosure is phosphoramidite-based synthesis of DNA. In some cases, the reagents for phosphoramidite-based synthesis include any one or combination of nucleoside phosphoramidites, oxidizing agents, activators, or deblockers, or the solvent includes acetonitrile. In some cases, the phosphoramidite-based synthesis method involves the controlled addition of phosphoramidite building blocks, i.e., nucleoside phosphoryl groups, to a growing polynucleotide chain in a coupling step that forms a phosphite triester bond between the phosphoramidite building block and the nucleoside bound to the substrate. In some cases, the nucleoside phosphoramidite is provided to an activated substrate. In some cases, the nucleoside phosphoramidite is provided to a substrate together with an activator. In some cases, the nucleoside phosphoramidite is provided to the substrate in a 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100-fold or greater excess over the substrate-bound nucleoside. In some cases, the addition of the nucleoside phosphoramidite is carried out in an anhydrous environment, for example, in anhydrous acetonitrile. Following the addition and conjugation of the nucleoside phosphoramidite in the coupling step, the substrate is optionally washed. In some cases, the coupling step is repeated one or more additional times, with optional washing steps between the addition of the nucleoside phosphoramidite to the substrate. In some cases, the polynucleotide synthesis method used herein includes one, two, three, or more consecutive coupling steps. Prior to coupling, the nucleoside attached to the substrate is often deprotected by removal of a protecting group, which functions to prevent polymerization. A protecting group can include any chemical group that prevents the extension of a polynucleotide chain. In some cases, the protecting group is cleaved (or removed) in the presence of acid. In some cases, the protecting group is cleaved in the presence of base. In some cases, the protecting group is removed using electromagnetic radiation, such as light, heat, or other energy source. In some cases, the protecting group is removed through an oxidation or reduction reaction (e.g., aIn some cases, the protecting group comprises a triarylmethyl group. In some cases, the protecting group comprises an allyl ether. In some cases, the protecting group comprises a disulfide. In some cases, the protecting group comprises an acid-labile silane. In some cases, the protecting group comprises an acetal. In some cases, the protecting group comprises a ketal. In some cases, the protecting group comprises an enol ether. In some cases, the protecting group comprises a methoxybenzyl group. In some cases, the protecting group comprises an azide. In some cases, the protecting group is 4,4'-dimethoxytrityl (DMT). In some cases, the protecting group is tert-butyl carbonate. In some cases, the protecting group is a tert-butyl ester. In some cases, the protecting group comprises a base-labile group.
[0102]
[0119] Following coupling, phosphoramidite polynucleotide synthesis optionally includes a capping step, in which the growing polynucleotide is treated with a capping agent. This step generally serves to block unreacted 5'-OH groups attached to the substrate after coupling from further chain elongation, thereby preventing the formation of polynucleotides with internal base deletions. Furthermore, lH-tetrazole-activated phosphoramidites often react only slightly with the O6 position of guanosine. Without being bound by theory, upon oxidation with I2 / water, this by-product undergoes depurination, presumably via O6-N7 migration. The non-purine base site may ultimately be cleaved during the final deprotection of the oligonucleotide, thereby reducing the yield of the full-length product. The O6 modification can be removed by treatment with a capping reagent prior to oxidation with I2 / water. In some cases, including a capping step during polynucleotide synthesis reduces the error rate compared to synthesis without capping. In one example, the capping step involves treating the substrate-bound polynucleotide with a mixture of acetic anhydride and 1-methylimidazole. Following the capping step, the substrate is optionally washed.
[0103]
[0120] Following the addition of the nucleoside phosphoramidite and optionally after capping and one or more washing steps, the substrates described herein include substrates bound to growing nucleic acids that can be oxidized. The oxidation step involves oxidizing a phosphite triester to a tetracoordinate phosphate triester, which is a protected precursor of the naturally occurring phosphodiester nucleoside linkage. In some cases, the phosphite triester is oxidized electrochemically. In some cases, oxidation of the growing polynucleotide is achieved by treatment with iodine and water, optionally in the presence of a weak base such as pyridine, lutidine, or collidine. Oxidation is sometimes performed under anhydrous conditions using tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is performed following oxidation. Because residual water from possible oxidation can inhibit subsequent coupling, the second capping step dries the substrate. Following oxidation, the substrate and growing polynucleotide are optionally washed. In some embodiments, the oxidation step is replaced with a sulfurization step to obtain oligonucleotide phosphorothioates, in which case any capping step can be performed after sulfurization. Many reagents allow for efficient sulfur transfer, including, but not limited to, 3-(dimethylaminomethylidene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide, also known as Beaucage reagent, and N,N,N',N'-tetraethylthiuram disulfide (TETD).
[0104]
[0121] To allow subsequent cycles of nucleoside incorporation through conjugation, the protected 5' end (or 3' end, if synthesis is performed in the 5' to 3' direction) of the growing substrate-bound polynucleotide must be removed so that the primary hydroxyl group can react with the next nucleoside phosphoramidite. In some cases, the protecting group is DMT, and deblocking occurs with trichloroacetic acid in dichloromethane. In some cases, the protecting group is DMT, and deblocking occurs with electrochemically generated protons. Detritylation for extended periods or using stronger acid solutions than recommended can lead to increased depurination of the solid support-bound oligonucleotide, thus reducing the yield of the desired full-length product. The methods and compositions described herein provide controlled deblocking conditions that limit undesired depurination reactions. In some cases, the substrate-bound polynucleotide is washed after deblocking. In some cases, efficient washing after deblocking contributes to synthesized polynucleotides with low error rates.
[0105]
[0122] The methods of synthesizing polynucleotides on a substrate described herein can include a repeated sequence of the following steps: applying a protected monomer to the surface of the substrate mechanism to link it to either the surface, a linker, or a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and applying another protected monomer for linking. One or more intermediate steps include oxidation and / or sulfurization. In some instances, one or more washing steps are performed before or after one or all of the steps.
[0106]
[0123] The methods for synthesizing polynucleotides on a substrate described herein can include an oxidation step. For example, the method can include a repeated sequence of the following steps: applying a protected monomer to the surface of the substrate mechanism to link it to either the surface, a linker, or a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; applying another protected monomer for linking; and an oxidation and / or sulfurization step. In some cases, this step occurs before or after one or all of the steps.
[0107]
[0124] The methods of synthesizing polynucleotides on substrates described herein can further include a repeating sequence of the following steps: applying a protected monomer to the surface of the substrate mechanism to link it to either the surface, a linker, or a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and oxidizing and / or sulfurizing steps, in some cases occurring before or after one or all of the steps.
[0108]
[0125] The methods of synthesizing polynucleotides on substrates described herein can further include a repeating sequence of the following steps: applying a protected monomer to the surface of the substrate mechanism to link it with either the surface, a linker, or a previously deprotected monomer, and an oxidation and / or sulfurization step, in some cases occurring before or after one or all of the steps.
[0109]
[0126] The methods of synthesizing polynucleotides on substrates described herein can further include a repeating sequence of the following steps: applying a protected monomer to the surface of the substrate mechanism to link it to either the surface, a linker, or a previously deprotected monomer; deprotecting the applied monomer so that it can react with a subsequently applied protected monomer; and oxidizing and / or sulfurizing steps, in some cases occurring before or after one or all of the steps.
[0110]
[0127] In some cases, polynucleotides are synthesized with photolabile protecting groups, and hydroxyl groups generated on the surface are blocked by the photolabile protecting groups. When the surface is exposed to UV light, such as through a photolithographic mask, a pattern of free hydroxyl groups can be generated on the surface. These hydroxyl groups can react with photoprotected nucleoside phosphoramidites according to phosphoramidite chemistry. A second photolithographic mask can be applied, and the surface exposed to UV light to generate a second pattern of hydroxyl groups, which can then be coupled with 5'-photoprotected nucleoside phosphoramidites. Similarly, patterns can be generated and oligomer chains can be extended. Without being bound by theory, the instability of the photocleavable groups depends on the wavelength and polarity of the solvent used, and the rate of photocleavage can be affected by the exposure duration and light intensity. This method can leverage several factors, such as the accuracy of mask alignment, the efficiency of photoprotective group removal, and the yield of the phosphoramidite coupling step. Furthermore, unintended light leakage to neighboring sites can be minimized. The density of synthesized oligomers per spot can be monitored by adjusting the packing of the leading nucleosides on the synthesis surface.
[0111]
[0128] The surface of the substrate described herein that provides support for polynucleotide synthesis can be chemically modified to allow cleavage of the synthesized polynucleotide chain from the surface. In some cases, the polynucleotide chain is cleaved simultaneously with deprotection of the polynucleotide. In some cases, the polynucleotide chain is cleaved after deprotection of the polynucleotide. In an exemplary method, a trialkoxysilylamine such as (CH3CHO)3Si-(CH2)2-NH2 reacts with the SiOH group surface of the substrate, followed by succinic anhydride reaction with the amine to generate an amide bond and a free OH, which then supports the growth of the nucleic acid chain. Cleavage includes gas cleavage with ammonia or methylamine. In some cases, cleavage includes linker cleavage with electrogenerated reagents such as acid or base. In some cases, once released from the surface, the polynucleotides are assembled into larger nucleic acids that are sequenced and decoded to extract the stored information.
[0112]
[0129] The surfaces described herein can be reused to support additional cycles of polynucleotide synthesis after polynucleotide cleavage. For example, the linker can be reused without additional treatment / chemical modification. In some cases, the linker is non-covalently attached to the substrate surface or polynucleotide. In some embodiments, the linker remains attached to the polynucleotide after cleavage from the surface. In some embodiments, the linker has a reversible covalent bond, such as esters, amides, ketals, β-substituted ketones, heterocycles, or other groups that can be reversibly cleaved. Such reversible cleavage reactions are controlled, in some cases, through the addition or removal of reagents or by an electrochemical process controlled by an electrode. Optionally, the chemical linker or surface-bound chemical group is regenerated after several cycles to restore reactivity and eliminate unwanted by-product formation at such linker or surface-bound chemical groups.
[0113]
[0130] Alternatively, polymer synthesis can be enzymatic DNA synthesis. In some cases, enzymatic DNA synthesis uses water as the solvent and the reagent is the enzyme terminal deoxynucleotidyl transferase (TdT) or a deblocker. In some cases, enzymatic synthesis of DNA uses a template-independent DNA polymerase, terminal deoxynucleotidyl transferase (TdT), which has evolved to rapidly catalyze the linkage of naturally occurring dNTPs. TdT indiscriminately adds to nucleotides, thus preventing uncontrolled synthesis by various techniques, such as tethering TdT, creating variant enzymes, and using nucleotides containing reversible terminators, to prevent chain elongation. TdT activity is maximized at approximately 37°C, and the enzymatic reaction is carried out in an aqueous environment.
[0114]
[0131] Polynucleotide storage device
[0132] The synthesized library of polynucleotides can be stored on a device. In some cases, the device includes a polynucleotide data storage system. In some cases, the library encoding the pool (e.g., multiple pools or) is stored in a compartment. In some cases, the compartment includes, by way of non-limiting example, an active surface (e.g., locus), a tube, a cell, a spot, or any other physical storage solution. In some examples, the compartment includes a location (e.g., spot) on a microfluidic chip, such as a digital microfluidic chip. In some examples, the compartment is labeled. In some examples, the label includes a barcode, a name (e.g., customer name, sample type, etc.), a timestamp, a list of stored objects, or any combination thereof.
[0115]
[0133] In some cases, the device that stores digital information in DNA includes one or more compartments. In some cases, each of the one or more compartments includes a library containing a plurality of polynucleotides. In some examples, the library encodes a pool (e.g., a pool of a plurality of pools described herein) containing digital information corresponding to one or more objects. In some examples, the pool includes a pool descriptor, one or more pool items, and a pool end descriptor, such as described herein. In some examples, the pool includes about 1 GB to about 1 TB of digital information, as described herein above.
[0116]
[0134] In some cases, each of the one or more compartments includes a medium for storing a plurality of polynucleotides. In some examples, the medium includes a solid, a liquid, a gas, or any combination thereof. In some examples, the medium includes a saline solution. In some examples, the molar ratio of salt to DNA can range from about 20:1 to about 2:1. In some examples, the molar ratio depends on the molecular weight of the salt used and the relative amounts of salt and DNA combined. In some examples, the molar ratio is calculated between the cations of the salt and the negatively charged phosphate groups of the DNA. In some examples, the saline solution includes a molar ratio of salt cations to phosphate groups in the DNA of less than 20:1. In some examples, the saline solution is dried to produce a dried product. In some cases, the saline solution includes, by way of non-limiting example, calcium chloride, calcium nitrate, calcium carbonate, calcium phosphate, magnesium chloride, magnesium sulfate, magnesium nitrate, magnesium carbonate, lanthanum chloride, lanthanum nitrate, lanthanum carbonate, lanthanum bromide, or a mixture thereof. In some cases, the saline solution includes, by way of non-limiting example, calcium chloride dihydrate, calcium chloride dihydrate, 1758521039658_· / A> , lanthanum trichloride, magnesium chloride hexahydrate, sodium chloride, or strontium chloride hexahydrate. In some cases, the concentration of the saline solution is from about 0.01 nM to about 0.1 nM.
[0117]
[0135] In some cases, the medium for storing the plurality of polynucleotides comprises nanoparticles. In some cases, the nanoparticles comprise silica nanoparticles. In some cases, a subset of the plurality of polynucleotides is encapsulated in the nanoparticles. In some cases, the nanoparticles encapsulating the polynucleotides are stored in an anhydrous or near-anhydrous environment. In some cases, the nanoparticles comprise a protective layer of silica (e.g., tetraethoxysilane). In some cases, the nanoparticles comprise a compound that co-interacts with the polynucleotides (e.g., N-[3-(trimethoxysilyl)propyl]-N,N,N-trimethylammonium chloride). In some cases, the nanoparticles encapsulating the polynucleotides are stored on a digital microfluidic chip. In some cases, the digital microfluidic chip enables fluid programmability. In some cases, the programmability enables automated storage and / or retrieval of polynucleotides. In some cases, each location on the digital microfluidic chip comprises approximately 100 GB, 500 GB, 1 TB, 2 TB, 10 TB, 20 TB, 30 TB, or 50 TB. In some cases, each location contains about 50 μg, 100 μg, 150 μg, 200 μg, 250 μg, 300 μg, 350 μg, 400 μg, 450 μg, 500 μg, 600 μg, 700 μg, 800 μg, 900 μg, or 1000 μg of nanoparticles.
[0118]
[0136] In some cases, one or more compartments are in communication with each other. In some cases, one or more compartments are in communication with each other through a medium. In some cases, one or more compartments are not in communication with each other. In some cases, one or more compartments are not in communication with each other through a medium.
[0119]
[0137] In some cases, the device further includes one or more second compartments. In some cases, each of the one or more second compartments includes a second library. In some examples, the second library encodes an index pool, such as described herein. In some cases, the one or more second compartments include a medium as described herein above. In some cases, the one or more second compartments include the same medium as the one or more compartments. In some cases, the one or more second compartments include a different medium than the one or more compartments. In some cases, each of the one or more second compartments is in communication (e.g., through a medium) with each other and / or with the one or more compartments. In some cases, each of the one or more second compartments is not in communication with each other and / or with the one or more compartments.
[0120]
[0138] In some cases, the device further comprises a solid support that forms a surface. Such devices are described herein as solid support-based nucleic acid synthesis and storage devices, where the solid support has various dimensions. In some cases, the size of the solid support is about 40-120 mm x about 25-100 mm. In some cases, the size of the solid support is about 80 mm x about 50 mm. In some cases, the width of the solid support is at least about 10 mm, 20 mm, 40 mm, 60 mm, 80 mm, 100 mm, 150 mm, 200 mm, 300 mm, 400 mm, 500 mm, or more than 500 mm. In some cases, the height ... solid support is at least about 100 mm. 2 ;200mm 2 ;500mm 2 ;1,000mm 2 ;2,000mm 2 ;4,500mm 2 ;5,000mm 2 ;10,000mm 2 ;12,000mm 2 ;15,000mm 2 ;20,000mm 2;30,000mm 2 ;40,000mm 2 and having a planar area of 50,000 mm or more. In some cases, the thickness of the solid support is about 50 mm to about 2000 mm, about 50 mm to about 1000 mm, about 100 mm to about 1000 mm, about 200 mm to about 1000 mm, or about 250 mm to about 1000 mm. Non-limiting examples of thicknesses of the solid support include 275 mm, 375 mm, 525 mm, 625 mm, 675 mm, 725 mm, 775 mm, and 925 mm. In some cases, the thickness of the solid support is at least about 0.5 mm, 1.0 mm, 1.5 mm, 2.0 mm, 2.5 mm, 3.0 mm, 3.5 mm, 4.0 mm, or greater than 4.0 mm.
[0121]
[0139] Described herein are devices in which two or more solid supports are assembled. In some cases, the solid supports are interfaced together on a larger unit. Interfacing can include the exchange of fluids, electrical signals, or other exchange media between the solid supports. The unit can interface with any number of servers, computers, or networked devices. For example, multiple solid supports are integrated into a rack unit, which can be conveniently inserted into or removed from a server rack. A rack unit can contain any number of solid supports. In some cases, a rack unit contains at least 1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 50,000, 100,000, or more than 100,000 solid supports. In some cases, two or more solid supports do not interface with each other. Nucleic acids (and the information stored therein) present on solid supports can be accessed from a rack unit. Accessing includes removal of polynucleotides from the solid support, direct analysis of polynucleotides on the solid support, or any other method that allows the information stored in the nucleic acids to be manipulated or identified. In some cases, the information is accessed from multiple racks, a single rack, a single solid support within a rack, a portion of a solid support, or a single locus on a solid support. In various cases, accessing includes interfacing the nucleic acids with an additional device, such as a mass spectrometer, HPLC, sequencing instrument, PCR thermocycler, or other device for manipulating nucleic acids. In some cases, accessing the nucleic acid information is achieved by cleavage of the polynucleotides from all or a portion of the solid support. In some cases, cleavage includes exposure to chemical reagents (ammonia or other reagents), electrical potential, radiation, heat, light, sound, or other forms of energy capable of manipulating chemical bonds. In some cases, cleavage is achieved by charging one or more electrodes in the vicinity of the polynucleotides. In some cases, electromagnetic radiation in the form of UV light is used to cleave the polynucleotides.In some cases, a lamp is used to cleave the polynucleotides and a mask controls the location of UV light exposure on the surface. In some cases, a laser is used to cleave the polynucleotides and the open / closed state of a shutter controls the exposure of UV light to the surface. In some cases, access to the nucleic acid information (including removal / addition of racks, solid supports, reagents, nucleic acids, or other components) is fully automated.
[0122]
[0140] The solid supports described herein include an active area. In some cases, the active area includes a region, cell, feature, or locus for nucleic acid synthesis. In some cases, the active area includes a region or locus for nucleic acid synthesis. In some examples, the region or locus includes one or more compartments. In some examples, the region or locus includes a second one or more compartments. In some cases, the region is addressable. In some examples, the region is addressable via an electrode.
[0123]
[0141] The active area can have a variety of dimensions. For example, the dimensions of the active area are from about 1 mm to about 50 mm by about 1 mm to about 50 mm. In some cases, the active area has a width of at least or about 0.5 mm, 1 mm, 1.5 mm, 2 mm, 2.5 mm, 3 mm, 5 mm, 5 mm, 10 mm, 12 mm, 14 mm, 16 mm, 18 mm, 20 mm, 25 mm, 30 mm, 35 mm, 40 mm, 45 mm, 50 mm, 60 mm, 70 mm, 80 mm, or greater than 80 mm. In some cases, the active area has a height of at least or about 0.5 mm, 1 mm, 1.5 mm, 2 mm, 2.5 mm, 3 mm, 5 mm, 5 mm, 10 mm, 12 mm, 14 mm, 16 mm, 18 mm, 20 mm, 25 mm, 30 mm, 35 mm, 40 mm, 45 mm, 50 mm, 60 mm, 70 mm, 80 mm, or greater than 80 mm.
[0124]
[0142] Described herein are solid support-based nucleic acid synthesis and storage devices, where the solid support has several sites (e.g., spots) or locations for synthesis or storage. In some cases, the solid support contains up to or about 10,000 x 10,000 locations within an area. In some cases, the solid support contains about 1,000-20,000 x about 1,000-20,000 locations within an area. In some cases, the solid support contains at least or about 10, 30, 50, 75, 100, 200, 300, 400, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 12,000, 14,000, 16,000, 18,000, 20,000 locations. The area may include at least or about 10, 30, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 12,000, 14,000, 16,000, 18,000, 20,000 locations. In some cases, the area may be up to 0.25 square inches, 0.5 square inches, 0.75 square inches, 1.0 square inches, 1.25 square inches, 1.5 square inches, or 2.0 square inches. In some cases, the solid support comprises loci with a pitch of at least or about 0.1 μm, 0.2 μm, 0.25 μm, 0.3 μm, 0.4 μm, 0.5 μm, 1.0 μm, 1.5 μm, 2.0 μm, 2.5 μm, 3.0 μm, 3.5 μm, 4.0 μm, 4.5 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, or more than 10 μm. In some cases, the solid support comprises loci with a pitch of about 5 μm. In some cases, the solid support comprises loci with a pitch of about 2 μm. In some cases, the solid support comprises loci with a pitch of about 1 μm. In some cases, the solid support comprises loci with a pitch of about 0.2 μm. In some cases, the solid support comprises loci having a pitch of about 0.2 μm to about 10 μm, about 0.2 μm to about 8 μm, about 0.5 μm to about 10 μm, about 1 μm to about 10 μm, about 2 μm to about 8 μm, about 3 μm to about 5 μm, about 1 μm to about 3 μm, or about 0.5 μm to about 3 μm.In some cases, the solid support comprises loci having a pitch of about 0.1 μm to about 3 μm.
[0125]
[0143] The solid supports for nucleic acid synthesis or storage described herein have high data storage capacities. For example, the capacity of the solid support is at least about 1 petabyte, 2 petabytes, 3 petabytes, 4 petabytes, 5 petabytes, 6 petabytes, 7 petabytes, 8 petabytes, 9 petabytes, 10 petabytes, 20 petabytes, 50 petabytes, 100 petabytes, 200 petabytes, 300 petabytes, 400 petabytes, 500 petabytes, 600 petabytes, 700 petabytes, 800 petabytes, 900 petabytes, 1000, or more than 1000 petabytes. In some cases, the capacity of the solid support is about 1 to about 10 petabytes or about 1 to about 100 petabytes. In some cases, the capacity of the solid support is about 100 petabytes. In some cases, the data is stored as an array of packets as droplets. In some examples, the array of packets is an addressable packet. In some examples, the packets are addressable using electrodes. In some cases, data is stored as an array of packets as droplets on spots. In some cases, data is stored as an array of packets as dry wells. In some cases, the array contains at least about 1 gigabyte, 2 gigabytes, 3 gigabytes, 4 gigabytes, 5 gigabytes, 6 gigabytes, 7 gigabytes, 8 gigabytes, 9 gigabytes, 10 gigabytes, 20 gigabytes, 50 gigabytes, 100 gigabytes, 200 gigabytes, or more than 200 gigabytes of data. In some cases, the array contains at least about 1 terabyte, 2 terabytes, 3 terabytes, 4 terabytes, 5 terabytes, 6 terabytes, 7 terabytes, 8 terabytes, 9 terabytes, 10 terabytes, 20 terabytes, 50 terabytes, 100 terabytes, 200 terabytes, or more than 200 terabytes of data. In some cases, the information item is stored in the background of the data, for example, the information item encodes about 10 to about 100 megabytes of data and is stored in the background of 1 petabyte of data.In some cases, an item of information encodes at least or more than about 1 megabyte, 10 megabytes, 20 megabytes, 30 megabytes, 40 megabytes, 50 megabytes, 60 megabytes, 70 megabytes, 80 megabytes, 90 megabytes, 100 megabytes, 150 megabytes, 200 megabytes, 300 megabytes, 400 megabytes, 500 megabytes, or 500 megabytes of data and is stored with more than 1 petabyte, 10 petabytes, 20 petabytes, 30 petabytes, 40 petabytes, 50 petabytes, 60 petabytes, 70 petabytes, 80 petabytes, 90 petabytes, 100 petabytes, 150 petabytes, 200 petabytes, 300 petabytes, 400 petabytes, 500 petabytes, or 500 petabytes of background data.
[0126]
[0144] Provided herein are solid support-based nucleic acid synthesis and storage devices in which, following synthesis, polynucleotides are collected in packets as one or more droplets. In some cases, the polynucleotides are collected and stored in packets as one or more droplets. In some cases, the number of droplets is at least or about 1, 10, 20, 50, 100, 200, 300, 500, 1,000, 2,500, 5,000, 75,000, 10,000, 25,000, 50,000, 75,000, 100,000, 1 million, 5 million, 10 million, 25 million, 50 million, 75 million, 100 million, 250 million, 50 million, 750 million, or more than 750 million droplets. In some cases, the droplet volume includes a diameter of 5 μm (micrometers), 10 μm, 15 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, 50 μm, 55 μm, 60 μm, 65 μm, 70 μm, 75 μm, 80 μm, 85 μm, 90 μm, 95 μm, 100 μm, or greater than 100 μm. In some cases, the droplet volume includes a diameter of 1 to 100 μm, 10 to 90 μm, 20 to 80 μm, 30 to 70 μm, or 40 to 50 μm.
[0127]
[0145] In some cases, the polynucleotides collected in a packet contain similar sequences. In some cases, the polynucleotides further contain non-identical sequences used as tags or barcodes. For example, the non-identical sequences are used to index the polynucleotides stored on the solid support and subsequently search for specific polynucleotides based on the non-identical sequences. Exemplary tag or barcode lengths include, but are not limited to, barcode sequences containing about 1 base, 2 bases, 3 bases, 4 bases, 5 bases, 6 bases, 7 bases, 8 bases, 9 bases, 10 bases, 15 bases, 20 bases, 25 bases, or more. In some cases, the tags or barcodes have a length of at least or about 10 base pairs, 50 base pairs, 75 base pairs, 100 base pairs, 200 base pairs, 300 base pairs, 400 base pairs, or more than 400 base pairs.
[0128]
[0146] Provided herein are solid support-based nucleic acid synthesis and storage devices in which polynucleotides are collected into packets containing redundancy. For example, the packets contain about 100 to about 1000 copies of each polynucleotide. In some cases, the packets contain at least or about 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, or more than 2000 copies of each polynucleotide. In some cases, the packets contain about 1000-fold to about 5000-fold synthetic redundancy. In some cases, the synthetic redundancy is at least about 500-fold, 1000-fold, 1500-fold, 2000-fold, 2500-fold, 3000-fold, 3500-fold, 4000-fold, 5000-fold, 6000-fold, 7000-fold, 8000-fold, or more than 8000-fold. The polynucleotides synthesized using the solid support-based methods described herein have various lengths. In some cases, the polynucleotides are synthesized and stored on a solid support. In some cases, the polynucleotide length is about 100 bases to about 1000 bases. In some cases, a polynucleotide has a length of at least or about 10 bases, 20 bases, 30 bases, 40 bases, 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 125 bases, 150 bases, 175 bases, 200 bases, 225 bases, 250 bases, 275 bases, 300 bases, 325 bases, 350 bases, 375 bases, 400 bases, 425 bases, 450 bases, 475 bases, 500 bases, 600 bases, 700 bases, 800 bases, 900 bases, 1000 bases, 1100 bases, 1200 bases, 1300 bases, 1400 bases, 1500 bases, 1600 bases, 1700 bases, 1800 bases, 1900 bases, 2000, or more than 2000 bases.
[0129]
[0147] Sequencing
[0148] Polynucleotides are extracted and / or amplified from the surface on which they are synthesized or stored. After extracting and / or amplifying the polynucleotides from the surface of the structure, suitable sequencing techniques can be employed to sequence the polynucleotides. In some cases, the DNA sequence is read on the substrate or within the features of the structure. In some cases, polynucleotides stored on the substrate are extracted, optionally assembled into longer nucleic acids, and then sequenced.
[0130]
[0149] Polynucleotides synthesized and stored on the structures described herein encode data that can be interpreted by reading the sequence of the synthesized polynucleotides and converting the sequence into computer-readable binary code. In some cases, the sequence requires assembly, and the assembly step may be required at the nucleic acid sequence stage or the digital sequencing stage.
[0131]
[0150] Provided herein is a detection system including a device capable of sequencing stored polynucleotides either directly on the substrate and / or after removal from the host structure. When the structure is a reel-to-reel tape of flexible material, the detection system includes a device for holding and advancing the structure through a detection location and a detector disposed near the detection location for detecting a signal generated from a section of the tape when it is at the detection location. In some cases, the signal indicates the presence of the polynucleotide. In some cases, the signal indicates the sequence of the polynucleotide (e.g., a fluorescent signal). In some cases, information encoded within the polynucleotides on the continuous tape is read by a computer as the tape is continuously advanced by a detector operably connected to the computer. In some cases, the detection system includes a computer system including a polynucleotide sequencing device, a database for storing and retrieving data related to the polynucleotide sequence, software for converting the DNA code of the polynucleotide sequence into binary code, a computer for reading the binary code, or any combination thereof.
[0132]
[0151] Provided herein are sequencing systems that can be integrated into the devices described herein. Various sequencing methods are well known in the art and include "base calling," in which the identity of bases in a target polynucleotide is identified. In some cases, polynucleotides synthesized using the methods, devices, compositions, and systems described herein are sequenced after cleavage from the synthesis surface. In some cases, sequencing is performed during or simultaneously with polynucleotide synthesis, and base calling is performed immediately after or immediately before the extension of nucleoside monomers into a growing polynucleotide chain. Base calling methods involve measuring the current / voltage generated by the polymerase-catalyzed addition of bases to a template strand. In some cases, the synthesis surface includes an enzyme, such as a polymerase. In some cases, such an enzyme is tethered to an electrode or synthesis surface. In some cases, the enzyme includes a terminal deoxynucleotidyl transferase or a variant thereof.
[0133]
[0152] Digital information retrieval system and method
[0153] Provided herein are methods and systems for digital information retrieval. In some cases, the digital information includes one or more objects as described herein. In some cases, each of the one or more objects is about 1 GB to about 1 TB as described herein. In some cases, the one or more objects include items of information such as, but not limited to, those described herein.
[0134]
[0154] In some cases, the systems and methods decode nucleic acid sequences (e.g., polynucleotides, oligonucleotides, multiple polynucleotides, etc.). In some cases, the method of retrieving digital information stored in multiple polynucleotides includes one or more steps.
[0135]
[0155] In some cases, retrieving the digital information stored in the plurality of polynucleotides includes accessing an index pool. In some cases, accessing the index pool includes fully or partially sequencing a library encoding the index pool. In some examples, the index pool is encoded in a library using the systems and methods described herein. In some examples, polynucleotides in the library encoding the index pool are sequenced using the systems and methods described herein. In some cases, two or more index pools are accessed. In some cases, polynucleotides in two or more libraries are sequenced. In some cases, the sequenced library is temporarily stored in a memory storage system (e.g., a flash drive). In some cases, to retrieve the index pool, the sequenced library is converted into digital information. In some cases, the index pool is temporarily stored in a memory storage system (e.g., a flash drive). In some cases, the digital information in the index pool is used to search for one or more objects of interest. In some examples, the one or more objects of interest are stored in a library containing a plurality of polynucleotides encoding the one or more objects. In some examples, the one or more objects of interest are searched for using metadata associated with the one or more objects. In some cases, accessing an index pool identifies multiple pools corresponding to one or more objects.
[0136]
[0156] In some cases, one or more objects of interest are retrieved from a compartment in the storage device. In some cases, retrieving the digital information stored in the plurality of polynucleotides includes sequencing the plurality of polynucleotides corresponding to one or more objects in the plurality of pools. In some cases, the plurality of polynucleotides are in a library. In some cases, the library is in a compartment of a device, as described herein above. In some cases, the plurality of polynucleotides in the library encoding the pool are sequenced using the systems and methods described herein. In some cases, the pool is encoded into a library using the systems and methods described herein. In some cases, the plurality of polynucleotides in two or more compartments are sequenced to retrieve one or more objects.
[0137]
[0157] In some cases, retrieving the digital information stored in the plurality of polynucleotides further includes applying a decoding scheme. In some cases, the decoding scheme decodes digital information in a plurality of pools. In some cases, the decoding scheme is applied to a sequenced library comprising a plurality of polynucleotides. In some cases, the decoding scheme includes an internal codec, an external codec (e.g., ECC), or a combination thereof. In some cases, the decoding scheme decodes the plurality of nucleotide sequences to generate an output comprising digital information (e.g., an object). In some cases, the decoding scheme includes undoing an operation in the encoding scheme. In some examples, the operation includes splitting, shuffling, concatenating, transposing, translating, duplicating, labeling (e.g., using an index), or any combination thereof, of data or portions of data.
[0138]
[0158] Provided herein are decoding methods and systems. In some cases, the methods and systems decode a nucleotide sequence (e.g., a polynucleotide, an oligonucleotide, a plurality of polynucleotides, etc.). In some cases, the nucleotide sequence is encoded using a method described herein. In some cases, the methods and systems include an internal codec, an external codec, or a combination thereof. In some cases, the method of decoding a plurality of nucleotide sequences may include identifying the plurality of nucleotide sequences. In some cases, identifying the plurality of nucleotide sequences includes sequencing the nucleotides. In some cases, the nucleotides are sequenced using a method described herein.
[0139]
[0159] After sequencing the plurality of nucleotides, the encoded binary data is decoded. In some cases, the plurality of nucleotides are decoded using, as a non-limiting example, the schematic diagram shown in Figure 9. The output from sequencing includes an unordered list of reads (e.g., nucleotide sequences), as shown in Figure 9.
[0140]
[0160] In some cases, sequenced polynucleotides, such as an unordered list of reads, are clustered after sequencing. In some cases, clustering is performed before applying an internal codec. In some cases, sequenced polynucleotides are clustered based on an index, such as a frame index, a lane index, or a combination thereof. In such cases, the sequenced polynucleotides are partially decoded to obtain the frame index, the lane index, or a combination thereof. In some cases, clustering is performed using a hash function, as described above in this specification. In some cases, a hash function is used when bases in a nucleotide sequence are identified using a hash in an encoding scheme, as described above in this specification.
[0141]
[0161] In some cases, the sequenced polynucleotides (e.g., reads) are aligned. In some cases, the sequenced polynucleotides are aligned after being clustered. In some cases, the sequenced polynucleotides are aligned before applying an internal codec. In some cases, aligning includes analyzing the consensus of the reads (e.g., nucleotide sequences) using an alignment algorithm. In some examples, the alignment algorithm includes a pairwise alignment algorithm, a multiple sequence alignment algorithm, or a combination thereof.
[0142]
[0162] In some cases, the pairwise alignment algorithm includes initializing the position of each read. Initializing includes aligning the nucleotide sequence to position 0. A consensus of one or more next bases is analyzed between the reads. In some cases, about 3 to about 10 reads are analyzed for a consensus. In some cases, about 3 to about 4, about 3 to about 5, about 3 to about 6, about 3 to about 7, about 3 to about 8, about 3 to about 9, about 3 to about 10, about 4 to about 5, about 4 to about 6, about 4 to about 7, about 4 to about 8, about 4 to about 9, about 4 to about 10, about 5 to about 6, about 5 to about 7, about 5 to about 8, about 5 to about 9, about 5 to about 10, about 6 to about 7, about 6 to about 8, about 6 to about 9, about 6 to about 10, about 7 to about 8, about 7 to about 9, about 7 to about 10, about 8 to about 9, about 8 to about 10, or about 9 to about 10 reads are analyzed for consensus. In some cases, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 reads are analyzed for consensus. In some cases, at least about 3, about 4, about 5, about 6, about 7, about 8, or about 9 reads are analyzed for a consensus. In some cases, at most about 4, about 5, about 6, about 7, about 8, about 9, or about 10 reads are analyzed for a consensus. In some cases, the next one or more bases include the next 2 to 10 bases. In some cases, the next one or more bases are about 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases. In some cases, the next one or more bases are at least about 2, 3, 4, 5, 6, 7, 8, or 9 bases. In some cases, the next one or more bases are at most about 3, 4, 5, 6, 7, 8, 9, or 10 bases. In some cases, the next one or more bases are about 2, 3, 4, or 5 bases. A consensus is analyzed between the reads to determine whether the next one or more bases are correct. If there is a consensus between all reads, for example, between bases at position x, the subsequent base, for example, x+1, can then be analyzed. If there is a discrepancy between reads, for example, between bases at position x, it is determined whether the read containing the discrepancy contains an error. In some cases, the error is an insertion, deletion, or substitution. Then, given the determination in each read (e.g., correct or erroneous), the position is incremented, for example, to x+1.In some cases, the process is repeated until the end of the lead is reached.
[0143]
[0163] In some cases, the decoding scheme includes an inner codec. In some cases, the inner codec is applied to multiple nucleotide sequences. The inner codec is used to convert the nucleotide sequences into digital or binary data. In some cases, the inner codec is capable of correcting deletion, substitution, or insertion errors, or any combination thereof. In some further embodiments, the inner codec is used to validate oligos and discard any suspect oligos to avoid contamination of the outer decoding. In some cases, the inner codec allows for efficient decoding using indexes (frame index and lane index).
[0144]
[0164] An internal codec comprising a decoding scheme is applied to the plurality of nucleotide sequences. In some cases, the internal codec may convert each of the plurality of nucleotide sequences into a lane of binary data. In some cases, the internal codec is applied to the sequenced plurality of nucleotides. In some cases, the internal codec is applied to the unordered reads. In some cases, the internal codec is applied to the reads or the plurality of polynucleotides once they have been clustered, as described herein. In some cases, the internal codec is applied to the reads or the plurality of polynucleotides once they have been aligned, as described herein.
[0145]
[0165] In some cases, the inner codec includes a greedy algorithm. In some cases, the inner codec includes a maximum likelihood (ML) algorithm. In some cases, the inner codec includes a greedy ML hybrid algorithm.
[0146]
[0166] An internal codec including a greedy algorithm (e.g., a greedy decoder) is illustratively shown in FIG. 10. As shown, the greedy algorithm considers transitions from only the most likely state as each bit position in the array is decoded. In some cases, each bit is guessed one at a time using the greedy algorithm. In some cases, more than one bit is guessed at a given time using the greedy algorithm. In some cases, the x-axis includes bit positions and the y-axis includes states. In some cases, the states include one or more valid coding states S that are analyzed at each bit position. In some cases, each state S is assigned a probability. In some cases, state S is defined as the bits coded from each lane, bit history, and bit position. In some cases, state S is defined as the bit history and bit word. The greedy algorithm iteratively finds the most likely state at each position until it reaches the most likely final state. In some cases, the decoded bits are backtracked by following the most likely state at each bit position. In some cases, this produces fully decoded bits. In some cases, the greedy decoder finds a local optimum solution. In some cases, the local optimum solution is an approximation of the global optimum solution. The greedy decoder provides a solution (or final state) in a reasonable amount of time compared to other intra codecs such as those described herein.
[0147]
[0167] In some cases, the performance of the internal codec improves by knowing where the oligonucleotide sequence ends. In some cases, the oligonucleotide length is determined during sequencing, for example, through paired-end sequencing. In some cases, a drift term is introduced into the greedy algorithm. The drift term includes an integer associated with the total number of insertions and deletions. Each insertion is represented as a +1 value, and each deletion is represented as a -1 value. For example, if there are no insertions and two deletions, the total drift is -2. In such an example, the greedy algorithm discards all decoding end states that do not match the oligo length as invalid. Thus, the drift term allows the greedy algorithm to know which decoding end states are valid, further improving performance. Thus, in some cases, the internal codec further includes a z-axis corresponding to drift, as shown in Figures 10 and 11.
[0148]
[0168] An inner codec including an ML algorithm is illustratively shown in FIG. 11. As shown, the ML algorithm takes into account transitions from all states as it decodes each bit position in the sequence. The states are defined as described above. In some cases, each bit is guessed one at a time using the ML algorithm. In some cases, more than one bit is guessed at a given time using the ML algorithm. In some cases, the ML algorithm iteratively finds all transition states at each position until a final candidate state is identified. In some cases, as described above, the x-axis includes bit positions and the y-axis includes states. In some cases, as described above, a drift term is used to filter the final candidate states. In some cases, the ML algorithm provides a global optimum by tracking all state transitions. In some cases, the ML algorithm is computationally intensive compared to other decoding schemes such as those described herein.
[0149]
[0169] In some cases, the inner codec includes a greedy ML blending algorithm that considers transitions from multiple states as each bit position in the array is decoded. In some cases, the multiple states are about 100 to about 1000 states as each bit in the array is decoded. In some cases, the plurality of states may be from about 100 to about 200, from about 100 to about 300, from about 100 to about 400, from about 100 to about 500, from about 100 to about 600, from about 100 to about 700, from about 100 to about 800, from about 100 to about 900, from about 100 to about 1,000, from about 200 to about 300, from about 200 to about 400, from about 200 to about 500, from about 200 to about 600, from about 200 to about 700, from about 200 to about 800, from about 200 to about 900, from about 200 to about 1,000, from about 300 to about 400, from about 300 to about 500, from about 300 to about 600, from about 300 to about 700, from about 300 to about 800, to about 900, about 300 to about 1,000, about 400 to about 500, about 400 to about 600, about 400 to about 700, about 400 to about 800, about 400 to about 900, about 400 to about 1,000, about 500 to about 600, about 500 to about 700, about 500 to about 800, about 500 to about 900, about 500 to about 1,000, about 600 to about 700, about 600 to about 800, about 600 to about 900, about 600 to about 1,000, about 700 to about 800, about 700 to about 900, about 700 to about 1,000, about 800 to about 900, about 800 to about 1,000, or about 900 to about 1,000. In some cases, the plurality of states is about 100 states, about 200 states, about 300 states, about 400 states, about 500 states, about 600 states, about 700 states, about 800 states, about 900 states, or about 1,000 states. In some cases, the plurality of states is at least about 100 states, about 200 states, about 300 states, about 400 states, about 500 states, about 600 states, about 700 states, about 800 states, or about 900 states. In some cases, the plurality of states is at most about 200 states, about 300 states, about 400 states, about 500 states, about 600 states, about 700 states, about 800 states, about 900 states, or about 1,000 states. States are defined as described herein above.In some cases, each bit is guessed one at a time using a greedy ML mixing algorithm. In some cases, two or more bits are guessed at a given time using the greedy ML mixing algorithm. In some cases, the greedy ML mixing algorithm iteratively finds about 100 to about 1000 transition states at each position until a final candidate state is identified. In some cases, a drift term, as described herein above, is used to filter the final candidate states. In some cases, the greedy ML mixing algorithm provides a globally optimal solution while being computationally inexpensive compared to other inner codecs, such as the ML algorithms described herein.
[0150]
[0170] In some cases, the inner codec includes a beam search decoder or a random sampling decoder (e.g., a pure sampling decoder, a top-K sampling decoder, etc.). In some cases, the beam search decoder or the random sampling decoder provides a greater variety of candidate states than a greedy decoder.
[0151]
[0171] In some cases, the internal codec further includes a checksum. In some cases, the checksum is used to verify data integrity, detect errors, or a combination thereof. In some cases, the checksum is generated using a checksum function or algorithm (e.g., parity byte or parity work (row-wise parity check), sum complement, position-dependent, fuzzy checksum, etc.). Examples of checksum functions or algorithms include, but are not limited to, BSD checksum (Unix), SYSV checksum (Unix), sum4, sum8, sum16, sum32, fletcher-4, fletcher-8, fletcher-16, fletcher-32, Adler-32, xor8, Luhn algorithm, Verhoeff algorithm, or Damm algorithm. In some cases, instead of taking only the most likely path, a small number of most likely paths are considered and tested against the checksum. In some cases, the checksum includes an RS code (e.g., a small RS code). In such cases, the decoder provides a list of possibilities (eg, "list decoding"), assuming the user can determine which exactly it is.
[0152]
[0172] In some cases, the decoding method further includes arranging the lanes into frames. In some cases, the lanes decoded from the internal codec are arranged into frames based on lane indexes and frame indexes. In some cases, one or more lanes are missing from a frame, as shown in FIG. 9 . In some cases, the lanes are missing due to errors occurring during nucleotide synthesis or sequencing. In some cases, about 1% to about 10% of the lanes are missing from a frame. In some cases, about 1% to about 2%, about 1% to about 4%, about 1% to about 6%, about 1% to about 8%, about 1% to about 10%, about 2% to about 4%, about 2% to about 6%, about 2% to about 8%, about 2% to about 10%, about 4% to about 6%, about 4% to about 8%, about 4% to about 10%, about 6% to about 8%, about 6% to about 10%, or about 8% to about 10% of the lanes are missing from a frame. In some cases, about 1%, about 2%, about 4%, about 6%, about 8%, or about 10% of the lanes are missing from the frame. In some cases, at least about 1%, about 2%, about 4%, about 6%, or about 8% of the lanes are missing from the frame. In some cases, at most about 2%, about 4%, about 6%, about 8%, or about 10% of the lanes are missing from the frame.
[0153]
[0173] In some cases, the internal codec has a "format." In some cases, there is no a priori information about the size of the data (e.g., binary data) during decoding. Thus, in some cases, frame index 0 contains the size of the data. In some cases, after placing lanes into frames and / or lining up frames, frame 1 is decoded first. Data is then extracted from frame 0 to reject frames outside the expected data size (e.g., from incorrectly decoded oligos).
[0154]
[0174] In some cases, the internal codec includes a hash (e.g., SHA-256). In some cases, the hash verifies that the data is decoded correctly. In some cases, encoding and decoding is performed as a stream, using the hash at the end (e.g., after ECC). In some cases, this can limit memory usage to only temporary buffers.
[0155]
[0175] The method for decoding the plurality of nucleotide sequences can include an external codec (e.g., ECC). In some cases, the plurality of nucleotide sequences are decoded into digital or binary data. In some cases, the external codec (e.g., ECC) is applied to the digital or binary data. In some examples, the ECC is applied to each frame. In some cases, the ECC is applied to lanes from an internal codec. In some cases, the ECC is applied after lanes from an internal codec are arranged into frames.
[0156]
[0176] In some cases, the outer codec includes an ECC used to code the data (e.g., binary data), and in some cases, the ECC includes a Reed-Solomon (RS) code, an LDPC code, a polar code, a turbo code, or any combination thereof.
[0157]
[0177] In some cases, the ECC includes a Reed-Solomon (RS) code. In such cases, the RS decoder receives a codeword r(x), which is the original codeword c(x) plus an error e(x) (e.g., r(x) = c(x) + e(x)). In some cases, the error e(x) is zero. In some cases, the RS decoder attempts to identify the location and magnitude of up to t errors (or 2t erasures). The RS code then attempts to correct these identified errors and / or erasures.
[0158]
[0178] In some cases, the RS decoder includes a syndrome calculation. In some cases, the syndrome calculation includes receiving input symbols and dividing them into a generator polynomial g(x), as described herein above. In some cases, the syndrome is calculated by substituting 2t roots of the generator polynomial g(x) (or the syndrome of the RS codeword c(x)) into r(x). In some cases, the generator polynomial g(x) is a known parameter of the decoder. In some cases, the RS codeword c(x) has 2t syndromes that are error sensitive.
[0159]
[0179] In some cases, the RS decoder includes finding the symbol error location. In some cases, the parity or check symbol t causes the syndrome calculation to zero if there is no error. In some cases, the parity or check symbol t includes the remainder in the RS encoder. If there is an error, the generated polynomial g(x) is passed to a Euclidean algorithm. In some cases, the remainder factors are found using the Euclidean algorithm. In some cases, the result is evaluated over iterations at each input symbol. In some cases, the error is discovered and corrected. In some cases, a corrected codeword c(x) is output from the RS decoder. In some cases, there are more errors in the codeword than can be corrected by the RS code (e.g., e(x)>2t). In such cases, the received codeword r(x) is output from the RS decoder. In some cases, the received codeword r(x) is output with an indication (e.g., a flag) that error correction failed. In some cases, the received codeword r(x) (e.g., a lane or frame containing binary data as described herein) is discarded.
[0160]
[0180] In some cases, frames from the ECC are merged to generate an output including binary data. In some cases, the binary data includes a byte stream or a byte array, as described above. The decoding methods described herein can be used to recover data when an error is present in at least one nucleotide sequence in the stored plurality of nucleotide sequences. In some cases, the error includes an insertion, deletion, substitution, or any combination thereof. In some cases, data is recovered when errors are present in about 0.001% to about 30% (e.g., an error rate) of the nucleotide sequences in the plurality of nucleotides.In some cases, the data may be from about 0.001% to about 0.01%, from about 0.001% to about 0.1%, from about 0.001% to about 0.5%, from about 0.001% to about 1%, from about 0.001% to about 2%, from about 0.001% to about 5%, from about 0.001% to about 10%, from about 0.001% to about 15%, from about 0.001% to about 20%, from about 0.001% to about 25%, from about 0.001% to about 30%, from about 0.01% to about 0.1%, from about 0.01 to about 0.5%, from about 0.01 % to about 1%, about 0.01% to about 2%, about 0.01% to about 5%, about 0.01% to about 10%, about 0.01% to about 15%, about 0.01% to about 20%, about 0.01% to about 25%, about 0.01% to about 30%, about 0.1% to about 0.5%, about 0.1% to about 1%, about 0.1% to about 2%, about 0.1% to about 5%, about 0.1% to about 10%, about 0.1% to about 15%, about 0.1% to about 20%, about 0.1% to about 25%, about 0.1% to about 30%, about 0.5% to about 1%, about 0.5% to about 2%, about 0.5% to about 5%, about 0.5% to about 10%, about 0.5% to about 15%, about 0.5% to about 20%, about 0.5% to about 25%, about 0.5% to about 30%, about 1% to about 2%, about 1% to about 5%, about 1% to about 10%, about 1% to about 15%, about 1% to about 20%, about 1% to about 25%, about 1% to about 30%, about 2% to about 5%, about 2% to about 10%, about 2% to about 15%, about 2% to about 20%, about 2% % to about 25%, about 2% to about 30%, about 5% to about 10%, about 5% to about 15%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 20% to about 25%, about 20% to about 30%, or about 25% to about 30% error rate. In some cases, data is recovered in the presence of errors at an error rate of about 0.001%, about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%.In some cases, data is recovered when errors are present at an error rate of at least about 0.001%, about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, or about 25%. In some cases, data is recovered when errors are present at an error rate of at most about 0.01%, about 0.1%, about 0.5%, about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, or about 30%.
[0161]
[0181] In some cases, the decoding scheme includes soft decoding. Soft decoding generally refers to decoding by considering a range of possible values (e.g., using probability estimation). As an example, sequencing has a quality for each base that can be considered during probability calculation. In such an example, each state includes a final probability, which can be used as a log-likelihood in the outer decoder if the outer decoder supports soft decoding. Furthermore, clustering and alignment can provide soft information about alignment reliability. As a further example, LDPC ECC includes an iterative decoder, which offers the possibility to iteratively switch back and forth between the inner and outer decoders instead of a single pass. However, in some cases, this comes at the cost of higher computational requirements.
[0162]
[0182] The hashes of the present disclosure can enable verification of the digital information during retrieval. In some cases, retrieving the digital information stored in the plurality of polynucleotides further includes verifying at least one or more objects. In some cases, the one or more objects are verified using a first one or more hashes in the plurality of pools. In some cases, retrieving the digital information stored in the plurality of polynucleotides further includes verifying one or more pool items. In some cases, the one or more pool items are verified using a second one or more hashes in the plurality of pools. In some examples, when an object is stored across two or more pools of a plurality of pools, two or more pool items are assembled into the object. In such examples, the first one or more hashes of the one or more objects, the second one or more hashes of the one or more pool items, or a combination thereof, enable proper assembly verification.
[0163]
[0183] Verifying a hash generally involves generating a hash (e.g., a cryptographic hash). Verifying can further involve comparing the generated hash to a previously determined hash. In some cases, the previous hash and the new hash are determined using the same hash function. In some cases, the hash function includes a cryptographic hash function. In some cases, the hash function includes MD-5, SHA-1, SHA-2, SHA-3, RIPEMD-160, Whirlpool, BLAKE, BLAKE2, BLAKE3, or variations thereof. In some cases, the hash function includes SHA-2. In some examples, SHA-2 includes SHA-224, SHA-256, SHA-384, SHA-512, SHA-512 / 224, or SHA-512 / 256. In some cases, if the new hash and the previous hash match, the integrity of the information item (e.g., an object) is verified. In some cases, if the new hash and the previous hash do not match, the verification fails. In some cases, if the verification fails, the integrity of the information item is not verified. In some cases, if the verification fails, the information item has been altered and / or corrupted.
[0164]
[0184] Retrieving the digital information can include combining information stored in pool items and / or across multiple pools. In some cases, retrieving digital information stored in multiple polynucleotides further includes combining digital information in multiple pools. In some cases, data payloads in one or more pool items are combined. In some cases, data payloads in one or more pool items across multiple pools are combined. In some cases, the combined data payloads include digital information. In some cases, the retrieved digital information is further stored in memory.
[0165]
[0185] In some cases, the retrieved digital information is presented to a user. In some cases, the information is presented to a user on an interface. In some cases, the interface is an interface of an electronic device (e.g., a personal electronic device). In some cases, the electronic device includes an application configured to communicate with a system described herein over a computer network to access the information.
[0166]
[0186] The method for retrieving digital information in DNA (or polynucleotides) can be performed on a system. In some cases, such a system includes one or more processing units, memory, instructions, a sequencing device, or a combination thereof. In some cases, the memory is in communication with the one or more processing units. In some cases, the instructions are stored in the memory. In some cases, the sequencing device is in communication with the memory, the one or more processing units, or a combination thereof. In some cases, the one or more processing units and memory are distributed across one or more physical or logical locations.
[0167]
[0187] In some cases, the memory is used to store digital information, polynucleotide sequences (e.g., partially or fully decoded sequences), or a combination thereof. In some cases, the memory is used to store information related to algorithms described herein (e.g., software code, parameters, executable instructions, etc.). In some examples, the memory can include any suitable memory described herein. In some examples, the memory can be configured according to embodiments described herein. In some examples, the sequencing device is configured to identify multiple nucleotide sequences using the methods described herein.
[0168]
[0188] In some cases, one or more processing units include a central processing unit (CPU), a graphical processing unit (GPU), a single-core processor, a multi-core processor, a processor cluster, an application-specific integrated circuit (ASIC), a programmable circuit such as a field-programmable gate array (FPGA), an AI accelerator, and any combination thereof. In some cases, one or more of the processing units include a single-instruction, multiple-data (SIMD) or single-program, multiple-data (SPMD) parallel architecture. As an example, one or more processing units include one or more GPUs or CPUs implementing SIMD or SPMD. In some cases, the AI accelerator includes Google-TPU, Graphcore, Cerebras, SambaNova, or a combination thereof. In some embodiments, one or more of the processing units are implemented in software and / or firmware in addition to a hardware implementation. The software or firmware implementation of the processing unit may include computer- or machine-executable instructions written in any suitable programming language to perform the various functions described herein. The software implementation of the one or more processing units may be stored in whole or in part in memory. Alternatively or additionally, the system may include one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc. In some cases, decoding is performed with compute-on-memory technology, such as, but not limited to, UpMem.
[0169]
[0189] In some cases, the one or more processing units are configured to perform one or more decoding steps. In some cases, the processing device is configured to perform one or more steps including applying a decoding scheme to decode the digital information in the multiple pools, verifying at least one or more objects using a first one or more hashes in the multiple pools, combining the digital information in the multiple pools to retrieve one or more objects, and storing the digital information in a memory. In some cases, the one or more processing units are configured to perform one or more steps including applying an internal codec to the multiple polynucleotides or applying an ECC to the multiple polynucleotides. In some cases, the internal codec converts each of the multiple polynucleotides into digital information. In some cases, the internal codec includes a hybrid decoding algorithm including a greedy algorithm and a maximum likelihood (ML) algorithm. In some cases, outputs from the ECC are merged to generate an output including the digital information.
[0170] Specific Definitions
[0190] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter belongs.
[0171]
[0191] Throughout this disclosure, numerical characteristics are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Accordingly, the description of a range should be considered to have specifically disclosed each individual value within that range, along with all possible subranges, to the tenth of the unit of the lower limit, unless the context clearly indicates otherwise. For example, description of a range such as 1 to 6 should be considered to have specifically disclosed each individual value within that range, such as 1.1, 2, 2.3, 5, and 5.9, along with subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specific excluded limits in the stated ranges. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention, unless the context clearly indicates otherwise.
[0172]
[0192] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of any embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It is further understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0173]
[0193] References throughout this specification to "in some instances," "further instances," or "particular instances" mean that the particular feature, structure, or characteristic described in connection with that instance is included in at least one instance. Thus, the appearances of the phrase "in some instances," or "in further instances," or "in particular instances" in various places throughout this specification are not necessarily all referring to the same instance. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more instances.
[0174]
[0194] Unless specifically stated or clear from the context, as used herein, the term "about" in reference to a number or range of numbers is understood to mean the stated number, + / - 10% of the stated number, or 10% below the recited lower limit and 10% above the recited upper limit of the values recited in relation to the range.
[0175]
[0195] As used herein, the terms "preselected sequence," "predefined sequence," or "predetermined sequence" are used interchangeably. These terms mean that the sequence of a polymer is known and selected prior to the synthesis or assembly of the polymer. In particular, various aspects of the invention are described herein primarily in connection with the preparation of nucleic acid molecules, and the sequence of a polynucleotide is known and selected prior to the synthesis or assembly of the nucleic acid molecule.
[0176]
[0196] As used herein, the term "hash" or "hashes" may generally refer to a fixed-length string output from a hash function. A hash function may generally include a function that takes an input of any length and produces a fixed-length output. In some cases, the input may be one or more terms of a transaction or contract, which may be passed through a hash function to produce a hash. In some cases, the hash function may be deterministic, and it may be infeasible to reverse-engineer the output from the hashed output. The act of feeding an input to a hash function may be referred to as "hashing."
[0177]
[0197] The polynucleotide sequences described herein may include DNA, RNA, or analogs or derivatives thereof, unless otherwise specified. As used herein, the terms nucleic acid, polynucleotide, oligonucleotide, oligo, and oligonucleic acid are used interchangeably throughout to refer to a polymer of nucleoside monomers. In some cases, nucleic acids are linked via phosphate- or sulfur-containing linkages. In some cases, nucleic acids include DNA, RNA, non-orthogonal nucleic acids, non-naturally occurring nucleic acids, or other nucleosides. In some cases, nucleotides include non-orthogonal bases, sugars, or other moieties. In some cases, nucleotides include terminators configured to prevent extension reactions. In some cases, such terminators are removed prior to subsequent addition of nucleotides to the growing chain.
[0178] Computing System
[0198] 12, a block diagram illustrating an exemplary machine including a computer system 1200 (e.g., a processing or computing system) within which a set of instructions may be executed to cause a device to perform or implement any one or more of the static code scheduling aspects and / or methodologies of the present disclosure is shown. The components in FIG. 12 are merely examples and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or combination of two or more such components, implementing a particular embodiment.
[0179]
[0199] Computer system 1200 may include one or more processors 1201, memory 1203, and storage 1208, which communicate with each other and with other components via a bus 1240. Bus 1240 may also link a display 1232, one or more input devices 1233 (which may include, for example, a keypad, keyboard, mouse, stylus, etc.), one or more output devices 1234, one or more storage devices 1235, and various tangible storage media 1236. All of these elements may interface with bus 1240 directly or through one or more interfaces or adapters. For example, the various tangible storage media 1236 may interface with bus 1240 through storage media interface 1226. Computer system 1200 may have any suitable physical form, including, but not limited to, one or more integrated circuits (ICs), a printed circuit board (PCB), a mobile handheld device (such as a cell phone or PDA), a laptop or notebook computer, a distributed computer system, a computational grid, or a server.
[0180]
[0200] Computer system 1200 includes one or more processors 1201 (e.g., a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a quantum processing unit (QPU)) that perform functions. Processor 1201 optionally includes a cache memory unit 1202 for temporary local storage of instructions, data, or computer addresses. Processor 1201 is configured to support the execution of computer-readable instructions. Computer system 1200 may perform functions illustrated in FIG. 12 as a result of processor 1201 executing non-transitory processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 1203, storage 1208, storage device 1235, and / or storage medium 1236. The software may provide functionality to components that are part of the computer-readable medium. A computer-readable medium may store software implementing particular embodiments, and the processor 1201 may execute the software. The memory 1203 may read the software from one or more other computer-readable media (such as mass storage devices 1235, 1236) or from one or more other sources through a suitable interface, such as the network interface 1220. The software may cause the processor 1201 to perform one or more processes or one or more steps of processes described or illustrated herein. Execution of such processes or steps may include defining data structures stored in the memory 1203 and modifying the data structures as directed by the software.
[0181]
[0201] The memory 1203 may include various components (e.g., magnetically readable media) including, but not limited to, random access memory components (e.g., RAM 1204) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase change random access memory (PRAM), etc.), read-only memory components (e.g., ROM 1205), and any combination thereof. The ROM 1205 may operate to communicate data and instructions uni-directionally to the processor 1201, and the RAM 1204 may operate to communicate data and instructions bi-directionally to the processor 1201. The ROM 1205 and RAM 1204 may include any suitable tangible computer-readable media, as described below. In one example, a basic input / output system 1206 (BIOS), including basic routines that help to transfer information between elements within the computer system 1200, such as during start-up, may be stored in the memory 1203.
[0182]
[0202] Persistent storage 1208 is optionally coupled bidirectionally to processor 1201 through storage control unit 1207. Persistent storage 1208 provides additional data storage capacity and may include any suitable tangible computer-readable media described herein. Storage 1208 may be used to store operating system 1209, executable files 1210, data 1211, applications 1212 (application programs), etc. Storage 1208 may also include an optical disk drive, a solid-state memory device (e.g., a flash-based system), or any combination of the above. Information in storage 1208 may, where appropriate, be incorporated as virtual memory within memory 1203.
[0183]
[0203] In one example, storage device 1235 may removably interface with computer system 1200 via storage device interface 1225 (e.g., via an external port connector (not shown)). In particular, storage device 1235 and associated machine-readable media may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 1200. In one example, software may reside, completely or partially, on the machine-readable media in storage device 1235. In another example, software may reside, completely or partially, within processor 1201.
[0184]
[0204] Bus 1240 connects a wide variety of subsystems. As used herein, reference to a bus may, where appropriate, include one or more digital signal lines that perform a common function. Bus 1240 may be any of several types of buses, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a MicroChannel Architecture (MCA) bus, a Video Electronics Standards Association (VESA) local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express bus (PCI-X), an Accelerated Graphics Port (AGP) bus, a HyperTransport (HTX) bus, a Serial Advanced Technology Attachment (ATA) (SATA) bus, and any combination thereof.
[0185]
[0205] Computer system 1200 may also include input devices 1233. In one example, a user of computer system 1200 may input commands and / or other information into computer system 1200 via input devices 1233. Examples of input devices 1233 include, but are not limited to, an alphanumeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touchscreen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combination thereof. In some embodiments, input devices are Kinect, Leap Motion, etc. Input devices 1233 may interface with bus 1240 via any of a variety of input interfaces 1233 (e.g., input interface 1233), including, but not limited to, serial, parallel, gameport, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0186]
[0206] In particular embodiments, when computer system 1200 is connected to network 1230, computer system 1200 may communicate with other devices connected to network 1230, including mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, among others. Communications to and from computer system 1200 may be transmitted through network interface 1220. For example, network interface 1220 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 1230, and computer system 1200 may store the incoming communications in memory 1203 for processing. Computer system 1200 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 1203, which may be communicated from network interface 1220 to network 1230. Processor 1201 may access these communication packets stored in memory 1203 for processing.
[0187]
[0207] Examples of network interface 1220 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of network 1230 or network segment 1230 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus, or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combination thereof. Networks such as network 1230 may employ wired and / or wireless communication modes. In general, any network topology may be used.
[0188]
[0208] Information and data can be displayed through display 1232. Examples of display 1232 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display, a plasma display, and any combination thereof. Display 1232 can interface with other devices, such as processor 1201, memory 1203, and fixed storage 1208, as well as input device(s) 1233, via bus 1240. Display 1232 is linked to bus 1240 via video interface 1222, and data transfer between display 1232 and bus 1240 can be controlled via graphics control 1221. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD), such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting example, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, FreeflyVR headsets, etc. In further embodiments, the display is a combination of devices such as those disclosed herein.
[0189]
[0209] In addition to the display 1232, the computer system 1200 may include one or more other peripheral output devices 1234, including, but not limited to, audio speakers, printers, storage devices, and any combination thereof. Such peripheral output devices may be connected to the bus 1240 via an output interface 1224. Examples of the output interface 1224 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combination thereof.
[0190]
[0210] Additionally or alternatively, computer system 1200 may provide functionality as a result of logic hardwired or otherwise embodied in circuitry that may operate in place of or in conjunction with software to perform one or more processes or one or more steps of one or more processes described or illustrated herein. References to software in this disclosure may encompass logic, and references to logic may encompass software. Furthermore, references to computer-readable media may encompass, where appropriate, circuitry (such as an IC) that stores software for execution, circuitry that implements logic for execution, or both. This disclosure encompasses any suitable combination of hardware, software, or both.
[0191]
[0211] Those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0192]
[0212] The various illustrative logic blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0193]
[0213] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processors, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a user terminal.
[0194]
[0214] In accordance with the description herein, suitable computing devices include, by way of non-limiting example, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those skilled in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the systems described herein. In various embodiments, suitable tablet computers include those having booklet, slate, and convertible configurations known to those skilled in the art.
[0195]
[0215] In some embodiments, the computing device includes an operating system configured to execute executable instructions. An operating system is software, including programs and data, that manages the device's hardware and provides services for running applications, for example. Those skilled in the art will recognize that non-limiting examples of suitable server operating systems include FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those skilled in the art will recognize that non-limiting examples of suitable personal computer operating systems include Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that non-limiting examples of suitable mobile smartphone operating systems include Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.Those skilled in the art will also recognize that non-limiting examples of suitable media streaming device operating systems include Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Those skilled in the art will also recognize that non-limiting examples of suitable video game console operating systems include Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.
[0196] Non-transitory computer-readable storage medium
[0216] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of an optionally networked computing device. In further embodiments, the computer-readable storage medium is a tangible component of the computing device. In still further embodiments, the computer-readable storage medium is optionally removable from the computing device. In some embodiments, computer-readable storage media include, by way of non-limiting example, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In some cases, the programs and instructions are encoded on the medium persistently, substantially persistently, semi-persistently, or non-transitoryly.
[0197] computer program
[0217] In some embodiments, the platforms, systems, media, and methods described herein include the use of at least one computer program. A computer program comprises a sequence of instructions that is executable by one or more processors of a computing device and that is written to perform specified tasks. The computer-readable instructions may be implemented as program modules, such as functions, objects, application program interfaces (APIs), computational data structures, etc. that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those skilled in the art will recognize that computer programs can be written in a variety of languages and in a variety of versions.
[0198]
[0218] The functionality of the computer-readable instructions may be combined or distributed in various environments as desired. In some embodiments, a computer program includes one instruction sequence. In some embodiments, a computer program includes multiple instruction sequences. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, add-ons, or combinations thereof.
[0199] Web Applications
[0219] In some embodiments, the computer program comprises a web application. In light of the disclosure provided herein, those skilled in the art will recognize that web applications, in various embodiments, utilize one or more software frameworks and one or more database systems. In some embodiments, the web application is created on a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, the web application utilizes one or more database systems, including, by way of non-limiting example, relational, non-relational, object-oriented, associative, XML, and document-oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting example, Microsoft® SQL Server, mySQL™, and Oracle®. Those skilled in the art will recognize that web applications, in various embodiments, are written in one or more versions of one or more languages. Web applications may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written in part in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or Extensible Markup Language (XML). In some embodiments, a web application is written in part in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written in part in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® Actionscript, JavaScript, or Silverlight®.In some embodiments, the web application is written in part in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tcl, Smalltalk, WebDNA®, or Groovy. In some embodiments, the web application is written in part in a database query language such as Structured Query Language (SQL). In some embodiments, the web application integrates with enterprise server products such as IBM® Lotus Domino®. In some embodiments, the web application includes a media player element. In various further embodiments, the media player element utilizes one or more of a number of suitable multimedia technologies, including, by way of non-limiting example, Adobe® Flash®, HTML5, Apple® QuickTime®, Microsoft® Silverlight®, Java®, and Unity®.
[0200] Mobile Applications
[0220] In some embodiments, the computer program comprises a mobile application provided to the mobile computing device. In some embodiments, the mobile application is provided to the mobile computing device at the time of manufacture. In other embodiments, the mobile application is provided to the mobile computing device over a computer network as described herein.
[0201]
[0221] In light of the disclosure provided herein, mobile applications are created using techniques known to those skilled in the art, using hardware, languages, and development environments known in the art. Those skilled in the art will recognize that mobile applications may be written in a variety of languages. Suitable programming languages include, by way of non-limiting example, C, C++, C#, Objective-C, Java, Javascript, Pascal, Object Pascal, Python, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or a combination thereof.
[0202]
[0222] Suitable mobile application development environments are available from a variety of sources. Commercially available development environments include, but are not limited to, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available free of charge, including, but not limited to, Lazarus, MobiFlex, MoSync, and Phonegap. Mobile device manufacturers also distribute software development kits, including, but not limited to, the iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows Mobile SDK.
[0203]
[0223] Those skilled in the art will recognize that various commercial forums are available for the distribution of mobile applications, including, by way of non-limiting example, the Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, the App Store for Palm devices, the App Catalog for webOS, Windows® Marketplace for Mobile, the Ovi Store for Nokia® devices, Samsung® Apps, and the Nintendo® DSi Shop.
[0204] Standalone Applications
[0224] In some embodiments, the computer program comprises a stand-alone application, which is a program that runs as an independent computer process rather than as an add-on, e.g., a plug-in, to an existing process. Those skilled in the art will recognize that stand-alone applications are often compiled. A compiler is a computer program that converts source code written in a programming language into binary object code, such as assembly language or machine code. Suitable compiled programming languages include, but are not limited to, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB.NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, the computer program comprises one or more executable compiled applications.
[0205] Web browser plugin
[0225] In some embodiments, the computer program includes a web browser plug-in (e.g., an extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Software application manufacturers support plug-ins to allow third-party developers to create features that extend the application, facilitate easy addition of new features, and reduce the size of the application. When supported, plug-ins allow customization of the software application's functionality. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display specific file types. Those skilled in the art are familiar with various web browser plug-ins, including Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar includes one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar includes one or more explorer bars, tool bands, or desk bands.
[0206]
[0226] In light of the disclosure provided herein, one of ordinary skill in the art will recognize that a variety of plug-in frameworks are available that allow for the development of plug-ins in a variety of programming languages, including, by way of non-limiting example, C++, Delphi, Java, PHP, Python, and VB.NET, or combinations thereof.
[0207]
[0227] A web browser (also called an Internet browser) is a software application designed for use with networked computing devices to retrieve, present, and traverse information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting example, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, minibrowsers, and wireless browsers) are designed for use on mobile computing devices, including, by way of non-limiting example, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting example, Google® Android® browser, RIM BlackBerry® browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® mobile, Amazon® Kindle® Basic Web, Nokia® browser, Opera Software® Opera® mobile, and Sony® PSP™ browser.
[0208] Software Module
[0228] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or the use thereof. In light of the disclosure provided herein, software modules are created by techniques known to those skilled in the art using machines, software, and languages known in the art. The software modules disclosed herein are implemented in numerous ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or a combination thereof. In further various embodiments, a software module comprises multiple files, multiple sections of code, multiple programming objects, multiple programming structures, multiple distributed computing resources, multiple cloud computing resources, or a combination thereof. In various embodiments, one or more software modules include, by way of non-limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, a software module is in one computer program or application. In other embodiments, a software module is in two or more computer programs or applications. In some embodiments, a software module is hosted on one machine. In other embodiments, a software module is hosted on two or more machines. In further embodiments, the software modules are hosted on a distributed computing platform, such as a cloud computing platform. In some embodiments, the software modules are hosted on one or more machines in one location. In other embodiments, the software modules are hosted on one or more machines in two or more locations.
[0209] Database
[0229] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases or the use thereof. In light of the disclosure provided herein, one of ordinary skill in the art will recognize that many databases are suitable for storing and retrieving information. In various embodiments, suitable databases include, by way of non-limiting example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, the database is internet-based. In further embodiments, the database is web-based. In still further embodiments, the database is cloud computing-based. In certain embodiments, the database is a distributed database. In other embodiments, the database is based on one or more local computer storage devices. [Example]
[0210] Example
[0230] The following illustrative examples represent embodiments of the software applications, systems, and methods described herein and are not meant to be limiting in any way.
[0211] Example 1 - Quality control of digital information after DMA synthesis
[0231] Digital information is encoded into data polynucleotides using the methods described herein (FIGS. 5-7). A voltage is applied to the synthesis surface and the current is measured to determine if the synthesis surface is defective. Approximately 100,000 data polynucleotides are synthesized using the methods described herein, with continuous quality control by current sensing, optical imaging, and flow sensing. After polynucleotide synthesis, the polynucleotide length distribution and polynucleotide mass are estimated using absorbance and fluorescence measurements.
[0212]
[0232] Quality control (QC) polynucleotides are synthesized on a synthesis surface along with data polynucleotides. The qc polynucleotides account for approximately 1% of the total polynucleotides on the surface (e.g., data polynucleotides + QC polynucleotides). The QC polynucleotides comprise a first primer sequence and are amplified based on the first primer sequence. The data polynucleotides comprise a second primer sequence that is different from the first primer sequence. The QC polynucleotides are fully sequenced. The QC polynucleotides are then aligned with a reference. The reference is a known, preselected sequence. By aligning the QC polynucleotides with the reference, a relative read count is generated, identifying the number of QC polynucleotides with the same sequence as the reference sequence.
[0213]
[0233] The QC polynucleotides are aligned with the standards to estimate the error rate in the polynucleotides. The error rate in the QC polynucleotides serves as a proxy for the error rate in the data polynucleotides. The error rate is estimated to be less than 5%. The QC polynucleotides are also aligned with the standards to estimate the synthesis uniformity in the data polynucleotides. The synthesis uniformity of the QC polynucleotides is analyzed across the synthesis surface locations. The synthesis uniformity is estimated to be greater than 95%.
[0214]
[0234] A subset of data polynucleotides, comprising approximately 0.1% of the data polynucleotides, is also selected and amplified. The subset is randomly selected across the synthesis surface and sequenced. The subset is partially decoded by an internal codec to determine an index, as shown in FIG. 9. The index is used to estimate the relative distribution of the subset of multiple data polynucleotides, which are then arranged according to lane and frame. Because the subset comprises approximately 0.1% of the 100,000 data polynucleotides, the relative distribution of the subset should be centered around every 100th decoded index. The relative distribution is used to estimate synthesis uniformity, which is estimated to be greater than 95%.
[0215]
[0235] An internal codec including a greedy ML blending algorithm is further applied to generate a likelihood. The likelihood is based on the number of steps required for blending and the probability associated with each step in the algorithm. The likelihood is associated with an error rate in the data polynucleotides, with a high likelihood being associated with a low error rate and a low likelihood being associated with a high error rate. The error rate is estimated to be less than 5%.
[0216]
[0236] The error and uniformity estimates from the QC polynucleotides and the error and uniformity estimates from the subset of data polynucleotides are then combined to determine a final pass or fail for the synthesized data polynucleotides. Because the error rate from the quality control for both the QC polynucleotides and the subset of polynucleotides is less than 5% and the uniformity is greater than 95%, the data polynucleotides pass the quality control.
[0217]
[0237] While preferred embodiments of the present subject matter have been shown and described, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the present subject matter. It will be understood that various alternatives to the embodiments of the present subject matter described herein may be employed in practicing the present subject matter.
[0218]
[0238] The present disclosure is further illustrated by the following non-limiting sections.
[0219]
[0239] Item 1. A method for quality control (QC) of data polynucleotides, comprising: a. providing a plurality of QC polynucleotides on a surface, wherein the plurality of QC polynucleotides comprises a first primer sequence; b. amplifying a plurality of QC polynucleotides based on a first primer sequence; c. sequencing a plurality of QC polynucleotides; d. aligning a plurality of QC polynucleotides with a reference to estimate an error rate in the data polynucleotides, a composite uniformity in the data polynucleotides, or a combination thereof for the data polynucleotide QC; A method comprising:
[0220]
[0240] Item 2. The method of item 1, wherein the error rate, synthetic uniformity, or a combination thereof is based at least in part on the relative read counts of a plurality of QC polynucleotides.
[0221]
[0241] Item 3. The method of items 1 or 2, wherein the plurality of QC polynucleotides is about 1% or less than 1% of the polynucleotides on the surface.
[0222]
[0242] Item 4. The method according to any one of items 1 to 3, wherein a plurality of QC polynucleotides are provided on a portion of a surface.
[0223]
[0243] Item 5. The method according to any one of items 1 to 4, wherein the plurality of QC polynucleotides are uniformly provided on the surface.
[0224]
[0244] Item 6. The method according to any one of Items 1 to 5, wherein the data polynucleotide comprises a second primer sequence.
[0225]
[0245] Item 7. The method of Item 6, wherein the first primer sequence is different from the second primer sequence.
[0226]
[0246] Item 8. The method according to Item 6 or 7, wherein the first primer sequence and the second primer sequence are of different lengths.
[0227]
[0247] Item 9. The method according to any one of items 1 to 8, wherein the quality control is performed after synthesis of the data polynucleotide.
[0228]
[0248] Item 10. The method according to item 9, wherein QC is performed before cleavage of the data polynucleotides from the surface.
[0229]
[0249] Item 11. The method according to any one of items 1 to 10, wherein each of the QC polynucleotides is about 50 to 200 nucleic acid bases in length.
[0230]
[0250] Item 12. The method of any one of Items 1 to 11, wherein each of the data polynucleotides is about 100 to about 300 nucleic acid bases in length.
[0231]
[0251] Item 13. A method for quality control (QC) of data polynucleotides, comprising: a. selecting a subset of the plurality of data polynucleotides; b. applying an inner codec to a subset of the plurality of data polynucleotides, the inner codec comprising probabilistic decoding; c. estimating an error rate in the plurality of polynucleotides based, at least in part, on a likelihood associated with each decoded sequence in the subset of the plurality of data polynucleotides; and A method comprising:
[0232]
[0252] Item 14. The method of item 13, in which a lower error rate is associated with a higher likelihood.
[0233]
[0253] Item 15. The method of Item 13 or 14, wherein a higher error rate is associated with a lower likelihood.
[0234]
[0254] Item 16. The method according to any one of Items 13 to 15, further comprising decoding an index of the subset of data polynucleotides.
[0235]
[0255] Item 17. The method of item 16, wherein the index is decoded using an internal codec, an external codec, or a combination thereof.
[0236]
[0256] Item 18. The method of Item 16, wherein the index is used to estimate the relative distribution of a subset of a plurality of data polynucleotides.
[0237]
[0257] Item 19. The method according to any one of Items 13 to 18, wherein QC is performed during synthesis of the polynucleotide, during QC of the stored polynucleotide, or a combination thereof.
[0238]
[0258] Item 20. The method of any one of Items 13 to 19, wherein the subset of the plurality of data polynucleotides is randomly selected.
[0239]
[0259] Item 21. The method of any one of items 13 to 19, wherein the subset of the plurality of data polynucleotides is selected based at least in part on their position on the surface.
[0240]
[0260] Item 22. The method of any one of Items 13 to 21, wherein the plurality of data polynucleotides comprises about 100,000 polynucleotides.
[0241]
[0261] Item 23. The method of Item 22, wherein the subset of the plurality of data polynucleotides is about 0.1% of the plurality of data polynucleotides.
[0242]
[0262] Item 24. The method according to any one of items 1 to 23, wherein the method is used in combination with current sensing, optical imaging, flow sensing, size estimation, quality estimation, mass estimation, or any combination thereof.
[0243]
[0263] Item 25. The method of item 24, wherein current sensing includes measuring current in the tip or a section of the tip.
[0244]
[0264] Item 26. The method of Item 25, wherein the current is compared to a reference value.
[0245]
[0265] Item 27. The method of item 26, wherein a difference between the current and a reference value indicates a chip failure, a deblocking failure, or a combination thereof.
[0246]
[0266] Item 28. The method according to any one of Items 24 to 27, wherein the current detection is performed before synthesis of the plurality of data polynucleotides.
[0247]
[0267] Item 29. The method according to Item 28, wherein current sensing is used to detect chip defects, adjust polynucleotide synthesis locations on the chip, or a combination thereof.
[0248]
[0268] Item 30. The method according to any one of Items 24 to 29, wherein mass estimation is performed using fluorescence.
[0249]
[0269] Item 31. The method of Item 30, wherein fluorescence is used to detect the yield of multiple polynucleotides.
[0250]
[0270] Item 32. The method of any one of Items 24 to 31, wherein the optical imaging includes detecting chip defects, non-uniformities, or a combination thereof.
[0251]
[0271] Item 33. A method for performing QC of a plurality of cells on a surface, comprising: a. Measuring the current of each cell in a plurality of cells on a surface; b. determining whether one or more cells in the plurality of cells have a defect based at least in part on the current; c. synthesizing and / or storing the polynucleotide in a second one or more cells in the plurality of cells, wherein the second one or more cells do not have the defect; A method comprising:
[0252]
[0272] Item 34. The method of Item 33, wherein the defects include physical defects.
[0253]
[0273] Item 35. The method of items 33 or 34, wherein the surface is a synthetic surface, a storage surface, or a combination thereof.
[0254]
[0274] Item 36. The method according to any one of Items 33 to 35, further comprising blocking one or more cells having the defect.
[0255]
[0275] Item 37. The method according to Item 36, wherein blocking is carried out by a protecting group on the surface.
[0256]
[0276] Item 38. The method according to item 37, wherein blocking is carried out by a photolabile protecting group on the surface.
[0257]
[0277] Item 39. The method according to any one of Items 36 to 38, wherein blocking is carried out by selectively applying energy to one or more cells.
[0258]
[0278] Item 40. The method according to any one of Items 36 to 39, wherein the blocking is performed by a masking material.
[0259]
[0279] Item 41. The method according to any one of Items 36 to 40, wherein the blocking is carried out by addressable control of each cell in the plurality of cells.
Claims
1. 1. A method for quality control (QC) of data polynucleotides, comprising: a. providing a plurality of QC polynucleotides on a surface, said plurality of QC polynucleotides comprising a first primer sequence; b. amplifying the plurality of QC polynucleotides based on the first primer sequence; c. sequencing the plurality of QC polynucleotides; d. aligning the plurality of QC polynucleotides with a reference to estimate an error rate in the data polynucleotides, a composite uniformity in the data polynucleotides, or a combination thereof for the data polynucleotide QC; A method comprising:
2. 2. The method of claim 1, wherein the error rate, the composite uniformity, or a combination thereof is based at least in part on relative read counts of the plurality of QC polynucleotides.
3. 3. The method of claim 1 or 2, wherein the plurality of QC polynucleotides is about 1% or less than 1% of the polynucleotides on the surface.
4. The method of any one of claims 1 to 3, wherein the plurality of QC polynucleotides are provided on a portion of the surface.
5. The method of any one of claims 1 to 4, wherein the plurality of QC polynucleotides are uniformly provided on the surface.
6. The method of any one of claims 1 to 5, wherein the data polynucleotide comprises a second primer sequence.
7. The method of claim 6 , wherein the first primer sequence is different from the second primer sequence.
8. The method of claim 6 or 7, wherein the first primer sequence and the second primer sequence are of different lengths.
9. The method of any one of claims 1 to 8, wherein the quality control is performed after synthesis of the data polynucleotides.
10. 10. The method of claim 9, wherein the QC is performed prior to cleavage of the data polynucleotides from the surface.
11. The method of any one of claims 1 to 10, wherein each of the QC polynucleotides is about 50 to 200 nucleobases in length.
12. 12. The method of any one of claims 1 to 11, wherein each of the data polynucleotides is from about 100 to about 300 nucleobases in length.
13. 1. A method for quality control (QC) of data polynucleotides, comprising: a. selecting a subset of a plurality of data polynucleotides; b. applying an inner codec to the subset of the plurality of data polynucleotides, the inner codec comprising stochastic decoding; c. estimating an error rate in the plurality of polynucleotides based, at least in part, on a likelihood associated with each decoded sequence in the subset of the plurality of data polynucleotides; A method comprising:
14. The method of claim 13 , wherein a lower error rate is associated with a higher likelihood.
15. 15. The method of claim 13 or 14, wherein a higher error rate is associated with a lower likelihood.
16. The method of any one of claims 13 to 15, further comprising decoding an index for said subset of said data polynucleotides.
17. The method of claim 16 , wherein the index is decoded using the inner codec, the outer codec, or a combination thereof.
18. 17. The method of claim 16, wherein the index is used to estimate the relative distribution of the subset of the plurality of data polynucleotides.
19. 19. The method of any one of claims 13 to 18, wherein the QC is performed during synthesis of the polynucleotide, during QC of a stored polynucleotide, or a combination thereof.
20. The method of any one of claims 13 to 19, wherein the subset of the plurality of data polynucleotides is randomly selected.