Barcode sequences and related systems and methods
By designing and using sample distinction codes or barcodes, the problems of low throughput and high cost of sample identification in synthesis and sequencing technology are solved, and efficient and accurate multi-sequencing and sample identification are achieved.
Patent Information
- Application Number
- CN202210354669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-05-14
- Filing Date
- 2016-05-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2036-05-13
AI Technical Summary
The existing synthesis and sequencing technology has problems with low sequencing throughput and high cost during sample identification, especially when dealing with highly complex samples, it is difficult to effectively identify sample nucleic acids or other biological molecules.
Design and use sample discrimination codes or barcodes, and by generating stream space codewords and inserting fill characters, satisfying the predetermined minimum distance criterion, preparing corresponding barcode sequences, combining fault-tolerant codes to improve identification accuracy and efficiency.
It improves the throughput and accuracy of multiple sequencing, reduces the error rate during the sequencing process, and enhances the flexibility and customization capabilities of sample identification.
Smart Images

Figure CN114540475B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention with the application date of May 13, 2016, application number 201680027931.8, and invention name "Barcode Sequences and Related Systems and Methods".
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 161,309, filed on May 14, 2015, which is hereby incorporated by reference in its entirety.
[0003] Sequence Listing
[0004] This application contains a sequence listing, which has been electronically submitted in ASCII format and is hereby incorporated by reference in its entirety. The name of the ASCII copy created on May 12, 2016 is LT01016_SL.txt and its size is 18,815 bytes. Technical Field
[0005] The present disclosure generally relates to methods, systems, and kits for sample identification, and more particularly to methods, systems, and kits for the design and / or production and / or use of sample discrimination codes or sample discrimination barcodes for identifying sample nucleic acids or other biomolecules or polymers. Background Art
[0006] Each instrument, device, and / or system performs nucleic acid sequencing by a method of sequencing-by-synthesis, such as including the Genome Analyzer / HiSeq / MiSeq platforms (Illumina, Inc.; see, e.g., U.S. Patent Nos. 6,833,246 and 5,750,341); the GS FLX, GS FLX Titanium, and GS Junior platforms (Roche / 454 Life Sciences; see, e.g., Ronaghi et al., Science, 281:363-365 (1998) and Margulies et al., Nature, 437:376-380 (2005)); and the Ion PGM™ Sequencer and Ion Proton™ Sequencer (Life Technologies / Ion Torrent; see, e.g., U.S. Patent No. 7,948,015 and U.S. Patent Application Publication Nos. 2010 / 0137143, 2009 / 0026082, and 2010 / 0282617, which are hereby incorporated by reference in their entireties). To increase sequencing throughput and / or reduce the cost of sequencing-by-synthesis (and other sequencing methods, such as sequencing-by-hybridization, sequencing-by-ligation, etc.), new methods, systems, machine-readable media, and kits are needed that allow for the efficient preparation and / or identification of potentially highly complex samples. SUMMARY OF THE INVENTION
[0007] The present disclosure generally relates to methods, systems, and kits for sample identification, and more particularly to methods, systems, and kits for the design and / or fabrication and / or use of sample discriminator codes or sample discriminator barcodes for identifying sample nucleic acids or other biomolecules or polymers. One embodiment provides a method for designing barcode sequences corresponding to flow space codewords. A plurality of flow space codewords consisting of a string of characters can be generated. The position of at least one padding character within the flow space codeword can be determined. The padding character can be inserted into the determined position within the flow space codeword. After insertion, based on meeting a predetermined minimum distance criterion, a plurality of flow space codewords can be selected, where the selected codewords correspond to valid base space sequences in a predetermined flow order. And barcode sequences corresponding to the selected codewords can be prepared.
[0008] In several embodiments, after inserting the padding character, at least one codeword can be filtered in a predetermined flow order, including an invalid base space shift. In several embodiments, the set of selected codewords includes an error-correcting code that meets a predetermined minimum distance criterion.
[0009] In some embodiments, determining the position of the padding character within the flow space codeword may further include iterating over multiple positions of the padding character within the codeword. Additionally, during each iteration, the number of codewords corresponding to a certain valid base space sequence in a predetermined flow order may be calculated. Then, the position with the highest calculated number of codewords corresponding to a certain valid base space sequence may be selected among the multiple positions.
[0010] In some embodiments, determining the position of the padding character within the flow space codeword may further include, during each iteration, determining the base space sequence corresponding to the flow space codeword, such that when the padding character is inserted at the iterative position of the codeword, the base space sequence corresponds to a valid base space sequence. During each iteration, based on at least one length criterion of the determined sequence, the determined base space sequences may be filtered. And the number of valid base space sequences at the filtered iterative positions may be calculated. In some embodiments, the filtering during each iteration further includes filtering the determined base space sequences according to a nucleotide percentage content criterion.
[0011] In some embodiments, after inserting at least one padding character, the codewords of the error-correcting code are synchronized within the flow space.
[0012] In some embodiments, the generated flow space codewords include an initial distance between the codewords, such that the minimum distance between the selected codewords is greater than the minimum distance between the generated codewords. After inserting the padding characters, this initial distance between the codewords may be maintained.
[0013] In some embodiments, the selection of multiple codewords further includes grouping the codewords, such that the minimum intra-group distance between the codewords within each group consists of a first value, and the minimum inter-group distance between the codeword groups of different groups consists of a second value, with the first value being greater than the second value.
[0014] In some embodiments, a subset of the selected codewords may be determined, including a termination flow that does not represent a merge. A subset of the barcode sequences corresponding to the selected codeword subset may be prepared, such that an adaptor of the subset of the barcode sequences is selected according to the termination flow corresponding to the codeword subset that does not represent a merge.
[0015] In some embodiments, preparing the barcode sequence further includes appending a series of key bases to the barcode sequence, wherein, for the first segment of this barcode sequence, the appended key bases are terminated by a repeating base. For example, the first segment may contain half of the barcode sequence. In some embodiments, for the second segment of the barcode sequence, the appended key bases may be terminated by a non-repeating base. In some embodiments, all of the selected codewords include an error-correcting code composed of the minimum distance between the codewords, such that the change in the termination key bases appended to the generated barcodes corresponding to the selected codewords increases the minimum distance between the codewords.
[0016] One embodiment provides a method for sequencing a polynucleotide sample comprising a barcode sequence. At least some of a plurality of barcodes can be incorporated into a plurality of target nucleic acids to form a polynucleotide, wherein the plurality of barcodes are designed so that the barcodes correspond to a flow space codeword in a predetermined flow order, the flow space codeword being composed of one or more error-tolerant codes, and the plurality of barcodes includes at least 1,000 barcodes. According to the predetermined flow order, a series of nucleotides can be introduced into the polynucleotide. As a result of the introduction of nucleotides into the target nucleic acid, a series of signals can be obtained. The series of signals can be parsed within the barcode range, presenting a flow space string such that the presented flow space string matches a codeword, wherein, in the presence of one or more errors, at least one presented flow space string matches at least one codeword. In some embodiments, in the presence of one or more errors, at least one presented flow space string that matches at least one flow space codeword is used to identify a signal obtained from one of the plurality of target nucleic acid sequences, associated with the codeword corresponding to the matched flow space codeword.
[0017] In some embodiments, a kit for use with a nucleic acid sequencing instrument is provided. The kit may include a plurality of barcode sequences that meet the following criteria: the barcode sequences correspond to flow-space codewords in a predetermined flow order, such that the corresponding codewords include an error-tolerant code with a minimum distance of at least three; the lengths of the barcode sequences are within a predetermined length range; the barcode sequences are synchronized in the flow space; and the plurality of barcode sequences is at least 500 different barcode sequences. In some embodiments, the plurality of barcode sequences is at least 1000 different barcode sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate one or more exemplary embodiments and serve to explain the principles of various exemplary embodiments. The drawings are merely exemplary and explanatory and should not be construed as limiting or restrictive in any way.
[0019] Figure 1 A block diagram illustrating components of an exemplary nucleic acid sequencing system.
[0020] Figure 2A Illustrated are cross-sectional and detailed views of an exemplary nucleic acid sequencing flow cell.
[0021] Figure 2B An exemplary uniform flow front between successive reagents flowing through a portion of an exemplary reaction chamber array is illustrated.
[0022] Figure 3 An exemplary label-free, pH-based sequencing process is described.
[0023] Figure 4 A block diagram illustrating an exemplary system for obtaining, processing, and / or analyzing multiplex nucleic acid sequencing data.
[0024] Figure 5 An exemplary ionization plot showing a signal indicative of base response.
[0025] Figures 6A and 6B illustrate the relationship between a base space sequence and a flow space vector.
[0026] Figure 7 An exemplary method for designing barcode sequences corresponding to flow space codewords is illustrated.
[0027] Figure 8 An exemplary method for sequencing a polynucleotide sample containing a barcode sequence is illustrated.
[0028] Figure 9 A set of distinct polynucleotide strands, each having a unique barcode sequence, is illustrated.
[0029] Figures 10A-10C An exemplary workflow for preparing a multiplex sample is illustrated.
[0030] Figure 11 An exemplary bead template containing a barcode sequence is illustrated.
[0031] Figure 12 Another exemplary bead template containing a barcode sequence is illustrated. DETAILED DESCRIPTION
[0032] The following description and the various embodiments described herein are merely exemplary and explanatory and should not be construed as limiting or constraining in any way. Other embodiments, features, objects, and advantages of the present disclosure will be apparent from the specification, drawings, and claims.
[0033] According to the various embodiments, methods, systems, and kits are provided that allow for the efficient preparation and / or identification of samples. In several instances, the methods, systems, and kits can facilitate increased throughput by allowing for the simultaneous sequencing and / or analysis of multiple samples (e.g., multiplex sequencing), and promoting the sequencing analysis with sample discrimination codes or coded molecular constructs. Multiplex sequencing can allow for the substantially simultaneous analysis of multiple coded samples (e.g., different samples or samples from different sources) in a single sequencing run (e.g., on a common slide, chip, substrate, or other sample holding device) or in substantially simultaneous sequencing runs (e.g., on multiple slides, chips, substrates, or sample holders).
[0034] In some embodiments, the disclosed methods, systems, and kits can be used to identify a source of a sample used in multiplexed sequencing. The identification can involve the analysis of sample sequencing data. The sequencing data source can be uniquely labeled, encoded, or identified (e.g., to distinguish a particular nucleic acid species associated with a particular sample population). The identification is facilitated by using a unique sample discrimination code or sequence (also called a barcode, such as a synthetic nucleic acid barcode) that can be incorporated into or associated with the sample. The use of sample discrimination codes is still limited by errors or misreads that can occur during the sequencing process. For example, an incorrect barcode read can alter the interpretation of the barcode information, rendering the barcode unrecognizable and preventing the correct identification of the sample. An incorrect barcode read can also result in a sample being associated with an incorrect sample source or source population.
[0035] However, the disclosed embodiments can mitigate the problem of detecting and / or correcting errors that can occur during the sequencing of barcoded samples. For example, methods are provided for sample discrimination codes or sequences or barcodes and for developing robust sample discrimination codes or sequences or barcodes, which incorporate a fault tolerance code (e.g., an error correction code or an error detection code).
[0036] The disclosed embodiments can also generate a large number of potential barcodes, such as barcodes that can be used to distinguish samples from one another, which can also correspond to codewords that incorporate a fault tolerance code (e.g., an error correction code or an error detection code). For example, when sequencing the generated barcodes, a sequencing instrument can receive signals, and the final signals received can represent a codeword of a particular fault tolerance code. In some embodiments, a large number of potential barcodes incorporating a barcode fault tolerance design can improve the efficiency (e.g., the number of simultaneous targets that can be sequenced), accuracy (e.g., fault tolerance), flexibility, and customization of multiplexed assays.
[0037] Unless specifically specified otherwise herein, the terms, techniques, and symbols used herein in the areas of biochemistry, cell biology, cell and tissue culture, genetics, molecular biology, nucleic acid chemistry, and organic chemistry (including the chemical and physical analysis of polymeric particles, enzyme reactions and purification, nucleic acid purification and preparation, nucleic acid sequencing and analysis, polymerization techniques, preparation of synthetic polynucleotides, recombinant techniques, etc.) follow the standard protocols and texts of the relevant fields.See Kornberg and Baker, DNA Replication, 2nd ed. (W.H. Freeman, New York, 1992); Lehninger, Biochemistry, 2nd ed. (Worth Publishers, New York, 1975); Strachan and Read, Human Molecular Genetics, 2nd ed. (Wiley-Liss, New York, 1999); Birren et al. (eds.), Genome Analysis: A Laboratory Manual Series (Vols. I-IV), Dieffenbach and Dveksler (eds.), PCR Primer: A Laboratory Manual, and Green and Sambrook (eds.), Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press); and Hermanson, Bioconjugate Techniques, 2nd ed. (Academic Press, 2008).
[0038] As used herein, "amplification" generally refers to performing one amplification reaction. As used herein, "amplicon" generally refers to the product of a polynucleotide amplification reaction, including a clonal population of polynucleotides. The amplicon can be single-stranded or double-stranded and can be replicated from one or more starting sequences. In one example, the one or more starting sequences can be one or more copies of the same sequence or a mixture of different sequences containing a common amplified region, such as a specific exon sequence present in a mixture of DNA fragments extracted from a sample. An amplicon can also be formed by amplification of a single starting sequence. An amplicon can be produced by multiple amplification reactions, and the reaction product includes replicas of one or more starting nucleic acids or target nucleic acids. The amplification reaction that produces the amplicon may be "template-driven", based on the base pairing of reactants (nucleotides or oligonucleotides) having complements on a template polynucleotide, which is a prerequisite for the formation of the reaction product. The template-driven reaction can be the extension of a primer with a nucleic acid polymerase or the ligation of oligonucleotides with a nucleic acid ligase. Examples of such reactions are polymerase chain reaction (PCR), linear polymerase reaction, nucleic acid sequence-based amplification (NASBA), rolling circle amplification or forming a monomer using rolling circle amplification that can specifically occupy a micro-well, as disclosed in the U.S. Patent to Drmanac et al. Application Publication No. 2009 / 0137404, which is incorporated herein by reference in its entirety. As used herein, "solid-phase amplicon" generally refers to a solid-phase support, such as a particle or bead, to which a clonal population of nucleic acid sequences is attached, and the population can be produced by methods such as emulsion PCR.
[0039] As used herein, an "analyte" generally refers to a molecule or biological sample that can directly affect an electronic sensor in a region (e.g., a defined space or reaction confinement zone or micro-well), or indirectly affect such an electronic sensor through a by-product of a reaction involving the molecule or biological cell located in the region. In one embodiment, the analyte can be a sample nucleic acid or template nucleic acid that can undergo a sequencing reaction, which in turn can generate a reaction by-product, such as one or more hydrogen ions, that can affect an electronic sensor. The term "analyte" can also encompass multiple copies of analytes such as proteins, peptides, nucleic acids, etc., which are attached to a solid support such as beads or particles. In one embodiment, an analyte can be a nucleic acid amplicon or a solid-phase amplicon. The sample nucleic acid template can associate with a surface through covalent bonding or a specific binding or coupling reaction, and can be derived from a shotgun fragmented DNA amplicon library (an example of a library fragment discussed later herein) or a sample emulsion PCR process to form a clonally amplified sample nucleic acid template on microparticles such as IonSphere™. Analytes can include microparticles attached to a clonal population of DNA fragments such as genomic DNA fragments, cDNA fragments.
[0040] As used herein, a "primer" generally refers to a natural or synthetic oligonucleotide that, when forming a duplex with a polynucleotide template, can serve as a starting point for nucleic acid synthesis and extend, such as along the template from its 3' end to form an extended duplex. Primer extension can be carried out using a nucleic acid polymerase such as DNA or RNA polymerase. The nucleotide sequence added during the extension process can depend on the sequence of the template polynucleotide. The primer length can range from 14 to 40 nucleotides, such as from 18 to 36 nucleotides, or from N to M nucleotides, where N is an integer greater than 18 and M is an integer greater than N and less than 36. Other suitable primer lengths can be applied in various embodiments. Primers can be used in multiple amplification reactions, for example, a single primer is used in linear amplification reactions and two or more primers are used in polymerase chain reactions. Guidelines for primer length and sequence selection can be found in Dieffenbach and Dveksler (eds.), PCR Primer: A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, New York, 2003).
[0041] As used herein, "polynucleotide" or "oligonucleotide" generally refers to a linear polymer of nucleotide monomers, which can be DNA or RNA. Through a regular pattern of monomer-monomer interaction, the monomers that make up a polynucleotide can specifically bind to a native polynucleotide, and such interactions include Watson-Crick base pairing, base stacking, Hoogsteen or reverse Hoogsteen base pairing. Such monomers and their internucleoside bonds can be naturally occurring or their analogs (e.g., naturally occurring or non-naturally occurring analogs). Examples of non-natural analogs include PNA, phosphorothioate internucleoside bonds, bases containing a bonding group that allows attachment of a label such as a fluorophore or a hapten. In one embodiment, an oligonucleotide refers to a (relatively) smaller polynucleotide, such as a polynucleotide having 5 - 40 monomer units. In several examples, a polynucleotide includes natural deoxynucleosides linked by phosphodiester bonds (e.g., deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine of DNA, or the ribose counterparts of RNA). However, they can also include non-natural nucleotide analogs (e.g., modified bases, sugars, or internucleoside bonds). In one embodiment, a polynucleotide can be represented by a series of letters (uppercase or lowercase), such as "ATGCCTG", which is understood to be in the 5' → 3' order from left to right, and "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, "T" represents deoxythymidine, "I" represents deoxyinosine, and "U" represents deoxyuridine, unless otherwise noted or indicated by the context. Whenever the use of an oligonucleotide or polynucleotide is related to enzymatic treatment, such as by polymerase extension or by ligase ligation, the oligonucleotide or polynucleotide in such examples should not contain some analogs of internucleoside bonds, sugar units, or bases at any or several positions. Unless otherwise noted, the terminology and atomic numbering conventions follow the publicly available information in the following reference: Strachan and Read, Human Molecular Genetics, 2nd ed. (Wiley-Liss, New York, 1999). The size of a polynucleotide can range from a few monomer units (e.g., 5 - 40) to several thousand monomer units.
[0042] For example, as used herein, a "confined space" (or "reaction space", used interchangeably with "confined space") generally refers to any space or region (which can be one-dimensional, two-dimensional, or three-dimensional) in which at least several portions of a molecule, liquid, and / or solid can be confined, retained, and / or positioned. In various embodiments, the space can be a predetermined area (flat region) or volume, and can be defined within or associated with a well in a microplate, microtiter plate, microarray reader, or chip, or a recess or a microfabricated hole. Depending on the amount of liquid or solid, the area or volume can also be determined. For example, a liquid or solid deposited on an area or volume that additionally defines a space. For example, an isolated hydrophobic region on a generally hydrophobic surface can provide a confined space. In one embodiment, the confined space can be a reaction chamber, such as a well or a microhole, which can be in a chip. In one embodiment, the confined space can be a substantially flat region on a non-porous substrate. The confined space can contain or be exposed to enzymes and reagents used in nucleotide binding.
[0043] As used herein, a "reaction confinement region" or "reaction chamber" generally refers to any region in which a reaction is confined, and includes a "reaction chamber", a "well", or a "microhole" (all used interchangeably). The reaction confinement region can include a region in which a physical or chemical property of a solid substrate allows the localization of a target reaction. In several embodiments, the reaction confinement region can include a discontinuous region on the surface of a substrate that can specifically bind a target analyte (such as a discontinuous region having an oligonucleotide or an antibody covalently bonded to the surface). The reaction confinement region can be hollow or can have a well-defined shape and volume and can be formed in a substrate. In several embodiments, the latter types of reaction confinement regions can refer to microholes or reaction chambers in this document, which can be fabricated using any suitable microfabrication technique and can have a volume, shape, aspect ratio (such as the ratio of the bottom width to the hole depth), and other dimensional characteristics that can be selected according to a specific application, including the nature of the reaction occurring and the reagents, by-products, and labeling techniques (if any) used. For example, the reaction confinement region can also be a substantially flat region on a non-porous substrate. In various embodiments, microholes can be fabricated using any suitable fabrication technique known in the art. The following patents disclose exemplary configurations (such as spacing, profile, and volume) of microholes or reaction chambers: Rothberg et al., U.S. Patent Publication Nos. 2009 / 0127589 and 2009 / 0026082; Rothberg et al., U.K. Patent Application Publication No. GB2461127; and Kim et al., U.S. Patent No. 7,785,862, which are hereby incorporated by reference in their entirety and made a part of this document.
[0044] A defined space or reaction confinement region can be arranged as an array, which is substantially a one - or two - dimensional planar arrangement of units such as sensors or pores. The number of columns (or rows) in a two - dimensional array can be the same or different. In several embodiments, the array consists of at least 100,000 cavities. For example, a reaction cavity can have a lateral width and a vertical depth, and the width - to - depth ratio is about 1:1 or less. In several embodiments, the spacing between reaction cavities is no greater than about 10 microns, and the volume of each reaction cavity is no greater than 10 cubic microns (i.e., 1 picoliter), or no greater than 0.34 picoliter, or no greater than 0.096 picoliter, or in several instances, no greater than 0.012 picoliter. For example, the cross - sectional area at the top of a reaction cavity can be 2 2 2 3 2 4 2 5 2 6 2 7 2 8 2 or 9 2 or 10 2 square microns. In several embodiments, the array can have at least 3 10 4 10 5 10 6 10 7 10 8 10 9 or more reaction cavities. The reaction cavities can be coupled to chemFETs.
[0045] A defined space or reaction confinement zone, whether arranged in an array or in some other configuration, can be in electrical contact with at least one sensor to detect or measure one or more detectable or measurable parameters or characteristics. The sensor can convert the presence, concentration, or change in the amount of a reaction byproduct (or a change in the ionic characteristics of a reactant) into an output signal that can be electronically recorded as a change in voltage or current, which in turn can be processed to extract information about a chemical reaction or a desired association event, an example of which is a nucleotide binding event. The sensor includes at least one chemically sensitive field effect transistor (“chemFET”) configured to generate at least one output signal related to the characteristics of a chemical reaction or a neighboring target analyte. The characteristics can include the concentration (or change in concentration) of a reactant, product, or byproduct, or the value of a physical property such as ionic concentration (or a change in that value). For example, an initial measurement or detection of the pH of a defined space or reaction confinement zone can be represented as an electrical signal or a voltage that can be digitized (e.g., converted into a digital representation of the electrical signal or voltage). In various embodiments, these measurements and representations can be considered raw data or a raw signal.
[0046] As used herein, a “nucleic acid template” (or “sequencing template,” which may be used interchangeably with “nucleic acid template”) generally refers to a nucleic acid sequence that is the target of one or more nucleic acid sequencing reactions. A nucleic acid template sequence can include a natural or synthetic nucleic acid sequence. A nucleic acid template sequence can also include a known or unknown nucleic acid sequence from a target sample. In various embodiments, the nucleic acid template can be attached to a solid support, such as a bead, particle, flow cell, or any other surface, scaffold, or object.
[0047] As used herein, a “fragment library” generally refers to a collection of nucleic acid fragments, one or more of which are used as a sequencing template. There are many ways to generate a fragment library (e.g., by cleavage, shearing, restriction, or breaking down a larger nucleic acid into smaller fragments). Fragment libraries can be generated from or obtained from natural nucleic acids, such as from bacteria, cancer cells, normal cells, solid tissues, etc. Libraries consisting of synthetic nucleic acid sequences can also be generated to form a synthetic fragment library.
[0048] A “molecular sample discrimination code” (or “molecular barcode,” interchangeable with “molecular sample discrimination code”) as used herein generally refers to a recognizable or distinguishable molecular marker that can be uniquely resolved and attached to a sample nucleic acid, biomolecule, or polymer. The molecular sample discrimination code can be used to track, sort, separate, and / or identify sample nucleic acids, biomolecules, or polymers, and can be designed to have properties useful for manipulating nucleic acids, biomolecules, polymers, or other molecules. The molecular sample discrimination code can be composed of the same type / class of substance or subunit as the nucleic acid, biomolecule, or polymer it is intended to identify, or can be composed of one or more different substances or subunits. A molecular sample discrimination code can be composed of a short-chain nucleic acid that includes a known, predetermined, or designed sequence. A molecular sample discrimination code can be a nucleic acid sample discrimination code (or nucleic acid barcode), which can be a recognizable or distinguishable nucleotide sequence (e.g., an oligonucleotide or polynucleotide sequence). A number of molecular sample discrimination codes can include one or more restriction endonuclease recognition sequences or cleavage sites, overhangs, linker sequences, primer sequences, etc. (including combinations of features or properties). A molecular sample discrimination code can be a biopolymer sample discrimination code, which can include one or more antibody recognition sites, restriction sites, intra- or intermolecular binding sites, etc. (including combinations of features or properties). Multiple different molecular sample discrimination codes can be used to discriminate or characterize samples belonging to a common group and can be attached to, coupled with, or associated with a library of nucleic acids, biomolecules, polymers, or other molecules (such as a fragment library). In various embodiments, a sample discrimination code or sequence or barcode can represent a molecular sample discrimination code or molecular barcode and can include a set of symbols, components, or characters for representing or defining a molecular sample discrimination code or barcode. For example, a sample discrimination code or barcode can be composed of a sequence of letters that defines a sequence of known or predetermined nucleic acid bases or other biomolecule or polymer components. Other embodiments can employ any other suitable symbols and / or alphanumeric characters that are not letters. The sample discrimination code or barcode can be used in multiple sets, subsets, and groupings, such as as part of a sequencing cycle or for performing multiplexed sequencing. The sample discrimination code or barcode can be interpreted, recognized, discriminated, or construed as a function of a sequence or other arrangement or relationship of subunits that together form a single code. In a number of embodiments, the sample discrimination code can be composed of a series of signals that are output by a sequencing instrument during barcoded sequencing according to a predetermined flow order (such as a flow space corresponding to a barcode), as described in more detail below. In a number of embodiments, the sample discrimination code or barcode can also include one or more additional functional elements, including key sequences for quality control and sample detection, primer sites, ligation junctions, substrate attachment linkers, inserts, and any other suitable elements.
[0049] Figure 1Illustrates the components of an exemplary nucleic acid sequencing system that can be implemented with various embodiments. The components include a flow cell consisting of a sensor array 100, a reference electrode 108, various reagents 114, a valve bank 116, a cleaning solution 110, a valve 112, a jet controller 118, pipelines 120 / 122 / 126, channels 104 / 109 / 111, a waste liquid container 106, an array controller 124, and a user interface 128. The flow cell and sensor array 100 include an inlet 102, an outlet 103, a reaction chamber array 107, and a flow chamber 105, defining a flow path of the reagent over the reaction chamber array 107. The reference electrode 108 can be of any suitable type or shape, including concentric cylinders with fluid channels or a wire inserted into the inner cavity of the channel 111. The reagent 114 can be driven through the fluid channels, valves, and flow cell by a pump, gas pressure, or other suitable means, and can be discarded into the waste liquid container 106 after flowing out of the flow cell and sensor array 100.
[0050] For example, in some embodiments, the reagent 114 can contain dNTP and will flow through the channel 130 and the valve bank 116, which can control the flow of the reagent 114 through the channel 109 to the flow chamber 105. For example, the system can include a reservoir 110 containing a cleaning solution, which can be used to wash away the previously flowed out dNTP. The reaction chamber array 107 can include an array of defined spaces or reaction confinement regions such as holes or micro-wells, which can be operatively associated with a sensor array such that each reaction chamber has a sensor suitable for detecting an analyte or a target reaction characteristic. The reaction chamber 107 can be integrated with the sensor array as a single device or chip. The flow cell can have various designs for controlling the path and flow rate of the reagent within the reaction chamber array 107 and can be a microfluidic device. The array controller 124 can provide bias voltage, timing, and control signals to the sensors and collect and / or process the output signals. The user interface 128 can display information from the flow cell and sensor array 100 as well as instrument settings and controls, and allow the user to input or set the instrument settings and controls.
[0051] In some embodiments, the system can be configured to allow a single fluid or reagent to contact the reference electrode 108 during a multi-step reaction. Valve 112 can be closed to prevent the wash fluid 110 from flowing into channel 109 while the reagent is flowing. Although the flow of the wash fluid can be stopped, there may still be an uninterrupted fluid and electrical connection between the reference electrode 108, channel 109, and the micro-well array 107. The distance between the reference electrode 108 and the junction points between channels 109 and 111 can be selected such that the reagent that may diffuse into channel 111 from the flow in channel 109 hardly or completely does not reach the reference electrode 108. In one embodiment, the wash fluid 110 can be selected to continuously contact the reference electrode 108. In one example, such a configuration can be used for multi-step reactions with frequent wash steps. In each embodiment, using any suitable instrument control software, such as LabView (National Instruments Corporation, Austin, Texas, USA), the jet controller 118 can programmatically control the driving force of the flow of reagent 114 and the operation of valve 112 and valve set 116 to deliver the reagent to the flow cell and sensor array 100 in a predetermined reagent flow sequence. The reagent can be delivered at a predetermined flow rate for a predetermined duration, and physical and / or chemical parameters can be measured to provide information about the state of one or more reactions occurring in a defined space or reaction confinement zone, such as within a well or micro-well.
[0052] Figure 2A A cross-sectional view and a detailed view of an exemplary nucleic acid sequencing flow cell 200 in each embodiment are illustrated. The flow cell 200 can include an array of reaction chambers 202, a sensor array 205, and a flow chamber 206 through which a reagent flow 208 can flow across a surface of the array of reaction chambers 202 and through the open end of a reaction chamber. The reagent flow (such as nucleotide species) can be in any suitable manner, including pipette delivery or through a pipe or channel connected to a reaction chamber. The duration, concentration, and / or other flow parameters of each reagent flow can be the same or different. Similarly, the duration, composition, and / or concentration of each wash flow can be the same or different.
[0053] The reaction chamber 201 in the reaction chamber array 202 can have any suitable volume, shape, and aspect ratio, which can be selected according to one or more reagents, by-products, and labeling techniques used, and the reaction chamber 201 can be formed in the layer 210 by using any suitable fabrication or microfabrication technique. The shape of the reaction chamber can be a pore, a micropore, a through-hole, a surface portion that is more hydrophilic to liquids and serves as a confinement region, or any other suitable confinement structure. The sensor 214 in the sensor array 205 can be an ion-sensitive (ISFET) or a chemical-sensitive (chemFET) sensor, with a floating gate 218, having a sensing plate 220, separated from the interior of the reaction chamber by a passivation layer 216, and can respond (and generate a related output signal) to the amount of charge 224 present on the passivation layer 216 opposite the sensing plate 220. A change in the amount of charge 224 causes a change in the current between the source 221 and the drain 222 of the sensor 214, and this change can be directly used to provide a current output signal or indirectly used via additional circuitry to provide a voltage output signal. Reactants, cleaning fluids, and other reagents can be moved into the reaction chamber, for example, by diffusion 240. One or more analytical reactions can be carried out in one or more reaction chambers of the reaction chamber array 202 to identify or determine the characteristics or properties of a target analyte.
[0054] In some embodiments, the reaction directly or indirectly generates a by-product that affects the amount of charge 224 in the sensing proximity region (such as the adjacent region) of the sensing plate 220. In one embodiment, the reference electrode 204 can be connected to the flow chamber 206 via a fluid through the flow channel 203. The reaction chamber array 202 and the sensor array 205 together can form an integrated unit that forms the bottom wall or floor of the flow cell 200. In one embodiment, for example, one or more copies of an analyte can be attached to the solid-phase carrier 212, and the carrier can include micron particles, nano particles, microbeads, gels, and can be a porous solid. The analyte can include a nucleic acid analyte, containing one or more copies, which can be prepared by rolling circle amplification (RCA), exponential RCA method, or other suitable techniques to produce an amplicon without a solid carrier.
[0055] Figure 2BAn exemplary uniform flow front between successive reagent movements is illustrated, which movements pass through a cross-section 234 of an exemplary array of reaction chambers in various embodiments. A "uniform flow front" between a first reagent 232 and a second reagent 230 can mean that, as the reagents move, the reagents are not or are hardly mixed, such that the boundary 236 therebetween is narrow. For a flow cell having inlets and outlets at opposite ends of its flow chamber, the boundary can be a straight line, and for a flow cell having a central inlet (or outlet) and a peripheral outlet (or inlet), the boundary can be a curve. In one embodiment, the design of the flow cell and the reagent flow rates can be selected such that, during the switching of reagents, each newly introduced reagent has a uniform flow front as it passes through the flow chamber.
[0056] Figure 3 An exemplary label-free, pH-based sequencing process in various embodiments is illustrated. A template 682 containing a sequence 685 and a primer binding site 681 is attached to a solid support 680. The template 682 can be attached to a solid support such as a microparticle or a microbead as a clonal population, and can be prepared according to the disclosure of the following U.S. Patent to Leamon et al.: Patent No. 7,323,305, which is incorporated herein by reference in its entirety. In one embodiment, the template can be associated with the surface of a substrate or can be present in a liquid phase that is coupled or not coupled to a carrier. By operation, a primer 684 and a DNA polymerase 686 can bind to the template 682. As used herein, "bind by operation" generally means that a primer anneals to a template such that the 3' end of the primer can be extended by a polymerase, and that a polymerase binds to such a primer-template duplex (or in the vicinity thereof) such that, upon addition of dNTPs, binding and / or primer extension can occur.
[0057] In step 688, dNTP (shown as dATP) is added, and DNA polymerase 686 incorporates a nucleotide "A" (since "T" is the next nucleotide in template 682 and is complementary to the flowing dATP nucleotide). In step 690, a wash is performed as described herein. In step 692, the next dNTP (shown as dCTP) is added, and DNA polymerase 686 incorporates a nucleotide "C" (since "G" is the next nucleotide in template 682). Base incorporation in pH-based nucleic acid sequencing can be determined by measuring hydrogen ions generated as a natural byproduct of the polymerase-catalyzed extension reaction, and one or more features of the following documents can be used at least in part for such base incorporation: Anderson et al., A System for Multiplexed Direct Electrical Detection of DNA Synthesis, Sensors and Actuators B: Chem., 129:79-86 (2008); Rothberg et al., US Patent Application Publication No. 2009 / 0026082; and Pourmand et al., Direct Electrical Detection of DNA Synthesis, Proc. Natl. Acad. Sci., 103:6466-6470 (2006), which are hereby incorporated by reference in their entirety. In one embodiment, after each addition of dNTP, a step can be added to treat the reaction chamber with a dNTP-destroying agent (such as apyrase) to eliminate any dNTP remaining in the chamber, and dNTP residues may cause false extensions in subsequent cycles.
[0058] In one embodiment, a primer-template-polymerase complex can be exposed to a series of different nucleotides in a predetermined or known sequence or order. When one or more nucleotides are incorporated, a signal generated by the incorporation reaction can be detected, and after several cycles of adding nucleotides, extending the primer, and collecting the signal, the nucleotide sequence of the template strand can be determined. In one example, the output signal measured throughout this process depends on the number of nucleotide incorporations. In particular, in each additional sequencing step, when the next base in the template is complementary to the added dNTP, the polymerase extends the primer by incorporating the added dNTP. For one complementary base, there is one incorporation; for two complementary bases, there are two incorporations; for three complementary bases, there are three incorporations, and so on. With each incorporation, hydrogen ions are released, and the population of released hydrogen ions collectively changes the local pH of the reaction chamber.
[0059] In one embodiment, the generation of hydrogen ions is monotonically related to the number of consecutive complementary bases in the template (and the total number of template molecules having a primer and polymerase participating in the extension reaction). Thus, when there is a series of consecutive identical complementary bases in the template (which can represent a homopolymer region), the number of hydrogen ions generated and the magnitude of the local pH change are proportional to the number of consecutive identical complementary bases (and the corresponding output signal is sometimes referred to as a "monomer", "dimer", "trimer" output signal, etc.). If the next base in the template is not complementary to the added dNTP, no incorporation occurs and no hydrogen ions are released (and the output signal at this time is sometimes referred to as a "zero-mer" output signal). In several examples, in each wash step of the cycle, a non-buffered wash solution of a predetermined pH can be used to remove the dNTP from the previous step to prevent misincorporation in subsequent cycles. The delivery of nucleotides to a reaction vessel or reaction chamber can be referred to as the "flow" of nucleotide triphosphates (i.e., dNTPs). For convenience, the flow of dATP is sometimes referred to as the "flow of A" or "A flow", and a series of flows can be represented as a sequence of letters, such as "ATGT" representing "a flow of dATP, followed successively by a flow of dTTP, a flow of dGTP, and a flow of dTTP".
[0060] In one embodiment, four different dNTPs are added to the reaction chamber sequentially such that each reaction is exposed to the four different dNTPs, one at a time. In one embodiment, the four different dNTPs are added in the following order: dATP, dCTP, dGTP, dTTP, dATP, dCTP, dGTP, dTTP, etc., and after the exposure, incorporation, and detection steps is a single wash step. Exposure to the nucleotide, followed by a single wash, can be considered one "nucleotide flow". In several instances, four consecutive nucleotide flows can be considered one "cycle". For example, a two-cycle nucleotide flow sequence can be represented as follows: dATP, dCTP, dGTP, dTTP, dATP, dCTP, dGTP, dTTP, and the subsequent step for each exposure is a wash. Different flow sequences can be implemented, as described in more detail below. In each embodiment, the predetermined sequence or ordering can be based on the repetitive cycle of a predetermined reagent flow order (such as the repetitive sequence of four nucleotide reagents, such as "TACG TACG..."), or on a random reagent flow order, or on an ordering that consists of all or part of a phase-protecting reagent flow order as described in the following document: Hubbell et al., U.S. Patent Application No. 13 / 440,849, published as U.S. Patent Publication No. 2012 / 0264621 on October 28, 2012, entitled PHASE-PROTECTING REAGENT FLOW ORDERINGS FOR USE IN SEQUENCING-BY-SYNTHESIS, incorporated herein by reference in its entirety or by some combination thereof. In other embodiments, labeled, pH-based sequencing can be implemented in a similar manner.
[0061] Figure 4 illustrates an exemplary system for obtaining, processing, and / or analyzing multiplexed nucleic acid sequencing data in each of the exemplary embodiments. The system includes a sequencing instrument 601, a server 402, and one or more end-user computers 405. The sequencing instrument 401 is configured to process barcoded samples or deliver reagents in a predetermined order as detailed herein. The predetermined order can be based on the repetitive cycle of a predetermined reagent flow order (such as the repetitive sequence of four nucleotide reagents, such as "TACG TACG..."), or on a random reagent flow order, or on an ordering that consists of all or part of a phase-protecting reagent flow order, or on some combination thereof. In one embodiment, the barcode is determined at least in part as a function of the ordering. For example, the barcode can consist of barcodes designed for the flow space, designed according to a predetermined flow order, as described in more detail below. Exemplary sequencing instruments that can be used with the barcodes of the present disclosure include, but are not limited to, Ion PGMTM 、Ion Proton TM 、Ion S5 TM and Ion S5 XL Next Generation TM sequencing systems. Those of ordinary skill in the art will appreciate the fact that other sequencing instruments and platforms, such as various fluorophore-labeled nucleotide sequencing platforms, can also be used with the barcodes of this disclosure.
[0062] The server 402 may include a processor 403 and a memory and / or database 404. The sequencing instrument 401 and the server 402 may include one or more machine-readable media for obtaining, processing, and / or analyzing multiplexed nucleic acid sequencing data. In one embodiment, the instrument and the server or other computing means or resources may be configured as a single component. One or more of these components may be used to perform all or part of the embodiments described herein.
[0063] In several embodiments, according to this disclosure, the barcode consists of the codewords of a certain error-correcting code, where the codewords are represented in a flow space (such as including numbers, characters, or several other symbols corresponding to the number of nucleotide incorporations, which is a response to a predetermined nucleotide flow) rather than a base space.
[0064] In each exemplary embodiment, a method of sequencing by synthesis can be used to determine a sequence and / or identify one or more nucleic acid samples. During the sequencing-by-synthesis process, the sequence of a target nucleic acid can be determined by stepwise synthesis of a complementary nucleic acid strand on a target nucleic acid (whose sequence and / or identity is to be determined), and the nucleic acid strand is used as a template for the synthesis reaction. For example, by means of a polymerase extension reaction, which generally includes forming a complex containing a template (or target polynucleotide), an annealed primer, and a polymerase, and by operation, the polymerase is coupled or associated with the primer-template mixture so as to be able to incorporate a nucleotide (such as, nucleoside triphosphate, nucleotide triphosphate, precursor nucleoside, or nucleotide) into the primer. During the sequencing-by-synthesis process, nucleotides can be sequentially added to the growing polynucleotide molecule or strand at positions complementary to the template polynucleotide molecule or strand. The growing complementary strand can be detected by various methods (such as, pyrosequencing, fluorescence detection, label-free electronic detection, etc.), and adding nucleotides thereto can be used to identify the sequence composition of the template nucleic acid. This process can be iterated until a full-length or selected-length complementary sequence of the template has been synthesized.
[0065] As mentioned above, in each of the embodiments, data and signals that can be generated, processed, and / or analyzed can be obtained by electronic or charge-based nucleic acid sequencing. In electronic or charge-based sequencing (e.g., pH-based sequencing), nucleotide incorporation events can be determined by detecting ions (e.g., hydrogen ions) that are produced as natural by-products of polymerase-catalyzed nucleotide extension reactions. This can be used to sequence a sample or template nucleic acid, which can be a fragment of a target nucleic acid sequence and can be directly or indirectly attached to a solid support, such as particles, microparticles, microbeads, etc., as a clonal population. By operation, the sample or template nucleic acid can be associated with a primer and a polymerase and can undergo repeated cycles or "flows" of deoxynucleoside triphosphate ("dNTP") addition and washing. The primer can anneal to the sample or template so that the 3' end of the primer can be extended by a polymerase whenever a dNTP complementary to the next base in the template is added. Based on the known sequence of the nucleotide flow and the measured indication signals of the ion concentration in each nucleotide flow, the type, sequence, and degree of consistency of the nucleotides associated with the sample nucleic acid in the reaction chamber can be determined.
[0066] Figure 5 An exemplary ionization plot showing a signal that can achieve base response is shown. In this example, the x-axis shows the nucleotides flowing out, and by rounding the y-axis values to the nearest integer, the corresponding number of nucleotide incorporations can be estimated. The signals used in establishing the base response and determining the sequencing data (such as the flow space vector) can come from any suitable time point during the acquisition or processing of the data signals received during the sequencing operation. For example, the signal can be raw acquisition data or data that has been processed (e.g., by background filtering, normalization, signal attenuation correction, and / or phase difference or phase effect correction, etc.). The base response can be established by analyzing any suitable signal characteristics (such as signal amplitude, intensity, etc.).
[0067] In each embodiment, if the identity of the nucleotides flowing out and the order in which the signals are obtained are known, the output signals generated by nucleotide incorporation can be further processed to establish a base call for the flowing bases and to compile the consecutive base calls associated with a particular sample nucleic acid template into a read. A base call refers to the identification of a particular nucleotide, e.g., dATP ("A"), dCTP ("C"), dGTP ("G"), or dTTP ("T"). The base call can include performing one or more signal normalizations, estimating signal phase and signal softening (e.g., enzyme inactivation), and correcting the signal, and can identify or estimate each flowing base call for each defined space. The base call can include performing or implementing one or more of the teachings disclosed in the following: Davey et al., U.S. Patent Application No. 13 / 283,320, published as U.S. Patent Publication No. 2012 / 0109598 on May 3, 2012, entitled PREDICTIVE MODEL FOR USE IN SEQUENCING-BY-SYNTHESIS, which is incorporated herein by reference in its entirety. Other aspects of signal processing and base calling can include performing or implementing one or more of the teachings disclosed in the following: Davey et al., U.S. Patent Application No. 13 / 340,490, published as U.S. Patent Publication No. 2012 / 0173159 on July 5, 2012, entitled METHOD,SYSTEM, AND COMPUTER READABLE MEDIA FOR NUCLEIC ACID SEQUENCING; Sikora et al., U.S. Patent Application No. 13 / 588,408, published as U.S. Patent Publication No. 2013 / 0060482 on March 7, 2013, entitled METHOD, SYSTEM, AND COMPUTERREADABLE MEDIA FOR MAKING BASE CALLS IN NUCLEIC ACID SEQUENCING, each of which is incorporated herein by reference in its entirety.
[0068] Figure 6AFIGS. 6A and 6B illustrate the relationship between a base space sequence and a flow space vector. A series of signals (e.g., generated by flowing dNTPs in the presence of a polynucleotide) representing the number of incorporations (e.g., incorporation of out-flowing dNTPs into the polynucleotide) or lack of incorporation (e.g., zero-mers, mono-mers, di-mers, etc.) can be referred to as a flow space vector, sequence, or string. In one embodiment, the flow space vector, sequence, or string can include a series of symbols representing incorporations (e.g., 0, 1, 2, 3, etc.). When a predetermined flow order and a flow space vector are known, a translation to the base space can be generated. For example, given the number of incorporations (e.g., 0, 1, 2, or 3) and a particular out-flowing dNTP (e.g., A, G, T, C), the translated base space can include a base complementary to the out-flowing and incorporated dNTP, where the number of consecutive repeated bases can be consistent with the number of incorporations indicated by the flow space vector (e.g., 2 or more).
[0069] In one embodiment, a flow space vector can be generated using any suitable nucleotide flow order, which includes a predetermined order based on a continuous repeating cyclic pattern of a predetermined reagent flow order, a random reagent flow order, or an order based on all or part of a phased reagent flow order, or some combination thereof. Figure 6A In FIGS. 6A and 6B, the exemplary base space AGTCCA undergoes a sequencing operation using a TACG cyclic flow order. The flow generates a series of signals, the signal amplitudes (e.g., signal intensities) of which are related to the number of nucleotide incorporations (e.g., zero-mers, mono-mers, di-mers, etc.). This series of signals generates the flow space vector 101001021. As Figure 6A shown, under the TACG-cyclic ordering, the base space sequence AGTCCA can be translated into the flow space vector 101001021. As Figure 6B and detailed herein, the flow space vector can be mapped back to the base space sequence, which is associated with the sample under a predetermined flow order.
[0070] Barcode
[0071] In various embodiments, a sample discrimination code or barcode can include or correspond to (either directly or indirectly) a nucleotide sequence, a sequence of biomolecular components and / or subunits, or a sequence of polymeric components and / or subunits. In one embodiment, a sample discrimination code or barcode can correspond to a sequence of a single nucleotide in a nucleic acid, or a subunit of a biomolecule or polymer, or to a collection, group, or continuous or discontinuous sequence of these nucleotides or subunits. In one embodiment, a sample discrimination code or barcode can also correspond to (either directly or indirectly) a conversion between nucleotides, biomolecular subunits, or polymeric subunits, or other relationships between the subunits (e.g., linkers, key bases, etc.) used to form the sample discrimination code or barcode.
[0072] In various embodiments, the sample discrimination code or barcode can have properties such that it can be sequenced, identified, characterized, or decoded with higher accuracy and / or a lower error rate for a given pattern, length, or complexity. In one embodiment, the sample discrimination code or barcode can be designed as a set (which can include subsets) of single sample discrimination codes or barcodes. In several embodiments, the selection of one or more sample discrimination codes or barcodes in a set (or a subset of the set) can be based on one or more criteria to increase accuracy and / or reduce the error rate in the reading order, identification, characterization, discrimination, or decoding of these codes.
[0073] In various embodiments, the sample discrimination code or barcode can be designed to exhibit a high-fidelity read order, which can be evaluated based on empirical sequencing metrics. The fidelity can be based on a prediction of the read order accuracy of a sample discrimination code or barcode having a specific nucleotide sequence. Nucleotide sequences that are known to cause unclear reads, errors, or sequencing biases can be avoided. The design can be based on an accurate response to the sample discrimination code or barcode (and associated sample or nucleic acid population), even in the presence of one or more errors. In various embodiments, the fidelity can be based on the probability of correctly sequencing the sample discrimination code or barcode, which is at least 82%, 85%, 90%, 95%, 99%, or greater.
[0074] In various embodiments, the sample discrimination code or barcode can be designed to exhibit a relatively high read order accuracy for sequencing using a sequencing-by-synthesis platform (as previously discussed), such platforms can include fluorophore-labeled nucleotide sequencing platforms or label-free sequencing platforms, such as the Ion PGM TM and Ion Proton TM sequencers, Ion S5 TM and Ion S5 XL NextGeneration TM sequencing systems. However, the design of the sample discrimination code or barcode and the specific sequences is not limited to any particular instrument platform or sequencing technology. For non-nucleic acid codes, the sample discrimination code or barcode can be sequenced, identified, decoded, or recognized using methods known in the art, such methods include amino acid sequencing of protein sample discrimination codes.
[0075] In various embodiments, the design means may include applying a series of sample discrimination codes or barcode constraints or criteria to achieve desired characteristics or performance. Such constraints or criteria may include the uniqueness of one or more nucleic acid barcode sequences and their separation from other nucleic acid barcode sequences. A barcode set may be a nested set of barcodes, which may be based on one or more design criteria. In one embodiment, the design of the nested barcode set may be similar to Matryoshka nesting, such that the characteristics of a certain subset are fully contained within the characteristics of a genus set. For example, the first barcode subset meeting certain characteristics (such as higher sequencing fidelity) may be selected from a larger barcode set meeting the same characteristics. For example, if a barcode set consists of 96 uniquely distinguishable barcodes, for a sequencing experiment with only 16 multiplexed samples, a subset of 16 barcodes may be selected from the 96 available barcodes. Thus, this subset of 16 barcodes can be optimized to be similar to a larger subset of 32 or 48 barcodes selected from the full set of 96 barcodes. In one embodiment, the barcodes can be designed as an ordered column of nested barcodes. In one embodiment, the barcodes (such as a 96-code set) can be sorted as follows: having a first code, a second code (one of the remaining 95 codes) that is the farthest from the first code under a suitable distance metric, a third code (one of the remaining 94 codes) that is the farthest from the first and second codes under a suitable distance metric, and so on until all the barcodes are sorted.
[0076] In various embodiments, the sample discrimination code or barcode can be combined with a target sequence, in which case it helps to uniquely identify or distinguish different target sequences. For example, the target sequence can be any type of sequence from any target source, including amplicons, candidate genes, mutation hotspots, single nucleotide polymorphisms, genomic library fragments, etc. For example, at any time point during sample preparation, techniques such as PCR amplification, DNA ligation, bacterial cloning, etc. can be used to operably couple the sample discrimination code or barcode sequence to the target sequence through operations. The sample discrimination code or barcode sequence can be included in an oligonucleotide and ligated to a genomic library fragment using any suitable DNA ligation technique.
[0077] In various embodiments, the lengths of the sample discrimination codes or barcodes can vary. For example, based on the number of samples to be identified, the length of the sample barcodes can be selected. In various embodiments, for a multiplexed sequencing experiment with 16 samples, 16 uniquely distinguishable barcodes may be sufficient to uniquely identify each sample. Similarly, for a multiplexed sequencing experiment with 64 or 96 samples, 64 or 96 barcodes may be sufficient respectively.
[0078] Several configurations can utilize longer codes or larger barcodes in order to obtain a larger number of multiplexed sequences. While longer barcodes would allow identification of more samples, in some cases, these longer barcodes may have drawbacks. For example, in sequencing-by-synthesis, longer barcodes require additional nucleotide flow, which can reduce accuracy if sequencing in an earlier flow tends to be most accurate. Additionally, if a sequencing system has a length criterion (such as 200 base pairs), longer barcodes can occupy more sequencing space. Thus, it may be required that the target fragments appended to the barcode conform to a smaller length criterion (e.g., longer barcodes are more suitable for sequencing shorter targets).
[0079] In various embodiments, a sample discrimination code or barcode can be designed based on one or more of the above criteria (which can be selected singly or in combination). Based on the sequencing experiment, different combinations of criteria can be selected. For example, if fewer samples are to be barcoded, it may not be necessary to design the barcode to have nested subsets. Design criteria can be selected based on the number of samples, target accuracy, single-sample detection sensitivity of the sequencing instrument, accuracy of the sequencing instrument, and so on.
[0080] In various embodiments, the sample discrimination codes or barcodes described herein can be used in any suitable manner to assist in identifying or differentiating samples. For example, a barcode can be used alone, or two or more barcodes can be used in combination. In one embodiment, a single barcode can identify one or more target sequences. For example, a single barcode can identify a set of target sequences. The read of a barcode can be separate from the target sequence, or can be part of a larger read operation that encompasses the barcode and the target sequence. The barcode can be located at any suitable position within the sample, including before and after a target sequence.
[0081] Barcode Design and Flow Space
[0082] In various embodiments, a sample discrimination code or barcode can be designed based on a flow space. In other words, the barcode can be designed based at least in part on a flow space vector (such as as a function of a flow order). For example, the design of a sample discrimination code or barcode can be based on a projection within the flow space as a flow space vector under a selected or predetermined nucleotide flow order. In another example, a series of flow space vectors can be generated, and then these vectors can be translated into base space (such as in a predetermined flow order) in order to generate a barcode sequence.
[0083] In one embodiment, a barcode stream space vector can be composed of a string of symbols (such as a string of numbers or characters, such as 0, 1, 2, etc., representing no incorporation, monomer incorporation, dimer incorporation, etc., respectively), and is a response to a nucleotide stream flowing out or being introduced in a predetermined order. In various embodiments, the stream space string or vector can represent or correspond to a codeword of a certain error - tolerant code (such as an error - correcting code). In an error - correcting code, a string of characters can enable errors introduced into the string (such as during sequencing) to be detected and / or corrected based on the remaining characters in the string. An error - correcting code can be composed of a set of different strings on a finite alphabet Σ of given character units, which can be called codewords. A codeword can be regarded as including a message plus some redundant data or parity - check data, allowing the decoder to correctly decode a codeword containing one or more errors. Codewords can be designed to be distinct enough from each other to allow the detection of an allowable number of errors during the transmission of a codeword, and in several cases, to correct these errors by calculating which actual codeword is closest to the received codeword.
[0084] In various embodiments, any suitable type of error - correcting code design can be used to design sample - differentiating codes or barcodes. The error - correcting code can be a linear block code using an alphabet Σ of character units, and each codeword has n coded character units. Redundant and / or parity - check data can be added to a message (such as a subset of the codeword) to allow the receiver to detect and / or correct errors in a transmitted codeword and to use a suitable decoding algorithm to recover the original message. For example, in sequencing - by - synthesis, when a barcode has been sequenced and projected into the flow space as a stream space string, a message string can be considered to have been "transmitted".
[0085] In various embodiments, different numbers of character units in the code alphabet can be used to design sample - differentiating codes or barcodes, and this alphabet can vary depending on the specific application. The error - correcting code can be a binary code using an alphabet of two character units. The error - correcting code can be a ternary code using an alphabet of three character units. In one embodiment, the number of character units can depend on the length of the longest homopolymer sequence allowed in the barcode sequence. For example, if a barcode has only monomers (no repeated bases), the error - correcting code can be a binary code with one character representing no incorporation and another representing single - base incorporation (e.g., the alphabet Σ of such a binary code can be {0, 1}). In another example, if the barcode has only monomers and dimers, the error - correcting code can be a ternary code with the same character representing no incorporation and the character representing single - base incorporation and a third character representing double - base incorporation (e.g., the alphabet Σ of such a ternary code can be {0,1,2}). If the barcode sequence has trimers, tetramers, etc., the size and set of characters used for other code alphabets can be modified appropriately.
[0086] In various embodiments, an error-correcting code design for a sample discrimination code or barcode can be employed that is at least partially based on Hamming codes, Gray codes, and / or tetracode codes. In one embodiment, the error-correcting code can be a binary Hamming code, a binary Gray code, a ternary Hamming code, a ternary Gray code, and / or any other suitable code. See Hoffman et al., Coding Theory: The Essentials, Marcel Dekker, Inc. (1991); and Lin et al., Error Control Coding: Fundamentals And Applications, Prentice Hall, Inc. (1983).
[0087] In various embodiments, the sample discrimination code or barcode can be designed to have fault tolerance characteristics expressed in the flow space. In other words, sequencing errors can be related to incorrect digits or characters in the flow space representation of the barcode in a predetermined flow order (e.g., an incorrect "0" or "zero-mer" instead of the expected "1" or "monomer" in the flow space representation). For example, a set of single-base (in the flow space) fault-tolerant barcodes can be designed such that if a sequencing error is encountered at any position in the flow space representation of one or more barcodes in the set, each barcode can still be distinguished from the other barcodes in the set because their flow space representations differ from the flow space representation of the incorrect barcode in at least two digit positions; thus, if one of these two digit positions is in error, the other digit positions are still available to allow for barcode discrimination. The set can also be designed such that barcodes can be distinguished even if there are multiple errors (e.g., 2, 3, etc.) in the flow space. Such fault tolerance characteristics can help provide a high level of confidence (e.g., accuracy) when resolving complex multiplexed sequencing samples in the presence of potential sequencing errors. Candidate barcodes in a set can be compared to confirm fault tolerance in the flow space. For example, such barcodes can be compared (e.g., through computer analysis or simulation) to determine if the codes can still be distinguished if any error (or 2, 3, etc., depending on the criteria) occurs in the flow space. In another example, candidate flow space codewords (e.g., codewords translated into candidate barcode sequences in a predetermined flow order) can be compared to confirm fault tolerance.
[0088] A variety of algorithms and / or software tools can be used to assist in generating error-correcting codes. A series of different design ideas can be incorporated into the development of coding strategies. As explained herein, for a given stream order, there is a mutual mapping relationship between the barcode sequence and the stream space codeword. Thus, the design or selection criteria related to the barcode sequence can be translated into corresponding stream space coding design / selection criteria. Similarly, the design / selection criteria related to stream space coding can be translated into corresponding barcode sequence design / selection criteria.
[0089] In various embodiments, one or more distance metrics capable of evaluating the codeword distance can be used to design sample discrimination codes or barcodes. In one embodiment, the distance metric can be the Hamming distance, corresponding to the number of positions where two codewords differ. Mathematically, if the Hamming distance between each codeword in a codeword set and all other codewords in the set is at least d, then the code can correct at most (d - 1) / 2 errors. Conversely, the Hamming distance d for decoding at most x errors is 2x + 1. The quantity d can be referred to as the minimum code distance. The notation [n, k, d] can be used to characterize an error-correcting code of length n bits that encodes k information bits and has a minimum distance d. For example, other distance metrics can be used, including the Euclidean distance metric, the sum of the absolute values of the differences between the corresponding input terms of two codewords, and the sum of the squared differences between the corresponding input terms of two codewords. In various embodiments, by using such distance metrics, the codeword distance can be evaluated within the stream space.
[0090] In various embodiments, a sample discrimination code or barcode can be designed to have an error-correcting code with a minimum distance of five, capable of correcting at most two errors within the codeword. In another embodiment, the error-correcting code can have a minimum distance of three and be capable of correcting single errors within the codeword. In several instances, a software algorithm or method compares each codeword in a candidate group with all other codewords in the group to construct a maximum codeword set that maintains an ideal minimum distance and has an ideal error-correcting ability. Such an algorithm or method can be used to select or combine codewords including error-correcting codes. The codewords (or corresponding barcodes) can be further divided into subsets that, when used alone, can correct multiple stream space errors (e.g., two or more errors). This enables a barcode set to correct at least two stream space errors. A ternary coding scheme can be used to generate a barcode set within the stream space (e.g., in a given stream within the stream space, a visual barcode can have 0, 1, or 2 incorporations).
[0091] In various embodiments, sample discrimination codes or barcodes can be designed to discriminate reads in the flow space rather than the base space, which is effective for sequencing by synthesis and helps avoid excessive flows, thereby reducing error accumulation and loss of sequencing capacity. In some cases, the Hamming distance works well in the base space. For example, a single-base insertion at the start of a sequence (e.g., from ACGT to AACGT) will result in a Hamming distance of 3 (although the insertion / deletion distance is only 1). Also, when the paired bits translate a binary code into 4 letters, an error automatically affects two bits simultaneously, and it is not guaranteed that a single error correction of 1 bit can correct 1 error in the base read. Moreover, conventional barcode designs may not adequately address sequencing error motifs. In one embodiment, codewords can be selected for useful biological properties.
[0092] In some embodiments, sample discrimination codes or barcodes can be designed around a ternary Hamming code that is mapped to a specific predetermined flow order. For example, this code can be an [n = 13, k = 10, d = 3] ternary Hamming code; this mapping can take the first 10 "ternaries" (e.g., the ternary code symbols 0, 1, and 2) and assign them to several flows of a predetermined flow order (e.g., flows 9 - 18), and then take three "parity check" ternaries and assign them to other flows (e.g., 19 - 21). In some embodiments, the final synchronization flow is a monomer (e.g., a 'C' at flow 22), and as a result, if designated as zero, the terminating flow of the codeword is zero. Some codewords generated under the Hamming code may not be admissible representations in the flow space (e.g., they can be valid mathematical codewords in the flow space that, according to a certain predetermined flow order, do not correspond to a possible nucleic acid sequence in the base space). These codewords can be filtered. In some embodiments, the codewords can be further filtered to include only codewords of an ideal length (e.g., 9 - 15 bases).
[0093] In some configurations, a multiplex sequencing application utilizes a set of 96 barcodes that can correct two errors in a flow space string, has some space to allow for potential losses due to problematic barcodes, and has a predetermined flow order TACG, then in sequence TACG, TCTG, AGCA, TCGA, TCGA, TGTA, CAGC; for such an application, a 13 - bit - long ternary Hamming code can be used to generate a set of barcode sequences, where ten bits of the code are treated as data and three bits of the code are treated as codeword parity checks. This particular coding scheme results in approximately 140 codewords that can correct up to two errors.
[0094] In one example, barcodes can be selected to be 9 - 11 bases long and designed for the Ion PGM TMOligonucleotides used for multiplex sequencing in a sequencer. In this example, the oligonucleotide contains components in the following order: a primer site, a TCAG key sequence (such as key bases) for quality control and sample detection, a unique barcode sequence, a synchronous common C base at the 3' end of the barcode sequence (to ensure that, if designated as zero, the codeword termination stream is zero), and a GAT buffer between the barcode and the insert (to minimize the impact of the variable barcode region on adapter ligation). Similar to the P1 adapter for the Ion PGM TM sequencer, the GAT buffer is the last three bases. The information in Table 1 below is organized by the serial numbers of the barcodes generated. The second column shows the key sequence, the barcode sequence, and the common C base. The third column shows the barcode sequence and the common C base. The fourth column shows the projection of the combined sequence unit into the flow space. In the table, the bases and the flow space vector units corresponding to the barcodes are shown in bold. In the flow space mapping, flows 1 - 8 are assigned to the key sequence (i.e., flow 1 = T, flow 2 = A, flow 3 = C, flow 4 = G, flows 5 - 8 repeat flows 1 - 4), flows 9 - 18 are assigned to the data bits of the barcode (i.e., flow 9 = T, flow 10 = C, flow 11 = T, flow 12 = G, flow 13 = A, flow 14 = G, flow 15 = C, flow 16 = A, flow 17 = T, flow 18 = C), and flows 19 - 21 are assigned to the parity bits (i.e., flow 19 = G, flow 20 = A, flow 21 = T). Since in this example, all barcodes are followed immediately by a common C base, flow 22 (i.e., flow 22 = C) is used for synchronization. In one embodiment, the predetermined flow order may include these 22 flows and additional flows, such that the flow order consists of a repeating series of 32 flows (i.e., flow 23 = G, flow 24 = A, flow 25 = T, flow 26 = G, flow 27 = T, flow 28 = A, flow 29 = C, flow 30 = A, flow 31 = G, flow 32 = C). Other suitable flow orders as described herein may also be implemented. In other embodiments, the key, the synchronization base, and / or the buffer may be different.
[0095] Table 1 - Exemplary Barcodes and Projections in Space
[0096]
[0097] In each embodiment, a sample discrimination code or barcode can be designed around a [n = 11, k = 6, d = 5] ternary Gray code using values 0, 1, 2. This code has 729 (i.e., 3 6( ) different codewords, with a length of 11, a codeword spacing of 5, and capable of correcting 2 errors. The codewords can be generated linearly, cyclically, or by any suitable method, such as through a generator matrix or a generator polynomial. In other embodiments, changes in the key bases (or streams) for barcodes (or stream space codewords) and changes in the termination bases (such as the termination "C" base and / or the corresponding termination "1" stream) can generate additional barcodes for multiplex reactions. For example, if the use of shared key bases (or streams) and a termination static base limits the number of eligible barcode sequences for multiplex sequencing (e.g., at most 96 or 384 barcode sequences), then changes in these characteristics can generate a larger number of barcode sequences for multiplex sequencing (e.g., at least 1000 barcode sequences).
[0098] Figure 7 An exemplary method for designing barcode sequences corresponding to stream space codewords is illustrated. In step 7002, a plurality of potential stream space codewords can be generated. For example, a generator function can generate a set of potential stream space codewords using a ternary [n = 13, k = 10, d = 3] Hamming code with a length of 13 bits, where ten bits are regarded as data and three bits are regarded as codeword parity checks. The number of potential codewords generated can include n^k, which is 3^10 in this example. The generated stream space codewords can include an ordered series of characters, such as alphanumeric mixed characters or other symbols. Other configurations of Hamming codes or Gray codes can be implemented similarly. The codewords can be generated linearly, cyclically, or by any suitable method, such as through a generator matrix or a generator polynomial.
[0099] In step 7004, a position can be determined for a padding character within the potential stream space codeword. For example, in several configurations, a base, such as a padding base, can be attached to the end of the barcode to assist in sequencing. Within the stream space, according to a predetermined stream order, the padding base can correspond to a padding stream. For example, the padding base can include a "C" base, and according to the predetermined stream order, the corresponding padding stream (or character) can include a "1". In these configurations, a termination "1" can be attached to the end of the stream space codeword, and similarly, a termination C base can be attached to the end of the corresponding barcode sequence. For example, if the generated stream space codeword consists of 13 characters, then after adding the padding character, the codeword can include 14 characters. In one embodiment, the predetermined stream order can include an order based on a stream order cycling pattern, based on a random stream order, or based on an ordering that includes all or part of a phase-preserving stream order, or several combinations thereof.
[0100] In some embodiments, the padding character may be moved so that it is not in the termination stream of a codeword (and the corresponding padding base is not the termination base of a barcode sequence). This flexibility in repositioning the padding character / base will result in a range of potential position choices, enabling design benefits to be obtained. For example, inserting a padding character into a 13-character stream space codeword at a selected position can increase the number of codewords that map to a valid base space sequence in a predetermined stream order. Based on the known dNTP reagents flowing out and the reagent incorporation status (e.g., 0, 1, 2, etc.), the corresponding base space sequence for the stream space codeword can be determined in a predetermined stream order. Although the translated sequences may have different base lengths, they can still be synchronized within the stream space based on the corresponding synchronous stream space codewords.
[0101] In one embodiment, in a predetermined stream order, a number of codewords fail to translate into a valid base space sequence. In one example, given a stream order of "A", "G", "A", a stream space string (or vector) 101 will not translate into a valid base space. Here, the second incorporation (e.g., the second occurrence of 1 in 101) does not translate into a valid base sequence. For example, if two "A"s appear in a barcode without an intervening base, the first stream of "A" will exhibit a dimer (e.g., will present a stream space value of 2), and thus it is impossible to present the second incorporation. Other methods may be implemented to determine an invalid base space translation.
[0102] In one embodiment, inserting padding characters at different positions in a codeword can provide an adjustment such that a number of codewords that previously failed to map to a valid base space sequence successfully map to a valid base space sequence after the insertion. In several examples, the position that results in the most codewords corresponding to a valid base space sequence may be selected. Here, for the generated codewords, the selected padding base insertion positions may be consistent in order to preserve the distance of the codewords (e.g., preserve the distance used to generate the codewords).
[0103] In one embodiment, determining the position of a padding character within a flow space codeword may include iterating through multiple positions of the padding character within the codeword such that, based on inserting the padding character at an iterative position into the codeword, the number of codewords for each position can be calculated, and in a predetermined flow order, the codewords correspond to valid base space sequences. For example, given a codeword of length 13, there are 14 possible padding character positions (i.e., before the first character, between the first and second characters, and so on). Here, an algorithm can be designed such that the padding character is iteratively inserted into these 14 possible positions, and for each iteration, the number of codewords after inserting the padding stream at the iterative position can be calculated, and the codewords are mapped to valid base space sequences. For example, this algorithm can determine whether, after insertion at the iterative position, the codeword is mapped to a valid base space sequence in a predetermined flow order. Then, for each iterative position, the number of these codewords mapped to valid base space sequences can be calculated.
[0104] In one embodiment, calculating the number of flow space codewords (mapped to a valid base space sequence in a predetermined flow order) for a given iterative position further includes determining the base space sequence corresponding to the flow space codeword (mapped to a valid base space sequence after inserting the padding character at the iterative position). In some embodiments, the calculated number of codewords may include the number of determined base space sequences. In other embodiments, the determined base space sequences can be further filtered. For example, according to one or more sequence length criteria and nucleotide percentage content (such as GC content) and other suitable criteria, the determined base space sequences can be filtered. For example, a sequence length criterion may include 9 - 11 or 9 - 14 bases, and base space sequences greater than the criterion requirements can be filtered. Other suitable length ranges can also be implemented.
[0105] In another example, barcode sequences can be designed or selected to avoid some nucleotide sequences that are known to cause sequencing misreads or sequencing biases. This can enhance PCR and / or sequencing performance. In some embodiments, since the GC (guanine / cytosine) content of a sequence can affect sequencing quality, the filtering criterion may include a GC content range of 40 - 60%. The AT content can be processed similarly. In one example, if the determined base space sequence does not meet the GC and / or AT content criteria, it can be filtered.
[0106] In one embodiment, the number of flow space codewords (mapped to a valid base space sequence in a predetermined flow order) at a certain iterative position that is calculated may include the number of determined base space sequences after filtering.
[0107] In one example, after iterating over the possible positions for padding characters, such as iterating over a 13-character flow space, the position corresponding to the highest calculated quantity can be selected. This selected position results in a larger number of flow space codewords, which are mapped to a valid base space sequence, and thus a larger number of corresponding barcode sequences (e.g., useful for multiplex sequencing).
[0108] The following example is presented to illustrate the above principle. In this example, the selected position can include flow 5 of the flow space codeword (e.g., selected based on this position corresponding to the highest calculated quantity). This example will illustrate how to utilize the flexibility of padding character repositioning (from the termination flow to a selected position) to adjust a codeword (not translated to a valid base space) such that after inserting a padding character at the selected position, the codeword is successfully mapped to a valid base space sequence. Using a flow space codeword "20012220010121" generated from a sample and a sample flow order "T C T G A G C A T C G A T C" (SEQI.D.NO.19), the flow space codeword can include a padding terminator "1". According to the sample flow order, although the padding character is in the termination flow, the sample flow space codeword does not map to a valid base space sequence, at least because of two consecutive incorporations with no base incorporation in between. For example, when considering a hypothetical synchronous flow, the underlined flow space symbols in the codeword 20012220010121 would correspond to the underlined flows in the sample flow order T C T G A G C A T C G A T C (SEQI.D.NO.19). Here, the starting "2" represents a dimer incorporation when the starting "C" flow exits. The subsequent "00" represents two zero-mer incorporations when the subsequent "A" and "T" flows exit. However, the subsequent "1" in the flow space codeword represents a monomer that cannot be generated based on the corresponding "C" flow in the sample flow order. This is at least because if there were such a complementary base, the starting "2" incorporated due to the starting "C" flow would have been incorporated as a "3" or trimer. In this example, repositioning the padding character from flow 14 to flow 5 results in a valid base space sequence according to the sample flow order. This repositioning yields the codeword "20011222001012". Then, the adjusted flow space codeword maps to a valid base space translation according to the same sample flow order. Based on the teachings of this disclosure, one of ordinary skill in the art will appreciate that other possible repositioning positions can be similarly implemented and will also result in a valid base space translation.
[0109] Another example also illustrates the flexibility of using a padding character to reposition from a termination stream to a selected position (as above, also stream 5) to adjust a codeword (not yet translated into a valid base space) such that after inserting the padding character at the selected position, the codeword successfully maps to a sequence of valid base spaces. Using a stream space codeword "00000210220211" generated from a sample and a sample stream order "T C T G A G C A T C G A T C" (SEQ I.D. NO. 19), the stream space codeword can also include a padding terminator "1". In this example, a series of key streams can also be utilized. For example, a series of key bases can be placed in front of a barcode to track the barcode and / or the attached target. The key streams in the sample ordering (e.g., a series of streams directly preceding the sample ordering) can include "T A C G", followed by a repeating stream order "T A C G". Following the sample stream order and the key streams, although the padding character is in the termination stream, the sample stream space codeword does not map to a sequence of valid base spaces, at least because too many streams result in a zero-mer following the key streams. For example, for a hypothetical synchronous stream, the underlined stream space symbols in the codeword 00000210220211 would correspond to the underlined streams in the sample stream order T C T G A G C A T C G A T C (SEQ I.D. NO. 19). Here, the last key stream includes a "G". In the highlighted stream series, all three other possible dNTPs that are not G flow out in the 5-stream series, and thus at least one of these streams would have to result in an incorporation in order for the codeword to map to a valid stream space. In this example, repositioning the padding character from stream 14 to stream 5 results in a valid sequence of base spaces according to the sample stream order. This repositioning results in the codeword "0000121022021". The adjusted stream space codeword then maps to a valid base space translation according to the same sample stream order. Based on the teachings of this disclosure, one of ordinary skill in the art will appreciate the fact that other possible repositioning positions can be similarly implemented and will also result in a valid base space translation.
[0110] In this embodiment, after inserting the padding character, the stream space codeword remains synchronous. For example, the stream space codeword after insertion will include length X plus the inserted character (e.g., X + 1). Thus, relative to the stream length and the predetermined stream order, the codeword will still be synchronous.
[0111] In step 7006, padding characters can be inserted at determined positions of the flow space codewords. For example, a position can be selected for the padding character based on the number of possible positions calculated as described herein. Then, the flow space codewords can be adjusted by inserting the padding character at the selected positions of the codewords, as described herein. In one embodiment, the generated set of codewords can be inserted such that the error tolerance of the codewords is maintained (e.g., maintaining the minimum distance property).
[0112] In step 7008, potential flow space codewords can be filtered. For example, flow space codewords that are not mapped to a valid base space according to a predetermined flow order can be filtered. After inserting a padding base at a selected position of a codeword, filtering of potential flow space codewords can occur. Here, for a number of codewords that have not been mapped to a valid base space sequence (e.g., not filtered), they can be retained based on the insertion of the padding base at the selected position. For example, several instances are further described herein regarding how the repositioning of a padding base from a termination flow to a selected flow (e.g., flow 5) results in a codeword that is mapped to a valid base space sequence.
[0113] In some embodiments, flow space codewords can also be filtered according to the base space sequence length. For example, potential flow space codewords that include a base space translation following a predetermined flow order greater than a certain threshold length can be filtered. For example, a sequence length criterion can be 9 - 11 or 9 - 14 bases, and flow space codewords corresponding to base space sequences greater than the criterion requirements can be filtered. Other suitable length ranges can also be implemented.
[0114] In some embodiments, potential flow space codewords can also be filtered according to the minimum distance. For example, a classification algorithm can be implemented that selects a subset of potential flow space codewords having a predetermined minimum distance. In one instance, a group of codewords can be selected such that they achieve a minimum distance among themselves and a second minimum distance from other codewords in other groups. Potential codewords not selected by such a classification algorithm can be filtered similarly.
[0115] In some embodiments, potential codewords can be filtered according to the nucleotide percentage content (e.g., GC content). For example, barcode sequences can be designed or selected to avoid some nucleotide sequences that are known to cause sequencing misreads or sequencing biases. This can enhance PCR and / or sequencing performance. In some embodiments, since the GC (guanine / cytosine) content of a sequence can affect the sequencing quality, the filtering criterion can include a range of 40 - 60% GC content. The AT content can be treated similarly. In one instance, potential codewords translated into a base space sequence can be filtered if they do not meet the GC and / or AT content criteria.
[0116] In some embodiments, potential codewords can be filtered according to the secondary structure or properties in the experiment. For example, the experimental performance of a barcode sequence that is self-complementary or complementary to a certain primer sequence (coupled to the barcode) may be poor. Accordingly, potential codewords that are translated into a base-space sequence (self-complementary or complementary to a certain primer sequence coupled to the barcode) can be filtered out.
[0117] In step 7010, a key stream can be attached to the filtered codewords. For example, when tracking a barcode or a corresponding nucleic acid fragment (such as a target nucleic acid), a key stream can be used. The key stream can correspond to key bases in a predetermined stream order. In some embodiments, static key bases (such as "T", "C", "A", "G") can be used to attach to the barcode sequence. In this example, the stream-space string corresponding to the predetermined stream order can similarly include a static key stream (such as 10100101) that can be attached to the stream-space codewords.
[0118] In some embodiments, a variable key stream (or base) can be implemented to further separate the stream-space codewords from each other. For example, two possible sets of different key bases can be used, which differ based on a repeating termination base (such as "T", "C", "A", "G" and "T", "C", "A", "G", "G"). In this example, the change in the stream space corresponding to the predetermined stream order can include either a "1" or a "2" in the last key stream (such as 10100101 and 10100102). In one embodiment, two different key streams can be attached to the end of the stream-space codewords to further separate the codewords from each other.
[0119] In some embodiments, a set of filtered codewords can be replicated, where the first set of codewords is attached with the first key stream, and the replicated set of codewords is attached with the second key stream. Here, two versions of the same codeword can differ due to the key stream attached to the end of the codeword (such as because of "1" or "2" for the termination key stream). The change in the key stream can effectively increase the minimum distance between the codewords by at least one unit.
[0120] In other embodiments, other variable key streams (or bases) can be implemented similarly. For example, the termination stream can include a "1", "2", or "3", such that three different key streams can be generated and attached to the end of the codewords. In other embodiments, other differences in the key streams can be implemented similarly to increase the codeword spacing.
[0121] In step 7012, the filtered codewords can be selected and combined. For example, the codewords can be combined by minimum distance. In one embodiment, a classification algorithm can be implemented that selects a subset of codewords having a predetermined minimum distance. In one example, a group of codewords can be selected such that they achieve a minimum distance among themselves and a second minimum distance from other codewords in other groups. In several embodiments, the minimum distance used by the classification algorithm can include effectively increasing the minimum distance by appending a variable key stream to the end of the codeword. In one embodiment, the first minimum distance can be greater than the second minimum distance. For example, the first minimum distance can include 6, while the second minimum distance can include 4. In several embodiments, one or more groups of codewords can include a general group such that the codewords in the general group include a first minimum distance (e.g., within the group and between groups) relative to all other codewords. Other suitable minimum distance values can be implemented.
[0122] In one embodiment, a single group of codewords can include a first error-tolerant code, and the selected barcodes can collectively (e.g., jointly among the groups) include a second error-tolerant code. For example, the minimum distance between codewords within the same group can define the first error-tolerant code, and the minimum distance between codewords in different groups can define the second error-tolerant code. Based on the different minimum distances, the first error-tolerant code can distinguish and / or correct more sequencing errors than the second error-tolerant code.
[0123] In one embodiment, after filtering and classification, the grouped flow space codewords can include at least 500, 1000, 3000, 5000, 7000, or 9000 codewords. In accordance with the present disclosure, a representative list of barcodes corresponding to these grouped codewords using the above techniques with reference to FIG. 7 can be seen in Table 2 below and Appendix A of the following U.S. Patent: Application No. 62 / 161,309, filed May 14, 2016, to which this application claims priority and which is hereby incorporated by reference in its entirety.
[0124] In step 7014, barcodes corresponding to the grouped flow space codewords can be fabricated or caused to be fabricated. For example, barcodes corresponding to the grouped flow space codewords in a predetermined flow order can be fabricated according to the details provided herein. In several examples, fabricating can include causing the barcodes to be fabricated. In one embodiment, the grouped and fabricated barcodes can include at least 500, 1000, 3000, 5000, 7000, or 9000 barcodes.
[0125] In one embodiment, the barcodes corresponding to the grouped codewords may be organized by a plate. For example, as described herein, the codewords selected for a certain group may correspond to a group of barcodes. The group of barcodes may be organized by a plate (such as a structure storing grouped barcodes). Thus, the barcodes of a particular plate may include the fault tolerance (such as the minimum distance property) of the grouped codewords corresponding to those barcodes, and the barcodes between plates may include the fault tolerance (such as the minimum distance property) of the non-grouped codewords corresponding to those barcodes.
[0126] In each exemplary embodiment, when manufacturing barcodes, a plurality of barcode adapters may be attached after the barcode sequence. However, when sequencing in a predetermined flow order, the barcode corresponding to the adjusted (such as in the case of a terminated static or filled flow, the corresponding base has been relocated) codeword no longer includes a static termination output signal (such as within the flow space). Here, the termination output signal (such as within the flow space) may be predicted based on the predetermined flow order and the flow space codewords. In one embodiment, the barcodes may be divided into two categories. For example, the first category includes barcodes ending with a positive incorporation signal in a predetermined flow order (such as a "1" or "2" within the flow space), while the second category includes barcodes not ending with a positive incorporation signal (such as a "0" within the flow space). In several embodiments, the barcodes of the first category may use any suitable adapter (such as a universal adapter). However, the barcodes of the second category may use an adapter starting with a specific base (such as G) due to the lack of an incorporation signal, in order to mitigate potential sequencing errors. According to the predetermined flow order, the specific base may include a predetermined base. For example, the specific base may be predetermined such that an expected dNTP flow based on the predetermined flow order results in an incorporation (such as generating an incorporation signal).
[0127] In one embodiment, barcode manufacturing may include the manufacturing of forward barcodes, forward primers (P1a), reverse barcodes, and reverse primers (P1b). In one embodiment, an initial step may be to purify these oligonucleotides, with all nucleotides normalized to 100 - 400 μM of TE or low TE buffer. In one embodiment, non-ligated oligonucleotides (such as reverse barcodes and P1b) may be purified by high performance liquid chromatography (HPLC), while ligated oligonucleotides (such as forward barcodes and P1a) may be purified by a desalting technique. Those of ordinary skill in the art are familiar with the various desalting techniques that can be used for barcode manufacturing.
[0128] For example, using HPLC for reverse barcodes and P1b helps reduce sequencing errors. Oligonucleotides are synthesized from 3' to 5', so synthesis failures from reverse barcodes and P1b may be truncated at the 5' end. Lack of HPLC treatment for these strands can increase adapter dimers (e.g., from roughly 0% to roughly 5 - 15%). Additionally, forward barcodes and P1a are directly ligated to amplicons, and any cross - contamination can lead to false base responses. Moreover, due to the large number of sequences, HPLC is both costly (or cost - ineffective) and vulnerable to cross - contamination. Desalting these strands without HPLC is less expensive and does not require the use of these strands on specialized laboratory equipment (i.e., an HPLC instrument), thus eliminating a source of cross - contamination. Also, during nick translation, using the forward barcode and P1a as a template, DNA polymerase rewrites the reverse barcode and P1b, thereby removing any contamination from HPLC - contaminated P1b and reverse barcode sequences. This further reduces the risk of contamination of the strands for HPLC analysis.
[0129] In one embodiment, after purification, equal amounts of forward and reverse barcode oligonucleotides can be combined with P1a and P1b oligonucleotides and annealed in separate tubes using certain annealing conditions. For example, the annealing conditions can include: denaturing at 95 ºC for 5 minutes; starting at 89 ºC for 2 minutes and then decreasing by 1 ºC every 2 minutes for 64 cycles; incubating at 4 ºC for 1 hour, up to overnight (e.g., between 6 and 12 hours).
[0130] After annealing, equal amounts of annealed barcode adapters and P1 adapters can be combined. The sample can be diluted 5 - fold with a low TE buffer. 2 μL of the diluted mixture / AmpliSeq reaction can be added. Other variations of barcode fabrication can be implemented similarly.
[0131] In one embodiment, the barcode fabrication step can include synthesizing a polynucleotide. Using conventional polynucleotide synthesis techniques known in the art, a polynucleotide containing a barcode sequence can be fabricated.
[0132] According to various exemplary embodiments, the manufactured barcodes can be combined to form a barcode kit for sequencing. For example, grouped barcodes can be allocated to one or more plates or other platforms for nucleic acid sequencing (including multiplexed sequencing). For example, the codewords corresponding to the barcodes within a particular plate can have a first minimum distance from the corresponding codewords of the barcodes within that particular plate, and a second minimum distance from the corresponding codewords of the barcodes of other plates. In several embodiments, a barcode plate can be constituted by a flexible plate such that the codewords corresponding to the barcodes within the flexible plate include a first minimum distance from the corresponding codewords of all other barcodes of the package. Here, considering the minimum distance characteristics of the corresponding flow space codewords of the barcodes, the barcodes of the flexible plate can be used as alternative barcodes for all other plates. By selecting a number of valid barcodes, the barcode kit can be customized based on the target application, and the barcode kit can also include a complete set of barcodes.
[0133] The sequencing kit can also include a polymerase. The sequencing kit can also include a plurality of containers for accommodating different polynucleotides, and each different polynucleotide can be stored in each different container. The polynucleotides can be oligonucleotides with a length of 5 - 40 bases. The sequencing kit can also include a plurality of different nucleotide monomers. The sequencing kit can also include a ligase.
[0134] In several embodiments, the sequencing kit can include a plurality of different polynucleotides (e.g., which can be stored in vials), and each different polynucleotide includes a different barcode sequence as described herein. The polynucleotides can be oligonucleotides with 5 - 40 bases. The polynucleotides themselves can be barcode sequences, and they can also include other units such as primer sites, adapters, ligation sites, linkers, etc. The sequencing kit can also include a collection of precursor nucleotide monomers (to perform sequencing-by-synthesis operations) and / or various other reagents involved in a particular workflow for sample preparation and / or sequencing.
[0135] In one embodiment, a barcode or a group of barcodes can be used for multiplexed sequencing. For example, unique barcodes can be attached to multiple target nucleic acids such that the unique barcode sequences (or their flow space representations) can be identified after sequencing to identify the target nucleic acids.
[0136] Figure 8 A method for sequencing a polynucleotide sample with barcode sequences according to an exemplary embodiment of the present disclosure is illustrated. For example, according to the exemplary embodiments described herein, a plurality of barcodes corresponding to the flow space codewords in a predetermined flow order.
[0137] In step 8002, multiple barcodes can be incorporated into multiple target nucleic acids to create polynucleotides. For example, the barcodes can be attached to the target nucleic acids by any conventional means such that the signals obtained from the barcodes during sequencing can identify the specific target nucleotides attached to the barcodes.
[0138] In one embodiment, multiple different target nucleic acids are provided for multiplexed sequencing by a predetermined nucleotide flow, each different target nucleic acid being attached to a respective supplied barcode sequence that corresponds to a different flow space string, and each different flow space string being a different codeword of a fault-tolerant code or an error-correcting code. In one embodiment, the barcodes utilized can include at least 500, 1000, 3000, 5000, 7000, or 9000 barcodes. Similarly, in one embodiment, the different target nucleic acids can be at least 500, 1000, 3000, 5000, 7000, or 9000.
[0139] In step 8004, a series of nucleotides can be introduced into the polynucleotide in a predetermined flow order. For example, a dNTP reagent flow can exit in a predetermined flow order such that the polynucleotide is exposed to the reagent flow and incorporation events can occur.
[0140] In step 8006, a series of signals can be obtained due to the introduction of the series of nucleotides. For example, hydrogen ions released due to nucleotide incorporation into the polynucleotide can be detected, wherein the signal amplitude can be related to the amount of hydrogen ions detected. In another example, inorganic pyrophosphate released due to nucleotide incorporation into the polynucleotide can be detected, wherein the signal amplitude is related to the amount of inorganic pyrophosphate detected.
[0141] In step 8008, a series of signals of the barcode sequence can be resolved to present a flow space string such that the presented flow space string matches the codeword, wherein in the presence of one or more errors, at least one presented flow space string at least matches one codeword. In one embodiment, the series of signals can include a flow space vector or a string of symbols (such as 0, 1, 2, etc.) representing the number of incorporations of a certain flow (such as, zero-mer, monomer, dimer, etc.).
[0142] In one embodiment, any suitable decoding algorithm and / or software tool can be used to decode the flow space string of the barcode sequence to correct and / or detect errors. For example, an exhaustive algorithm can be employed for decoding, which compares a codeword with one error to all other members of the code and decodes it to the nearest matching codeword. If the erroneous codeword is equidistant from two codewords or farther from any codeword than half the minimum distance, the algorithm indicates the detection of an error without correction. In another example, decoding can involve performing the encoding operation in reverse. In another example, the decoding algorithm can use linear algebra techniques to decode the codeword.
[0143] In one embodiment, once at least one error-containing codeword matches a codeword of a fault-tolerant code or an error-correcting code, a signal obtained from one of the target nucleic acid sequences (associated with the corresponding barcode of the matched flow space codeword) can be identified. For example, for a target nucleic acid based on the matched codeword, a presented flow space string and a corresponding base space sequence can be identified.
[0144] In several embodiments, the scale of multiplexed sequencing facilitated by the large number of barcodes provided can facilitate some sequencing applications. For example, using a large number of barcodes that enable highly multiplexed sequencing, genotyping while sequencing, clone verification, and other detection synthesis verifications (such as verifying that a synthetic sequence is correct) can be performed more effectively. In several embodiments, a non-transitory machine-readable storage medium including instructions is also provided, which when executed by a processor, cause the processor to perform the methods and their variations detailed herein. A system is also provided, including: a machine-readable memory; a processor configured to execute the machine-readable instructions, which when executed by the processor, cause the system to perform the methods and their variations detailed herein.
[0145] According to an exemplary embodiment, a set of different polynucleotide chains is provided, with different barcode sequences for the polynucleotide chains; wherein each barcode sequence gives different flow space strings according to the flow space projection of a predetermined flow order, and these strings are codewords of a certain fault-tolerant code detailed herein. Figure 9 Seven different polynucleotide chains each associated with a unique barcode sequence are illustrated. Each embodiment includes a larger number of barcode sequences and polynucleotide chains, and the seven polynucleotides are representative examples. Each polynucleotide chain can have a primer site, a standard key sequence, and a unique barcode sequence. Each polynucleotide chain can also have a different target sequence. For such a set of polynucleotide chains, multiplexed sequencing can be performed, and the barcodes help identify the sequence data sources derived from a multiplexed sample.
[0146] According to an exemplary embodiment, a sample identification kit is provided, including: a plurality of sample discrimination codes, wherein: a) each sample discrimination code consists of a sequence of a single subunit; b) the subunit sequence of each sample discrimination code can be distinguished from the sequence of a single subunit of each other member among the plurality of sample discrimination codes; c) each sample discrimination code tolerates one or more errors so that other sample discrimination codes can be independently distinguished.
[0147] According to an exemplary embodiment, a sample identification kit is provided, including: a plurality of sample discrimination codes, wherein: a) each sample discrimination code consists of a sequence of a single subunit; b) a detectable signal is associated with each subunit or subunit pair or subunit set, such that each sample discrimination code is associated with a sequence of a detectable signal; c) each detectable signal sequence can be distinguished from the detectable signal sequences of every other member among the plurality of sample discrimination codes; and d) the detectable signal sequence of each sample discrimination code tolerates at least one error so as to be able to independently distinguish other sample discrimination codes.
[0148] Figures 10A-10C An exemplary workflow for preparing a multiplex sample is illustrated. Figure 10A An example of constructing a library of genomic DNA fragments is shown. A bacterial genomic DNA 10 can be fragmented into many DNA fragments 12 by any suitable technique, such as sonication, mechanical shearing or enzymatic digestion. Then platform-specific adapters 14 can be ligated to the fragments 12. Referring Figure 10B , each fragment sample 18 can then be isolated and combined with a bead 16. For the identification of the fragment 18, a barcode sequence (not shown in the figure) can be ligated to the fragment 18. Then the fragment 18 can be clonally amplified onto the bead 16, obtaining many cloned copies of the fragment 18 on the bead 16. This process can be repeated for different fragments 12 in the library, obtaining many beads, each having the product of multiple amplifications of a single library fragment 12. Referring Figure 10C , then the beads 16 can be loaded onto a reaction chamber array (e.g., a microwell array). Figure 10C A partial view of a DNA fragment undergoing a sequencing reaction in a reaction chamber is shown. A template strand 20 can pair with a growing complementary strand 22. In the left panel, an A nucleotide is added to the reaction chamber, generating a single-base incorporation event and producing a hydrogen ion. In the right panel, a T nucleotide is added to the reaction chamber, generating a double-base incorporation event and producing two hydrogen ions. The signal generated by the hydrogen ions is shown as a peak 26 in the ionization map. In various embodiments, a sequencing kit can contain one or more of the substances required for the above sample preparation and sequencing processes, including DNA fragmentation reagents, adapters, primers, ligases, beads or other solid-phase carriers, polymerases or precursor nucleotide monomers for the incorporation reaction.
[0149] According to an exemplary embodiment, a system is provided that consists of multiple distinguishable nucleic acid barcodes. The nucleic acid barcodes can be attached or associated with target nucleic acid fragments to form barcoded target fragments (such as polynucleotides). A library of barcoded target fragments can include multiple first barcodes attached to target fragments from a first source. A library of barcoded target fragments can also include different distinguishable barcodes attached to target fragments from different sources to establish a multiplex library. For example, a multiplex library can include a mixture of multiple first barcodes and multiple second barcodes, where the first barcodes are attached to target fragments from a first source and the second barcodes are attached to target fragments from a second source. In this multiplex library, the first and second barcodes can be used to identify the sources of the first and second target fragments, respectively. Any number of different barcodes can be attached to target fragments from any number of different sources. In a library of barcoded target fragments, the barcode portion can be used to identify: a single target fragment; a single source of a target fragment; a group of target fragments; target fragments from a single source; target fragments from different sources; target fragments from a user-defined group; or any other group that requires or benefits from identification. The barcode-labeled portion sequence of a barcoded target fragment can be sequenced separately from the target fragment or as part of a larger sequence read that encompasses the barcode and the target fragment. In a sequencing experiment, the nucleic acid barcodes can be sequenced using the target fragments, and then parsed using an algorithm during the processing of the sequencing data. In various embodiments, a nucleic acid barcode can include a synthetic or natural nucleic acid sequence, DNA, RNA, or other nucleic acids and / or derivatives. For example, a nucleic acid barcode can include the nucleotide bases adenine, guanine, cytosine, thymine, uracil, inosine, or analogs thereof. These barcodes can be used to identify a polynucleotide chain and / or distinguish it from other polynucleotide chains (such as a polynucleotide chain containing a different target sequence of interest), and can be used for various purposes, such as tracking, sorting, and / or identifying samples. Since different barcode sequences can be associated with different polynucleotide chains, these barcode sequences can be applicable to multiplex sequencing of different samples.
[0150] Multiplex library
[0151] In various embodiments, sample discrimination codes or barcodes (such as nucleic acid barcodes) are provided, which can be attached or associated with a target (such as a nucleic acid fragment) to generate a barcoded library (such as a barcoded nucleic acid library). One or more suitable nucleic acid or biomolecular manipulation procedures can be used to prepare such a library, and the procedures include: fragmentation; size selection; end repair; tailing; adapter ligation; nick translation; purification. In various embodiments, using one or more suitable procedures, including ligation, sticky-end hybridization, nick translation, primer extension or amplification, the nucleic acid barcode can be attached or associated with a fragment of a target nucleic acid sample. In several embodiments, using an amplification primer having a specific barcode sequence, the nucleic acid barcode can be attached to a target nucleic acid.
[0152] In various embodiments, a sample of a target nucleic acid or biomolecule (such as proteins, polysaccharides, nucleic acids and their polymeric subunits, etc.) can be isolated from any suitable source, examples of which include solid tissue, tissue, cells, yeast, bacteria or similar sources. Any suitable method can be used to isolate the sample from these sources. For example, the solid tissue or tissue can be weighed, cut, mashed, homogenized, and then the sample can be isolated from the homogenized sample. The isolated nucleic acid sample can be chromatin, which can be cross-linked with DNA-binding proteins in the ChIP (chromatin immunoprecipitation) procedure. In several embodiments, any suitable procedure can be used to fragment the sample, including enzymatic or chemical cleavage, or shearing. Enzymatic cleavage can include cleavage mediated by restriction endonucleases, endonucleases or transposases.
[0153] Fragment library
[0154] In various embodiments, a fragment library is provided, which can include: a first primer site (P1), a second primer site (P2), an insert, an internal adapter (IA) and a barcode (BC). In several embodiments, a fragment library can include a framework with a certain arrangement, such as: a P1 primer site, an insert, an internal adapter (IA), a barcode (BC) and a P2 primer site. In several embodiments, the fragment library can be attached to a solid-phase support, such as a bead.
[0155] Figure 11Describes an exemplary bead template according to a fragment library embodiment. It shows an exemplary nucleic acid attached to a solid support such as a bead. A bead template 700 includes a bead 710 with an adaptor sequence 720 that attaches the template 730 to the solid support. The template 730 may include a first or P1 primer site 740, an insert 750, and a second or P2 primer site 760. The template 730 may be a synthetic template. The template 730 may represent a fragment library. The template 730 may include a nucleic acid barcode BC that may be located between the P1 primer site 740 and the insert 750. An internal adaptor may be placed between the P1 primer site 740 and the barcode BC, or between the barcode BC and the insert 750, or between the insert 750 and the P2 primer site 760.
[0156] Figure 12 Describes another exemplary bead template according to a fragment library embodiment. The nucleic acid barcode BC may be located between the insert 750 and the P2 primer site 760. An internal adaptor may be placed between the P1 primer site 740 and the insert 750, or between the insert 750 and the barcode BC, or between the barcode BC and the P2 primer site 760.
[0157] In various embodiments, the lengths of the adaptor sequence 720 and the template 730 may vary. For example, the length of the adaptor sequence 720 may be between 10 and 100 bases, or 15 and 45 bases, and may be 18 bases (18b). The template 730, which consists of the P1 primer site 740, the insert 750, and the P2 primer site 760, may also have different lengths. For example, the respective lengths of the P1 primer site 740 and the P2 primer site 760 may be between 10 and 100 bases, or 15 and 45 bases, and may be 23 bases (23b). The length of the insert 750 may be between 2 bases (2b) and 20,000 bases (20kb), and may be 60 bases (60b). In one embodiment, the insert 750 may include more than 100 bases, such as 1,000 or more bases. In various embodiments, the insert may be in the form of a linker, in which case the insert 750 may consist of up to 100,000 bases (100 kb) or more bases.
[0158] In various embodiments, the position of barcode BC can be selected based on different considerations, such as the length of the insert, signal-to-noise issues, and / or sequencing bias issues. For example, when there are signal-to-noise problems (e.g., the signal-to-noise ratio can decrease during additional ligation cycles in the process of sequencing while ligating), barcode BC can be positioned near the P1 primer site 740 to mitigate potential errors caused by the decrease in signal-to-noise ratio. If the signal-to-noise ratio is not a significant issue, barcode BC can be placed near the P1 primer site 740 or the P2 primer site 760. In some cases, the interaction between the template sequence and the probe sequence used during the sequencing experiment can be different. Placing barcode BC in front of insert 750 can affect the sequencing results of insert 750. Placing barcode BC behind insert 750 can reduce sequencing errors caused by bias. In summary, the barcode position can be affected by sequencing or can affect sequencing, and the position that gives the best results based on the conditions of the sequencing process can be selected.
[0159] In various embodiments, a forward sequence read (e.g., along the 5’–3’ direction of the template) can be used to sequence and decode a nucleic acid barcode, such as reading barcode BC and insert 750 in a single read. In one embodiment, through an algorithm, the forward read can be parsed into a barcode portion and an insert portion.
[0160] In addition to the fragment library and the corresponding bead templates described herein, additional libraries and / or bead templates can be constructed using publicly available barcodes. For example, U.S. Patent Application No. 13 / 599,876, published as U.S. Patent Publication No. 2013 / 0053256 on February 28, 2015, with the title METHODS, SYSTEMS, AND KITS FOR SAMPLE IDENTIFICATION and the assignee Hubbell, which is hereby incorporated by reference in its entirety, also discloses Mate Pair libraries, Paired End libraries, SAGE TM libraries, yeast sequencing libraries, and ChIP-Seq libraries, which can be constructed using various publicly available embodiments.
[0161] In accordance with various embodiments, one or more functions of one or more of the materials and / or embodiments discussed above can be performed or implemented using appropriately configured and / or programmed hardware and / or software units. Determining whether a particular embodiment is implemented using hardware and / or software units can be based on any number of factors, such as the desired operating rate, power level, heat tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, etc., as well as other design constraints or performance constraints.
[0162] Examples of hardware units include processors, microprocessors, input and / or output (I / O) devices (or peripherals) interconnected through a local interface circuit, circuit elements (such as transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chip sets, etc. The local interface may include one or more buses or other wired or wireless connections, controllers, buffers (caches), drivers, repeaters, receivers, etc., to allow proper communication between hardware components. A processor is a hardware device for executing software, especially software stored in memory. The processor can be any custom or commercially available processor, a central processing unit (CPU), an auxiliary processor among several computer-related processors, a semiconductor-type microprocessor (such as a microchip or chip set), a macroprocessor, or any device commonly used to execute software instructions. The processor may also represent a distributed processing architecture. The I / O devices may include input devices, such as keyboards, mice, scanners, microphones, touchscreens, usage interfaces of medical devices and / or laboratory instruments, barcode readers, styli, laser readers, radio frequency device readers, etc. Moreover, the I / O devices may also include output devices, such as printers, barcode printers, displays, etc. Finally, the I / O devices may also include input / output communication devices, such as modulators / demodulators (modems; for accessing other devices, systems, or networks), radio frequency (RF) transceivers or other transceivers, telephone interfaces, bridges, routers, etc.
[0163] Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, program segments, software interfaces, application programming interfaces (APIs), instruction sets, operation codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Memory software may include one or more independent programs, which may include an ordered list of executable instructions for implementing logical functions. Memory software may include a system for identifying data streams according to existing data and any suitable custom or commercially available operating system (O / S), which may control the execution of other computer programs such as the system and provide scheduling, input / output control, file and data management, memory management, communication control, etc.
[0164] In accordance with various embodiments, one or more functions of any one or more of the materials and / or embodiments discussed above can be performed or implemented by a non-transitory machine-readable medium or article that is appropriately configured and / or programmed, which can store an instruction or set of instructions that, when executed by a machine, can cause the machine to perform a method and / or operation in an embodiment. The machine can include any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, scientific instrument, or laboratory instrument, etc., and can be implemented using any suitable combination of hardware and / or software. The machine-readable medium or article can include any suitable type of memory unit, memory device, memory item, memory medium, storage device, storage item, storage medium, and / or storage unit, such as memory, removable or non-removable medium, erasable or non-erasable medium, writable or rewritable medium, digital or analog medium, hard disk, floppy disk, compact disc read-only memory (CD-ROM), recordable compact disc (CD-R), rewritable compact disc (CD-RW), optical disc, magnetic medium, magneto-optical medium, removable memory card or disk, various types of digital versatile discs (DVD), magnetic tape, magnetic cassette, etc., including any medium suitable for a computer. The memory can include any volatile memory unit or any combination thereof (such as random access memory RAM, such as DRAM, SRAM, SDRAM, etc.) and non-volatile memory units (such as ROM, EPROM, EEROM, flash memory, hard disk drive, magnetic tape, CDROM, etc.). Moreover, the memory can be combined with electronic, magnetic, optical, and / or other types of storage media. The memory can have a distributed architecture where the various components are far apart but still accessible by the processor. The instructions can include any suitable type of code, such as source code, assembly code, interpreted code, executable code, static code, dynamic code, encrypted code, etc., implemented using any suitable high-level, low-level, object-oriented, intuitive, assembly, and / or interpreted programming language.
[0165] In accordance with various embodiments, one or more functions of any one or more of the materials and / or embodiments discussed above can be at least partially implemented using a distributed, clustered, remote, or cloud computing resource.
[0166] According to various embodiments, one or more functions of any one or more of the materials and / or embodiments discussed above can be implemented or carried out by a source program, an executable program (object code), a script, or any other entity consisting of an instruction set. The source program is translated by a compiler, assembler, interpreter, etc. (which may or may not be included in memory) to operate properly together with the O / S. Instructions can be written using the following tools: a) an object-oriented programming language with multiple classes of data and methods, or (b) a procedural programming language with routines, subroutines, and / or functions, which may include C, C++, Pascal, Basic, Fortran, Cobol, Perl, Java, and Ada.
[0167] According to various embodiments, one or more of the above-discussed embodiments may include transmitting, displaying, storing, printing, or outputting information related to any information, signal, data, and / or intermediate or final result generated, accessed, or used by the embodiments to a user interface device, a computer-readable storage medium, a local computer system, or a remote computer system. The information so transmitted, displayed, stored, printed, or output may take the form of a searchable and / or filterable running list, and reports, pictures, tables, data graphs, curves, spreadsheets, correlations, sequences, and combinations thereof.
[0168] Other embodiments can be derived by repeating, adding, or substituting any of the generally or specifically described functions and / or components and / or substances and / or steps and / or operating conditions set forth in one or more of the above embodiments. Moreover, it should be understood that as long as the goal of a step or action can still be achieved, the order or command of a certain step for performing a certain action is immaterial, unless otherwise specifically stated. Also, as long as the goal of a step or action can still be achieved, two or more steps or actions can be carried out simultaneously, unless otherwise specifically stated. Also, as long as the goal of any one of the other embodiments discussed above can still be achieved, any one or more functions, components, aspects, steps, or other features mentioned in one of the embodiments discussed above can be regarded as a potential alternative function, component, aspect, step, or other feature of any one of the other embodiments discussed above, unless otherwise specifically stated.
[0169] While embodiments of the present disclosure can be effectively used with synthesis - sequencing - by - synthesis methods, as described herein and in the following references: Rothberg et al., U.S. Patent Publication No. 2009 / 0026082; Anderson et al., Sensors and Actuators B Chem., Vol. 129, pp. 79 - 86, 2008; Pourmand et al., Proc. Nat1 Acad. Sci., Vol. 103, pp. 6466 - 6470, 2006, each of which is incorporated herein by reference in its entirety, the present disclosure can also be used in other ways, such as alternative synthesis - sequencing - by - synthesis methods, including methods of modifying nucleotide precursors or nucleotide triphosphate precursors to reversible terminators [sometimes called cyclic reversible termination (CRT) methods] and methods that do not modify nucleotide precursors or nucleotide triphosphate precursors [sometimes called cyclic single - base delivery (CSD)], or more general methods that include repeating steps of delivering nucleotides (to a polymerase - primer - template complex) and acquiring signals (either directly or indirectly detecting incorporation) (or extension in response to delivery).
[0170] While embodiments of the present disclosure can be effectively used with pH - based sequence detection, as described herein and in the following references: Rothberg et al., U.S. Patent Application Publication Nos. 2009 / 0127589 and 2009 / 0026082, and Rothberg et al., U.K. Patent Application Publication No. GB2461127, each of which is incorporated herein by reference in its entirety, the present disclosure can also be used with other detection methods, including detection of pyrophosphate ions (PPi) released from the incorporation reaction (see U.S. Patent Nos. 6,210,891, 6,258,568, and 6,828,100), various fluorescence sequencing instrument methods (see U.S. Patent Nos. 7,211,390, 7,244,559, and 7,264,929), several synthesis - sequencing - by - synthesis techniques [which can detect nucleotide - associated labels, such as mass tags, fluorescence, and / or chemiluminescence tags, in which case a passivation step can be incorporated into the workflow prior to the next synthesis - detection cycle (e.g., by chemical cleavage or photobleaching)], and more general methods in which an incorporation reaction produces or results in a product or component having a property that can be monitored and used to detect the incorporation event, including changes in magnitude (such as heat) or concentration (such as pyrophosphate ions and / or hydrogen ions) and signals (such as fluorescence, chemiluminescence, photogeneration), in which cases the amount of the detected product or component can be monotonically related to the number of incorporation events.
[0171] Although some embodiments are described in detail in this specification, other embodiments are possible and within the scope of the invention. For example, those skilled in the art may find it reassuring that this material can be implemented in various forms, such as using various sequencing instruments, and each embodiment can be implemented alone or in combination. For those skilled in the art, changes and modifications will be obvious in view of the description, figures, and patent practice set forth in the specification, the figures, and the claims.
[0172] Table 2 shows representative barcode sequences according to the various embodiments described herein
[0173] Seq.I.D. Barcode sequence Group (SEQ.ID.NO.20) TTCCGGAGGATGCC Plate_XX (SEQ.ID.NO.21) TTGAGGCCAAGTCC Plate_XX (SEQ.ID.NO.22) GACCACCGGTTC Plate_XX (SEQ.ID.NO.23) GTGGACCTCCGTTC Plate_XX (SEQ.ID.NO.24) TGGACCACGAATTC Plate_XX (SEQ.ID.NO.25) TTCTGGACATCCGC Plate_XX (SEQ.ID.NO.26) TTAGGCCTCCATTC Plate_XX (SEQ.ID.NO.27) GTTGAGGAACCACC Plate_XX (SEQ.ID.NO.28) CCGGACAAGAATTC Plate_XX (SEQ.ID.NO.29) CGGAGTTCCGGTTC Plate_XX (SEQ.ID.NO.30) GTCCACCAACCACC Plate_XX (SEQ.ID.NO.31) GTTCCAGCCATCTC Plate_XX (SEQ.ID.NO.32) GTTAGCGGATTC Plate_XX (SEQ.ID.NO.33) GCCACAACTTCC Plate_XX (SEQ.ID.NO.34) GTTCCTTAGAAGAC Plate_XX (SEQ.ID.NO.35) GCCAGCACCAATTC Plate_XX (SEQ.ID.NO.36) GCTTGGAGCCGTTC Plate_01 (SEQ.ID.NO.37) TCCAGGCACCTTCC Plate_01 (SEQ.ID.NO.38) GTTCCTACGTTC Plate_01 (SEQ.ID.NO.39) CCAGAACGGAATCC Plate_01 (SEQ.ID.NO.40) GTCAGGACCAAC Plate_02 (SEQ.ID.NO.41) CTTACCATCCTTCC Plate_02 (SEQ.ID.NO.42) GCTGACACCACC Plate_02 (SEQ.ID.NO.43) TCACCAACGGAC Plate_02 (SEQ.ID.NO.44) CTGAGAATCCAACC Plate_02 (SEQ.ID.NO.45) TTCCTACAATCTCC Plate_02 (SEQ.ID.NO.46) GTCTTGACAAGAAC Plate_02 (SEQ.ID.NO.47) GTTCTTAGAGAACC Flat plate_02 (SEQ.ID.NO.48) GTCCAGGAGGTC Flat plate_02 (SEQ.ID.NO.49) TCGGACCAATTGCC Flat plate_02 (SEQ.ID.NO.50) CCTTACCAATAACC Flat plate_03 (SEQ.ID.NO.51) TCGAGGCCATCGAC Flat plate_03 (SEQ.ID.NO.52) TTCCTTACCTTATC Flat plate_03 (SEQ.ID.NO.53) TTCTGAGCCGAC Flat plate_03 (SEQ.ID.NO.54) GTCCTACCAATGAC Flat plate_03 (SEQ.ID.NO.55) TAGCCAATTGAACC Flat plate_03 (SEQ.ID.NO.56) GCCTTAGCAACACC Flat plate_03 (SEQ.ID.NO.57) GTCCTGAGCAGAAC Flat plate_03 (SEQ.ID.NO.58) GTCTACCTCGGC Flat plate_03 (SEQ.ID.NO.59) GTCTGACCGGATCC Flat plate_03 (SEQ.ID.NO.60) CCAGAATTCGGACC Flat plate_04 (SEQ.ID.NO.61) TTCCGGAGTTCATC Flat plate_04 (SEQ.ID.NO.62) CCTTAGATCCTTCC Flat plate_04 (SEQ.ID.NO.63) GCCTTAGGATCGCC Flat plate_04 (SEQ.ID.NO.64) GCCAGGATTGGTCC Flat plate_04 (SEQ.ID.NO.65) GTCCGGAGATGAAC Flat plate_04 (SEQ.ID.NO.66) GCCTTATTCCAACC Flat plate_04 (SEQ.ID.NO.67) GTTCTAGGATTCAC Flat plate_04 (SEQ.ID.NO.68) TCCTAGTCCGGTCC Flat plate_04 (SEQ.ID.NO.69) GTCTTGGAGTTAAC Flat plate_04 (SEQ.ID.NO.70) GTTCTATCGTTC Flat plate_05 (SEQ.ID.NO.71) TTCGAGTGTTCC Flat plate_05 (SEQ.ID.NO.72) TCTTGATTGGTC Flat plate_05 (SEQ.ID.NO.73) GCTTACTCCGGTCC Flat plate_05 (SEQ.ID.NO.74) GATTCGGATTCC Flat plate_05 (SEQ.ID.NO.75) GTTCCTGAGTTCTC Flat plate_05 (SEQ.ID.NO.76) GTCGGACCATGAAC Flat plate_05 (SEQ.ID.NO.77) CAGATCCGTTCC Plate_05 (SEQ.ID.NO.78) GTTCTGACGTCC Plate_05 (SEQ.ID.NO.79) TCCGAGGATGAATC Plate_05 Sequence Listing <110> Life Technologies Corporation <120> Barcode Sequences and Related Systems and Methods <130> LT01064 PCT <140> <141> <150> 62 / 161,309 <151> 2015-05-14 <160> 79 <170> Patent Version 3.5 <210> 1 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / note="Artificial Sequence Description: Synthetic oligonucleotide" <400> 1 tcagtcctcg aatc 14 <210> 2 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / note="Artificial Sequence Description: Synthetic oligonucleotide" <400> 2 tcagcttgcg gatc 14 <210> 3 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 3 tcagtctaac ggac 14 <210> 4 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 4 tcagttctta gcgc 14 <210> 5 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 5 tcagtgagcg gaac 14 <210> 6 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 6 tcagttaagc ggtc 14 <210> 7 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 7 tcagctgacc gaac 14 <210> 8 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 8 tcagtctaga ggtc 14 <210> 9 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 9 tcagaagagg attc 14 <210> 10 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 10 tcctcgaatc 10 <210> 11 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 11 cttgcggatc 10 <210> 12 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 12 tctaacggac 10 <210> 13 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 13 ttcttagcgc 10 <210> 14 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 14 tgagcggaac 10 <210> 15 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 15 ttaagcggtc 10 <210> 16 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 16 ctgaccgaac 10 <210> 17 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 17 tctagaggtc 10 <210> 18 <211> 10 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 18 aagaggattc 10 <210> 19 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 19 tctgagcatc gatc 14 <210> 20 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 20 ttccggagga tgcc 14 <210> 21 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 21 ttgaggccaa gtcc 14 <210> 22 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 22 gaccaccggt tc 12 <210> 23 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 23 gtggacctcc gttc 14 <210> 24 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 24 tggaccacga attc 14 <210> 25 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 25 ttctggacat ccgc 14 <210> 26 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 26 ttaggcctcc attc 14 <210> 27 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 27 gttgaggaac cacc 14 <210> 28 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 28 ccggacaaga attc 14 <210> 29 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 29 cggagttccg gttc 14 <210> 30 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 30 gtccaccaac cacc 14 <210> 31 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 31 gttccagcca tctc 14 <210> 32 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 32 gttagcggat tc 12 <210> 33 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic [[ID=...]]oligonucleotide" <400> 33 gccacaactt cc 12 <210> 34 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 34 gttccttaga agac 14 <210> 35 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 35 gccagcacca attc 14 <210> 36 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 36 gcttggagcc gttc 14 <210> 37 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Explanation of artificial sequence: Synthetic oligonucleotide" <400> 37 tccaggcacc ttcc 14 <210> 38 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 38 gttcctacgt tc 12 <210> 39 <211> 14 <212> DNA <213> artificial sequence <220> <221> source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 39 ccagaacgga atcc 14 <210> 40 <211> 12 <212> DNA <213> artificial sequence <220> <221> source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 40 gtcaggacca ac 12 <210> 41 <211> 14 <212> DNA <213> artificial sequence <220> <221> source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 41 cttaccatcc ttcc 14 <210> 42 <211> 12 <212> DNA <213> artificial sequence <220> <221> source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 42 gctgacacca cc 12 <210> 43 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 43 tcaccaacgg ac 12 <210> 44 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 44 ctgagaatcc aacc 14 <210> 45 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 45 ttcctacaat ctcc 14 <210> 46 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 46 gtcttgacaa gaac 14 <210> 47 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 47 gttcttagag aacc 14 <210> 48 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 48 gtccaggagg tc 12 <210> 49 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 49 tcggaccaat tgcc 14 <210> 50 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 50 ccttaccaat aacc 14 <210> 51 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 51 tcgaggccat cgac 14 <210> 52 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 52 ttccttacct tatc 14 <210> 53 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 53 ttctgagccg ac 12 <210> 54 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 54 gtcctaccaa tgac 14 <210> 55 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 55 tagccaattg aacc 14 <210> 56 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 56 gccttagcaa cacc 14 <210> 57 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 57 gtcctgagca gaac 14 <210> 58 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 58 gtctacctcg gc 12 <210> 59 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Artificial sequence description: Synthetic oligonucleotide" <400> 59 gtctgaccgg atcc 14 <210> 60 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 60 ccagaattcg gacc 14 <210> 61 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 61 ttccggagtt catc 14 <210> 62 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 62 ccttagatcc ttcc 14 <210> 63 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 63 gccttaggat cgcc 14 <210> 64 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 64 gccaggattg gtcc 14 <210> 65 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 65 gtccggagat gaac 14 <210> 66 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 66 gccttattcc aacc 14 <210> 67 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 67 gttctaggat tcac 14 <210> 68 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 68 tcctagtccg gtcc 14 <210> 69 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 69 gtcttggagt taac 14 <210> 70 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 70 gttctatcgt tc 12 <210> 71 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 71 ttcgagtgtt cc 12 <210> 72 <211> 12 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 72 tcttgattgg tc 12 <210> 73 <211> 14 <212> DNA <213> Artificial sequence <220> <221> Source <223> / Note = "Artificial sequence description: synthetic oligonucleotide" <400> 73 gcttactccg gtcc 14 <210> 74 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 74 gattcggatt cc 12 <210> 75 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 75 gttcctgagt tctc 14 <210> 76 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 76 gtcggaccat gaac 14 <210> 77 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: synthetic oligonucleotide" <400> 77 cagatccgtt cc 12 <210> 78 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 78 gttctgacgt cc 12 <210> 79 <211> 14 <212> DNA <213> Artificial Sequence <220> <221> Source <223> / Note = "Description of artificial sequence: Synthetic oligonucleotide" <400> 79 tccgaggatg aatc 14
Claims
1. A kit for use with a nucleic acid sequencing instrument, the kit comprising: a plurality of barcoded polynucleotides, the plurality of barcoded polynucleotides comprising at least 16 different barcoded polynucleotides, each barcoded polynucleotide of the plurality of barcoded polynucleotides comprising a unique barcoded nucleic acid sequence of a plurality of barcoded nucleic acid sequences, each barcoded polynucleotide having a length within a predetermined length range of at least 5 nucleotide bases in length, wherein each barcoded nucleic acid sequence comprises: a codeword nucleic acid sequence corresponding to a flowspace codeword of a set of flowspace codewords such that: each flowspace codeword of the set comprises a different string and a padding character, wherein the padding character is not at the end of the flowspace codeword and the corresponding padding base is not the end base of one of the barcoded nucleic acid sequences of the plurality of barcoded nucleic acid sequences, wherein the padding character is configured for synchronization of the barcoded polynucleotide within the flowspace, and wherein the padding character configures the flowspace codeword to correspond to a valid base space sequence in a predetermined flow order; each flowspace codeword of the set satisfies a minimum distance; and each flowspace codeword of the set is configured to appear in the flowspace in a predetermined flow order; and a key nucleic acid sequence appended to the codeword nucleic acid sequence.
2. The kit according to claim 1, wherein the plurality of barcoded polynucleotides comprises at least 96 different barcoded polynucleotides.
3. The kit according to claim 1, wherein the plurality of barcoded polynucleotides comprises at least 500 different barcoded polynucleotides.
4. The kit according to claim 1, wherein the set of flowspace codewords comprises at least 500 different flowspace codewords.
5. The kit according to claim 1, wherein the length of the predetermined length range is 5 - 40 nucleic acid bases.
6. The kit according to claim 1, wherein the codeword nucleic acid sequence has a length of 9 - 14 nucleic acid bases.
7. The kit according to claim 1, wherein the length of the key nucleic acid sequence is at least 4 nucleotide bases.
8. The kit according to claim 1, wherein the key nucleic acid sequence is placed 5' to the codeword nucleic acid sequence in the barcoded nucleic acid sequence.
9. The kit according to claim 1, wherein each barcoded nucleic acid sequence of the plurality of barcoded nucleic acid sequences comprises a common 3' nucleotide base for configuring synchronization of each barcoded polynucleotide within the flowspace.
10. The kit according to claim 1, wherein the set of flowspace codewords does not comprise a flowspace codeword not configured to be valid in the base space.
11. The kit according to claim 1, wherein: for a first group of the set of flowspace codewords, the key nucleic acid sequence appended to each codeword nucleic acid sequence terminates with a repeating base; and for a second group of the set of flowspace codewords, the key nucleic acid sequence appended to each codeword nucleic acid sequence terminates with a non-repeating base.
12. The kit according to claim 1, wherein the set of flowspace codewords as a whole defines an error-tolerant code.
13. The kit according to claim 12, wherein the minimum distance is at least 3.
14. The kit according to claim 12, wherein the error-tolerant code is a ternary code using an alphabet of three-character units.
15. The kit according to claim 12, wherein the error-tolerant code is capable of correcting at least one stream space error.
16. The kit according to claim 12, wherein the error-tolerant code: provides a minimum distance between the stream space codewords of the set, and a change in the key nucleic acid sequence increases the minimum distance between the stream space codewords of the set.
Citation Information
Patent Citations
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090127589A1
Methods and apparatus for measuring analytes
US20100137143A1
Methods and apparatus for detecting molecular interactions using FET arrays
US20100282617A1
Predictive Model for Use in Sequencing-by-Synthesis
US20120109598A1