Systems and methods for decoding by sequencing
The method enhances DNA sequencing by using coded recognition elements and rolling circle amplification for sensitive and specific target detection, addressing the limitations of existing technologies.
Patent Information
- Application Number
- PCT/US2024/060535
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
Existing DNA sequencing technologies face challenges in achieving high sensitivity and specificity for target detection, often resulting in low signal levels and difficulties in identifying the presence or absence of targets.
The method involves subjecting targets to a recognition event with coded recognition elements, followed by ligation and rolling circle amplification to generate immobilized amplification products. These products are then extended with labeled and unlabeled nucleotides, allowing for iterative detection of sequence incorporation and subsequent decoding of the target-specific code.
This approach enables sensitive and specific detection of targets by amplifying and decoding the recognition element codes, improving signal output and accuracy in identifying the presence of targets.
Smart Images

Figure US2024060535_26062025_PF_FP_ABST
Abstract
Description
[0001] SYSTEMS AND METHODS FOR DECODING BY SEQUENCING
[0002] The present application claims the benefit of United States Provisional Patent Application Serial No. 63 / 611,696, filed December 18, 2023, and United States Provisional Patent Application Serial No. 63 / 611,686, filed December 18, 2023, both of which are incorporated herein by reference in their entireties.
[0003] BACKGROUND
[0004] Rolling circle amplification (RCA) and sequencing by synthesis (SBS) are two distinct techniques that can be used together in DNA sequencing processes. RCA is a DNA amplification method that generates multiple copies of DNA, forming single -stranded DNA amplification products (e.g., concatemers) including multiple tandem repeats of a template sequence. SBS determines the DNA sequence by synthesizing a complementary DNA strand to the amplification products one base at a time, using detectable nucleotides or nucleotide analogs.
[0005] SUMMARY
[0006] In various embodiments disclosed herein, a method is provided for conducting an assay for determining the presence of one or more targets from a set of targets, the method comprising: subjecting the set of targets to a recognition event, wherein each target of the set of targets is uniquely recognized by and bound to at least one recognition element from a set of recognition elements, wherein each recognition element comprises a target-specific binding site and a code, wherein each code is associated with the target of the set of targets bound to at least one recognition element, thereby generating a set of bound target recognition elements, subjecting the set of target bound recognition elements to a ligation reaction to yield modified recognition elements, wherein the modified recognition elements comprise ligated circularized recognition elements that are immobilized onto a solid surface, performing rolling circle amplification on the modified and immobilized recognition elements thereby generating amplified recognition elements, introducing the amplified recognition elements to a detectable nucleotide or nucleotide analog from a set of labeled nucleotides or nucleotide analogs under conditions for incorporating the labeled nucleotide or nucleotide analog into an extension reaction using an amplified recognition element of the amplified recognition elements as a template, introducing the amplified modified recognition elements to an unlabeled nucleotide or nucleotide analog from a set of unlabeled nucleotides or nucleotide analogs under conditions for incorporating the unlabeled nucleotide or nucleotide analog into the extension reaction using the amplified recognition element as the template, detecting a signal from the incorporation of the labeled nucleotide or nucleotide analog into the extension reaction by imaging the solid surface; and iteratively performing (d), (e), and (f) , thereby determining the sequence of the code of the amplified modified recognition elements and the presence of one or more targets from the set of targets.
[0007] In some embodiments, the method comprises introducing the labeled and unlabeled nucleotides or nucleotide analogs substantially simultaneously. In some embodiments, the method comprises introducing the labeled and unlabeled nucleotides or nucleotide analogs sequentially. In some embodiments, the labeled and unlabeled nucleotides or nucleotide analogs comprise a reversible terminator moiety. In some embodiments, the reversible terminator moiety comprises a 3’-O-alkyl hydroxylamino group, a 3 ’ -phosphorothioate group, a 3’-O-malonyl group, a 3 ’-O-benzyl group, a 3’-O-Azidomethyl group, a 3 ’-O-Azide group, a 3 ’-O-Azido group, or a 3 ’-O-methyl group. In some embodiments, the reversible terminator moiety prevents incorporation of an additional labeled nucleotide or nucleotide analog or an additional unlabeled nucleotide or nucleotide analog at an N+l position, wherein N is the position of the labeled nucleotide or nucleotide analog or the unlabeled nucleotide or nucleotide analog in an extension reaction.
[0008] In some embodiments, the method further comprises subjecting the amplified modified recognition elements to a deblocking reaction sufficient to remove the reversible terminator moiety from the labeled nucleotide or nucleotide analog and remove the reversible terminator moiety from the unlabeled nucleotide or nucleotide analog. In some embodiments, the method further comprises subjecting the amplified modified recognition elements to a deblocking reaction sufficient to remove the reversible terminator from the detectable nucleotide or nucleotide analog and remove the reversible terminator from the unlabeled nucleotide or nucleotide analog; and removing the label from the labeled nucleotide or nucleotide analog.
[0009] In some embodiments, the method further comprises subjecting the amplified modified recognition elements to a deblocking reaction sufficient to remove the reversible terminator from the labeled nucleotide or nucleotide analog and remove the reversible terminator moiety from the unlabeled nucleotide or nucleotide analog, removing the label from the labeled nucleotide or nucleotide moiety; and iteratively performing the steps of the method for a plurality of cycles.
[0010] In some embodiments, the deblocking reaction comprises a buffer having one or more of phosphine compound comprising Tris(2 -carb oxy ethyljphosphine, bis-sulfo triphenyl phosphine or Tri(hydroxyproyl)phosphine, tetrakis(triphenylphosphine)palladium(0) (Pd(P(C6H5)3)4), (CH2)SNH, or2,3-Dichloro-5,6-dicyano-l,4-benzo-quinone, palladium on carbon, a thiol group comprising beta -mercaptoethanol or dithiothritol, potassium carbonate in MeOH, triethylamine in pyridine, or with Zn in acetic acid, or tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0011] In some embodiments, detecting the incorporation occurs prior to removal of the reversible terminator and removal of the label. In some embodiments, detecting the incorporation occurs simultaneously with removal of the reversible terminator and removal of the label. In some embodiments, detecting the incorporation occurs simultaneously with removal of the label and prior to removal of the reversible terminator. In some embodiments, detecting the incorporation occurs simultaneously with removal of the reversible terminator and prior to removal of the label.
[0012] In some embodiments, the code of the recognition elements comprises a soft decodable code. In some embodiments, decoding the codes of the amplified modified recognition elements comprises recording a signal produced in response to interrogation of each segment of the codes, and upon completion of the interrogation, determining a probability of the presence of each of the codes by applying a soft-decision probabilistic decoding algorithm to the recorded signal, wherein the presence of the code is indicative of the presence of the target. In some embodiments, decoding the codes comprises decoding the codes by a soft decision decoding method. In some embodiments, a percentage of the amplified modified recognition elements that comprise a plurality of codes that are successfully decoded by the soft decoding algorithm is greater than 10% of the amplified modified recognition elements. In some embodiments, a percentage of the amplified modified recognition elements that comprise a plurality of codes that are successfully decoded by the soft decoding algorithm is greater than 50% of the amplified modified recognition elements.
[0013] In some embodiments, the label comprises a detectable optical label. In some embodiments, the label comprises a detectable fluorescent label. In some embodiments, the set of labeled nucleotides or nucleotide analogs comprises a plurality of species, wherein the fluorescent label of each of the plurality of species emits a different wavelength than other fluorescent labels of the other species of the plurality. In some embodiments, the plurality of species comprises adenine (A), cytosine (C), Guanine (G), or Thymine (T). In some embodiments, the set of labeled nucleotides or nucleotide analogs comprises A, C, G and T, and wherein the detectable label for each of A, C, G and T is different. In some embodiments, the set of labeled nucleotides or nucleotide analogs comprises two or more species of A, C, G and T, wherein the detectable label for each of the two or more species is different. In some embodiments, the set of labeled nucleotides or nucleotide analogs comprises one species comprising A, C, G or T. In some embodiments, the labeled nucleotides or nucleotide analogs of the set of labeled nucleotides or nucleotide analogs comprises the same detectable label. In some embodiments, the set of unlabeled nucleotides or nucleotide analogs comprises a plurality of species, wherein the plurality of species comprises A, C, G or T.
[0014] In some embodiments, the set of targets comprises nucleic acid targets. In some embodiments, the nucleic acid targets comprise DNA, bisulfite converted DNA, RNA or methylated nucleic acid targets. In some embodiments, the set of targets comprises amino acid targets or polypeptide targets.
[0015] In some embodiments, the recognition elements are padlock probes or molecular inversion probes. In some embodiments, the recognition elements further comprise one or more target recognition element sequences, one or more sequencing primer binding site sequences, one or more amplification primer binding site sequences, a unique molecular identifier sequence, a sample index sequence, a restriction enzyme site sequence, or any combination thereof. In some embodiments, the one or more amplification primer binding site sequences comprise a universal primer sequence that is common to all the recognition elements.
[0016] In some embodiments, the codes in the set of recognition elements are the same length. In some embodiments, at least a sub set of the set of recognition elements has codes of the same length. In some embodiments, each code from the set of codes comprises a nucleic acid sequence having a length that is 25 or fewer nucleotides, 20 or fewer nucleotides, 15 or fewer nucleotides, 10 or fewer nucleotides, or 5 or fewer nucleotides.
[0017] In some embodiments, the rolling circle amplification is performed on a solid surface, wherein the solid surface does not comprise a covalent attachment molecule. In some embodiments, the solid surface is a charged surface, for example a cation-coating layer bound to the solid surface. In some embodiments, the cation-coating layer comprises poly-L-lysine.
[0018] In some embodiments, the rolling circle amplification generates a concatemer comprising multiple copies of the modified recognition element.
[0019] In some embodiments, the assays as described herein are conducted in vitro.
[0020] The present disclosure further provides a method for identifying the presence of a target nucleic acid from a sample, comprising providing a nucleic acid sample and a plurality of recognition elements, wherein a recognition element of the plurality of recognition elements comprises a code, a 5’ recognition region, a 3’ recognition region, and one or more sequencing primer binding sites, hybridizing the 5’ recognition region and the 3 ’ recognition region to complementary sequences in the target nucleic acid from the sample, ligating and circularizing the hybridized recognition element thereby generating circularized recognition elements, immobilizing the circularized recognition elements on a sequencing substrate, determining the code of the circularized recognition element by sequencing and using the determined code sequence to identify the presence of the target nucleic acid from the sample. In some embodiments, the target nucleic acid is DNA, RNA or cDNA. In some embodiments, the plurality of recognition elements comprises a subset comprising recognition regions complementary to a wild type sequence in the target nucleic acid from the sample. In some embodiments, the plurality of recognition elements comprises a subset comprising recognition regions complementary to a variant sequence in the target nucleic acid from the sample. In some embodiments, the recognition element of the plurality of recognition elements further comprises one or more of a universal primer binding site sequence, a unique molecule identifier sequence, a restriction endonuclease recognition sequence, and a cleavage sequence. In some embodiments, the sequencing is next generation sequencing. In some embodiments, the next generation sequencing is sequence by synthesis. In some embodiments, the variant sequence is a single nucleotide polymorphism, an indel, or a copy number variant. In some embodiments, the 5’ recognition region and the 3’ recognition region of the recognition element hybridize to adjacent sequences of the target nucleic acid. In some embodiments, the 5’ recognition region and the 3 ’ recognition region of the recognition element hybridize to non-adjacent sequences of the target nucleic acid.
[0021] The present disclosure further provides computer-implemented systems comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device to create an application comprising a software module configured to perform a soft decision decoding of a plurality of the signals detected in a method as described herein.
[0022] The present disclosure further provides systems comprising a computer processor, wherein the processor is programmed to execute a soft decision decoding of a plurality of the signals detected in a method as described herein.
[0023] The present disclosure further provides systems for conducting an assay for a set of targets or target analytes, comprising a reaction vessel, a reagent dispensing module, and software to execute a method comprising a soft decision decoding of a plurality of the signal detected in a method as described herein, wherein the method is executed robotically. The present disclosure further provides kits for conducting an assay for a set of targets, the kit comprising a set of recognition elements, wherein each recognition element of the set comprises a 5’ target- specific binding site and a 3 ’ target specific binding site and a code associated with a target from the set of targets, a plurality of the labeled nucleotides or nucleotide analogs, a plurality of the unlabeled nucleotide or nucleotide analogs, and instructions for use of the kit in an assay, wherein the instructions comprise steps for performing a method as described herein or using a system as described herein. In some embodiments, the labeled nucleotide or nucleotide analog of a kit or the unlabeled nucleotide or nucleotide analog of a kit comprises a reversible terminator moiety at a 3 ’ OH position. In some embodiments, the reversible terminator moiety comprises a 3 ’-O-alkyl hydroxylamino group, a 3 ’- phosphorothioate group, a 3 ’-O-malonyl group, a 3 ’-O-benzyl group, a 3 ’-O-Azidomethyl group, a 3’-O-Azide group, a 3’-O-Azido group, or a 3’-O-methyl group. In some embodiments, the plurality of labeled nucleotides or nucleotide analogs of a kit comprises a plurality of species, wherein a fluorescent label of each of the plurality of species emits a different wavelength than other fluorescent labels of the other species of the plurality, wherein the plurality of species comprises adenine (A), cytosine (C), Guanine (G), or Thymine (T). In some embodiments, the plurality of labeled nucleotides or nucleotide analogs comprises A, C, G and T, and wherein the detectable label for each of A, C, G and T is different. In some embodiments, the plurality of labeled nucleotides or nucleotide analogs comprises two or more species of A, C, G and T, wherein the detectable label for each of the two or more species is different. In some embodiments, the plurality of labeled nucleotides or nucleotide analogs comprises one species comprising A, C, G or T. In some embodiments, the plurality of unlabeled nucleotides or nucleotide analogs comprises a plurality of species, wherein the plurality of species comprises A, C, G or T.
[0024] In some embodiments, a kit further comprises one or more of a well plate, a 96-well plate, a glass-bottomed plate, a poly-L-lysine-coated plate, a polymerase, a DNA-polymerase, a Bst DNA polymerase, a Bst-like DNA polymerase, a Therminator X polymerase, a Bst3.0 polymerase, a ArcticZymes Bst polymerase, a Bsm DNA polymerase, and a cleaving reagent. In some embodiments, the cleaving reagent cleaves the detectable label from a subset of the plurality of labeled nucleotides or nucleotide analogs and / or the cleaving reagent cleaves the reversible terminator from a subset of the plurality of labeled nucleotides or nucleotide analogs or from a subset of the plurality of unlabeled nucleotides or nucleotide analogs. In some embodiments, the cleaving reagent subjects the modified recognition elements to a deblocking reaction sufficient to remove the reversible terminator from the labeled nucleotide or nucleotide analog and remove the reversible terminator from the unlabeled nucleotide or nucleotide analog. In some embodiments, the cleaving reagent subjects the modified recognition elements to a deblocking condition sufficient to remove the detectable label from the labeled nucleotide or nucleotide moiety. In some embodiments, the cleaving reagent comprises a buffer having one or more of a phosphine compound comprising Tris(2 -carboxy ethyljphosphine (TCEP), bis-sulfo triphenyl phosphine (BS-TPP) or Tri(hydroxyproyl)phosphine (THPP), tetrakis(triphenylphosphine)palladium(0) (Pd(P(C6H5)3)4), (CEEjsNH (piperidine), 2,3- Dichloro-5,6-dicyano-l,4-benzo-quinone (DDQ), Palladium on carbon (Pd / C), a thiol group comprising beta-mercaptoethanol or dithiothritol (DTT), potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or with Zn in acetic acid (AcOH), or tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0025] In some embodiments, a kit further comprises a scan reagent used during the detection of the nucleotides. In some embodiments, the scan reagent improves detection of a binding event between a labeled nucleotide or nucleotide analog and an extension product using an amplified modified recognition element as the template. In some embodiments, a scan reagent comprises buffer and radical scavengers. In some embodiments, a kit further comprises a solution comprising a plurality of detectable labeled nucleotides or nucleotide analogs, wherein the label is an optical label such as a fluorescent label that is detectable by imaging such as fluorescent imaging. In some embodiments, a kit comprises a solution comprising the plurality of unlabeled nucleotides or nucleotide analogs.
[0026] In some embodiments, a kit comprises one or more wash solutions, wherein the one or more wash solutions comprises Tris(hydroxymethyl)aminomethane, 2-Amino-2- (hydroxymethyl)-l,3-propanediol (TRIS), ethylenediaminetetraacetic acid (edetic acid; EDTA), and Polyoxyethylene(20)sorbitan monolaurate (Tween 20). In some embodiments, the wash solution comprises sodium chloride at a concentration greater than 0.9 M. In some embodiments, the wash solution comprises sodium chloride at a concentration of about 1 M. In some embodiments, the wash solution comprises sodium chloride at a concentration of about 10 mM to about 30 mM. In some embodiments, the wash solution comprises sodium chloride at a concentration of about 20 mM. In some embodiments, the wash solution further comprises ethylenediaminetetraacetic acid (EDTA). In some embodiments, the wash solution comprises sodium chloride at a concentration of about 1 M in combination with TRIS, EDTA, and Tween 20 (TET). INCORPORATION BY REFERENCE
[0027] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
[0028] BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present disclosure provides inventive concepts set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments and the accompanying drawings.
[0030] FIG. 1 provides a schematic diagram of an example of a coded recognition element according to some embodiments herein.
[0031] FIG. 2A- FIG. 2B shows a schematic diagram illustrating an example of capturing an amplification product on a cation-coated surface according to some embodiments herein.
[0032] FIG. 3A provides a schematic diagram of a transformation process for circularizing a linear recognition element to form a circular modified recognition element according to some embodiments herein, and FIG. 3B provides a schematic diagram showing RCA amplification of a circular modified recognition element to yield an amplification product according to some embodiments herein.
[0033] FIG. 4 provides a schematic diagram of an example of a process for capturing an unknown region of a target for sequencing according to some embodiments herein.
[0034] FIG. 5 provides a schematic diagram illustrating an example for generating a sequencing library from an amplification product that may be used to identify the sequence of the codes associated with the target of interest according to some embodiments herein.
[0035] FIG. 6 provides a schematic diagram illustrating an example of a process for directly sequencing an amplification product set to identify the codes associated with the target set of interest according to some embodiments herein.
[0036] FIG. 7 provides a schematic diagram illustrating a process for determining a code from an amplification product according to some embodiments here.
[0037] FIG. 8 provides a schematic diagram illustrating a process for determining the sequence of a code from a recognition element according to some embodiments here. FIG. 9 provides a schematic diagram illustrating bridge amplification on a circular recognition element that is immobilized on a sequencing substrate as shown in FIG. 8.
[0038] FIG. 10 provides a flow diagram of an example of an assay workflow according to some embodiments herein.
[0039] FIG. 11 provides a flow diagram of an example of an amplification product library preparation, sequencing, and imaging workflow according to some embodiments herein.
[0040] FIG. 12A provides an image of amplification products at the end of a sequencing flow, and FIG. 12B provides an additional image of amplification products at the end of a sequencing flow.
[0041] FIG. 13A provides images of amplification products in four different color channels over eight sequencing cycles, and FIG. 13B provides images of amplification products in four different color channels over eight sequencing cycles.
[0042] FIG. 14 shows a plot of the intensity measured in each of five color channels in the first imaging cycle for a plurality of amplification product codes, including ACACGTCTCA (SEQ ID NO: 1), AGATCGATCGCA (SEQ ID NO: 2), ATCACGACGACA (SEQ ID NO: 3), CATCGATCAG (SEQ ID NO: 4), CGACATCGTA (SEQ ID NO: 5), CTGAGACA, GAGACGAG, GATCGCTA, GTCGATAGATA (SE ID NO: 6), TATCACGATA (SEQ ID NO: 7), TCGAGCGTCA (SEQ ID NO: 8), and TGACGATCATCA (SEQ ID NO: 9).
[0043] FIG. 15 shows a plot of the signal intensity readout in three color channels in each of eight cycles for a plurality of amplification products all comprising the same code.
[0044] FIG. 16 provides a plot showing attainable codespace sizes for different numbers of cycles, number of codes, and hamming distances between codes.
[0045] FIG. 17 provides the number of counts of decoded sequences assigned to each of a plurality of targets (top) and the percentage of amplification products whose code sequences were successfully decoded (decode rate) in a decoding experiment (bottom).
[0046] FIG. 18 is an exemplary plot showing the number of counts of decoded sequences assigned to each of a plurality of wildtype (WT) and variant targets, the top row shows counts from a sample comprising WT and variant recognition elements (REs) and targets (Ts), the bottom shows counts from a sample with only WT REs and Ts.
[0047] FIG. 19 is an exemplary plot demonstrating the specificity of amplification and detection of wildtype (WT) targets in amplification reactions containing both WT and variant recognition elements in comparison to amplification reactions containing only WT recognition elements. FIG. 20 provides additional exemplary plots demonstrating the specificity of amplification and detection of WT targets in amplification reactions containing both WT and variant recognition elements.
[0048] FIG. 21 provides a plot showing the similarity of detection rates of wildtype (WT) targets in amplification reactions that comprises WT recognition elements and only WT targets and in amplification reactions comprising WT recognition elements and both variant targets and WT targets.
[0049] FIG. 22 provides a plot showing the similarity of detection rates of wildtype (WT) targets in amplification reactions that comprise WT targets and only WT recognition elements and in amplification reactions comprising WT targets and both variant recognition elements and WT recognition elements.
[0050] FIG. 23 shows the similarity of read counts in duplicate amplification reactions that contained WT and variant targets and WT and variant recognition elements .
[0051] FIG. 24 provides a plot showing the signal intensity readout when the expected state of the code is either ON or OFF in reactions using three different polymerase enzymes (Bst3.0; Bsm; AZ Bst, from left to right) in each of seven cycles in a 2-state amplification reaction.
[0052] FIG. 25 shows a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface.
[0053] FIG. 26 shows a non-limiting example of a web / mobile application provision system; in this case, a system providing browser-based and / or native mobile user interfaces.
[0054] FIG. 27 shows a non-limiting example of a cloud-based web / mobile application provision system; in this case, a system comprising an elastically load balanced, auto -scaling web server and application server resources as well synchronously replicated databases.
[0055] FIG. 28 provides a schematic diagram illustrating an example of a process of using a bisulfite conversion reaction in combination with a coded recognition element to detect a methylated target site of interest.
[0056] FIG. 29A- FIG. 29C provide a schematic diagram illustrating an example of a process for detecting a target sequence using a third oligonucleotide recognition element to produce a modified recognition element comprising the code.
[0057] FIG. 30A- FIG. 30C provide a schematic diagram illustrating an example of a process for detecting a target sequence using a circular third oligonucleotide recognition element to produce a modified recognition element comprising the code. FIG. 31 A - FIG. 31C provides a schematic diagram illustrating an example of a process for detecting a target of interest using a pre-circularized encoded modified recognition element and a PCR amplification / 5' endonuclease cleavage reaction.
[0058] DETAILED DESCRIPTION
[0059] Assays exist for identifying whether targets, such as target nucleic acids, are present in a sample. Such assays require high levels of sensitivity and specificity; however, such assays are sometimes limited by their levels of detection, resulting in low signal levels and difficulties in identifying the presence or absence of the target. As such, there is a need for assays that are sensitive, specific, and able to identify the presence or absence of a target from a sample without ambiguity. The present disclosure provides a solution which is sensitive, specific and provides robust signal output for target identification.
[0060] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0061] The term “about” when referring to a number refers to that number plus or minus 10% of that number. The term “about” in reference to a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.
[0062] The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” “identifying” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context. “Identify,” “determine” and the like with respect to codes, targets or analytes disclosed herein are intended to include any or all of : (A) an indication of the presence or absence of the relevant code, target or analyte, (B) an indication of the probability of the presence or absence of the relevant code, target or analyte, and / or (C) quantification of the relevant code, target or analyte.
[0063] The terms “decoding” or “detecting” with respect to a code includes determining the presence of a known code or the probability of the presence of a known code. Decoding and detecting the code can be done in a number of ways, for example by determining the codes using next generation sequencing chemistries to determine the sequence of the code in a base by base fashion, or by using detection oligonucleotide complexes that comprise a detectable moiety such as a fluorescent moiety. The data generated can be decoded by applying hard or soft decision detecting algorithms which is then used to identify the presence of the code and, by proxy, the presence or absence of the target molecule from the sample. As such, “sequencing” as used herein refers to the determination of the code sequence from a recognition element by using sequencing chemistries, such as SBS chemistries.
[0064] The terms "hard decision detecting" or “hard decision” refers to a method or model that includes making a call for each nucleotide in a nucleic acid segment (commonly referred to as a “base call”) in order to determine the sequence of nucleotides in the nucleic acid segment. Models of the inventive concepts incorporate hard decision detecting models. The particular nucleic acid being detected may be or include a code of the inventive concepts described herein.
[0065] The terms “soft decision detecting" or “soft detection” refers to a method or a model that uses data collected during a sequencing or detecting process to calculate a probability that a particular nucleic acid or nucleic acid segment is present. The probability may optionally be calculated without making a base call for each nucleotide in a nucleic acid segment. In another example, a probability is calculated without making a hard call that a string of nucleic acids in a segment of a code is present. Instead of making a hard call for each nucleotide or nucleotide segment, a probabilistic detecting algorithm is applied to the recorded signal upon completion of signal collection. A probability of the presence of each of the codes may be determined without discarding signal in contrast to hard decision detecting method in which hard calls are made during the signal collection process. In soft decision detecting, the data may, for example, include or be calculated from, intensity readings in spectral bands for signals produced by the SBS chemistry. In one embodiment, soft decision detecting uses data collected during a sequencing / detecting process to calculate a probability that a particular nucleic acid segment from a known set of sequences is present.
[0066] The terms “coded” and “encoded” are, for the purposes of this disclosure, considered equivalent.
[0067] The term “target” can be any molecule from a sample, for example a protein, peptide, nucleic acid, etc. In preferred embodiments, that target molecule is a nucleic acid, either DNA or RNA. A target can be a nucleic acid analyte (e.g., mRNA, cfDNA, etc.) or a proxy for the target analyte of interest (e.g., an antibody conjugated with oligonucleotide). Thus, in some instances, the term “target” and the term “target analyte” are used interchangeably. “Target” with respect to a nucleic acid includes wild-type and mutated nucleic acid sequences, including for example, point mutations (e.g., substitutions, insertions and deletions), chromosomal mutations (e.g., inversions, deletions, duplications), and copy number variations (e.g., gene amplifications). “Target” with respect to a nucleic acid may also include the presence or absence of one or more methyl groups on the nucleic acid target. “Target” with respect to a polypeptide includes wildtype and mutated polypeptides of any length, including proteins and peptides.
[0068] The term “sample” comprising a target molecule can be from any species, human or nonhuman primate, mammals, aves, bacteria, viral, etc. For example, a “sample” can be a set of nucleic acids fortesting, a sample preparation produced may be used to provide a sequencing ready sample from a raw sample or partially processed sample. One or more samples may be combined for sample preparation and / or sequencing and may be distinguished post-sequencing using sample specific codes of the recognition elements. The present disclosure is not limited by the source of the sample. Examples of samples include biological samples, such as whole blood, lymphatic fluid, serum, plasma, sweat, tear, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, fluidized tissues, fluidized organisms, liquids containing multi-celled organisms, biological swabs and biological washes. Samples may be from any organism (e.g., prokaryotes, eukaryotes, plants, animals, humans) or other sample (e.g., environmental or forensic samples).
[0069] The terms “bind”, “binding”, “bound” and the like refer to covalent and non -covalent interactions, for example, “bind” can include any degree of interaction including hybridization between two nucleic acid sequences. Alternatively, “bind” can include covalent binding, where the sharing of electrons between atoms occurs. A covalent bond can be reversible or irreversible, depending on the need. Hydrogen bonding, van der Waals interactions and other weak interactions between molecules are also considered “bonds” for the present disclosure.
[0070] The terms “subject,” “individual,” or “patient” are often used interchangeably herein. A “subject” can be a biological entity containing expressed genetic materials. The biological entity can be a plant, animal, or microorganism, including, for example, bacteria, viruses, fungi, and protozoa. The subject can be tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro. The subject can be a mammal. The mammal can be a human. The subject may be diagnosed or suspected of being at high risk for a disease. In some cases, the subject is not necessarily diagnosed or suspected of being at high risk for the disease.
[0071] The term “in vivo” is used to describe an event that takes place in a subject’s body. The term “zw vitro” is used to describe an event that takes place contained in a container for holding laboratory reagents such that it is separated from the biological source from which the material originated. In vitro assays can encompass cell-based assays in which living or dead cells are employed. In vitro assays can also encompass a cell-free assay in which no intact cells are employed.
[0072] The term “linked” with respect to two nucleic acids means not only a fusion of a first moiety to a second moiety at the C-terminus orthe N-terminus, but also includes insertion of the first moiety to the second moiety into a common nucleic acid. Thus, for example, the nucleic acid A may be linked directly to nucleic acid B such that A is adjacent to B (-A-B-), but nucleic acid A may be linked indirectly to nucleic acid B, by intervening nucleotide or nucleotide sequence C between A and B (e.g., -A-C-B- or -B-C-A-). The term “linked” is intended to encompass these various possibilities.
[0073] The term “set” includes sets of one or more elements or objects. A “subset” of a set includes any number elements or objects from the set, from one up to all of the elements of the set.
[0074] The terms “phasing” or “signal phasing” means misalignment of SBS cycles during an SBS process caused by the non-incorporation of a nucleotide during a cycle or by the incorporation of two or more nucleotides during an SBS cycle.
[0075] The term “crosstalk” refers to the situation in which a signal from one nucleotide addition reaction may be picked up by multiple channels (referred to as “color crosstalk”) or the situation in which a signal from an amplification product or sequencing cluster interferes with an adjacent or nearby cluster or amplification product (referred to as “cluster crosstalk” or “nanoball crosstalk”).
[0076] The term “color channel” means a set of optical elements for sensing and recording an electromagnetic signal from a sequencing reaction. Examples of optical elements include lenses, filters, mirrors, and cameras.
[0077] The terms “spectral band” or “spectral region” means a continuous wavelength range in the electromagnetic spectrum.
[0078] Provided herein are systems, methods, and kits related to multiplexed target detection from a sample. The methods and systems disclosed herein utilize encoded recognition elements (e.g., probes) that undergo a molecular transformation in the presence of a given target (but not in the absence of the given target) such that the code of the encoded recognition element may be detected using one or more nucleic acid sequencing techniques. The code may be a known sequence that is detected, and thus serves as a surrogate or proxy for detecting the target. In some embodiments, the encoded recognition elements that have undergone the molecular transformation (referred to herein as “modified recognition elements”) are amplified using rolling circle amplification (RCA) to produce an amplification product. In some embodiments, the amplification products comprise concatemers that can immobilize to a surface for subsequent detection by nucleic acid sequencing. In some embodiments, the concatemer is a nucleic acid molecule comprising multiple tandem repeats of the modified recognition element including the code.
[0079] Methods disclosed herein comprise numerous formats of the encoded assays disclosed herein, such as, for example the encoded assays provided in United States Patent Application Nos. 18 / 253,803, 18 / 150,669, and 18 / 150,661 ; and International Application Nos. PCT / US2022 / 037778, PCT / US2022 / 037781, PCT / US2022 / 037785, each of which are incorporated by reference herein in their entirety. In some embodiments, such formats encompass a decoding step in which the presence of a code, or a portion of a code, or a probability of the presence of a code is determined. The present disclosure provides methods of determining the sequence of the code using nucleic acid sequencing. In some embodiments, the identity of nucleotides in the code, or a portion of the code, are determined using nucleic acid sequencing. Non-limiting examples of nucleic acid sequencing methods of the present disclosure include those provided in Slatko BE, Gardner AF, Ausubel FM. Overview of Next-Generation Sequencing Technologies. Curr Protoc Mol Biol. 2018 Apr;122(l):e59, which is hereby incorporated by reference in its entirety. For example, the sequencing method may be sequencing by hybridization, sequencing-by-synthesis (SBS), SMRT (Singe Molecule Real Time) sequencing, nanopore-based DNA sequencing, or sequencing-by-binding (SBB), such as Avidity™ sequencing, or other sequencing workflows.
[0080] In some embodiments, decoding comprises performing a sequencing-by-synthesis method that involves detection of the stepwise incorporation of detectable complementary nucleotides into a primed template strand. In some embodiments, the primed template strand comprises the code or a segment or portion within the code. In some embodiments, the nucleotides or nucleotide analogs that are incorporated are detectable, which enable detection of incorporation. In some embodiments, a detectable nucleotide or nucleotide analog is labeled. In some embodiments, the nucleotides or nucleotide analogs are labeled with a fluorescent label. In some cases, the incorporation of sequential nucleotides or nucleotide analogs into the primed template is monitored in real-time. In such sequencing-by-synthesis methods, the nucleotide or nucleotide analogs are detectable but unblocked, meaning they are incorporated into the growing primed template upon binding to a complementary nucleotide in the template at the N+l position. In some embodiments, the detectable nucleotides or nucleotide analogs have a blocking group (also referred to here as “reversible terminators”) that pause the synthesis process during detection. Subsequent to, or substantially simultaneously with, detection, the blocking group and / or the fluorescent label maybe cleaved off of the nucleotide or nucleotide analog, thereby enabling incorporation of the next complementary nucleotide or nucleotide analog into the growing primed template strand. While sequencing, a template strand which has failed to incorporate a nucleotide or nucleotide analog in a sequencing cycle will continue to lag behind, which is referred to as “phasing,” which can cause errors in sequencing base call accuracy. To avoid phasing, the sequencing-by-synthesis method may also utilize unlabeled nucleotides or nucleotide moieties that incorporate into the primed template at places where the template strand fails to incorporate a detectable or previously detectable nucleotide.
[0081] Systems disclosed herein may include computer systems with one or more processors capable of implementing instructions for decoding the codes of the amplification products. In some embodiments, the instructions comprise soft decision decoding. In some embodiments, the instructions comprise hard decision decoding. The systems disclosed herein may also include an instrument for performing the encoded assays, a sequencer, or a combination thereof. In some embodiments, the instrument for performing the encoded assays may include one or more reaction vessels, a reagent dispensing module, a fluidic control system, an imaging system, or any combination thereof. The system may also include assay components of the encoded assays disclosed here, such as the recognition elements, sequencing primers, amplification primers, or additional oligonucleotide probes (e.g., synthetic splint oligonucleotides, synthetic bridge elements, etc.). In some embodiments, the systems also include reagents for conducting the encoded assay, the nucleic acid sequence reactions, or any combination thereof.
[0082] Also provided are kits comprising one or more components of the system or encoded assays disclosed herein. In some embodiments, the kits comprise instructions for performing an encoded assay disclosed herein. In some embodiments, the kits comprise instructions for performing the decoding step (e.g., nucleic acid sequencing and code detection) disclosed herein. In some embodiments, the kits comprise instructions for utilizing the readout of the encoded assay and decoding steps to predict a presence of a plurality of targets from a sample in a multiplexed fashion. Such multiplexed encoded assays are capable of detecting thousands of targets in a single reaction.
[0083] Disclosed herein, in some embodiments, are methods of detecting one or more target molecules from a sample. In some embodiments, the method is a multiplexed method involving the detection of a plurality of target molecules from a sample involving a multiplex nucleic acid sequence reaction. In some embodiments, the methods comprise introducing the target molecules from a sample to a plurality of recognition elements having target recognition elements complementary to flanking regions of a target nucleic acid sequence, under conditions that the recognition element binds to the target nucleic acid molecule. In some embodiments, the recognition element or the target nucleic acid molecule comprise a code that is associated with the target molecule. In some embodiments, the recognition elements bound to the target nucleic acid molecule undergoes a molecular transformation (and recognition elements that are not bound to a target nucleic acid molecule are not expected to undergo the molecular transformation). In some embodiments, the transformation is a ligation reaction in the presence of a ligating enzyme to produce a circular modified recognition element. In some embodiments, the methods comprise amplifying the modified recognition elements to produce a plurality of amplification products associated with multiple copies of the code that is associated with the target nucleic acid molecule sequence.
[0084] In some embodiments, a sequencing library is produced from the amplification products as further detailed herein (e.g., fragmentation, ligation to a UMI, an index, one or more adapters, etc.). In some embodiments, the sequencing library is sequenced to determine or predict by soft decision decoding the presence of the code, thereby predicting the presence of the target in the sample. In some embodiments, the presence of the target is predicted with an accuracy of at least about 99%. Nucleic acid sequencing may be carried out on one or more devices, or by one or more instruments or systems, of the present disclosure.
[0085] The disclosure provides encoded assays that make use of recognition elements comprising codes that are amplified and detected as a surrogate for detecting the presence of the target. In some embodiments, the code in a recognition element may be a soft decodable code (e.g., a trellis code). A recognition element may include target-specific regions that may be used for target recognition and binding (e.g., a target recognition region). A recognition element may include a 5 ' terminal phosphate and a 3 ’ terminal hydroxyl that may be used to facilitate ligation (e.g., circularization) after target recognition andbinding. In some embodiments, the recognition element is a linear oligonucleotide probe that has a 5’ target recognition region probe arm and a 3 ’ target recognition region probe arm. In some embodiments, the 3’ probe arm comprises a nucleotide that is the complement to a nucleotide at a target site of interest, such as a single nucleotide polymorphism (SNP), indel or copy number variant. In some embodiments, the 5’ probe arm comprises the nucleotide that is the complement to a nucleotide at a target sequence of interest. A recognition element may configure as a padlock probe. A recognition element may configure as a molecular inversion probe. In some embodiments, the recognition element may include regions at the 3 ' probe arm and 5 ' probe arm that are complementary to adjacent regions of a target nucleic acid molecule. In some embodiments, the recognition element may include regions at the 3’ probe arm and the 5’ probe arm that are complementary to non -adjacent regions of a target nucleic acid molecule. When the recognition element regions hybridize to the target nucleic acid molecule, the recognition element may be circularized by a ligation reaction (e.g., for adjacently hybridized probe arms) or gap-fill ligation reaction (e.g., for non-adjacently hybridized probe arms) in the presence of one or more of a ligase and reagents necessary for performing a gap fill extension / ligation reaction. As described elsewhere in this disclosure, the target nucleic acid molecule may be a nucleic acid analyte (e.g., mRNA, cfDNA, DNA, etc.) or a proxy for the target analyte of interest if the analyte of interest is a peptide or protein or other non-nucleic acid target (e.g., an oligonucleotide conjugated to an antibody specific to a protein target, or a portion thereof).
[0086] The recognition elements of the present disclosure are selectively amplified in the presence of the target where a ligation of the recognition element has occurred (but not in the absence of the target). In some embodiments, the amplification is an isothermal amplification reaction. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification (TMA), Loop -mediated isothermal amplification (LAMP), Helicase-dependent amplification (HD A), Nucleic Acid Sequence Based Amplification (NASBA), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Fakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci. 2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety.
[0087] In some embodiments, the amplification is RCA. A recognition element may include an amplification primer binding site. In some embodiments, the amplification primer binding site is for a universal amplification primer. In some embodiments, the amplification primer binding site is suitable for any amplification method disclosed herein (e.g., RCA). When RCA is employed in sequencing-by-synthesis, the high-density amplification of RCA enhances the signal detection during the sequencing process, ultimately improving the accuracy and sensitivity of the results. This combined approach is valuable in various scientific applications that require robust DNA amplification and accurate sequencing. In some embodiments, the recognition element comprises a code (referred to herein as “coded or encoded recognition element”). In other embodiments, the recognition element does not have a code. In such embodiments, the recognition element may associate with, or bind to, another oligonucleotide molecule that comprises the code.
[0088] FIG. 1 provides a schematic diagram of an example of a coded or encoded recognition element 100. Coded recognition element 100 may include a 5 ' target specific region 110a and a 3' target specific region 110b that are complementary to regions of a target nucleic acid sequence (not shown). The target may be a nucleic acid analyte (e.g., mRNA, cfDNA, DNA, etc.) or a proxy for the target analyte of interest (e.g., an antibody conjugated with an oligonucleotide that serves as a proxy for the target analyte of interest). Target specific region 110a may include a 5' terminal phosphate to facilitate ligation and circularization after target recognition and hybridization. Target specific region 110b may include one or more terminal 3' nucleotides (N) complementary to a nucleotide at a target site of interest. In this example, the 3 ' target site specific nucleotide “N” may be a SNP specific nucleotide.
[0089] Target specific regions 110a and 110b may hybridize to the target, and the recognition element may be circularized. For example, when the complementary nucleotide is present in the target, the 3’ SNP specific nucleotide hybridizes to the target, enabling circularization, e.g., by ligation or gap-fill ligation. Other types of features or mutations may be detected by varying the terminal nucleotide (N) or nucleotides of target specific region 110a and / or target specific regions 110b to hybridize when the target feature is present and not hybridize when the target feature is not present. In some embodiments, the 3’ nucleotide and the target nucleotide are not variants of interest.
[0090] Coded recognition element 100 may include an RCA priming site 115 for priming an RCA reaction. In this example, RCA priming site 115 is downstream from target specific region 115b. However, other locations are possible, as long as the positioning of the primer site doesn’t interfere with the other functions of the recognition element, e.g., the recognition element hybridization function and the encoding function.
[0091] A coded recognition element may optionally include other functional sequences 120. For example, the recognition element may include index sequences which are unique oligo identifiers present in the recognition element sequence or inserted as part of the assay. Index sequences, such as sample barcodes, allow differentiation among different samples, experiments, etc. during the decoding event (e.g., reading (decoding) the code).
[0092] The coded recognition element may include one or more unique molecular identifiers (UMIs). UMIs may be inserted anywhere within the recognition element to address downstream readout and data analysis. For example, UMIs may be introduced to distinguish unique recognition events with single-molecule resolution during the readout. UMI’s may facilitate error correction and / or individual molecule counting.
[0093] A coded recognition element may include other primer binding sites in addition to the priming region required for RCA amplification. Other primer binding regions may, for example, be present to facilitate the readout of an index, a UMI, a restriction endonuclease site, a cleavage site or other oligonucleotide sequences present in the recognition element which may be of use for downstream applications. Primer binding regions may allow parallel or serial reading schemes. They may also be used to increase the amount of multiplexing or allow sequential readout. For instance, if a plurality of recognition elements or amplified objects are present, only those containing a specific primer binding site will be amplified or read. Primer binding sites may also be used to facilitate the capture and immobilization of a recognition element or amplified object onto a surface (e.g., via DNA-DNA hybridization).
[0094] A coded recognition element may include one or more sequences recognizable by enzymes, such as endonucleases. Various sequences may be selected and used to facilitate additional transformations, such as digestion, nick or gap formation, phosphorylation etc . In one embodiment, the recognition element includes one or more restriction sites.
[0095] A coded recognition element may include one or more non-natural nucleic acid components. Examples include phosphorothioate groups, locked DNA (LNA), peptide DNA (PNA) and others, which may be included to improve certain features of the recognition element, such as melting temperature for target recognition, or primer recognition, or resistance to degradation. Additionally, abasic nucleotides (“wobble bases”) may be included in the recognition element sequence to add degeneracy to targeting or priming regions and extend the ability to recognize a broader number of complementary sequences.
[0096] A coded recognition element may include one or more chemical moieties. Such chemical moieties may be included in the recognition element structure or added at any stage of the workflow to enable additional transformations or properties. Examples include cleavable groups to open or linearize the recognition element, reactive groups to add additional components such as dyes, and groups to facilitate immobilization on surfaces.
[0097] A coded recognition element may include CRISPR recognition sequences, oligo sequences designed to be recognized by CRISPR enzymes and replaced with other arbitrary sequences. The recognition element may optionally include one or more oligo sequences designed to be recognized by transposases and replaced with other arbitrary sequences. A coded recognition element may optionally include one or more adapter primers for compatibility with sequencing-by-synthesis. A coded recognition element may optionally include one or more adapter primers for compatibility with non-SBS platforms. The adapter primers may be included in the recognition element sequence or added at any stage as part of the workflow. Such adapter primers may be used directly to immobilize, cluster, extend, and amplify as precursor activities to a decoding run by sequencing-by-synthesis or another sequencing method.
[0098] In one embodiment, a recognition element assay workflow may include:
[0099] (i) hybridizing the recognition element to a target;
[0100] (ii) optionally, extending the hybridized recognition element to fill any singlestranded gap remaining between the two recognition element arms;
[0101] (iii) circularizing the recognition element when the target is present;
[0102] (iv) removing (e.g., by exonuclease or other mean) non -circularized recognition elements remaining after ligation (and other linear nucleic acid molecules);
[0103] (v) amplifying the circularized recognition element by RCA or other method;
[0104] (vi) capturing of the amplified product on a surface;
[0105] (vii) preparing the library for sequencing, using sequencing sample preparation workflows suitable for a desired sequencing platform; and
[0106] (viii) reading out or decoding the sequence of the code.
[0107] The recognition elements or amplification products thereof may comprise one or more index sequences. An index sequence may be added to the amplification products during tagmentation to produce a nucleic acid sequence library. The term, “tagmentation” as used herein may refer to the initial step in library prep where unfragmented DNA is cleaved and tagged for analysis via a transpositional system comprising a transposase and transposon complex (i.e., transposome). Index sequences, such as sample barcodes, that can identify the sample source from which the target nucleic acids are derived may allow differentiation among different samples, batches, spatial locations, or experiments, during the detection event. Indexes may be added to a recognition element using a variety of strategies. Indexes may be added during the synthesis of a recognition element. In this case, for every recognition element manufactured, the number of recognition elements is N x P, where N is the number of indices and P is the plexity of the recognition element pool.
[0108] Indexes may be added after recognition element synthesis as part of manufacturing or at a site of use as a step prior to performing an encoded assay . In this case, only one synthesis is required for each recognition element and additional functional elements. Additional functional elements may be added to a recognition element to enable insertion of an index. Examples of functional elements that may be added include (i) non-natural nucleotides (e.g., biotin, amine, etc.) and (ii) polynucleotides that enable biochemical transformation of the recognition element to contain an index sequence such as adapters for ligations or extension ligations, restriction endonuclease recognition sites, and transposome binding sites.
[0109] Indexes may be added during an encoded assay. For example, a ligation reaction to insert an index can occur at the same time as ligation of the recognition element at the target site of interest to generate a circularized recognition element (e.g., the transformation event). In some cases, the ligation reaction may be a gap-fill extension / ligation reaction.
[0110] Indexes may be added after ligation of the recognition element and amplification (RCA) by including modified nucleotides during the amplification reaction. The modified nucleotides may be bound to an index sequence. In cases where there is a covalent or non-covalent interaction, either moiety can be linked to the index sequence or incorporated during amplification.
[0111] Examples of binding strategies include: (i) ligand protein pairs such as biotinstreptavidin, antigen-antibody, CLIP tag and SNAP tag pair (e.g., O6-benzylguanine derivatives binding to O6-alkylguanine-DNA-alkyltransferase, wherein either the protein or the substrate may be bound to the recognition element), carbohydrate-protein pairs (e.g., lectins), and digoxigenin-DIG-binding protein; (ii) peptide-protein pairs (e.g., SpyTag - SpyCatcher); and (iii) hybridizing indexes to a common sequence on the RCA product.
[0112] Indexes may be added to amplification products by restriction endonuclease cleavage followed by index ligation. Indexes may be added to amplification products using a transposome with transposon sequences modified to include index sequences that fragments and indexes the amplification products. In some embodiments, the amplification products are amplification products of an RCA reaction. Indexes may be added to the fragmented amplification products resulting from tagmentation of the amplification products when producing a nucleic acid sequencing library. The methods and systems disclosed herein may include amplification of a nucleic acid molecule. In some embodiments, the nucleic acid molecule comprises the recognition element. In some embodiments, the nucleic acid molecule comprises the target nucleic acid molecule or a fragment thereof. In some embodiments, the amplification is selective. For example, the recognition element may be amplified only in the presence of the target (but not in the absence of the target) such that the recognition elements binds to the target and the bound recognition element is circularized by ligation. In another example, the target nucleic acid, or a fragment thereof may be selectively amplified using target-specific primers. In some embodiments, the amplification is not selective. For example, when immobilized primed templates in sequencing reaction are clonally amplified to form amplified colonies for solid -phase sequence detection.
[0113] The amplification methods disclosed herein may include polymerase chain reaction (PCR). In some embodiments, the PCR is multiplexed PCR. The amplification methods disclosed herein may include isothermal amplification. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification (TMA), Loop-mediated isothermal amplification (LAMP), Helicasedependent amplification (HD A), Nucleic Acid Sequence Based Amplification (NASB A), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Fakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci. 2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety.
[0114] The amplification methods may be performed in solution. The amplification methods may be performed on a surface. Surface-based amplification may be performed with surface- anchored primers (e.g., Illumina® bridge amplification technology) or recombinase polymerase amplification (RPA) (e.g., ExAmp technology). Clonally amplified material may result in an amplification product or a DNA cluster (e.g., Illumina® surface-based amplification).
[0115] In some embodiments, the methods comprise adding a second surface adapter to a recognition element. The second surface adapter may be complementary to a second primer on a flow cell surface (e.g., a bridge amplification primer). The second surface adapter may, for example, be added to a recognition element during the ligation or gap -fill ligation event or added separately by PCR or through its own ligation to a recognition element. For example, an amplification strategy may include using a splint ligation approach to add a second surface adapter to a surface bound recognition element to facilitate bridge amplification. Bridge amplification may be used to create clusters of amplification products for sequencing.
[0116] In some embodiments, the methods comprise adding a restriction enzyme site to a recognition element sequence. For example, the recognition element may include a restriction enzyme sequence site that when hybridized with a complementary oligonucleotide provides a double-stranded site for a restriction endonuclease to cleave the recognition element, rendering a linear recognition element. The linear recognition element may be amplified for downstream processing, e.g., for sequencing. For example, the linear recognition element may be captured on a flow cell and amplified by bridge amplification (e.g., Illumina® bridge amplification technology) or recombinase polymerase amplification (RPA) (e.g., Ex Amp technology).
[0117] The recognition element may include surface primer binding sequences or surface adapter sequences that are complementary to surface bound primers on a flow cell. The adapter sequences may be linked to or adjacent to the restriction site, so that when the site is cut by a restriction enzyme the linear recognition element is ready for sequencing. As noted, other forms of cleavage are possible, such as CRISPR mediated cleavage or any other double -stranded break inducing protein.
[0118] Similarly, an amplification product may include surface primers or sequencing adapters linked to or adjacent to a restriction site, so that when the site is cut by a restriction enzyme the linear recognition elements are released and potentially ready for sequencing. As noted, other forms of cleavage are possible, such as CRISPR mediated cleavage.
[0119] In another embodiment, an amplification product with adapter sequences complementary to surface bound primers may be seeded directly onto the surface without cleaving.
[0120] Amplification may proceed through bridge amplification (e.g., Illumina® bridge amplification technology) or recombinase polymerase amplification (RPA) (e.g., ExAmp technology) initiated directly.
[0121] Rolling circle amplification (RCA) may be used to produce amplification products as part of the assays disclosed herein. An RCA reaction may be performed as a surface-bound reaction. For example, RCA may be initiated by an oligonucleotide bound to a surface (e.g., beads, flow cells, microwells, or nanowells). Any method may be used to bind the oligonucleotide to the surface, either directly or indirectly. In one example, the oligonucleotide may be covalently bound to the surface. An oligonucleotide may be covalently attached to a surface. An oligonucleotide may include an RCA primer sequence that is complementary to an RCA primer binding site on a recognition element. An oligonucleotide may be used to capture a recognition element by hybridization of the complementary sequences and initiate the RCA reaction. When the oligonucleotide is covalently bound to the surface, the surface-bound RCA reaction can generate an amplification product that is also covalently attached to the surface.
[0122] In another example, a cation-coated surface (e.g., beads, flow cells, microwells, or nanowells) may be used to capture amplification products indirectly on a flowcell or other substrate. In one example, the cation-coated surface may be a polylysine-coated surface. In one example, the cation-coated surface may be a poly-L-lysine-coated surface. Other options for cation coated surface treatments may be equally amenable; the current example is not limited to any specific coating and the use of poly-L-lysine is exemplary only. FIGs. 2A-B illustrate schematic diagrams illustrating an example of capturing an amplification product on a cation- coated surface. A surface 215 may be coated with a cation 230. By way of example, a surface 215 may be coated with a poly-L-lysine coating 230. An RCA reaction maybe performed in the presence of the coated surface, resultingin simultaneous immobilization and amplification of an amplification product 235. RCA primers 232 may be supplied in solution (A) or bound to the cation-coated surface prior to performing the RCA reaction (B).
[0123] In another example, a streptavidin-coated surface (e.g., beads, flow cells, microwells, or nanowells) may be used to capture amplification products. In this approach, biotin-linked deoxynucleotides may be incorporated into the amplification products during RCA. A surface may be coated with a streptavidin coating. The amplification products may be bound to the surface by a biotin-streptavidin linkage. An RCA reaction may be performed in the presence of the streptavidin coated surface using biotin-linked deoxynucleotides to produce an amplification product that includes biotin moieties resulting in simultaneous immobilization and amplification of amplification product.
[0124] In another embodiment, biotin linked RCA primers may be bound to a surface by a streptavidin - biotin linkage and used to initiate an RCA reaction as described above. A surface may be coated with a streptavidin coating. An oligonucleotide that includes a biotin moiety may be attached to surface through a biotin-streptavidin linkage. The oligonucleotide may include an RCA primer sequence that is complementary to an RCA primer binding site on a recognition element. The oligonucleotide may be used to capture a recognition element by hybridization of the complementary sequences and initiation of the RCA reaction to produce an amplification product. Amplification in the presence of the streptavidin coated surface further anchors the amplification product to the surface.
[0125] Following the formation of an amplification product, a determination may be made with respect to the identity of the code. Prior to making the determination, various secondary processing steps are possible within the scope of the assays described herein. The recognition element may include various elements that facilitate secondary processing steps. Examples include restriction endonuclease sites and CRISPR sites.
[0126] The amplification product may be converted to double-stranded DNA (dsDNA) prior to fragmentation. The dsDNA amplification product may be fragmented. In one embodiment, the recognition element includes restriction sites which are replicated in the amplification product, and the amplification product is fragmented using a restriction enzyme having specificity for the restriction sites.
[0127] The fragmentation may be performed by physical shearing or enzymatic means. Enzymatic means may be performed by an endonuclease or a mixture of different types of endonuclease enzymes. The enzyme may be a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) system. The enzyme may be a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) system. The enzyme may be a CRISPR-associated (CAS) protein. In some embodiments, CRISPR may be used to fragment the amplification product at specific sites. Random fragmentation of amplification products may be performed using physical shearing methods, such as sonication, nebulization, or acoustic shearing. The fragmented DNA may be end-repaired and A-tailed. Fragmentation may be tagmentation, in which the DNA is simultaneously fragmented and tagged for analysis. Tagmentation may be performed on the amplification product. The tagmentation may be used to add sequencing adapters.
[0128] Methods and system disclosed herein may provide methods and systems for preparing a sequencing library for nucleic acid sequencing. In some embodiments, the nucleic acid sequencing is multiplexed nucleic acid sequencing. Sequencing library preparation may involve shearing the recognition element or the amplification products thereof, or size selection and clean-up of the DNA fragments, or a combination thereof. In some embodiments, in which the recognition element does not already contain one or more sequencing-specific elements (e.g., UMI, index, adaptor(s)), the sequencing-specific elements may be added to the DNA fragments.
[0129] Size selection and clean-up may be performed to enrich for DNA fragments of a defined length. Magnetic beads, columns or gels may be used for size selection and clean-up. Suitable techniques for DNA fragment size selection and clean up include use of solid phase reversible immobilization beads, such as those described in DeAngelis et al., Solid-phase reversible immobilization for isolation of PCR products. Nucleic Acids Res., 23(22)(1995), pp.4742 -4743, and Xu et al., Solid-phase reversible immobilization in microfluidic chips for the purification of dye-labeled DNA sequencing fragments. Anal. Chem., 75(13)(2003), pp. 2975-2984, each of which is incorporated by reference in its entirety. Spin-column clean-up kits can also be used to perform DNA fragment size selection and clean-up. Alternatively, size selection may be performedby gel electrophoresis. Instruments useful for size selection for this purpose include the BluePippin™ System (Sage Science, 2019). The DNA fragments that are selected may be about 50 base pairs to about 50,000 bp in length. The fragments that are selected may be about 50 base pairs to about 10,000 bp in length. The fragments that are selected may be about 50 base pairs to about 1,000 bp in length. The fragments that are selected may be about 50 bp to about 200 bp in length. The fragments that are selected may be about 60 bp to about 150 bp in length. The fragments that are selected may be about 70 bp to about 100 bp in length. The fragments that are selected may be about 80 bp to about 100 bp in length.
[0130] In some embodiments, the recognition elements or amplification products may be fragmented with restriction endonucleases (RE) to yield a multitude of code -containing single stranded nucleic acids. The single-stranded nucleic acids may be prepared for sequencing by ligation to adapter sequences. Amplification products may be fragmented enzymatically with a CRISPR or a transposome methodology.
[0131] In some embodiments, the sequencing library is produced from modified (e.g., circularized) recognition elements. In some embodiments, the sequencing library is produced from amplification products of an amplification reaction involving the modified (e.g., circularized) recognition element. In certain embodiments, amplification of the modified (e.g., circularized) recognition elements and preparation of the amplification products for sequencing may be performed in a single reaction (e.g., adapter addition via PCR).
[0132] Sequencing adapters may be added to the modified recognition element prior to amplification or the amplification product thereof. Alternatively, sequencing adaptors may be added to DNA fragments produced from shearing the modified recognition element or the amplification product thereof. A sequencing adaptor may comprise a surface-binding element that binds to an immobilized capture oligonucleotide on a nucleic acid sequencing flow cell. The surface-binding element may be a nucleic acid sequence complementary to a sequence of the capture oligonucleotide. Two different adaptors may be added, one adaptor that binds to the 5’ end of a flow cell capture oligonucleotide, and a second adaptor that binds to the 3’ end of a flow cell capture oligonucleotide. Sequencing adaptors may comprise a sequencing primer binding site that allows the tagged DNA fragment immobilized to a solid surface (e.g., flow cell surface) to be primed and recognized by DNA polymerase to initiate a DNA synthesis reaction. The sequencing primer binding sites may be added for paired-end sequencing, or single-end sequencing. The sequencing adaptors may have an index comprising a barcode. A barcode may permit differentiation among different samples, batches, spatial locations, or experiments, in a multiplexed nucleic acid sequence reaction. The sequencing adaptors may have a unique molecular identifier (UMI). The sequencing adaptor may be ligated to A-tailed library fragments. The sequencing adaptor may be a Y-adaptor.
[0133] Sequencing adapters may be added in a polymerase chain reaction (PCR). In this case, amplification and preparation for sequencing may be a single step . Depending on the recognition element design, the code, UMI, and index may be read in a single step or in two separate reads with a dehybridization step. For example, sequencing adapters may be added by transposomes that simultaneously fragment double-stranded DNA and add adapters.
[0134] As discussed elsewhere in the application, the assays disclosed herein include a transformation step. The transformation may involve circularization of a recognition element when a target is present (e.g., by ligation or gap-fill ligation).
[0135] FIG. 3 A is a schematic diagram of a transformation process 300 for circularizing a linear recognition element to form a circular modified recognition element according to some embodiments herein. In this example, a linear recognition element 310 includes a UMI sequence 312, a code 314, an SBS primer binding site 316, and an index primer sequence 318 all located between a 5' target recognition region 320a and a 3 ' target recognition region 320b. In the presence of a target (not shown), recognition element 310 is hybridized and circularized in a ligation reaction to yield a circular modified recognition element 325. The ligation reaction may be followed by an exonuclease digestion step to remove liner recognition elements 310 and linear targets.
[0136] The circular modified recognition element 325 may, in some cases, be amplified in a rolling circle amplification reaction to form a concatemeric amplification product. FIG. 3B is a schematic diagram showing RCA amplification of the circular modified recognition element to yield a concatemeric amplification product 330. For example, in an RCA reaction an SBS primer 316b that is complementary to the SBS primer binding site sequence 316 may be hybridized to the circular modified recognition element 325 and used to initiate the RCA reaction to generate a concatemeric amplification product 330. Concatemeric amplification product 330 is a polymeric molecule (e.g., concatemer) that includes multiple repeated copies of the circular modified recognition element 325, wherein each copy includes a SBS primer binding site 316, a code 314, a UMI sequence 312, target recognition regions 320a and 320b, and index primer 318 and other sequences that might be present in the modified recognition element. In this example, the complement (e.g., copy) of modified recognition element 325 is indicated by the dashed line. The encoded assays disclosed herein are capable of multiplex target detection. The readout of the encoded assays can be measured alongside the readout of various molecular assays that may be performed in parallel, thereby enabling a multiomic platform for the analysis of different target molecules in a sample.
[0137] Examples of target molecules include, but are not limited to, proteins, nucleic acids (e.g., DNA and RNA), metabolites, glycosylation, exosomes, viruses, bacteria, and cells (e.g., circulating tumor cells). In some embodiments, the target molecule is a fragment or component of a protein, a nucleic acid (e.g., DNA and RNA), a metabolite, glycosylation, an exosome, a virus, a bacteria, or a cell. A DNA target may include one or more single nucleotide variants (SNVs), insertion / deletions (indels), copy number variant, or methylated nucleotides, or any combination thereof. In some embodiments, the DNA target is a cell -free DNA (cfDNA), such as maternal cfDNA, fetal cfDNA, or cfDNA from a solid tumor. From a DNA target may be a synthetic DNA target, such as a product of a polymerase chain reaction (PCR). The DNA target may be transcribed from single-stranded RNA templates, such as complementary DNA (cDNA). An RNA target may be messenger RNA (mRNA). The mRNA may be a splice variant. The RNA target may be a microRNA (miRNA), a pre-miRNA, a pri-miRNA, a mRNA, a pre- mRNA, a viral RNA, a viroid RNA, a virusoid RNA, circular RNA (circRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), a pre-tRNA, a long non-coding RNA (IncRNA), a small nuclear RNA (snRNA), a circulating RNA, a cell-free RNA, an exosomal RNA, a vector- expressed RNA, an RNA transcript, a synthetic RNA, and combinations thereof. In one embodiment, an encoded assay may be performed for the analysis of a set of nucleic acid targets from a sample.
[0138] In one embodiment, the target molecule is DNA. In an encoded assay, a set of DNA targets may be targeted for detection of a single nucleotide difference relative to a reference nucleotide. A single nucleotide difference may be a change in the methylation status of a nucleotide at a target site of interest. In another example, a single nucleotide difference may be a change in nucleotide usage at a target site of interest, e.g., a single nucleotide polymorphism (SNP).
[0139] In one embodiment, the target molecule is RNA. In an encoded assay, an RNA sample may, for example, be processed in a reverse transcription reaction to generate cDNA molecules for detection of a set of targets of interest. An encoded RNA assay may, for example, be used to detect and count RNA targets of interest in a sample. In another example, an encoded RNA assay may be used to detect alternative splicing variants for a target of interest. FIG. 10 is a flow diagram of an example of an assay workflow 1000 according to some embodiments herein. Assay workflow 1000 may include, but is not limited to, the following steps.
[0140] At 1010, a sample is collected. The sample maybe whole blood, lymphatic fluid, serum, plasma, sweat, tear, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, fluidized tissues, fluidized organisms, liquids containing multi-celled organisms, biological swabs or biological washes, tissue samples, cell samples and biopsy samples. For example, a blood or saliva sample may be collected. In one example, a whole blood sample maybe collected and processed to separate the plasma fraction from the cellular components of whole blood.
[0141] At 1015, target extraction, concentration, conversion, and / or purification processes are performed. In some embodiments, depending on the sample type, if a specific target nucleic acid is desired an amplification step to enrich that target nucleic acid could be performed after extracting. In this example, the target is DNA. DNA (e.g., cell-free DNA) in the plasma sample may be extracted, purified, and concentrated for analysis. A proteinase K (ThermoFisher, Waltham, MA) digestion step may be used to digest proteins present in the plasma sample . In some cases, a heat denaturation step (e.g., 94-98°C for 20-30 seconds) may be used to denature double-stranded DNA into single-stranded nucleic acid. A bead-based extraction and concentration protocol may be used to capture single -stranded DNA in the plasma sample. In some embodiments, the bead-based extraction protocol uses magnetically responsive nucleic acid capture beads. The bead-bound DNA may be released from the capture beads using an elution buffer (or other elution means suitable to the capture bead used) to produce a processed DNA sample for analysis. In one embodiment, the DNA sample may be further processed in a bisulfite conversion reaction for analysis of the methylation status of a set of targets in the sample. In some embodiments, the sample is a tissue or cell sample wherein the DNA is extracted from tissues or cells for example by protease digestion of tissue and cellular proteins . Once extracted, the tissue or cell derived DNA can be purified away from cellular proteins and components using methods known in the art.
[0142] At 1020, the processed DNA sample is transferred into, for example, an analysis cartridge according to some embodiments herein. The analysis cartridge may comprise a reaction vessel. Non-limiting examples of reaction vessels include a plate, a well, a container, a tube, a flow cell, a microfluidic chip, or the like. The plate may be a welled plate, such as a 96- well plate. The reaction vessel, or a reaction surface thereof, may be optically clear to enable optical target detection in the reaction vessel. The reaction vessel may comprise a glass surface. The reaction vessel may comprise a glass-bottomed, 96-well plate. The reaction vessel may comprise a cationic coating. The reaction vessel may comprise a polylysine coating. The reaction vessel may comprise a poly-L-lysine coating. The reaction vessel may comprise a glass- bottomed, poly-L-lysine-coated 96-well plate. In some embodiments, a reaction vessel is a glass sheet of ~ 1mm thickness that has been coated with poly-L-lysine and is attached to a plastic well plate with an adhesive bottom. In some embodiments, a reaction vessel is a glass sheet of ~lmm thickness that has been coated with poly-L-lysine and is attached to a plastic well plate by compression of a silicon gasket matching the bottom structure of the well plate. The well plate may be composed of an array of circular or square features of varying depth and diameters of ~3 mm to ~6 mm. In some embodiments, the processed DNA sample is transferred to a reaction tube.
[0143] At 1025, a recognition event for each target in a set of targets is performed. For example, each target is uniquely recognized by and bound to a recognition element associated with a code (and optionally other elements). In one example, the recognition event for the set of targets uses a panel of coded recognition elements. In another example, the recognition event for the set of targets uses a panel of molecular inversion probes. The recognition event yields a set of coded targets comprising the target and the recognition element.
[0144] The recognition event may include sequence-specific binding between a 5’ probe arm and a 3’ probe arm of the recognition element to the target nucleic acid molecule under conditions sufficient to form a binding complex comprising the recognition element and the target nucleic acid molecule. In embodiments where the recognition element is configured to form a padlock probe, the 5’ probe arm or the 3 ’ probe arm comprises a target recognition region that binds to the target. In another embodiment, the target region is unknown, and the 5 ’ probe arm and the 3 ’ probe arm bind to the target nucleic acid molecule at 3 ’ and 5 ’ regions flanking the target region leaving a gap between the 5’ probe arm and the 3 ’ probe arm of the padlock probe.
[0145] The recognition event may include sequence-specific binding between a 5’ probe arm and a 3’ probe arm of the recognition element to 3 ’ region and a 5’ region of a bridge oligonucleotide having a target-specific element complementary to the target nucleic acid molecule interposedbetween the 3 ’ region and the 5’ region. In some embodiments, a cleavage product from a flap endonuclease cleavage reaction between a dual -probe recognition element and the target molecule can be used as a proxy for the presence of the target nucleic acid molecule. In some embodiments, the bridge oligonucleotide and the recognition element are introduced to the target nucleic acid molecule under conditions sufficient to form a ternary binding complex comprising the recognition element, the bridge oligonucleotide and target nucleic acid molecule.
[0146] The recognition element may include sequence-specific binding between the target nucleic acid molecule and a target-binding region of a pre-circularized recognition element. In some embodiments, the target nucleic acid molecule serves as a primer for an amplification reaction, such as a rolling circle amplification (RCA) reaction.
[0147] At 1030, a transformation event for each recognition element of the set of coded targets is performed. The transformation event may comprise introducing the binding complex to one or more enzymes under conditions sufficient to circularize the recognition element. The transformation event may include a ligation reaction between the 3 ’ probe arm and the 5’ probe arm of the recognition element by a ligating enzyme if the probe arms adjacently hybridize to the target nucleic acid molecule. The ligating enzyme may be a DNA ligase or catalytically active portion thereof. Non-limiting examples of ligases include Taq DNA Ligase, HiFi Taq DNA ligase (NEB), Ampligase Thermostable DNA ligase (LGC), Pfu DNA Ligase, Tth DNA Ligase, Tfi DNA Ligase, Tsc DNA Ligase, 9°N DNA Ligase (NEB), T4 RNA Ligase 1 and SplintR® Ligase (NEB). The transformation even may include a gap-fill ligation reaction in embodiments where there is a gap between the hybridized 3’ probe arm and the hybridized 5’ probe arm of the recognition element following target nucleic acid binding. In addition to the ligase, any gap between the 3’ hybridized probe arm and the 5’ hybridized from arm may first be filled by extension of the 3 ’ probe arm, using dNTPs and a DNA polymerase, until the extended 3’ probe arm is adjacent to the 5’ probe arm. The DNA polymerase or a catalytically active portion thereof may be one or more of a Q5 high-fidelity DNA polymerase (NEB), T4 DNA Polymerase, Sulfolobus DNA Polymerase IV, and T7 polymerase. Transformation of a recognition element in a ligation or gap -fill ligation reaction generates a circular molecule, or a modified recognition element.
[0148] An exonuclease cleanup step may be used following the transformation event to digest any remaining single stranded nucleic acid, such as unreacted coded recognition elements, amplification primers, and single stranded target sequences. Single stranded or double stranded targets may also be reduced or eliminated by exonuclease digestion. Non-limiting examples of exonucleases useful for digesting remaining nucleic acid without digesting modified recognition elements include thermolabile Exonuclease I, Exonuclease I, Exonuclease VII, Lambda exonuclease, Red, RecJf, Exonuclease VIII truncated, or Msz Exonuclease I. The transformation event yields a set of modified recognition elements comprising the code. At 1035, an amplification event for each code of the set of modified recognition elements is performed. In one example, the amplification event may be a rolling circle amplification (RCA) reaction to generate a set of target-specific concatemeric amplification products. The amplification event yields a set of concatemeric amplified recognition elements including codes (among other elements).
[0149] At 1040, a detection and decoding event for each amplified code of the set of amplified codes is performed to identify the code which can be used as a proxy for the presence of the target nucleic acid molecule. In one example, the code may be detected and decoded by sequencingthe code (and optionally other elements). The detection event detects the code as a surrogate for detection of the target. Decoding by sequencing may in some cases make use of soft decision decoding.
[0150] At 1045, using the code information (and optionally other elements) from step 1040, bioinformatics is performed. The bioinformatic analysis pipeline may be performed by one or more computer systems of the present disclosure.
[0151] In some embodiments, the amplification products may be sequenced directly. In some embodiments, sequencing adapters may be added by PCR amplification, followed by clustering and sequencing.
[0152] In some embodiments, an unknown region of a target sequence may be captured by a recognition element transformation reaction and sequenced along with the code. FIG. 4 is a schematic diagram of an example of a process 400 for capturing an unknown region of a target for sequencing.
[0153] In 410, a recognition element is hybridized to a target 420 and circularized by ligation after a gap -fill extension reaction that captures an unknown region of the target sequence. For example, a recognition element 410 that includes a code 412 (among other elements not shown) and a pair of target recognition regions 414a and 414b is hybridized to a target 420. Target 420 may include region 422 comprising an unknown sequence. Target recognition regions 414a and 414b recognize and bind to target 420 at sites flanking the unknown region 422. A gap-fill extension and ligation reaction (indicated by dashed arrow) is performed to copy region 422 into the recognition element to yield a circular modified recognition element (not shown) comprising the complement of the unknown region 422 of target 420. The ligation reaction may be followed by an exonuclease digestion step to remove unligated and linear recognition elements 410 and targets 420.
[0154] In 415, the circular modified recognition element is amplified in an RCA reaction to form an RCA amplification product 425 comprising multiple copies of the unknown region 422 and the code 412 (among other sequences). The RCA product 425 may be sequenced directly or sequencing adapters may be added by PCR amplification, followed by clustering and sequencing.
[0155] FIG. 28 is a schematic diagram illustrating an example of a process 2800 of using a bisulfite conversion reaction in combination with a coded recognition element to detect a methylated target site of interest. In this example, a DNA sample may include a target sequence of interest 2810 that may be methylated (e.g., 2810a “Methylated Target”) or unmethylated (e.g., 2810b “Unmethylated Target”) at a CpG site of interest. A bisulfite conversion reaction is used to convert non-methylated cytosine to thymine (C —> T) in the target sequence 2810b.
[0156] In the recognition event, target sequence 2810 is recognized and bound by a recognition element comprising a code, e.g., padlock probe configuration 2815. Recognition element 2815 includes a 3 '-terminal G nucleotide that base pairs with the target C at the CpG site of interest.
[0157] In some embodiments, in a transformation event, ligation of recognition element 2815 occurs when the 3 '-terminus of the recognition element (e.g., a guanine “G”) is matched to the target site “C” of interest in target sequence 2810a to generate a circularized modified recognition element 2820. No ligation occurs at the target site “T” in the bisulfite converted target sequence 2810b and consequently, transformation of recognition element 2815 hybridized to target sequence 2810b to a circular modified recognition element does not occur. As described above with reference to FIG. 10, modified recognition element 2820 may be amplified in an amplification reaction to generate an amplification product (step 1035) comprising many copies of the code (among other elements) and the code may be decoded (step 1040)
[0158] In one embodiment of process 2800, the recognition element (e.g., encoded probe) may be a molecular inversion probe that includes a 3 '-terminal single base gap at a target site of interest. A gap-fill ligation event using only a single added nucleotide may be used to generate the modified recognition element comprising the code only when the nucleotide corresponding to the target site of interest is incorporated. This approach provides two forms of specificity to the assay: (i) the 3 '-terminus of the probe recognizes and binds the interrogated site; and (ii) a single base extension reaction that incorporates the nucleotide corresponding to the target site of interest occurs.
[0159] FIGs. 29A-B illustrate a schematic diagram illustrating an example of a process 2900 for detecting a target sequence using a linear third oligonucleotide probe 2935 to produce a hybrid complex comprising a recognition element fragment and an oligonucleotide probe. The steps of process 2900 may, for example, be used in a methylation assay or a genotyping assay. Sample preparation for input into process 2900 may, for example, be performed starting from a whole blood sample, performing nucleic acid extraction, concentration, and / or purification, and transferring the nucleic acid sample to the analysis cartridge. Process 2900 may include, but is not limited to, the following steps.
[0160] At FIG. 29A, a recognition event is performed for each target in a set of targets to yield a set of released recognition element fragments. For example, an upstream probe 2910 and a downstream probe 2920 are combined in a binding reaction with a target sequence 2915 and a flap endonuclease (not shown). Downstream probe 2920 may include a target-specific sequence 2922 and a mismatch sequence 2924. In this example, target sequence 2915 includes a target site of interest that is a “C” nucleotide.
[0161] Hybridization of upstream probe 2910 and downstream probe 2920 to target sequence 2915 with no mismatches forms a ternary nucleic acid complex that may be recognized and cleaved (indicated by the dashed arrow) by a flap endonuclease to release a recognition element fragment 2930. Recognition element fragment sequence 2930 includes mismatch sequence 2924 and the base complementary to the target site of interest, e.g., “G” in this example.
[0162] Multiple rounds of target recognition and fragment release may be performed to increase the number of recognition element fragment 2930 released in the recognition event.
[0163] At FIG. 29B and FIG. 29C, a transformation event is performed to produce a set of modified recognition elements comprising hybrid complexes that include target-associated codes. In the transformation event, a bridge oligonucleotide 2935 may be used to mediate the ligation of a recognition element fragment to a coded third oligonucleotide probe 2940 to form a circular hybrid complex comprising the recognition element fragment and the third oligonucleotide probe. For example, a bridge oligonucleotide 2935 that includes sequences complementary to a coded third oligonucleotide probe 2940 and recognition element fragment 2930 may be used in a hybridization reaction to bring the ends of the third oligonucleotide probe 2935 and the recognition element fragment 2930 into proximity for ligation. In this example, a single set of recognition element fragments 2930, a coded third probe 2940, and a bridge oligonucleotide 2935 are shown, but any number of released fragment sets, coded third probes, and bridge oligonucleotides may be used.
[0164] The ligation of recognition element fragment 2930 to coded third probe 2940 may yield a circularized hybrid complex 2950 comprising the code 2945.
[0165] A detecting and decoding event (not shown) for circularized hybrid complex 2950 may include, for example, a rolling circle amplification event to generate an amplification product. In one embodiment, the third oligonucleotide probe may be a circular probe that includes a target-specific code and sequences for recognizing and hybridizing to a target- specific recognition element fragment (e.g., a mismatch sequence).
[0166] FIGs. 30A-C illustrate a schematic diagram illustrating an example of a process 3000 for detecting a target sequence using a circular third oligonucleotide probe to produce a circularized hybrid complex (e.g., modified recognition element) comprising the code. The steps of process 3000 may, for example, be used in a methylation assay or a genotyping assay.
[0167] At FIG. 30A, a recognition event is performed for each target 3015 in a set of targets to yield a set of released recognition element fragments 3030. For example, an upstream probe 3010 and a downstream probe 3020 are combined in a binding reaction with a target sequence 3015 and a flap endonuclease (not shown). Downstream probe 3020 may include a targetspecific sequence 3022 and a mismatch sequence 3024. Mismatch sequence 3024 may include a sequence that is complementary to a pre -circularized third oligonucleotide probe comprising a target-associated code. In this example, target sequence 3015 includes a target site of interest that is a “C” nucleotide.
[0168] Hybridization of upstream probe 3010 and downstream probe 3020 to target sequence 3015 with no mismatches forms a ternary nucleic acid complex that may be recognized and cleaved (indicated by the dashed arrow) by a flap endonuclease to release a recognition element fragment 3030. Recognition element fragment sequence 3030 includes mismatch sequence 3024 and the base complementary to the target site of interest, e.g., “G” in this example.
[0169] Multiple rounds of target recognition and fragment release may be performed to increase (e.g., amplify) the number of recognition element fragment 3030 released in the recognition event.
[0170] At FIG. 30B and FIG. 30C, a transformation event for the set of recognition element fragments 3030 may be performed to produce a set of circular modified recognition elements comprising hybrid complexes that include target-associated codes. In the transformation event, the recognition element fragment may be hybridized to a pre -circularized third oligonucleotide probe comprising a target-associated code and used to prime an RCA reaction to generate a nanoball detection product comprising the amplified code. For example, recognition element fragment 3030 may be hybridized to a pre-circularized third oligonucleotide probe 3040. The pre-circularized third oligonucleotide probe 3040 includes, for example, a code sequence 3042 and a hybridization sequence 3044 that is complementary to recognition element fragment 3030. An RCA reaction using recognition element fragment 3030 as a primer sequence is performed to generate an amplification product (not shown) comprisingthe amplified target-associated code. In one example, Phi29 DNA polymerase may be used in the RCA reaction.
[0171] Unreacted (e.g., full-length) downstream probe 3020 that includes mismatch sequence 3024 may also hybridize to pre-circularized third probe 3040. In this case, the 3 ' probe overhang of the unreacted probe may prevent priming of the RCA reaction . To prevent Phi29 exonuclease activity from degrading the 3 ' terminus of any unreacted probes different strategies may be used. In one example, an exo(-) Phi29 polymerase may be used in the RCA reaction. In another example, probes with 3 ' termini that are resistant to exonuclease degradation may be used (e.g., by including phosphorothioated nucleotides, alkyl linkers, or inverted bases).
[0172] In this example, a single set of recognition element fragments 3030 and pre-circularized third probe 3040 are shown, but any number of released fragment sets and encoded third probes may be used to generate a set of amplification products (e.g., concatemers) for detection of the set of targets.
[0173] FIGs. 31 A-C illustrate a schematic diagram illustrating an example of a process 3100 for detecting a target of interest using a pre-circularized single probe recognition element and a PCR amplification / 5' nuclease cleavage reaction. The steps of process 3100 may, for example, be used in a methylation assay or a genotyping assay.
[0174] Sample preparation for input into process 3100 may, for example, be performed starting from a whole blood sample, performing the nucleic acid extraction, concentration, and / or purification processes, and transferring the nucleic acid sample to the analysis cartridge . Process 3100 may include, but is not limited to, the following steps.
[0175] At FIG. 31 A, a recognition event is performed for each target in a set of targets to yield a set of released recognition element fragments. For example, a single probe 3110 is combined in an amplification reaction with a forward primer 3120a and a reverse primer 3120b that are specific for a target sequence 3125 of interest, and a DNA polymerase having 5' nuclease activity (e.g., Taq DNA polymerase). Single probe 3110 may include a target-specific sequence 3112 and a mismatch sequence 3114. In this example, target sequence 3125 includes a target site of interest that is a “C” nucleotide.
[0176] Hybridization of single recognition element 3110 to target sequence 3125 may form a ss- ds forked structure that includes a double-stranded (e.g., hybridized) region comprising target sequence 3125 and a single-stranded region that includes the mismatch sequence 3114.
[0177] During amplification, the structure-specific 5 ' nuclease activity of the DNA polymerase may cleave the 5' terminus of the hybridized probe and releases the non -complementary mismatch sequence 3114 to yield a recognition element fragment 3130 that is associated with the target. Recognition element fragment 3130 may include mismatch sequence 3114 and the base that is the complement of the target site of interest, e.g., a “G”. The site of cleavage may also be 5' or 3' of the matched base.
[0178] Multiple cycles of PCR amplification / 5' nuclease cleavage may be performed to increase the number of recognition element fragments 3130 released in the recognition event.
[0179] At FIG. 31B and FIG. 31C, a transformation event for the set of recognition element fragments may be performed to produce a set of circular modified recognition elements comprising hybrid complexes that include target-associated codes. In the transformation event, the recognition element fragment may be hybridized to a pre -circularized coded oligonucleotide probe comprising a target-associated code and used to prime an amplification reaction to generate an amplification product detection product comprising the amplified code as describe above with reference to FIG. 30. For example, recognition element fragment 3130 may be hybridized to a pre-circularized oligonucleotide probe 3140. Oligonucleotide probe 3140 includes, for example, a code sequence 3142 and a hybridization sequence 3144 that is complementary to recognition element fragment 3130. An RCA reaction using recognition element fragment 3130 as a primer sequence can be performed to generate the amplification detection product (not shown) comprising the amplified target-associated code.
[0180] In one embodiment of process 3100, the coded oligonucleotide probe may be a linear probe that includes a target- specific code and sequences for recognizing and hybridizing to a target-specific recognition element fragment (e.g., a mismatch sequence). In this case, in the transformation event, a bridge oligonucleotide may be used to mediate the ligation of the recognition element fragment to the coded oligonucleotide probe to form a circular hybrid complex (e.g., modified recognition element) comprisingthe recognition element fragment and the coded oligonucleotide probe as described above with reference to FIG. 29.
[0181] In some embodiments, a single probe recognition element may include a mismatch sequence comprising a target-specific code (among other elements). In this case, a recognition element fragment may be released from the single probe and the transformation event may include a hybridization and ligation reaction.
[0182] In some embodiments, the recognition element fragment 3130 comprises the code and the recognition element fragment 3130 binds to a linear bridge oligonucleotide instead of the pre-circularized oligonucleotide probe 3140. In this case, the linear bridge oligonucleotide binds to a 5’ region and a 3’ region of the recognition element fragment to bring them in close proximity under conditions sufficient to ligate the 5 ’ region to the 3 ’ region thereby generating modified recognition element. In some embodiments of workflow 1000 in FIG. 10, a sequencing library comprisingthe codes (among other elements) may be generated. The library may be sequenced to determine a presence of a code or a probability of the presence of a code that is associated with a target of interest. In one embodiment, a sequencing library may be generated from a modified (e.g., circularized) recognition element (step 1030). In another embodiment, a sequencing library may be generated from the amplification products of the modified recognition element.
[0183] In some embodiments, the identity of nucleotides in the code are determined using nucleic acid sequencing. Non -limiting examples of nucleic acid sequencing methods of the present disclosure include those provided in Slatko BE, Gardner AF, Ausubel FM. Overview of Next-Generation Sequencing Technologies. CurrProtoc Mol Biol. 2018 Apr;122(l):e59, which is hereby incorporated by reference in its entirety. For example, the sequencing method may be sequencing by hybridization, sequencing-by-synthesis (SBS), SMRT (Singe Molecule Real Time) sequencing, Nanopore-based DNA sequencing, or sequencing-by-binding (SBB), such as Avidity™ sequencing, or other sequencing methodologies.
[0184] In some embodiments, decoding comprises performing a sequencing-by-synthesis method that involves detection of the step -wise incorporation of detectable complementary nucleotides into a primed template strand. In some embodiments, the primed template strand is the code or a segment within the code. In some embodiments, the nucleotides or nucleotide analogs that are incorporated are labeled, which enable detection of incorporation . In some embodiments, the nucleotides or nucleotide analogs are labeled with a fluorescent label. In some cases, the incorporation of sequential nucleotides or nucleotide analogs into the primed template is monitored in real-time. In such sequencing-by-synthesis methods, the nucleotide or nucleotide analogs are detectable but unblocked, meaning they are incorporated into the growing primed template upon binding to a complementary nucleotide in the template at the N+l position. In some embodiments, the detectable nucleotides or nucleotide analogs have a blocking group (also referred to here as “reversible terminators” when the block can be reversed) that pause the synthesis process during detection. Subsequent to, or substantially simultaneously with, detection, the blocking group and / or the fluorescent label may be cleaved off of the nucleotide or nucleotide analog, thereby enabling incorporation of the nucleotide or nucleotide analog into the growing primed template strand. While sequencing, a template strand which has failed to incorporate a nucleotide or nucleotide analog in a sequencing cycle will continue to lag behind, which is referred to as “phasing,” and can cause errors in base call accuracy. To avoid phasing, the sequencing-by-synthesis method may also utilize unlabeled nucleotides or nucleotide moieties that incorporate into the primed template at places where the template strand failed to incorporate a detectable or previously detectable nucleotide or nucleotide.
[0185] A variety of sequencing workflows are possible within the scope of the assays disclosed. A sequencing workflow may use a recognition element. The code of the recognition element may be a soft decodable code.
[0186] In some embodiments, a sequencing workflow may include: i. hybridizing recognition elements to a set of targets, yielding modified recognition elements; ii. amplifying the modified recognition elements, e.g., amplifying the modified recognition element using, for example, rolling circle amplification (RCA); iii. binding (e.g., binding or incorporating) detectable (e.g., labeled) nucleotides or nucleotide analogs to a subset of the amplified, modified recognition elements; iv. binding (e.g., binding or incorporating) unlabeled nucleotides or nucleotide analogs to a subset of the amplified, modified recognition elements; v. detecting (e.g., via imaging) a signal from the binding of the detectable nucleotides or nucleotide analogs to the subset of the amplified, modified recognition elements; and vi. iteratively performing (iii), (iv), and (v) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements.
[0187] In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least twice. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least three times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least five times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least eight times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least ten times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 12 times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 15 times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 17 times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 20 times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 25 times. In some embodiments, a sequencing workflow comprises iteratively performing (iii), (iv), and (v) at least 30 times.
[0188] In some embodiments, a sequencing workflow does not comprise binding (e.g., binding or incorporating) unlabeled nucleotides or nucleotide analogs to a subset of the amplified, modified recognition elements.
[0189] In some embodiments, a sequencing workflow comprises performing (iii) and (iv) at the same time (e.g., concurrently or simultaneously). In some embodiments, a sequencing workflow comprises performing (iii) and (iv) sequentially. In some embodiments, (iii) may occur before (iv). In some embodiments, (iv) may occur before (iii).
[0190] The modified recognition elements or amplification products thereof may have a sequencing primer binding site suitable for binding to a sequencing primer. Alternatively, the modified recognition elements or amplification product thereof may have been tagmented whereby the modified recognition elements or amplification products were fragmented and tagged with one or more adaptors comprising the sequencing primer binding site. In either embodiment, the workflow may include flowing into the reaction vessel a plurality of sequencing primers under conditions sufficient to bind the sequencing primer to the primer binding sites to produce primed nucleic acid molecules comprising the code that serve as a template for the nucleic acid sequencing reaction.
[0191] The sequencing primer may be blocked from a primer extension reaction. For example, the sequencing primer may have a nucleotide at the 3 ’ end that is reversibly terminated and does not allow for extension by a polymerase when blocked. Alternatively, the detectable nucleotide or nucleotide analog may comprise a reversible terminator moiety. In either embodiment, the workflow may include subjecting the primed templates bound to the detectable nucleotide or nucleotide analog to a deblocking condition sufficient to remove a reversible terminator from a detectable nucleotide or nucleotide analog or the sequencing primer, thereby enabling the incorporation of a detectable nucleotide or nucleotide analog into the primed template. The workflow may include removing a detectable label from a detectable nucleotide or nucleotide moiety. In some embodiments, a sequencing workflow includes: viii. removing a reversible terminator from a detectable nucleotide or nucleotide analog, and removing a reversible terminator from a nucleotide or nucleotide analog; and ix. removing a detectable label from a detectable nucleotide or nucleotide moiety. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 2 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 3 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 5 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 7 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 8 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 10 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 12 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 16 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 20 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 25 times. In some embodiments, a sequencing workflow includes iteratively performing (iv), (v), (vi), (viii), and (ix) at least 30 times.
[0192] In some embodiments, the detecting of the binding event occurs prior to removal of the reversible terminator and removal of the detectable label. The detecting of the binding event may occur simultaneously with removal of the reversible terminator and removal of the detectable label. The detecting of the binding event may occur simultaneously with removal of the detectable label and prior to removal of the reversible terminator. The detecting of the binding event may occur simultaneously with removal of the reversible terminator and prior to removal of the detectable label.
[0193] In some embodiments provided herein, a sequencing workflow may proceed as described in FIG. 11. In some embodiments provided herein, a sequencing workflow may proceed in steps substantially similar to those described in FIG. 11. In some embodiments provided herein, a sequencing workflow may proceed in steps substantially similar to those described in FIG. 11 comprising different reagents from those disclosed in FIG. 11, the different reagents having similar functions to the reagents disclosed in FIG. 11.
[0194] In some embodiments, one or more codes associated with a target of interest are identified in part by detection of binding events between a modified recognition element and a detectable nucleotide or detectable nucleotide analog. A binding event may be a binding between the modified recognition element and a detectable nucleotide or detectable nucleotide analog. For example, a binding event may be a binding between a nucleotide comprising a fluorescent label and a modified recognition element. Binding between a nucleotide and a modified recognition element may not include incorporation of the nucleotide into the modified recognition element. After a binding event between a detectable nucleotide or a nucleotide analog and a modified recognition element, the detectable nucleotide or nucleotide analog may be unbound (e.g., washed away) from the modified recognition element. For example, a destabilizing buffer may be introduced that destabilizes a binding complex between the detectable nucleotide or a nucleotide analog and a modified recognition element. The unbound detectable nucleotide or nucleotide analog may be washed way. The nucleotide or nucleotide analog may be unbound after the binding event is detected, or the binding event may be detected by detecting the unbinding of the nucleotide or nucleotide analog.
[0195] In some embodiments, the binding event may include an incorporation of a detectable nucleotide or detectable nucleotide analog into a modified recognition element. An incorporation of a detectable nucleotide or nucleotide analog into a modified recognition element may not be a binding between a detectable nucleotide or nucleotide analog and a modified recognition element. In some embodiments, a binding event may occur between an element (e.g., a nucleic acid) of a code sequence of a modified recognition element and a nucleotide or nucleotide analog. In some embodiments, a binding event may occur between an element (e.g., a nucleic acid) of an index sequence of a modified recognition element and a nucleotide or nucleotide analog. In some embodiments, a binding event may occur between an element (e.g., a nucleic acid) of a unique molecular identifier (UMI) sequence of a modified recognition element and a nucleotide or nucleotide analog.
[0196] The binding of the nucleotide or nucleotide analog to the modified recognition element may be a detectable event or may be detected. The binding of the nucleotide or nucleotide analog to the modified recognition element may produce an effect that is detected. The detection of a binding event may comprise an intensity readout (e.g., the intensity of light emitted in a certain range of frequencies or in a color channel), or a categorical read out (e.g., ON or OFF, whether light is detected in a certain range of frequencies or in a color channel.)
[0197] Provided herein are methods of identifying a target of interest in part by detection of binding events between a modified recognition element and a detectable nucleotide or detectable nucleotide analog. As contemplated herein, a nucleotide or nucleotide analog may serve as a monomeric unit of a nucleic acid polymer, e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). A nucleotide or a nucleotide analog may comprise a nucleobase (e.g., guanine, adenine, cytosine, thymine, or uracil) and a phosphate group.
[0198] In some embodiments, a nucleotide analog comprises a nucleobase (e.g., guanine, adenine, cytosine, thymine, or uracil) and a phosphate group. A nucleotide analog may comprise a nucleic acid analog, an analog nucleobase, or an artificial nucleic acid. A nucleotide analog may comprise a nucleotide. A nucleotide analog may comprise a Xeno Nucleic Acid (XNA) . In some embodiments, a nucleotide analog may be incorporated into a growing nucleic acid chain or strand. Incorporation of a nucleotide analog into a modified recognition element may terminate the growth of a nucleic acid strand or chain. For example, the nucleotide may have a 3 ’-OH reversible terminator moiety. The nucleotide analog may be a conjugate comprising one or more nucleotides conjugated to another molecule, such as a nanoparticle, a protein or peptide, or a polymer. The protein or peptide may be an antibody or antigen -binding fragment thereof. The protein may be streptavidin or avidin. The polymer may be a polyethylene glycol (PEG), poly dimethylsiloxane, polystyrene microporous polystyrene (MPPS), polymethylmethacrylate (PMMA), polycarbonate (PC), polypropylene (PP), polyethylene (PE), high density polyethylene (HDPE), cyclic olefin polymers (COP), cyclic olefin copolymers (COC), or polyethylene terephthalate (PET), or a combination thereof.
[0199] A nucleotide analog may be a deoxynucleotide triphosphate (dNTP) comprising adenine (dATP), cytosine (dCTP), guanine (dGTP), thymine (dTTP), or uracil (dUTP). A nucleotide analog may be a dideoxynucleotide triphosphate (ddNTP) comprising adenine (ddATP), cytosine (ddCTP), guanine (ddGTP), thymine (ddTTP), or uracil (ddUTP). A nucleotide analog may comprise a reversible terminator. A nucleotide analog may comprise a detectable label.
[0200] In some embodiments provided herein, a nucleotide or nucleotide analog comprises a reversible terminator moiety. The reversible terminator moiety may prevent incorporation of an additional nucleotide or nucleotide analog in an N+l position, wherein N is the position of the nucleotide or nucleotide analog.
[0201] In some embodiments, the detectable nucleotide or nucleotide analog comprises a reversible terminator moiety at a 3 ’ OH position. The reversible terminator may comprises a 3 ’ - O-alkyl hydroxylamino group, a 3’-phosphorothioate group, a 3 ’-O-malonyl group, a 3’-O- benzyl group, a 3’-O-Azidomethyl group, a 3’-O-Azide group, a 3’-O-Azido group, or a 3’-O- methyl group.
[0202] In some embodiments provided herein, a nucleotide or nucleotide analog comprises a detectable label, e.g., an optical label or a fluorescent label. A detectable label may emit a signal (e.g., a chemical signal or an optical signal). The signal emitted by a detectable label may be easily detected or distinguishable. The detectable label may be a dye. The detectable label may be a fluorescent dye. The fluorescent dye may emit red, far-red, near-red, yellow, green, or blue light. In some embodiments, the fluorescent dye comprises 6 -FAM (6-carboxy fluorescein). In some embodiments, the fluorescent dye comprises JOE (6-carboxy -4', 5'-dichloro-2', 7'- dimethoxyfluorescein). In some embodiments, the fluorescent dye comprises TAMRA (6- carboxytetramethylrhodamine). In some embodiments, the fluorescent dye comprises ROX (6- carboxy-X-rhodamine), 5-Cy5 (5 -carb oxy rhodamine). In some embodiments, the fluorescent dye comprises 5-Cy5.5 (5-carboxylic acid succinimidyl ester). In some embodiments, the fluorescent dye comprises 5-Cy7 (5 -carb oxyrhodamine). In some embodiments, the fluorescent dye comprises HEX (hexachlorofluorescein). In some embodiments, the fluorescent dye comprises Alexa Fluor 488 (AF488). In some embodiments, the fluorescent dye comprises Alexa Fluor 514 (AF514). In some embodiments, the fluorescent dye comprises Texas Red. In some embodiments, the fluorescent dye comprises Cyanine 3. In some embodiments, the fluorescent dye comprises Cyanine 5. In some embodiments, the fluorescent dye comprises Pacific Blue. In some embodiments, the fluorescent dye comprises Tetramethylrhodamine. In some embodiments, the fluorescent dye comprises Oxazole Yellow. In some embodiments, the fluorescent dye comprises Atto647N. In some embodiments, the fluorescent dye comprises Rhodamine 6G (R6G).
[0203] In some embodiments, the label is a fluorophore. Non-limiting examples of fluorescent moieties include, but are not limited to, fluorescein and fluorescein derivatives such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl -amino fluorescein, rhodamine and rhodamine derivatives such as TRITC, TMR, lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonyl hydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Yellow and derivatives such as Lucifer Yellow iodoacetamide, Lucifer Yellow CH, cyanine and derivatives such as indolium based cyanine dyes, benzo- indolium based cyanine dyes, pyridium based cyanine dyes, thiozolium based cyanine dyes, quinolinium based cyanine dyes, imidazolium based cyanine dyes, Cy 3, Cy5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT, BHHCT, BCOT, Europium chelates, Terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and derivatives thereof, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, Malachite green, stilbene, DEG dyes, NR dyes, nearinfrared dyes and others known in the art such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition; Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or Hermanson, Bioconjugate Techniques, 2nd Edition, or derivatives thereof, or any combination thereof. Cyanine dyes may exist in either sulfonated or non-sulfonated forms, and consist of two indolenin, benzo-indolium, pyridium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3, (which may comprise 1 -[6-(2,5-dioxopyrrolidin-l -yloxy)-6-oxohexyl]-2-(3 -{ 1 -[6-(2,5-dioxopyrrolidin-l- yloxy)-6-oxohexyl]-3,3-dimethyl-l,3-dihydro-2H-indol-2-ylidene}prop-l-en-l-yl)-3,3- dimethyl-3H-indolium or l -[6-(2,5-dioxopyrrolidin-l-yloxy)-6-oxohexyl]-2-(3-{ l-[6-(2,5- dioxopyrrolidin-l-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-l,3-dihydro-2H-indol-2- ylidene}prop-l-en-l-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may comprise 1 - (6-((2,5-dioxopyrrolidin-l-yl)oxy)-6-oxohexyl)-2-((lE,3E)-5-((E)-l-(6-((2,5-dioxopyrrolidin-l- yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-l,3-dien-l-yl)-3,3-dimethyl-3H- indol-l-ium or l-(6-((2,5-dioxopyrrolidin-l-yl)oxy)-6-oxohexyl)-2-((lE,3E)-5-((E)-l-(6-((2,5- dioxopyrrolidin- 1 -yl)oxy)-6-oxohexy l)-3 ,3 -dimethyl-5 -sulfoindolin-2-ylidene)penta- 1 , 3-dien- 1 - yl)-3,3-dimethyl-3H-indol-l-ium-5-sulfonate), and Cy7 (which may comprise l-(5- carboxypentyl)-2-[(lE,3E,5E,7Z)-7-(l-ethyl-l,3-dihydro-2H-indol-2-ylidene)hepta-l,3,5-trien- l-yl]-3H-indolium or l-(5-carboxypentyl)-2-[(lE,3E,5E,7Z)-7-(l-ethyl-5-sulfo-l,3-dihydro-2H- indol-2-ylidene)hepta-l,3,5-trien-l-yl]-3H-indolium-5-sulfonate), where “Cy” stands for 'cyanine', and the first digit identifies the number of carbon atoms between two indolenine groups. Cy2 which is an oxazole derivative rather than indolenin, and the benzo -derivatized Cy3.5, Cy5.5 and Cy7.5 are exceptions to this rule.
[0204] In some embodiments, the fluorescent or optical label of a nucleotide or nucleotide analog species emits a different wavelength than other fluorescent labels of other nucleotides or nucleotide analog species. In some embodiments, the fluorescent or optical label of a nucleotide or nucleotide analog species emits a wavelength that is differentiable from the wavelength of other fluorescent labels of other nucleotides or nucleotide analog species. For example, a nucleotide analog comprising guanine may further comprise a fluorescent label that emits in the red wavelengths, a nucleotide analog comprising adenine may further comprise a fluorescent label that emits in the green wavelengths, a nucleotide analog comprising cytosine may further comprise a fluorescent label that emits in the yellow wavelengths, a nucleotide analog comprising thymine may further comprise a fluorescent label that emits in the far -red wavelengths. In some embodiments, the fluorescent or optical label of a nucleotide or nucleotide analog species emits a wavelength that can be detected in a color channel. In some embodiments, the fluorescent or optical label of a nucleotide or nucleotide analog species emits a wavelength that can be detected in a different color channel than the fluorescent label of another nucleotide or nucleotide analog species. In some embodiments, the fluorescent or optical label of a nucleotide or nucleotide analog species is differentiable from another nucleotide or nucleotide analog species because the two species emit wavelength that can be detected in a different color channels. For example, a nucleotide analog comprising guanine may further comprise a fluorescent label that is detected in a red color channel, a nucleotide analog comprising adenine may further comprise a fluorescent label that is detected in a green color channel, a nucleotide analog comprising cytosine may further comprise a fluorescent label that is detected in a yellow color channel, a nucleotide analog comprising thymine may further comprise a fluorescent label that is detected in a far-red color channel. In some embodiments, the fluorescent or optical label of a first nucleotide or nucleotide analog species is differentiable from a second nucleotide or nucleotide analog species because the first species emits a wavelength that can be detected in two or more different color channels, and the second species emits light that can be detected in fewer color channels than the first. For example, a nucleotide analog comprising guanine may further comprise a fluorescent label that is detected in a red color channel and in the far-red color channel, while a nucleotide analog comprising thymine may comprise a fluorescent label that is detected only in the far-red color channel.
[0205] In some embodiments, the nucleotides or nucleotide analogs may comprise the same detectable label. In some embodiments, the nucleotides or nucleotide analogs may comprise the same fluorescent label. The nucleotides or nucleotide analogs provided in methods used herein may comprise different species (e.g., A, C, G, or T) and the different species may comprise the same fluorescent label.
[0206] Provided herein are methods of identifying a target of interest in part by introducing amplified modified recognition elements to a detectable nucleotide or nucleotide analog. In some embodiments, a detectable nucleotide or nucleotide analog couples with a recognition element. In some embodiments, a detectable nucleotide or nucleotide analog binds with a recognition element. In some embodiments, a detectable nucleotide or nucleotide analog integrates into a recognition element.
[0207] Provided herein are methods of identifying a target of interest in part by introducing amplified modified recognition elements to a nucleotide or nucleotide analog that is unlabeled. In some embodiments, an unlabeled nucleotide or nucleotide analog couples with a recognition element. In some embodiments, an unlabeled nucleotide or nucleotide analog binds with a recognition element. In some embodiments, an unlabeled nucleotide or nucleotide analog integrates into a recognition element.
[0208] In some embodiments, a nucleotide or nucleotide analog is unlabeled (e.g., a nucleotide or nucleotide analog that does not comprise a label). In some embodiments, a subset of nucleotides or nucleotides are unlabeled (e.g., a nucleotide or nucleotide analog that does not comprise a label). In some embodiments, a nucleotide or nucleotide analog or a subset of nucleotides or nucleotides do not comprise an optical label. In some embodiments, a nucleotide or nucleotide analog or a subset of nucleotides or nucleotides do not comprise a fluorescent label.
[0209] In some embodiments, a subset of the nucleotides or nucleotide analogs disclosed herein is detectable (e.g., labeled), and a subset of the nucleotides or nucleotide analogs disclosed herein is unlabeled. For example, a method disclosed herein may comprise the use of detectable A, C, G, and T nucleotides or nucleotide analogs and unlabeled A, C, G, and T nucleotides or nucleotide analogs.
[0210] A sequencing library comprising the codes (among other elements) may be generated from a set of target-specific amplification products (step 1035). The amplification product library may be sequenced to identify codes associated with targets of interest.
[0211] FIG. 5 is a schematic diagram illustrating an example of a process 500 for generating a sequencing library from an amplification product set that may be used to identify the codes associated with the target set of interest. Sample preparation for input into process 500 may, for example, be performed as described for FIG. 10 starting from a whole blood sample (step 1010), performing the nucleic acid extraction, concentration, and / or purification processes (step 1015), and transferring the nucleic acid sample to the analysis cartridge (step 1020). Process 500 may include, but is not limited to, the following steps.
[0212] In 510, recognition and transformation events (steps 1025 and 1030) for each target in a set of targets of interest is performed to yield a set of modified recognition elements comprising the code. For example, a set of coded recognition elements 512 that include target- specific recognition regions associated with a code may be used. The transformation event may include a ligation or a gap -fill ligation reaction to produce a plurality of circularized modified recognition elements comprising the code. In the transformation event, only the coded recognition elements 512 that hybridize to a target sequence of interest with no mismatches at the end of the 3 ’ region may be ligated to yield a circular modified recognition element comprising the code. In this example, a single modified recognition element 514 is shown, but any number of modified recognition elements 514 may be generated to yield a set of modified recognition elements 514.
[0213] In 515, an amplification event for each code of the set of modified recognition elements is performed. For example, modified recognition element 514 may be amplified in a rolling circle amplification (RCA) to generate an amplification product 516. In this example, a single amplification product 516 is shown, but any number of amplification products may be generated corresponding to the number of circular modified recognition elements present to yield a set of amplification products comprising the codes.
[0214] In 520, a sequencing library is generated from the amplification product. For example, 25 cycles of amplification may be used to add sequencing adapters and sample index sequences (among other optional sequences) to the code sequence generating a sequencing library 522 that includes a set of codes. Sequencing library 522 may be loaded onto a sequencing flow cell for next generation sequencing (NGS).
[0215] In 525, a detection event for each code of the set of codes is performed. For example, the library 522 is sequenced using an NGS sequencing protocol to identify the sequence of the amplification products used to generate the sequencing library including the codes (and other elements (e.g., sample index, UMIs)) associated with the set of targets of interest. The code data may then be used as a digital count of the target- specific detection events.
[0216] A set of amplification products (step 516) may be directly sequenced to identify codes associated with the set of targets of interest. The code data may then be used as a digital count of the target-specific detection events.
[0217] In one embodiment, the amplification products may be immobilized onto the surface of a sequencing flow cell for direct sequencing on the amplification products. The amplification products may be immobilized onto the flow cell surface using an immobilization agent. In one example, the immobilization agent is a surface bound oligonucleotide that is complementary to a sequence on the amplification product. In another example, the immobilization agent is a polypeptide.
[0218] To facilitate immobilization of an amplification product on a flow cell surface for direct sequencing, a recognition element associated with a code (e.g., an encoded recognition element) may include a palindrome sequence that is incorporated into the amplification product to create a secondary structure that compacts (collapses) the amplification product. The compacted amplification product provides a structure that may be more readily sequenced.
[0219] FIG. 6 is a schematic diagram illustrating an example of a process 600 for directly sequencing an amplification product set to identify codes associated with the target set of interest. Sample preparation for input into process 600 may, for example, be performed as described for FIG. 6 starting from a whole blood sample (step 1010), performing the nucleic acid extraction, concentration, and / or purification processes (step 1015), and transferring the nucleic acid sample to the analysis cartridge (step 1020). Process 600 may include, but is not limited to, the following steps.
[0220] In 610, recognition and transformation events (steps 1025 and 1030) for each target in a set of targets of interest is performed to yield a set of modified recognition elements comprising the code. For example, a set of coded recognition elements 612 that include target- specific recognition elements associated with a code may be used. The transformation event may include a ligation or a gap-fill ligation reaction to produce a circularized modified recognition element comprising the code. In the transformation event, only the coded recognition elements 612 that hybridize to a target sequence of interest with no mismatches may be ligated to yield a circular modified recognition element comprising the code. In this example, a single modified recognition element 614 is shown, but any number of modified recognition elements 614 maybe generated to yield a set of modified recognition elements 614.
[0221] In 615, an amplification event for each modified recognition element is performed. For example, modified recognition element 614 may be amplified in a rolling circle amplification (RCA) to generate an amplification product 616. In this example, a single amplification product 616 is shown, but any number of amplification products may be generated corresponding to the number of circular modified recognition elements present to yield a set of amplification products comprising the codes.
[0222] In 620, the amplification product is loaded onto the surface of a sequencing flow cell. For example, amplification product 616 is loaded onto a sequencing flow cell. The amplification products may be immobilized onto the flow cell surface using an immobilization agent. In one example, the immobilization agent is a surface bound oligonucleotide that is complementary to a sequence on the amplification product. In another example, the immobilization agent is a polypeptide.
[0223] In 625, a detection event for each amplified code of the set of amplified recognition elements is performed. For example, the amplification product is directly sequenced to identify codes associated with the set of targets of interest. The code data may then be used as a digital count of the target-specific detection events.
[0224] FIG. 7 shows another workflow 700 for using circularized recognition elements for determining the sequence of a code using SBS chemistry. In FIG. 7, at 710 and 715, recognition elements recognize, bind, ligate and circularize if a target nucleic acid is present in the sample. The target could be a known target or it could be an unknown sequence. If the target sequence, be it known or unknown, is present the recognition element ligates and circularizes 714, however in the absence of the target sequence there is no ligation or circularization 712. A recognition element that does not ligate or circularize 712, is not expected to participate in downstream applications. A nuclease can be used to remove uncircularized recognition elements, target sequences, and other linear nucleic acids from the workflow for increasing the sensitivity and specificity of downstream applications. In this particular example, a circularized recognition element is immobilized on a sequencing substrate 720 without the need to first generate concatemers of the circularized recognition element prior to sequencing 725.
[0225] FIG. 8 is an expanded view of how a circularized recognition element can be used in a sequencing workflow. In this example, recognition elements that are generated to hybridize to wild type versions of a target nucleic acid include a sequence that is specific to a wild type sequence 801. Alternatively, recognition elements that are generated to hybridize to variants of a target nucleic acid, for example a SNP or other sequence variant, include a different sequence that is specific to the variant sequence of interest 802. The additional sequences present in a recognition element to discern between two different sequence types, in this example wild type and variant, can be any sequence that can participate in the described workflow. Once circularized, the recognition elements can be captured on a sequencing flowcell 805.
[0226] In FIG. 8, the flowcell 805 comprises a plurality of capture oligonucleotides 803 and 804. The capture oligonucleotides, for example on an Illumina® flowcell, include the P5 and P7 primer sequences that can capture a nucleic acid for sequencing if that nucleic acid has complementary P5 and P7 sequences. In this example, 801 comprises a sequence complementary to a P5 primer and 802 comprises a sequence complementary to a P7 primer. As such, both the wild type circularized recognition elements and the variant circularized recognition elements are captured on the flowcell. Variant sequences, which are not always of great abundance in a sample relative to a wild type sequence, benefit from this scenario as they have dedicated immobilization primers which increases their relative abundance on the flowcell. Once immobilized, the P5 and P7 primers immobilized on the surface of the flowcell and hybridized to the circularized recognition elements can be used in extension reactions, the products of which can be used in bridge amplification and subsequent code determination using SBS chemistry.
[0227] In some embodiments, the circular recognition elements that are immobilized on a flowcell are linearized to participate in bridge amplification. For example, a recognition element is generated with a restriction site that is located immediately adjacent to the sequence used to immobilize the circular recognition element on the flowcell. Once captured on the flowcell, the recognition element is cleaved using a restriction enzyme that will cut at the engineered restriction site, thereby linearizing the circular recognition element such that the linear recognition element, or at least the code region of the circularized recognition element, will participate in bridge amplification for use in a sequence reaction. Once the sequence of the code is determined, the data can be used to identify the presence of the target nucleic acid from the sample.
[0228] FIG. 9 is an exemplary expanded view demonstrating how bridge amplification can be performed on an immobilized circular recognition element. The recognition element, as previously described, comprises sequencing specific sequences such as SBS primer binding sites, P5 and P7 complementary sequences, and the like. After target nucleic acid hybridization, ligation and circularization the circular recognition element is immobilized on the sequencing flowcell. Bridge amplification is performed by the addition of a DNA polymerase. Once the circular recognition element is copied via extension reaction from an immobilized probe sequence, the extended single stranded copy of the recognition element is able to participate in bridge amplification, thereby generating the colonies for determining the sequence of the codes, the data of which can be used in soft or hard decision decoding workflows to identify whether the target nucleic acid, either wild type or variant, was present in the original sample.
[0229] The encoded assays disclosed herein may be performed on a surface. For example, a target may be immobilized on a surface for conducting assays disclosed herein. The recognition elements disclosed herein may be immobilized on a surface for conducting assays disclosed herein. DNA amplification products disclosed herein may be immobilized on a surface for conducting assays disclosed herein. Various intermediate assemblies of molecules of the assays disclosed herein may be immobilized on a surface for conducting assays disclosed herein.
[0230] Various steps disclosed herein may be performed on a surface, such as target capture, recognition events, transformation events, amplification, and / or detection events, e.g., determination of the absence or presence of the code (e.g., by sequencing or hybridization -based detection). Thus, for example, the disclosure provides a surface having a recognition element as described herein immobilized on the surface. The disclosure provides a surface having an amplification product as described herein immobilized on the surface. The disclosure provides a surface having a target immobilized on the surface. The disclosure provides a surface having a target immobilized on the surface with a recognition element as described herein hybridized to the target. The disclosure provides a surface having a recognition element immobilized on the surface with a target as described herein hybridized to the recognition element. The disclosure provides a surface having a target nucleic acid immobilized on the surface, and a protein or peptide bound to the target nucleic acid. The disclosure provides a surface having a target nucleic acid immobilized on the surface, and an antibody, aptamer, binder, or antibody fragment bound to the target nucleic acid. The disclosure provides a surface having a ligand that has affinity for any of the foregoing immobilized on the surface. For example, the ligand may have affinity for a recognition element as described herein, an amplification product as described herein, or a target as described herein. The ligand may, for example, be a protein, peptide, antibody, aptamer, binder, or antibody fragment.
[0231] A variety of surfaces may be used for the surface attachments described herein . In various embodiments, the surface includes an oxide, a nitride, a metal, an organic or an inorganic polymer (e.g., hydrogel, resin, plastic or other).
[0232] The surface may take a variety of forms, e.g., it may be flat or curved. It may be heads or particles. In some cases, the surface is the surface of a flow cell. Beads or other particles may in some embodiments range in size from less than 100 nm up to several millimeters.
[0233] Various surface modifications may be used to permit attachment of various components of the assays disclosed herein to a surface. For example, various anchoring ligands may be used (e.g., streptavidin, biotin, aptamers, antibodies, etc.). Chemical handles, such as click chemistry handles, may be used. Examples include azides, alkynes, unsaturated bonds, amines, carboxylic acids, NHS, DBCO, BCN, tetrazine, epoxy and the like. Single- or double-stranded oligonucleotides may be used. Size ranges of the oligonucleotides may, in some cases, be from about 10 to about 200 nucleotides. Proteins or peptides may be used for surface attachment. Charge-based molecules or polymers may be used, e.g., polyethylenimine.
[0234] Various techniques may be used to prepare a surface for binding to a target or to a component of an assay disclosed herein. In one example, a flow cell with primers may be used. A splint DNA segment that comprises a segment complementary to the primer and a segment that is complementary to the target, or the component of the assay may be hybridized to the primer. A variety of splints may be used on a surface, with various subsets of the splints having different segments complementary to different components disclosed herein or different targets. Specific splints may be arranged on different regions of a surface. For example, splints may be arranged in a manner that permits the identification of distinct regions of a surface targeted to specific targets or components of the assays.
[0235] In various embodiments, amplification of a nucleic acid may occur on the surface. The nucleic acid may be a target or any nucleic acid component of an assay disclosed herein. For example, a target may be amplified on a surface, or a recognition element disclosed herein may be amplified on a surface, and / or a fragment of any of the foregoing may be amplified on a surface. The amplification maybe performed on a bead or particle, or on a flat surface, such as on the surface of a flow cell.
[0236] It should also be noted that DNA may be amplified in solution, e.g., in an aqueous suspension or emulsion, such as in microdroplets. Solution-based amplification may be performed, for example, in an open environment, such as the well of the microtiter plate, in a nanowell, or in an enclosed space, droplet in an emulsion, or on a flow cell or other microfluidic device.
[0237] Amplification may be by any method of amplification, including for example, PCR, isothermal amplification and / or other amplification methodology.
[0238] Attachment for immobilization of components of the assays or of targets may be covalent or non-covalent (e.g., Coulombic in nature), temporary or permanent, and / or rendered labile when subject to a particular stimulus.
[0239] Examples of mechanisms of lability include:
[0240] • Enzymatic - protease, restriction endonuclease, CRISPR-Cas9
[0241] • Chemical - reduction, hydrolysis, nucleophilic attack, displacement, reducing of a disulfide bond
[0242] • Temperature - melting of duplexed hybridized DNA, thermodynamically unfavorable conditions (Positive deltaG)
[0243] • pH - hydrazone, carbonate, etc.
[0244] • Light - O-nitrobenzyl or derivatives where absorption of light of a particular wavelength(s) can cause bond rearrangements or cleavage. Light sensitive groups include nitro-benzene derivatives
[0245] • Ligand mediated - competitive competition for binding site (see examples below) o Peptide-tagged oligos with protein interactions - e.g., Spy-catcher. The moiety may be the ligand or the protein. o Peptide-tagged oligo with heavy metal interactions - e.g., Hexa-histidine - to Cu. The moiety may be the ligand or the protein. o CLIP tag and SNAP tag pair - e.g., O6-benzylguanine derivatives binding to 06- alkylguanine-DNA-alkyltransf erase. Either the protein or the substrate may be bound to the oligo. o Carbohydrate-protein pairs, e.g., lectins o The moiety may be a ligand (e.g., biotin, digoxigenin) bound to a fluorescently - tagged protein (e.g., avidin, streptavidin, DIG-binding protein)
[0246] • Cleavage can be performed by cleaving a moiety dangling on a nucleotide, or a nucleotide ora nucleobase within the oligo sequence or the di-nucleotide linkage, e.g., uracil and USER cocktail (uracil-N-deglycosylase (UNG)) followed by Endonuclease VIII or FPG (Formamidopyrimidine DNA Glycosylase with bifunctional DNA glycosylase with DNA N-glycosylase and AP lyase activities)
[0247] • Cleavage can be performed by an enzyme
[0248] A variety of surface-based workflows are possible within the scope of the assays as disclosed. In some embodiments, a surface-based workflow may use a recognition element that includes a recognition element associated with a code. The code may be a soft decodable code, such as a trellis code. In some embodiments, a surface-based workflow may use a dual recognition element that includes a recognition element associated with a code (e.g., a trellis code).
[0249] In some embodiments, a surface-based workflow may include immobilizing a target on a surface and hybridizing a recognition element to the target. In one embodiment, a surface-based workflow may include:
[0250] (i) immobilizing the target on a surface;
[0251] (ii) hybridizing a recognition element to the immobilized target;
[0252] (iii) circularizing the recognition element to produce a circular modified recognition element; and
[0253] (iv) releasing the circular modified recognition element from the target.
[0254] In some cases, the RCA reaction may be performed in a solution that remains in contact with the surface on which the target is immobilized (e.g., in the same container, well, reservoir, liquid volume or droplet). In some cases, the solution comprising the released modified recognition element may be transferred to a separate container prior to performing the RCA reaction. In some cases, the solution comprising the released modified recognition element may be transferred to a different surface prior to performing the RCA reaction. In some embodiments, the immobilized target (e.g., DNA) may be used to prime the RCA reaction. In one embodiment, a surface-based workflow may include:
[0255] (i) immobilizing the target on a surface;
[0256] (ii) hybridizing a recognition element to the target;
[0257] (iii) circularizing the recognition element to produce a circular modified recognition element; and
[0258] (iv) using the target to prime an RCA reaction to generate an amplification product.
[0259] In some embodiments, a surface-based workflow may include immobilizing a recognition element (or a part thereof) on a surface and using the immobilized recognition element to capture a target. In one embodiment, a surface-based workflow may include:
[0260] (i) immobilizing the recognition element (or a part thereof) on a surface;
[0261] (ii) hybridizing a target to the recognition element;
[0262] (iii) circularizing the recognition element to produce a circular modified recognition element; and
[0263] (iv) using the target to prime an RCA reaction to generate an amplification product.
[0264] In some embodiments, the circular modified recognition element may be released from the surface prior to amplification. In some cases, the RCA reaction may be performed in a solution that remains in contact with the surface on which the recognition element was anchored (e.g., in the same container, well, reservoir, liquid volume or droplet). In some cases, the solution comprising the released modified recognition element may be transferred to a separate container prior to performing the RCA reaction.
[0265] In some embodiments, the solution comprising the released modified recognition element may be transferred to a different surface prior to performing the RCA reaction . In one embodiment, oligonucleotides bound to the new surface may be used as capture moieties to immobilize the circular modified recognition element on the surface and to initiate the amplification reaction. In one embodiment, the target may be immobilized on the new surface and used to initiate the amplification reaction.
[0266] A surface-based workflow may use a dual recognition element as a recognition element. In one embodiment, a surface-based workflow using a dual recognition element may include:
[0267] (i) hybridizing a target to a first recognition element;
[0268] (ii) hybridizing the target to a second recognition element; and
[0269] (iii) performing a ligation or a gap -fill ligation reaction to link the first recognition element and the second recognition element. In some embodiments, the first recognition element and the second recognition element may both be immobilized on the surface. In some embodiments, the first recognition element is immobilized on the surface and the second recognition element is in solution. The surface may, for example, be the surface of a flow cell.
[0270] Assays disclosed herein may be used to interrogate the methylation status of a target sequence of interest. In one embodiment, methylated cytosines in a target sequence of interest may be detected using assays that include a bisulfite conversion reaction to detect methylated cytosines. In another embodiment, methylated cytosines in a target sequence of interest may be detected using assays that do not use a conversion reaction (e.g., conversion-free).
[0271] In one embodiment of a conversion assay for detection of methylated cytosines, a bisulfite conversion reaction that converts non -methylated cytosines to thymine (C —> T) may be used.
[0272] For example, a methylated cytosine assay using encoded recognition elements may include: (i) a bisulfite conversion reaction to convert non -methylated cytosine to thymine (C —> T); (ii) a recognition event, in which a target nucleic acid is uniquely recognized and bound by a recognition element associated with a code (e.g., an encoded recognition element); (ii) a transformation event, in which a molecular transformation of the recognition element produces a modified recognition element comprising the code; and (iii) a detection event, that uses the code as a surrogate for detection of the target nucleic acid, e.g., by recognizing or decoding code (and optionally other elements).
[0273] In some embodiments, a methylated target site of interest may be interrogated using an encoded recognition element in combination with a transformation event that includes a ligation reaction to detect the methylation status of the target site.
[0274] In one embodiment, the recognition element (e.g., an encoded recognition element) may be a coded recognition element that includes a 3 '-terminal guanine (“G”). The transformation event (e.g., ligation) to generate the modified recognition element may only occur when the 3 guanine is matched to a cytosine at a target site of interest.
[0275] In one example, a DNA sample may include a target sequence of interest that may be methylated or unmethylated at a CpG site of interest. A bisulfite conversion reaction is used to convert non -methylated cytosine to thymine (C —> T) in the target sequence.
[0276] In the recognition event, a target sequence may be recognized and may be bound by a recognition element associated with a code, e.g., a recognition element. A recognition element may include a 3 '-terminal G nucleotide that base pairs with the target C at the CpG site of interest. In one embodiment, in the transformation event, a ligation of a recognition element occurs only when the 3 '-terminus of the recognition element (e.g., a guanine “G”) is matched to the target site “C” of interest in a target sequence to generate a circularized modified recognition element. If no ligation occurs at the target site “T” in the bisulfite converted target sequence, transformation of a recognition element hybridized to target sequence to a circular modified recognition element may not occur. As described above with reference to FIG. 10, a modified recognition element may be amplified in an RCA reaction to generate an amplification product (step 1035) comprising many copies of the code (among other elements) and the code may be decoded (step 1040).
[0277] In one embodiment, the recognition element (e.g., encoded recognition element) may be a molecular inversion probe that includes a 3 '-terminal single base gap at a target site of interest. A gap-fill ligation event using only a single added nucleotide may be used to generate the modified recognition element comprising the code only when the nucleotide corresponding to the target site of interest is incorporated. This approach provides two forms of specificity to the assay: (i) the 3 '-terminus of the probe must recognize and bind the interrogated site; and (ii) a single base extension reaction that incorporates the nucleotide corresponding to the target site of interest occurs.
[0278] In one example, a DNA sample may include a target sequence of interest that may be methylated or unmethylated at a CpG site of interest. A bisulfite conversion reaction may be used to convert non-methylated cytosine to thymine (C —> T) in the target sequence.
[0279] In the recognition event, two target sequences may be recognized and bound by a recognition element associated with a code, e.g., a molecular inversion probe. A molecular inversion probe may include a single 3 '-terminal base gap that spans a target site of interest.
[0280] In the transformation event, a single dGTP nucleotide (“G”) may be incorporated in a molecular inversion probe, thereby allowing ligation of the probe to generate a circularized modified probe. If no incorporation of dGTP occurs at the target site “T” in the bisulfite converted target sequence, transformation of the molecular inversion probe hybridized to the target sequence to a circular modified probe may not occur. As described above with reference to FIG. 10, a modified probe may be amplified in an RCA reaction to generate an amplification product (step 1035) comprising many copies of the code (among other elements) and the code may be decoded (step 1040).
[0281] In one embodiment, the recognition element (e.g., a molecular inversion probe) may be designed to target two methylated cytosine sites of interest in a target sequence of interest. A gap-fill ligation event using all dNTPs may be used to generate the modified recognition element comprisingthe code. In this approach, both methylated cytosines mustbe present in the target nucleic acid molecule for ligation to occur. The requirement for multiple matches has several advantages: (i) it provides enhanced specificity relative to a single match at a methylated cytosine; (ii) the ability to discriminate between a disease state (e.g., all CpG sites in a region are methylated) and a healthy state (e.g., only some CpG sites are methylated) is increased by requiring multiple methylated cytosines for detection; and (iii) multiple matches can be used to correct for incomplete bisulfite conversion of unmethylated cytosines at the target site of interest.
[0282] A recognition element may target two methylated cytosines in a target of interest. A DNA sample may include a target sequence of interest that may be methylated at multiple CpG sites. A bisulfite conversion reaction may be used to convert non-methylated cytosine to thymine (C —> T) in the target sequence (not shown).
[0283] In the recognition event, a target sequence may be recognized and bound by a recognition element associated with a code, e.g., a molecular inversion probe. The molecular inversion probe may include a 3 '-probe arm that terminates at a first methylated cytosine site and a 5 '-probe arm that terminates at a second methylated cytosine site. Both a 3 '-GC match and a 5'-GC match during the recognition event (hybridization) may be required for a transformation event to occur.
[0284] In the transformation event, a gap-fill ligation reaction using all dNTPs may be performed. The 3'-GC match may be required for polymerase extension in the gap-fill reaction. The 5'-GC match may be required for ligation of the gap-filled molecule. Gap-fill ligation may generate a circularized modified recognition element. No incorporation of dGTP occurs at the target site “T” in the bisulfite converted target sequence and consequently, transformation to a circular modified recognition element does not occur in non-methylated target sequences. As described above with reference to FIG. 10, a modified recognition element may be amplified in an RCA reaction to generate an amplification product (step 1035) comprising many copies of the code (among other elements) and the code may be decoded (step 1040).
[0285] The assays disclosed herein may be used in a genotyping assay. A target site of interest may be interrogated using an encoded recognition element in combination with a ligation reaction to detect a single nucleotide variant (SNV) of interest. In one example, the single nucleotide change may be a single nucleotide polymorphism (SNP).
[0286] In one embodiment, a genotyping assay using encoded recognition elements may include: (i) a recognition event, in which a target nucleic acid is uniquely recognized and bound by a recognition element associated with a code (e.g., an encoded recognition element); (ii) a transformation event, in which a molecular transformation of the recognition element produces a modified recognition element comprising the code; and (iii) a detection event, that uses the code as a surrogate for detection of the target nucleic acid, e.g., by recognizing or decoding code (and optionally other elements).
[0287] In one embodiment, the recognition element (e.g., an encoded recognition element) may be a coded recognition element that includes a 3 '-terminal nucleotide that is matched to a SNV of interest. The transformation event (e.g., ligation) to generate the modified recognition element may only occur when the 3 nucleotide is matched to the SNV at the target site of interest.
[0288] In one embodiment, the recognition element (e.g., an encoded recognition element) may be a molecular inversion probe that includes a 3 '-terminal single base gap at a target site of interest. A gap-fill ligation event using only a single added nucleotide may then be used to generate the modified recognition element comprising the code only when corresponding nucleotide is incorporated.
[0289] The assays disclosed herein may be used in an RNA analysis assay. In one embodiment, an RNA assay using encoded recognition elements may include: (i) a reverse transcription reaction to convert RNA (e.g., mRNA) to cDNA; (ii) a recognition event, in which a target nucleic acid cDNA is uniquely recognized and bound by a recognition element associated with a code (e.g., an encoded recognition element); (ii) a transformation event, in which a molecular transformation of the recognition element produces a modified recognition element comprising the code; and (iii) a detection event, that uses the code as a surrogate for detection of the cDNA target nucleic acid and therefore the mRNA, e.g., by recognizing or decoding code (and optionally other elements).
[0290] In some cases, the reverse transcription step (i) may be omitted and a ligase tolerant to DNA-RNA hybrid duplexes may be used in the transformation event. In one example, the ligase is a Chlorella virus ligase such as SplintR® ligase (New England BioLabs).
[0291] In one embodiment, the encoded recognition element may be a recognition element that comprises a recognition element associated with a code. In one embodiment, the encoded recognition element may be a molecular inversion probe that comprises a recognition element associated with a code. In one embodiment, the encoded recognition element may be a padlock probe that comprises a recognition element associated with a code.
[0292] Assays disclosed herein may be used to detect and count RNA targets of interest in a sample. Assays disclosed herein may be used to detect alternative splicing variants for a target of interest. In one example, splicing variants may be identified by placing one half of a recognition element (e.g., a coded recognition element) on either side of the splice site. The transformation event (e.g., ligation) to generate the modified recognition element may only occur when the 3 nucleotide is matched to the splice variant at the target site of interest. In another example, splicing variants maybe identified using a molecular inversion probe and an extension ligation reaction, wherein one probe arm spans the splice site.
[0293] In a sequencing workflow provided herein, detectable nucleotides or nucleotide analogs used in the sequencing workflow may comprise A, C, G, and T, and each of the nucleotides or nucleotide analogs may comprise a different, detectable fluorescent label. In some embodiments, a sequencing reaction may comprise N-states, each N-state corresponding to the detection of a binding event of a different, detectable nucleotide or nucleotide analog to a modified recognition elements, wherein N is a number of states. In some embodiments, a sequencing reaction may comprise 4-states, each state corresponding to the binding of a different, detectable nucleotide or nucleotide analog to a modified recognition elements. In some embodiments, a sequencing reaction may comprise N-states, each state corresponding to the binding of a different, detectable nucleotide or nucleotide analog to a modified recognition elements, wherein N is a number of states. In some embodiments provided herein, the number of states (N) is 8, 7, 6, 5, 4, 3, or 2. In some embodiments, an N-state sequencing reaction comprises nucleotides or nucleotide analogs comprising a reversible terminator. In some embodiments, an N-state sequencing reaction comprises nucleotides or nucleotide analogs that do not comprise a reversible terminator moiety.
[0294] In some embodiments, the detectable nucleotides or nucleotide analog present in a cycle differ between cycles. For example, in a first cycle, detectable A and G nucleotides or nucleotide analogs are present, and detectable C and T nucleotides or nucleotide analogs are present in a second cycle. The detectable nucleotides or nucleotide analogs used in a cycle may alternate or differ between cycles. A single detectable nucleotide or nucleotide analog may be present in a cycle. An unlabeled nucleotide or nucleotide analog may be present in a flow or cycle. A flow or cycle may comprise an unlabeled nucleotide or nucleotide analog. A flow or cycle may comprise a detectable nucleotide or nucleotide analog and an unlabeled nucleotide or nucleotide analog. For example, in a first cycle, a detectable A nucleotide or nucleotide analog is used; in a second cycle, a detectable C nucleotide or nucleotide analog is used; in a third cycle, a detectable G nucleotide or nucleotide analog is used; and in a fourth cycle, a detectable T nucleotide or nucleotide analog is used. In a sequencing reaction, the pattern of detectable nucleotide or nucleotide analogs used in a flow may repeat. For example, in an 8 -cycle sequencing reaction, the detectable nucleotide or nucleotide analogs used in each cycle may be A, C, G, T, A, C, G, and T. In a second example of an 8-cycle sequencing reaction, the detectable nucleotide or nucleotide analogs used in each cycle may be A+C, G+T, A+C, G+T, A+C, G+T, A+C, and G+T. In a third example of an 8 -cycle sequencing reaction, the detectable nucleotide or nucleotide analogs used in each cycle may be A+C+G+T, A+C+G+T, A+C+G+T, A+C+G+T, A+C+G+T, A+C+G+T, A+C+G+T, and A+C+G+T.
[0295] In some embodiments, a state comprises the intensity of a signal in a color channel. For example, a higher intensity signal may indicate the binding of a plurality of a species of nucleotides or nucleotide analogs to a modified recognition sequence. The intensity of a signal in a color channel may indicate the number of nucleotides or nucleotide analogs bound to a modified recognition element in a cycle of a sequencing reaction. For example, a relatively high signal may indicate two nucleotide or nucleotide analogs bound to a modified recognition element, and a relatively low signal may indicate a single nucleotide or nucleotide analog bound to a modified recognition element.
[0296] In some embodiments, a sequencing reaction may comprise 4 -states, each state corresponding to the detection of a binding event of a different, detectable nucleotide or nucleotide analog to a modified recognition elements. In some embodiments, a sequencing reaction may comprise 4-states, each state corresponding to the binding of a different, detectable nucleotide or nucleotide analog to a modified recognition elements. For example, in a sequencing workflow provided herein, a set of detectable nucleotides or nucleotide analogs may comprise A, C, G and T, and wherein the detectable label for each of A, C, G and T is different. In a first state, a detectable A nucleotide or nucleotide analog may be bound to a modified recognition element. In a second state, a detectable C nucleotide or nucleotide analog may be bound to a modified recognition element. In a third state, a detectable G nucleotide or nucleotide analog may be bound to a modified recognition element. In a fourth state, a detectable T nucleotide or nucleotide analog may be bound to a modified recognition element. In an embodiment, detectable A, C, T, and G nucleotides or nucleotide analogs are all present in a single sequencing flow or cycle. In some embodiments, binding events between detectable A, C, T, and G nucleotides and a plurality of coded recognition elements may be detected in a single cycle or simultaneously. In some embodiments, binding events between detectable A, C, T, and G nucleotides and a plurality coded recognition elements may occur or exist in a single cycle or simultaneously. In some embodiments, a 4-state sequencing reaction comprises nucleotides or nucleotide analogs comprising a reversible terminator. In some embodiments, a 4-state sequencing reaction comprises nucleotides or nucleotide analogs that do not comprise a reversible terminator. In some embodiments, a 4-state sequencing reaction comprises unlabeled nucleotides or nucleotide analogs. In some embodiments, a sequencing reaction may comprise 3 -states, each state corresponding to the detection (or lack of detection) of a binding event of a different, detectable nucleotide or nucleotide analog to a modified recognition elements. In some embodiments, the set of detectable nucleotides or nucleotide analogs used in a sequencing workflow may comprise two or more species of A, C, G and T, wherein the detectable label for each of the two or more species is different. For example, in a first state, a nucleotide or nucleotide analog that is unlabeled binds to a modified recognition element, and a binding event is not detected. In a second state, a detectable A nucleotide or nucleotide analog may be bound to a modified recognition element. In a second state, a detectable G nucleotide or nucleotide analog may be bound to a modified recognition element. In some embodiments, a first state comprises an unlabeled nucleotide or nucleotide analog bound to a modified recognition element, a second state comprises a detectable C nucleotide or nucleotide analog bound to a recognition element, and a third state comprises a detectable T nucleotide or nucleotide element bound to a recognition element. In some embodiments, a 3 -state sequencing reaction comprises nucleotides or nucleotide analogs comprising a reversible terminator. In some embodiments, a 3 -state sequencing reaction comprises nucleotides or nucleotide analogs that do not comprise a reversible terminator. In some embodiments, a 3 -state sequencing reaction comprises unlabeled nucleotides or nucleotide analogs.
[0297] In some embodiments, a sequencing reaction may comprise 2-states, each state corresponding to the detection (or lack of detection) of a binding event of a different, detectable nucleotide or nucleotide analog to a modified recognition element. In some embodiments, the set of detectable nucleotides or nucleotide analogs used in a sequencing workflow may comprise one species of nucleotides or nucleotide analogs (e.g., A, C, G, or T), wherein the detectable label for each of the species is the same. In some embodiments, the set of detectable nucleotides or nucleotide analogs used in a sequencing workflow may comprise a single species of nucleotides or nucleotide analogs (e.g., A, C, G, and T), wherein the detectable label for each of the species is different. In some embodiments, in a state (e.g., a first state or a second state), a binding event occurs between an unlabeled nucleotide or nucleotide analog and a modified recognition element, and the binding event is not detected. In a first state, in some embodiments, a binding event does not occur between a detectable nucleotide or a nucleotide analog and a modified recognition element (OFF-state). In a second state, a detectable nucleotide or nucleotide analog (e.g., a detectable A, C, G, or T) is bound to a modified recognition element (ON-state). In some embodiments, a 2-state sequencing reaction comprises nucleotides or nucleotide analogs comprising a reversible terminator. In some embodiments, a 2-state sequencing reaction comprises nucleotides or nucleotide analogs that do not comprise a reversible terminator. In some embodiments, a 2-state sequencing reaction comprises unlabeled nucleotides or nucleotide analogs.
[0298] The present disclosure provides systems and methods for detection of one or more target molecules from a sample. A sample may include nucleic acids from tissue(s) which nucleic acids may be extracted using the techniques described herein. Non-limiting examples of tissues include solid tissue, lysed solid tissue, fixed tissue samples, whole blood, plasma, serum, dried blood spots, buccal swabs, other forensic samples, fresh or frozen tissue, biopsy tissue, organ tissue, cultured or harvested cells, and bodily fluids.
[0299] The sample may include a biological sample, such as whole blood, lymphatic fluid, serum, plasma, sweat, tear, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, fluidized tissues, fluidized organisms, liquids containing multi-celled organisms, biological swabs or biological washes. Samples may be provided directly from biological sources, or may be processed samples, such as samples which are enriched for targets, nucleic acids, or proteins from any of the foregoing sources.
[0300] The methods disclosed herein may be used for screening or diagnosing a subject for a disease, such as cancer or for selecting a therapy for treating a disease, such as selecting a therapy for treating a cancer. In one embodiment, the methods disclosed herein may be used in a liquid biopsy application. In one example, a liquid biopsy assay may include determination of the methylation status and / or the variant usage of a set of target sequences. In one embodiment, the methods disclosed herein may be used in a pathogen detection application. In one example, pathogen detection may include detecting both a protein and nucleic acid (e.g., an RNA) associated with the pathogen. In one embodiment, the methods disclosed herein may be used to monitor and / or determine complications associated with a transplantation procedure.
[0301] Provided herein are encoded assays that determine a presence of a code to predict a presence of one or more target molecules. A recognition element provided herein may comprise a code. A code may comprise a plurality of segments, wherein each segment is unique from other segments in the code. A segment of a code may have a length of one nucleotide. The length of a segment may be greater than or equal to about 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. The length of a segment may be fewer than or equal to about 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. In some embodiments, codes in a set of coded recognition elements are the same length. In some embodiments, codes in a set of coded recognition elements are different lengths. A code may have a length of N-nucleotides or nucleotide analogs, where N is a positive integer. In some embodiments, a code has a length that is 30 or fewer nucleotides. In some embodiments, a code has a length that is 25 or fewer nucleotides. In some embodiments, a code has a length that is 24 or fewer nucleotides. In some embodiments, a code has a length that is 23 or fewer nucleotides. In some embodiments, a code has a length that is 22 or fewer nucleotides. In some embodiments, a code has a length that is 21 or fewer nucleotides. In some embodiments, a code has a length that is 19 or fewer nucleotides. In some embodiments, a code has a length that is 30 or fewer nucleotides. In some embodiments, a code has a length that is 18 or fewer nucleotides. In some embodiments, a code has a length that is 17 or fewer nucleotides. In some embodiments, a code has a length that is 16 or fewer nucleotides. In some embodiments, a code has a length that is 15 or fewer nucleotides. In some embodiments, a code has a length that is 14 or fewer nucleotides. In some embodiments, a code has a length that is 13 or fewer nucleotides. In some embodiments, a code has a length that is 12 or fewer nucleotides. In some embodiments, a code has a length that is 11 or fewer nucleotides. In some embodiments, a code has a length that is 10 or fewer nucleotides. In some embodiments, a code has a length that is 9 or fewer nucleotides. In some embodiments, a code has a length that is 8 or fewer nucleotides. In some embodiments, a code has a length that is 7 or fewer nucleotides. In some embodiments, a code has a length that is 6 or fewer nucleotides. In some embodiments, a code has a length that is 5 or fewer nucleotides. In some embodiments, a code has a length that is 4 or fewer nucleotides. In some embodiments, a code has a length that is 3 or fewer nucleotides. In some embodiments, a code has a length that is 2 or fewer nucleotides.
[0302] A code in a codespace may comprise symbols. The symbols comprised in a code in a codespace may correspond to or represent a sequence of nucleotides. A codespace may comprise codes comprising 5 or fewer unique symbols (e.g., digits; e.g., 0, 1, 2, 3, 4). A codespace may comprise codes comprising 4 or fewer unique symbols (e.g., 1, 2, 3, 4; e.g., A, C, G, T). A codespace may comprise codes comprising 3 or fewer unique symbols (e.g., 0, 1, 2; e.g., OFF, A, G). A codespace may comprise codes comprising 2 or fewer unique symbols (e.g., 0, 1; e.g., OFF, ON).
[0303] In some embodiments, a plurality of codes comprises a codespace. A codespace may comprise a set of codes. A codespace may be a snug codespace, e.g., a code may have a minimum-permissible hamming-distance from a maximum number of codes in a codespace. Codes in a codespace may be chosen based in part on a hamming-distance (e.g., a number of positions at which corresponding symbols are different) between codes. A hamming-distance may be 10. A hamming-distance may be 8. A hamming-distance may be 5. A hamming-distance may be 4. A hamming-distance may be 3. A hamming-distance may be 2. A hamming-distance may be 1 . A codespace may comprise a set of codes, such that the hamming-distance between any two codes in a code space is the same. A codespace may comprise a set of codes, such that there are a plurality of hamming-distances between sets of two codes in a codespace.
[0304] The methods and systems disclosed herein may comprise soft decoding to predict the presence of the code in a modified recognition element of amplification product thereof. Disclosed herein are methods and systems that record and obtain a signal produced in response to interrogation of each segment of a code from an encoded assay of the present disclosure. Upon completion of the interrogation of each segment, the methods and system disclosed herein determine a probability of the presence of each of the codes by applying a soft-decision probabilistic decoding algorithm to the recorded signal, wherein detecting the presence of the code is indicative of the presence of the target.
[0305] In some embodiments, a soft decoding algorithm decodes at least 5% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 10% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 15% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 20% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 25% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 30% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 35% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 40% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 45% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 50% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully. In some embodiments, a soft decoding algorithm decodes at least 55% of the codes in amplified recognition elements of a plurality of amplified recognition elements successfully.
[0306] In some embodiments, successfully decoding a code may comprise detecting the presence of a code (e.g., probabilistically assigning a code known to be present in a codespace to a signal) that is known to be part of a codespace. In some embodiments, a signal is not successfully decoded if a code is probabilistically assigned a sequence that is not part of a codespace (e.g., unused error).
[0307] The systems and methods of the present disclosure comprise computer systems. Referring to FIG. 25, a block diagram is shown depicting an exemplary machine that includes a computer system 2500 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in FIG. 25 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0308] Computer system 2500 may include one or more processors 2501, a memory 2503, and a storage 2508 that communicate with each other, and with other components, via a bus 2540. The bus 2540 may also link a display 2532, one or more input devices 2533 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 2534, one or more storage devices 2535, and various tangible storage media 2536. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 2540. For instance, the various tangible storage media 2536 can interface with the bus 2540 via storage medium interface 2526. Computer system 2500 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0309] Computer system 2500 includes one or more processor(s) 2501 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Processor(s) 2501 optionally contains a cache memory unit 2502 for temporary local storage of instructions, data, or computer addresses. Processor(s) 2501 are configured to assist in execution of computer readable instructions. Computer system 2500 may provide functionality for the components depicted in FIG. 25 as a result of the processor(s) 2501 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 2503, storage 2508, storage devices 2535, and / or storage medium 2536. The computer-readable media may store software that implements particular embodiments, and processor(s) 2501 may execute the software. Memory 2503 may read the software from one or more other computer-readable media (such as mass storage device(s) 2535, 2536) or from one or more other sources through a suitable interface, such as network interface 2520. The software may cause processor(s) 2501 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 2503 and modifying the data structures as directed by the software.
[0310] The processor may be comprised of any of a variety of suitable integrated circuits, microprocessors, logic devices, field-programmable gate arrays (FPGAs) and the like. In some instances, the processor may be a single core or multi core processor, or a plurality of processors may be configured for parallel processing. Although the disclosure is described with reference to a processor, other types of integrated circuits and logic devices are also applicable.
[0311] The memory 2503 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 2504) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase - change random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 2505), and any combinations thereof. ROM 2505 may act to communicate data and instructions unidirectionally to processor(s) 2501, and RAM 2504 may act to communicate data and instructions bidirectionally with processor(s) 2501. ROM 2505 and RAM 2504 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 2506 (BIOS), including basic routines that help to transfer information between elements within computer system 2500, such as during start-up, may be stored in the memory 2503.
[0312] Fixed storage 2508 is connected bidirectionally to processor(s) 2501, optionally through storage control unit 2507. Fixed storage 2508 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 2508 may be used to store operating system 2509, executable(s) 2510, data 2511, applications 2512 (application programs), and the like. Storage 2508 can also include an optical disk drive, a solid- state memory device (e.g., flash -based systems), or a combination of any of the above. Information in storage 2508 may, in appropriate cases, be incorporated as virtual memory in memory 2503.
[0313] In one example, storage device(s) 2535 may be removably interfaced with computer system 2500 (e.g., via an external port connector (not shown)) via a storage device interface 2525. Particularly, storage device(s) 2535 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 2500. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 2535. In another example, software may reside, completely or partially, within processor(s) 2501
[0314] Bus 2540 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 2540 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI -Express (PCLX) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0315] Computer system 2500 may also include an input device 2533. In one example, a user of computer system 2500 may enter commands and / or other information into computer system 2500 via input device(s) 2533. Examples of an input device(s) 2533 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 2533 may be interfaced to bus 2540 via any of a variety of input interfaces 2523 (e.g., input interface 2523) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0316] In particular embodiments, when computer system 2500 is connected to network 2530, computer system 2500 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 2530. Communications to and from computer system 2500 may be sent through network interface 2520. For example, network interface 2520 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 2530, and computer system 2500 may store the incoming communications in memory 2503 for processing. Computer system 2500 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 2503 and communicated to network 2530 from network interface 2520. Processor(s) 2501 may access these communication packets stored in memory 2503 for processing.
[0317] Examples of the network interface 2520 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 2530 or network segment 2530 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 2530, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0318] Information and data can be displayed through a display 2532. Examples of a display 2532 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 2532 can interface to the processors) 2501, memory 2503, and fixed storage 2508, as well as other devices, such as input device(s) 2533, via the bus 2540. The display 2532 is linked to the bus 2540 via a video interface 2522, and transport of data between the display 2532 and the bus 2540 can be controlled via the graphics control 2521. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non -limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0319] In addition to a display 2532, computer system 2500 may include one or more other peripheral output devices 2534 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 2540 via an output interface 2524. Examples of an output interface 2524 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof. In addition or as an alternative, computer system 2500 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0320] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0321] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0322] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is bound to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0323] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub -notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those of skill in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations, known to those of skill in the art.
[0324] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non -limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of nonlimiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX -like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non -limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Those of skill in the art will also recognize that suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Those of skill in the art will also recognize that suitable video game console operating systems include, by way of non -limiting examples, Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.
[0325] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non -transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semipermanently, or non -transitorily encoded on the media.
[0326] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages.
[0327] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.
[0328] In some embodiments, a computer program includes a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non -limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non -limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.
[0329] Referring to FIG. 26, in a particular embodiment, an application provision system comprises one or more databases 2600 accessed by a relational database management system (RDBMS) 2610. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBMDB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system further comprises one or more application severs 2620 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 2630 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 2640. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.
[0330] Referring to FIG. 27, in a particular embodiment, an application provision system alternatively has a distributed, cloud-based architecture 2700 and comprises elastically load balanced, auto-scaling web server resources 2710 and application server resources 2720 as well synchronously replicated databases 2730.
[0331] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.
[0332] In view of the disclosure provided herein, a mobile application is created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of non -limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, ObjectPascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0333] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non -limiting examples, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.
[0334] Those of skill in the art will recognize that several commercial forums are available for distribution of mobile applications including, by way of non -limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.
[0335] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non -limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.
[0336] In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Those of skill in the art will be familiar with several web browser plug-ins including, Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.
[0337] In view of the disclosure provided herein, those of skill in the art will recognize that several plug-in frameworks are available that enable development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.
[0338] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of nonlimiting examples, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, andKDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non- limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non -limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ browser.
[0339] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non- limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.
[0340] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of signals recorded in a method disclosed herein. In various embodiments, suitable databases include, by way of non -limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity -relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web -based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0341] Provided herein is a kit for conducting an assay for a set of targets. In some embodiments disclosed herein, a kit comprises a set of coded recognition elements, wherein each coded recognition element of the set comprises a target-specific binding site specific to a target of the set of targets and a code associated with the target. In some embodiments, a kit comprises a plurality of the detectable nucleotides or nucleotide analogs. In some embodiments, a kit comprises instructions for use of the kit in an assay, wherein the instructions comprise steps for performing any of the methods disclosed herein. In some embodiments, a kit comprises instructions for use of the kit in an assay, wherein the instructions comprise steps for performing any of the systems disclosed herein. In some embodiments, a kit comprises a plurality of the nucleotide or nucleotide analogs that are unlabeled. A kit disclosed herein may comprise: (a) a set of coded recognition elements, wherein each coded recognition element of the set comprises a target-specific binding site specific to a target of the set of targets and a code associated with the target; (b) a plurality of the detectable nucleotides or nucleotide analogs; (c) a plurality of the nucleotide or nucleotide analogs that are unlabeled; and (d) instructions for use of the kit in an assay, wherein the instructions comprise steps for performing any of the methods disclosed herein.
[0342] A kit disclosed herein may comprise a plurality of nucleotides or nucleotide analogs as described herein. In some embodiments, a kit disclosed herein comprises a nucleotide or nucleotide analog that may serve as a monomeric unit of a nucleic acid polymer, e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, a kit disclosed herein comprises a nucleotide or a nucleotide analog comprising a nucleobase (e.g., guanine (G), adenine (A), cytosine (C), thymine (T), or uracil (U)) and a phosphate group. A kit disclosed herein may comprise a plurality of nucleotides or nucleotide analogs comprising a plurality of species (e.g., A, C, G, T, or U). A kit provided herein may comprise a solution comprising a plurality of nucleotides or nucleotide analogs comprising a detectable label (e.g., an incorporation mix). A kit provided herein may comprise a solution comprising a plurality of nucleotides or nucleotide analogs that are unlabeled (e.g., an incorporation mix).
[0343] In some embodiments provided herein, a kit comprises a plurality of nucleotides or nucleotide analogs comprising a reversible terminator moiety. A reversible terminator moiety may be any reversible terminator moiety described herein. The reversible terminator moiety may prevent incorporation of an additional nucleotide or nucleotide analog in an N+l position, wherein N is the position of the nucleotide or nucleotide analog. A detectable nucleotide or nucleotide analog may comprise a reversible terminator moiety at a 3 ’ OH position. The reversible terminator may comprise a 3 ’-O-alkyl hydroxylamino group, a 3 ’-phosphorothioate group, a 3 ’-O-malonyl group, a 3’-O-benzyl group, a 3 ’-O-Azidomethyl group, a 3’-O-Azide group, a 3 ’-O-Azido group, or a 3 ’-O-methyl group.
[0344] In some embodiments provided herein, a kit comprises a plurality of nucleotides or nucleotide analogs comprising a detectable label, e.g., an optical label or a fluorescent label. A detectable label may be any detectable label described herein. A detectable label may emit a signal (e.g., a chemical signal or an optical signal). The signal emitted by a detectable label may be easily detected or distinguishable. The detectable label may be a fluorescent dye. The fluorescent dye may emit red, far-red, near-red, yellow, green, or blue light. In some embodiments provided herein, a kit comprises a plurality of nucleotides or nucleotide analogs that do not comprise a detectable label. In some embodiments provided herein, a kit comprises a plurality of nucleotides or nucleotide analogs that are unlabeled.
[0345] In some embodiments, the label is a fluorophore. Non-limiting examples of fluorescent moieties include, but are not limited to, fluorescein and fluorescein derivatives such as carboxyfluorescein, tetrachlorofluorescein, hexachlorofluorescein, carboxynapthofluorescein, fluorescein isothiocyanate, NHS-fluorescein, iodoacetamidofluorescein, fluorescein maleimide, SAMSA-fluorescein, fluorescein thiosemicarbazide, carbohydrazinomethylthioacetyl -amino fluorescein, rhodamine and rhodamine derivatives such as TRITC, TMR, lissamine rhodamine, Texas Red, rhodamine B, rhodamine 6G, rhodamine 10, NHS-rhodamine, TMR-iodoacetamide, lissamine rhodamine B sulfonyl chloride, lissamine rhodamine B sulfonyl hydrazine, Texas Red sulfonyl chloride, Texas Red hydrazide, coumarin and coumarin derivatives such as AMCA, AMCA-NHS, AMCA-sulfo-NHS, AMCA-HPDP, DCIA, AMCE-hydrazide, BODIPY and derivatives such as BODIPY FL C3-SE, BODIPY 530 / 550 C3, BODIPY 530 / 550 C3-SE, BODIPY 530 / 550 C3 hydrazide, BODIPY 493 / 503 C3 hydrazide, BODIPY FL C3 hydrazide, BODIPY FL IA, BODIPY 530 / 551 IA, Br-BODIPY 493 / 503, Cascade Blue and derivatives such as Cascade Blue acetyl azide, Cascade Blue cadaverine, Cascade Blue ethylenediamine, Cascade Blue hydrazide, Lucifer Y ellow and derivatives such as Lucifer Y ellow iodoacetamide, Lucifer Yellow CH, cyanine and derivatives such as indolium based cyanine dyes, benzo - indolium based cyanine dyes, pyridium based cyanine dyes, thiozolium based cyanine dyes, quinolinium based cyanine dyes, imidazolium based cyanine dyes, Cy 3, Cy5, lanthanide chelates and derivatives such as BCPDA, TBP, TMT, BHHCT, BCOT, Europium chelates, Terbium chelates, Alexa Fluor dyes, DyLight dyes, Atto dyes, LightCycler Red dyes, CAL Flour dyes, JOE and derivatives thereof, Oregon Green dyes, WellRED dyes, IRD dyes, phycoerythrin and phycobilin dyes, Malachite green, stilbene, DEG dyes, NR dyes, nearinfrared dyes and others known in the art such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition; Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or Hermanson, Bioconjugate Techniques, 2nd Edition, or derivatives thereof, or any combination thereof. Cyanine dyes may exist in either sulfonated or non-sulfonated forms, and consist of two indolenin, benzo-indolium, pyridium, thiozolium, and / or quinolinium groups separated by a polymethine bridge between two nitrogen atoms. Commercially available cyanine fluorophores include, for example, Cy3, (which may comprise 1 -[6-(2,5-dioxopyrrolidin-l -yloxy)-6-oxohexyl]-2-(3 -{ 1 -[6-(2,5-dioxopyrrolidin-l- yloxy)-6-oxohexyl]-3,3-dimethyl-l,3-dihydro-2H-indol-2-ylidene}prop-l-en-l-yl)-3,3- dimethyl-3H-indolium or l -[6-(2,5-dioxopyrrolidin-l-yloxy)-6-oxohexyl]-2-(3-{ l-[6-(2,5- dioxopyrrolidin-l-yloxy)-6-oxohexyl]-3,3-dimethyl-5-sulfo-l,3-dihydro-2H-indol-2- ylidene}prop-l-en-l-yl)-3,3-dimethyl-3H-indolium-5-sulfonate), Cy5 (which may comprise 1 - (6-((2,5-dioxopyrrolidin-l-yl)oxy)-6-oxohexyl)-2-((lE,3E)-5-((E)-l-(6-((2,5-dioxopyrrolidin-l- yl)oxy)-6-oxohexyl)-3,3-dimethyl-5-indolin-2-ylidene)penta-l,3-dien-l-yl)-3,3-dimethyl-3H- indol-l-ium or l-(6-((2,5-dioxopyrrolidin-l-yl)oxy)-6-oxohexyl)-2-((lE,3E)-5-((E)-l-(6-((2,5- dioxopyrrolidin- 1 -yl)oxy)-6-oxohexy l)-3 ,3 -dimethyl-5 -sulfoindolin-2-ylidene)penta- 1 , 3-dien- 1 - yl)-3,3-dimethyl-3H-indol-l-ium-5-sulfonate), and Cy7 (which may comprise l -(5- carboxypentyl)-2-[(lE,3E,5E,7Z)-7-(l-ethyl-l,3-dihydro-2H-indol-2-ylidene)hepta-l,3,5-trien- l-yl]-3H-indolium or l-(5-carboxypentyl)-2-[(lE,3E,5E,7Z)-7-(l-ethyl-5-sulfo-l,3-dihydro-2H- indol-2-ylidene)hepta-l,3,5-trien-l-yl]-3H-indolium-5-sulfonate), where “Cy” stands for 'cyanine', and the first digit identifies the number of carbon atoms between two indolenine groups. Cy2 which is an oxazole derivative rather than indolenin, and the benzo-derivatized Cy3.5, Cy5.5 and Cy7.5 are exceptions to this rule.
[0346] In some embodiments provided herein, a kit comprises a surface. A surface may be any surface described herein. During a sequencing reaction, an amplification product may bind to the surface. The surface may comprise a plate or a welled plate. The surface may comprise a 96- well plate. The surface may comprise a glass surface. The surface may comprise a glass- bottomed, 96-well plate. The surface may comprise a cationic coating. The surface may comprise a polylysine coating. The surface may comprise a glass-bottomed, cationic-coated 96- well plate.
[0347] In some embodiments, a kit comprises a reaction vessel. In some embodiments, a reaction vessel is a flow cell. In some embodiments, a reaction vessel is a glass sheet of ~lmm thickness that has been coated with poly-L-lysine and is attached to a plastic well plate with an adhesive bottom. In some embodiments, a reaction vessel is a glass sheet of ~lmm thickness that has been coated with poly-L-lysine and is attached to a plastic well plate by compression of a silicon gasket matching the bottom structure of the well plate. The well plate may be composed of an array of circular or square features of varying depth and diameters of ~3 mm to ~6 mm.
[0348] In some embodiments, a kit comprises a flow cell. For example, a kit may comprise a flow cell used in next generation sequencing. In some embodiments, a kit comprises a flow cell and a reaction vessel. In some embodiments, a kit comprises a reaction vessel. In some embodiments, a kit comprises a flow cell comprising a surface comprising a poly-L-lysine coating. In some embodiments, a kit comprises a reaction vessel comprising a surface comprising a poly-L-lysine coating.
[0349] Provided herein are kits comprising polymerases. In some embodiments, a polymerase is a DNA-polymerase. A polymerase may be a Bst polymerase. A polymerase may be a Bst-like polymerase. For example, a kit may comprise a Therminator X polymerase . A kit may comprise a Bst3.0 polymerase. A kit may comprise an ArcticZymes polymerase. A kit may comprise a Bsm DNA polymerase. A kit may comprise Klenow (-exo). A kit may comprise Sequenase. A kit may comprise SNPase. A kit may comprise Bsu DNA polymerase. A kit may comprise Bst 2.0 DNA polymerase. A kit may comprise ArcticZymes IsoPol DNA polymerase. A kit may comprise ArcticZymes IsoPol SD+ DNA polymerase.
[0350] Provided herein are kits comprising a first reagent. In some embodiments, a kit provided herein comprises a plurality of reagents. A first reagent may be a buffer or a buffered solution. In some embodiments, a first reagent may cleave a detectable label from a nucleotide or a nucleotide analog. A first reagent may subject a modified recognition element to a deblocking condition sufficient to remove a detectable label from the detectable nucleotide or nucleotide moiety. A first reagent may cleave a reversible terminator moiety from a nucleotide or nucleotide analog (e.g., a detectable nucleotide or nucleotide analog, or an unlabeled nucleotide or nucleotide analog). A first reagent may cleave a reversible terminator from a nucleotide or a nucleotide analog. A first reagent may subject a modified recognition element to a deblocking condition sufficient to remove a detectable label from the detectable nucleotide or nucleotide moiety. A first reagent may comprise a buffer having one or more of (1) a phosphine compound comprising Tris(2-carboxyethyl)phosphine (TCEP), bis-sulfo triphenyl phosphine (BS-TPP) or Tri(hydroxyproyl)phosphine (THPP); (2) tetrakis(triphenylphosphine)palladium(0) (Pd(P(C6H5)3)4), (CEFjsNH (piperidine), or 2,3-Dichloro-5,6-dicyano-l,4-benzo-quinone (DDQ); (3) palladium on carbon (Pd / C); (4) a thiol group comprising beta-mercaptoethanol or dithiothritol (DTT); (5) potassium carbonate (K2CO3) in MeOH, triethylamine in pyridine, or with Zn in acetic acid (AcOH); (6) or tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
[0351] In some embodiments, a kit provided herein comprises a second reagent. A second reagent may be used during the detection of a binding event between a nucleotide or nucleotide analog and a coded recognition element. The second reagent may improve detection of a binding event between a detectable nucleotide or nucleotide analog and a coded recognition element. The second reagent may protect a label bound to a nucleotide from photobleaching. A second reagent may comprise a buffer. A second reagent may comprise radical scavengers. In some embodiments, a kit provided herein comprises one or more wash solutions. A wash solution may wash the contents of a sequencing reaction from amplified recognition elements. For example, a wash solution may wash away, in part or in full, a first reagent from amplified recognition elements. In some embodiments, a wash solution comprises one or more of Tris(hydroxymethyl)aminomethane, 2-Amino-2-(hydroxymethyl)-l,3-propanediol (TRIS). In some embodiments, a wash solution comprises ethylenediaminetetraacetic acid (edetic acid; EDTA). In some embodiments, a wash solution comprises polyoxyethylene(20)sorbitan monolaurate (Tween 20). In some embodiments, a wash solution comprises TRIS, EDTA, and Tween 20. In some embodiments, a wash solution comprises sodium chloride (NaCl). In some embodiments, the concentration of NaCl in the wash solution is greater than 0.8 M. In some embodiments, the concentration of NaCl in the wash solution is greater than 0.9 M. In some embodiments, the concentration of NaCl in the wash solution is greater than 1.0 M. In some embodiments, the concentration of NaCl in the wash solution is greater than 1.1 M. In some embodiments, the concentration of NaCl in the wash solution is about 1 M. In some embodiments, the concentration of NaCl in the wash solution is about 5 mM to about 100 mM. In some embodiments, the concentration of NaCl in the wash solution is about 10 mM to about 40 mM. In some embodiments, the concentration of NaCl in the wash solution is about 10 mM to about 30 mM. In some embodiments, the concentration of NaCl in the wash solution is about 15 mM to about 25 mM. In some embodiments, the concentration of NaCl in the wash solution is about 20 mM.
[0352] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0353] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0354] One exemplary embodiment includes a method of conducting an assay for a set of targets, the method comprising: subjecting a set of targets to a recognition event, in which each target of the set of targets is uniquely recognized by and bound to at least one recognition element from a set of coded recognition elements, each coded recognition element comprising a target-specific binding site and a code, each code associated with one or more targets of the set of targets, to yield a set of coded targets comprising the target and the at least one recognition element; subjecting the set of coded recognition element to a ligation reaction to yield modified recognition elements comprising the codes, such that the codes of the modified recognition elements can be amplified in a rolling circle amplification (RCA) event, wherein the modified recognition elements comprise ligated circularized recognition elements that are immobilized to a solid surface; performing the RCA event on the modified recognition elements; introducing the modified recognition elements that were amplified in (iii) to a detectable nucleotide or nucleotide analog of a set of detectable nucleotides or nucleotide analogs under conditions sufficient that a first subset of the modified recognition elements couples to the detectable nucleotide or nucleotide analog, wherein the detectable nucleotide or nucleotide analog comprises a detectable label; introducing the modified recognition elements that were amplified in (iii) to a nucleotide or nucleotide analog that is unlabeled of a set of unlabeled nucleotides or nucleotide analogs under conditions sufficient that a second subset of the modified recognition elements couples to the nucleotide or nucleotide analog; detecting a signal from a binding event between a modified recognition element of the first subset of the modified recognition elements and the detectable nucleotide or nucleotide analog by imaging the solid surface; and iteratively performing (iv), (v), and (vi) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements.
[0355] A second exemplary embodiment includes a method of the first embodiment but further comprising subjecting the modified recognition elements to a deblocking condition sufficient to remove the reversible terminator from the detectable nucleotide or nucleotide analog and remove the reversible terminator from the nucleotide or nucleotide analog; and removing the detectable label from the detectable nucleotide or nucleotide moiety.
[0356] A third exemplary embodiment includes a method of the first embodiment but further comprising subjecting the modified recognition elements to a deblocking condition sufficient to remove the reversible terminator from the detectable nucleotide or nucleotide analog and remove the reversible terminator from the nucleotide or nucleotide analog; removing the detectable label from the detectable nucleotide or nucleotide moiety; and iteratively performing
[0357] (iv), (v), (vi), (viii), and (ix) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements.
[0358] A fourth exemplary embodiment includes a method of conducting an assay for a set of targets, the method comprising subjecting a set of targets to a recognition event, in which each target of the set of targets is uniquely recognized by and bound to at least one recognition element from a set of coded recognition elements, each coded recognition element comprising a target-specific binding site and a code, each code associated with one or more targets of the set of targets, to yield a set of coded targets comprising the target and the at least one recognition element; subjecting the set of coded recognition element to a ligation reaction to yield modified recognition elements comprising the codes, such that the codes of the modified recognition elements can be amplified in a rolling circle amplification (RCA) event, wherein the modified recognition elements comprise ligated circularized recognition elements that are immobilized to a solid surface; performing the RCA event on the modified recognition elements; introducing the modified recognition elements that were amplified in (iii) to a detectable nucleotide or nucleotide analog of a set of detectable nucleotides or nucleotide analogs under condition s sufficient that a first subset of the modified recognition elements couples to the detectable nucleotide or nucleotide analog, wherein the detectable nucleotide or nucleotide analog comprises a detectable label; detecting a signal from a binding event between a modified recognition element of the first subset of the modified recognition elements and the detectable nucleotide or nucleotide analogby imagingthe solid surface; and iteratively performing (iv) and
[0359] (v) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements.
[0360] A fifth exemplary embodiment includes the fourth embodiment but further comprising subjecting the modified recognition elements to a deblocking condition sufficient to remove the reversible terminator from the detectable nucleotide or nucleotide analog and remove the reversible terminator from the nucleotide or nucleotide analog; and removing the detectable label from the detectable nucleotide or nucleotide moiety.
[0361] A sixth exemplary embodiment includes the further embodiment but further comprising subjecting the modified recognition elements to a deblocking condition sufficient to remove the reversible terminator from the detectable nucleotide or nucleotide analog and remove the reversible terminator from the nucleotide or nucleotide analog; removing the detectable label from the detectable nucleotide or nucleotide moiety; and iteratively performing (iv), (v), (vii), and (viii) for at least two nucleotides in the nucleic acid sequence of the code of the modified recognition elements.
[0362] A seventh exemplary embodiment includes a computer-implemented system comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device to create an application comprising a software module configured to perform a soft decision decoding of a plurality of the signals detected in a method of any one of Embodiments 1, 2, 3, 4, 5, or 6.
[0363] An eighth exemplary embodiment includes a system comprising a computer processor wherein the processor is programmed to execute a soft decision decoding of a plurality of the signals detected in a method of any one of the exemplary embodiments 1, 2, 3, 4, 5, or 6.
[0364] A ninth exemplary embodiment includes a system for conducting an assay for a set of targets or target analytes, comprising: a reaction vessel, a reagent dispensing module, and software to execute a method comprising a soft decision decoding of a plurality of the signal detected in a method of any one of the exemplary embodiments 1, 2, 3, 4, 5, or 6, wherein the method is executed robotically.
[0365] A tenth exemplary embodiment includes a kit for conducting an assay for a set of targets, the kit comprising: a set of coded recognition elements, wherein each coded recognition element of the set comprises a target-specific binding site specific to a target of the set of targets and a code associated with the target, a plurality of the detectable nucleotides or nucleotide analogs, a plurality of the nucleotide or nucleotide analogs that are unlabeled; and instructions for use of the kit in an assay, wherein the instructions comprise steps for performing the method of the exemplary embodiments 1, 2, 3, 4, 5, or 6 orusingthe system of the exemplary embodiment 7, 8, or 9.
[0366] EXAMPLES
[0367] The following examples are included for illustrative purposes only and are not intended to limit the scope of the inventive concepts.
[0368] Example 1 - Decoding an Amplification Product Library by Incorporation of Fluorescent Nucleotides: 4-State Decoding
[0369] To evaluate detecting target sequences using an amplification product library, a 4-state decode-by-incorporation approach was implemented. As illustrated in FIG. 4, in the decode-by- incorporation approach, target sequences were detected using coded recognition elements 400. Recognition elements 410 comprising a code 412 were hybridized to a target sequence 420 and were amplified in a rolling circle amplification reaction 415, generating an amplification product 425 comprising a plurality of concatenated copies of the recognition element 410 and code 412 and the target region of interest 422. In a 4-state decode-by-incorporation approach, the codes were sequenced by incorporating a nucleotide with a reversible terminator and a dye label into a synthesized nucleic acid strand complementary to each code copy 412 in an amplification product 425. After imaging each incorporation event, the dye was chemically removed and the reversible terminator was cleaved, allowing for the next cycle of incorporation and imaging. This decode-by-incorporation scheme has four states in each cycle, each state corresponding to the incorporation of one of four dye-labeled nucleotides into a complementary nucleic acid strand for sequencing a code.
[0370] Recognition elements used in this experiment were designed as illustrated in FIG. 1. Each recognition element 100 comprised a 5 ’ target specific region 110a and a 3’ target specific region 110b that were complementary to a first and a second region of a target sequence, respectively. The recognition elements also comprised an RCA priming binding site 115, and a code 120.
[0371] An amplification product library was prepared, sequenced, and imaged as described in example workflow 1100 in FIG. 11. Amplification products were generated in a 64 well plate from the modified ligation and / or exonuclease product of a coded recognition element and a target sequence (FIG. 4) using rolling circle amplification. Briefly, the ligation mix used in the RCA reaction included InM per well of a recognition element and 0.1 pM per well of the target sequence. The RCA reaction used 0.5XDNA polymerase buffer 1110. luM of unlabeled RCA primer was hybridized to the amplification products using 1 X hybridization buffer with a surface blocker (e.g., blocking DNA that passivates the surface of the 64 well plate) 1115. The amplification product was washed multiple times with Tris(hydroxymethyl)aminomethane, 2- Amino-2-(hydroxym ethyl)- 1,3 -propanediol (TRIS), ethylenediaminetetraacetic acid (edetic acid; EDTA), and Polyoxy ethylene(20)sorbitan monolaurate (Tween 20) and IM NaCl (TETS) to remove any unhybridized RCA primers. Incorporation mixes (e.g., mixes comprising dye- labeled blocked nucleotides or non-labeled blocked nucleotides) and cleave mixes were prepared 1120. The amplification product-containing wells were washed 2 times with 25uL of TRIS, EDTA, and Tween 20 (TET) + 20mM sodium chloride (NaCl) wash buffer. After the second wash, the TET+20mM NaCl solution was left in the well, and the amplification products were incubated with the TET + NaCl wash buffer solution at 60°C for 10 minutes 1125. The incorporation mixes comprising either dye-labeled blocked nucleotides or non-labeled blocked nucleotides, both types of nucleotides with a reversible terminator, were added to strip tubes and incubated at room temperature for 10 minutes 1130. The TET+20mMNaCl wash buffer solution was removed from the wells, and 20ul of dye-labeled blocked nucleotide incorporation mix was added to the well 1135. The amplification products and the dye-labeled blocked nucleotide incorporation mix were incubated at 60°C for 15 minutesin the wells. The dye-labeled blocked nucleotide incorporation mix was removed, and 20ul of non -labeled blocked nucleotide incorporation mix was added to each well.
[0372] The amplification products and the non -lab eled blocked nucleotide incorporation mix were incubated at 60°C for 10 minutes in the wells 1140. The non-labeled blocked nucleotide incorporation mix was removed and 25ul of TET+1 OmM ethylenediaminetetraacetic acid (edetic acid; EDTA) was added to each well to stop the reaction 1145. A plate washer was used to wash the wells with TETS (TET and IM NaCl). The wash in the wells was replaced with 20ul of a scan reagent comprising buffer and radical scavengers 1150, and the amplification products in the wells were imaged using a cell imaging multimode reader 1155. The scan reagent was removed by washing with TET+20mM NaCl, and 35ul of a cleaving reagent was added to the wells. The amplification products were incubated with the cleaving reagent for 20 min at 60°C in the wells 1160. The cleave mix was removed, and the wells were washed with TETS 1165. The wells were washed once with wash buffer, and then 60ul of wash buffer solution comprising TET was added to each well. The amplification products were incubated with the wash buffer solution for 10 minutes at 60°C in the wells 1170. Steps 1125 through 1170 were repeated as many times as necessary to read the amplified codes comprisedin the amplification products. In the experiment described herein, 4 amplification cycles (steps 1110-1170) were run per day over a two day period, for a total of 8 cycles.
[0373] As described in step 1120, incorporation reagents and cleaving reagents were prepared each day. The dye-labeled blocked nucleotide incorporation mix included: premix (362.74uL), fluorescent nucleotides (2.85uL), polymerase (8.5 luL), and 2M Mg (0.90uL) for a total of 375uL. The non-labeled blocked nucleotide mix contained: premix (357 uL), fluorescent nucleotides (8.5 luL), polymerase (8.5 luL), and 2M Mg (0.90uL) for a total of 375uL. Nucleotides had reversible terminators. A polymerase was included in the incorporation reagents, and it was found that the Therminator™ X polymerase (NEB) was an effective polymerase. 20ul of the dye-labeled blocked nucleotide incorporation reagent or 20ul of the nonlabeled blocked nucleotide reagent was added to each well in steps 1135 and 1140, respectively. The volume of each component in the incubation reagents was scaled as appropriate for the number of amplification product-containing wells used in the experiment. The cleaving reagent contained 200mg of TCEP added to lOmL of IM Tris at pH 8. The cleaving reagent was pre- mixed, aliquoted, and stored at -20°C. 35ul of the cleaving reagent was added to each well in step 1160.
[0374] As described in step 1155, the amplification products in the well plate were imaged using a cell imagining multimode reader. The imaging was performed by enabling a 2x1 montage in manual mode on the reader, and the set-up file for the reader was edited for GFP, RFP, TRITC, and Cy5 imaging channels. The focal height was found and set in the RFP channel of each well before images were triggered. A +3um offset was used for Cy5 imaging, and a 0 offset was used for GFP and TRITC imaging. Exposure and gain were increased as necessary for each cycle. The first imaged well in the plate was used to establish the exposure and gain settings, and these settings were maintained for the other wells. FIG. 12A and FIG. 12B provide RFP, GFP, and Cy5 channel images of the same well collected on the first cycle and the final cycle (eighth cycle) (FIG. 12A and 12B, respectively) of a sequencing experiment. The contrast used in each image is the same, demonstrating that there is only a small amount of signal decay over an 8 - cycle run.
[0375] In 4-state decoding strategy, dye-labeled nucleotides were incorporated into the complementary strand of the codes of amplification products, imaged, and then cleaved. There were four states in each cycle, each state corresponding to four different dye molecules on each of the four nucleotides. Each nucleotide was labeled with a dye that emitted light at a wavelength distinguishable from the other used dyes. FIG. 13A and FIG. 13B provide examples of images captured during an 8-cycle run from two different wells each well containing a plurality of amplification products. Each row in FIG. 13A and FIG. 13B corresponds to a cycle and each column corresponds to a fluorescent imaging channel. A fluorescent dot in each picture represents an amplification product. The position of an example amplification product is circled in each image, with each circle identified by an arrow. High intensity signal in a channel corresponds to a plurality of nucleotides bound to a dye visible in that channel being incorporated into a complementary extension strand of an amplification product in that cycle. For example, in FIG. 13A, in the first cycle, a fluorescent signal is strongest in the first channel (ON), while the fluorescence signal in the second cycle is strongest in the fourth channel (ON). The intensity measured in each of five color channels (corresponding to dyes detectable in at least one of the GFP, RFP, TRITC, Cy5, and Cy5.5 channels) in the first imaging cycle for a plurality of amplification product codes is shown in FIG. 14. Amplification products comprising code sequences that began with a guanine (G) or thymine (T) had relatively high intensity signal in the RFP channel, code sequences that began with a thymine (T) had relatively high intensity signal in the TRITC channel, code sequences that began with an adenine (A) had high intensity signal in the Cy5 channel, and sequences that began with a cytosine (C) had relatively high intensity readout in the Cy5.5 channel.
[0376] In another experiment each nucleotide was labeled with a different dye that emitted light that was visible in one of three fluorescent channels (e.g., Cy5, GFP, or RFP) in the imaging instrument. Additionally, one of the four dyes emitted light that was visible in two of the channels (the RFP and the TRITC channels), while a second of the four dyes was visible in only one channel (the RFP channel) but not in a second channel (the TRICT channel). In this experiment, the dye linked to the guanine nucleotide was visible in the Cy5 channel; the dye linked to the thymine nucleotide was visible in the RFP channel; the dye linked to the cytosine nucleotide was visible in the GFP channel, and the dye linked to the adenine nucleotide was imaged in the RFP and TRICT channels. FIG. 15 provides examples of signal intensity readouts in the Cy5, RFP, and GFP channels in each of eight cycles of a run for a well comprising amplification products all labeled with the same code comprising the code sequence GTCGATGA. Images were not taken with the TRITC filter.
[0377] Using a soft decoding algorithm , intensity signals measured from amplification products can be assigned to a most likely code sequence. Code sequences were chosen as follows. First, a list was constructed of all legal codes. Legal codes are sequences of numbers containing the digits 1, 2, 3, and 4 that avoided certain problematic codes. Problematic codes included codes that use only one or two of the possible digits. Codes that include a run of the same digit repeated 4 or more times were rejected. The codespace was started by selecting a usable code at random, adding it to the list of chosen codes, and then removing from the available codes any code that was too close (e.g., a code that had a hamming distance less than a certain threshold hamming distance to the chosen starting code) to the first added code. For example, if the minimum hamming distance is set at 4, then adding 12341234 to the chosen codes meant that 12341243 was never chosen. In other embodiments, a plurality of color balance codes were added to the codespace, which were chosen to have maximum pairwise hamming distance and to exercise every state in every cycle.
[0378] After the initial code was selected, a second available code was selected, added to the list of chosen codes, and un-chosen codes were removed from the available list as needed (e.g., codes that were too close to the second-chosen code were removed). In each step, a snug code was chosen, e.g., a code that had minimum-permissible hamming-distance from the maximum number of chosen codes. For example, if the codespace contained three codes, and the minimum hamming distance was 3 : (1) the next code was chosen randomly from available codes with hamming-distance-3 from all 3 chosen codes; (2) if that was not possible, the next code was chosen randomly from available codes with hamming-distance-3 from any 2 chosen codes; (3) if that was not possible, the next code was chosen randomly from available codes with hamming- distance-3 from any 1 chosen code; (4) if that was not possible, the next code was chosen randomly from the available codes. The codespace-generation procedure was repeated several times, and the largest resulting codespace was used to create codes for the recognition elements. Using snug codespaces leads to larger codespaces than choosing available codes at random or maximizing hamming distances from already-chosen codes. FIG. 16 provides attainable codespace sizes for different cycle counts. Each line in FIG. 16 corresponds to a different decoding strategy, state count (see Example 2 for details of 2 -state decoding and Example 3 for details of 3 -state decoding), and minimum hamming distance (HD). The y-axis is log-scaled as codespace sizes increase exponentially with the number of cycles.
[0379] In a decode experiment, a codespace comprising six 8 -base sequences with a minimum pairwise hamming distance of 3 was generated to test the feasibility of application of the soft- decoding algorithm to a decode-by-incorporation approach. The sequences in the codespace were chosen to cover all four nucleotide bases in each of eight sequencing flows. The sequences used in this experiment comprised: ATATCTCA; AGACGTAG; GTCGATGA; TCGAGCGT; CATCGATC; and TCTCTGAG. As described above, RCA was used to amplify a plurality of modified target-bound recognition elements comprising code sequences into amplification products. In two wells of amplification products , the amplification products in one well comprising an ATATCTCA code sequence and the second comprising a TCTCTGAG code sequence, the percentage of decoded sequences that did not correspond to the code in the well (unused error rate; the percentage of sequences that were assigned one of the 5 unused codes in the 6-plex code space, rather than the used code) was 0% and 0.3%, respectively. In two wells that contained all six code sequences (6-plex pool), the percentage of amplification products whose code sequences were successfully decoded (decode rate) was between -55 -60%, and all six codes used were well-represented. As all six codes of the 6-plex code space were used in the 6-plex wells, calculating an unused error rate forthese wells using the 6-plex codespace was not performed. To further test the soft decoding algorithm in these 6-plex well, the codespace was augmented to a 2739-plex codespace with hamming distance of 2. The two 6-plex wells were still decodable with unused error rates of 3-4%, decode rates of 15-20%, and all 6 codes represented.
[0380] In a second decode experiment, a 69-plex codespace was used. In this experiment, wells contained either wild-type targets (WT) and recognition elements targeting the wild-type targets, variant targets and recognition elements targeting the mutated targets, or a combination of targets and recognition elements. Two wells included all targets (WT and variant) and all recognition elements (WT and variant), one well included only WT targets and both WT and variant recognition elements, and one well included only WT targets and WT recognition elements. Spot counts (the number of amplification products detected) and decode rates for each of 4 wells (2 wells with all targets and recognition elements, 1 well with WT targets and both types of recognition elements, and 1 well with WT recognition elements and WT targets) are provided in FIG. 17. There were two images taken in each well and these were analyzed individually. One of the two images in the last well developed an artifact and was excluded from analysis. Shown in FIG. 18 are the number of counts of decoded sequences assigned to each of the 69 targets. When a well included only WT recognition elements and targets, there was detection of the WT targets was specific, as almost none of the variant targets were detected in the wells (bottom row of FIG. 18 and circled variant dots of FIG. 19). The data indicates that the soft decoding algorithm accurately decoded recognition element sequences. Similarly, in wells that included both WT and variant recognition elements and only WT targets, very few variant targets were detected in the wells (circled dots of FIG. 20), indicating that recognition elements were selective in their target binding. In wells that included both WT and variant targets and recognition elements, both WT and variant targets were detected (top row of FIG. 18). Detection rates of WT targets were similar in wells that included only WT recognition elements and in wells that included both WT and variant recognition elements (comparison of the top and bottom rows of FIG. 18, and FIG. 19). In wells that included only WT recognition elements, detection rates ofWT targets were similar in wells that included either WT and variant targets or only WT targets (FIG. 21). In wells that included only WT targets, detection rates of WT targets were similar in wells that included either WT and variant recognition elements or only WT recognition elements (FIG. 22). Read counts in duplicate wells that included WT and variant targets and WT and variant recognition elements were also similar (FIG. 23).
[0381] Example 2 - Decoding an Amplification Product Library by Incorporation of Fluorescent Nucleotides: 2-State Decoding
[0382] To evaluate detecting target sequences using an amplification product library, a 2-state decode-by-incorporation approach was performed. As illustrated in FIG. 4, in the decode-by- incorporation approach, target sequences were detected using coded recognition elements 410. Recognition elements 410 comprising a code 412 hybridized a target sequence 420 and were amplified in a rolling circle amplification reaction 415, generating an amplification product 425 comprising a plurality of copies of the recognition element. In a 2-state decode-by-incorporation approach, the recognition elements were sequenced by incorporating a nucleotide with a dye label into an extension reaction using the amplification product as a template. This scheme had two states in each flow - each amplification product does or does not incorporate a dye-labeled nucleotide base. A different nucleotide was used in each cycle (e.g., A, T, C, G, A, T, C, G, etc.), and after imaging, and between cycles, dyes were removed from the labelled nucleotides.
[0383] Using a soft decoding algorithm , intensity signals measured from amplification products can be assigned to a most likely code sequence. Code sequences were chosen as follows. First, a list was constructed of all usable codes. Usable codes were sequences of numbers containing the digits 0 and 1 that avoided certain problematic codes. Problematic codes include codes that had a long run of zeros. For example, codeword 00001111 would not be useful, because the sequence cannot start with an A since the first digit is a 0, the sequence cannot start with a C since the second digit is a 0, the sequence cannot start with a G since the third digit is a 0, and the sequence cannot start with a T since the fourth digit is a 0. The codespace was started by selecting a usable code at random, adding it to the list of chosen codes, and then removing from the list of available codes any code that was too close (e.g., a code that had a hamming distance less than a certain threshold hamming distance to the chosen starting code) to the added code.
[0384] After the initial step, a second available code was selected, added to the list of chosen codes, and un-chosen codes were removed from the available list as needed (e.g., codes that were too close to the second-chosen code were removed). In each step, a snug code was chosen, e.g., a code that had minimum-permissible hamming-distance from the maximum number of chosen codes. For example, if the codespace contained three codes, and the minimum hamming distance was 3 : (1) the next code was chosen randomly from available codes with hamming - distance-3 from all 3 chosen codes; (2) if that was not possible, the next code was chosen randomly from available codes with hamming-distance-3 from any 2 chosen codes; (3) if that was not possible, the next code was chosen randomly from available codes with hamming- distance-3 from any 1 chosen code; (4) if that was not possible, the next code was chosen randomly from the available codes. The codespace-generation procedure was repeated several times, and the largest resulting codespace was used to create codes for the recognition elements. Using snug codespaces lead to larger codespaces than choosing available codes at random or maximizing hamming distances from already -chosen codes. FIG. 16 provides codespace sizes for different cycle counts. Each line in FIG. 16 corresponds to a different decoding strategy, state count (see Example 1 for details of 4-state decoding and Example 3 for details of 3 -state decoding), and minimum hamming distance (HD). The y-axis is log-scaled as codespace sizes increase exponentially with the number of cycles. Many different polymerases were screened for the ability to incorporate dye-labeled nucleotides over several cycles with high efficiency and specificity in the context of a single dye-labeled nucleotide. Bst3.0 DNA polymerase (NEB) was initially found to have the best performance, and incorporation appeared strong and accurate at 37°C as it did at 60°C. Other Bst-like enzymes, including ArcticZymes Bst DNA polymerase (AZ Bst) and Bsm DNA polymerase (ThermoScientific) displayed similar performance to Bst3.0 DNA polymerase in single-plex decode experiments performed in parallel. In these experiments, one of four cleavable dye-labeled dNTPs was provided with a polymerase over seven cycles in the order A- T-C-G-A-T-C. Each well included amplification products from only one code sequence. The signal intensity relative to the expected state (ON or OFF) was examined (see FIG. 24). The OFF state was expected to have a low relative intensity, while the ON state was expected to have a high relative intensity. Amplification products were subjected to a cleavage reagent that removed the dye on the incorporated nucleotide between each flow.
[0385] In one decode-by-incorporation experiment, 1,745 spots (localized amplification products) were decoded in a well, 4 of 4 used codes were identified in the well with good coverage, and 4.3% of the sequences decoded by the soft-decoding algorithm were sequences that were not present in the well (unused error rate).
[0386] Another decode-by-incorporation experiment was a 2-state decode of an 8-plex codespace. One well had a 3.6% unused error rate, with 2,002 spots decoded, and 4 of 4 used codes identified with good coverage.
[0387] In a 15 -cycle experiment, a well had 10,314 amplification products decoded with an unused error rate of 5.4%, and coverage for 22 of 24 expected codes.
[0388] Example 3 - Decoding an Amplification Product Library by Incorporation of Fluorescent Nucleotides: 3-State Decoding
[0389] To evaluate detecting target sequences using an amplification product library, a 3-state decode-by-incorporation approach is used. As illustrated in FIG. 4, in the decode-by- incorporation approach, target sequences are detected using coded recognition elements 410. Recognition elements 410 comprise a code 412 and are hybridized to a target sequence 420 and are amplified in a rolling circle amplification (RCA) reaction 415, generating an amplification product 425 comprising a plurality of copies of the code. In a 3-state decode-by-incorporation approach, the recognition elements are sequenced by incorporating a nucleotide with a dye label into an extension reaction using an amplification product as the template. This scheme has three states in each flow - each amplification product does or does not incorporate a dye-labeled nucleotide base. By addressing sequencesthat contain runs of 2, 3, or more occurrences of the same nucleotide in a row, more states could be encoded per cycle (e.g., low and high signal intensities resulting from single or multiple dyes being added in a cycle, respectively), but at the cost of a lower signal -to-noise ratio.
[0390] A two-dye, three-state system is used to detect target sequences, wherein: (i) Cycle 1 uses A and G nucleotides, and each code incorporates either A (color 1), or G (color 2), or both (colors 1 and 2 together); (ii) Cycle 2 uses C and T nucleotides, and each code incorporates either C (color 1), or T (color 2), or both (colors 1 and 2 together); (iii) Cycle 3 uses A and G nucleotides again, and each code incorporates either a (color 1), or G(color 2), or both (colors 1 and 2 together). In general, the odd-numbered cycles are purines (A and G), and the even- numbered cycles are pyrimidines (C and T).
[0391] A codespace is generated including usable codes for the 3 -state decode by incorporation approach. FIG. 16 provides attainable codespace sizes for different cycle counts. Each line in FIG. 16 corresponds to a different decoding strategy, state count (see Example 1 for details of 4- state decoding and Example 3 for details of 3 -state decoding), and minimum hamming distance (HD). The y-axis is log-scaled as codespace sizes increase exponentially with the number of cycles.
[0392] Example 4- Use of Circularized Recognition Elements and SBS Chemistry for Determining the Presence of a Code and Its Target
[0393] Target nucleic acids are extracted from a tissue sample. The extracted nucleic acids, in this example DNA, are purified and quantitated by any method known to a skilled artisan. Recognition elements are designed and generated; 1) a recognition element is designed with complementary 5’ and 3’ regions to adjacent sequences of a wild type nucleic acid sequence, while 2) a recognition element is designed with 5 ’ and 3 ’ target regions that would detect a SNP at the 3’ end of a recognition element sequence. The recognition element design also includes unique codes to identify whether a sequence is wild type or variant, sequencing primer binding site sequences, UMI sequences, and sequences that are complementary to one or more immobilized primers on a sequencing flowcell. In this example the flowcell is a SBS flowcell for use with an Illumina® sequencing instrument, as such sequences in the recognition elements are P5 or P7 oligonucleotide complementary sequences.
[0394] The recognition elements and extracted and purified DNA from a tissue sample are combined. The target sequences from the sample DNA, wild type and variant, hybridize to their complementary recognition element regions, and the hybridized recognition elements are ligated and circularized. After circularization, exonucleases are added to the reaction to remove any linear DNA, including linear and unhybridized recognition elements and genomic DNA, to decrease background in downstream applications.
[0395] The circularized recognition elements are flowed onto the flowcell and captured by the complementary P5 or P7 primers already immobilized on the flowcell. Bridge amplification is performed using the P5 and P7 primers on the flowcell and the hybridized recognition elements as the templates, thereby generating a plurality of amplicon clusters. The amplicons are sequenced using established SBS sequencing chemistry as defined by the Illumina ® instrument being used.
[0396] Sequencing data is generated and analyzed using soft decision decoding and / or hard decision decoding for the code sequences which correlate to the target nucleic acid sequence from the sample, and the presence or absence of the target nucleic acids from the sample are determined based on the presence or absence of the code associated with the target.
[0397] While preferred embodiments have been shown and described herein, it is obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the methods and kits disclosed herein be limited by the specific examples provided within the specification. While the methods and kits disclosed herein have been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from methods and kits disclosed herein. Furthermore, it shall be understood that all aspects disclosed herein are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments described herein may be employed in practicing the methods and kits disclosed herein. It is therefore contemplated that the methods and kits disclosed herein shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of methods and kits disclosed herein and that methods and kits within the scope of these claims andtheir equivalents be covered thereby.
Claims
CLAIMSWhat is Claimed is:1 . A method of detecting the presence of one or more targets, the method comprising:(a) providing a set of targets comprising the one or more targets and a set of recognition elements, wherein each target from the set of targets is uniquely recognized by and complementary to at least one recognition element from a set of recognition elements, wherein each recognition element from the set of recognition elements comprises a target- specific binding site and a code from a set of codes;(b) hybridizing each target from the set of targets to the target- specific binding site of at least one recognition element from the set of recognition elements, wherein the code of the recognition element is associated with the target that is hybridized to the target- specific binding site of the recognition element, thereby generating hybridized recognition elements;(c) performing a ligation reaction on the hybridized recognition elements to yield modified recognition elements, wherein the modified recognition elements are circularized;(d) immobilizing the modified recognition elements onto a solid surface to produce immobilized recognition elements;(e) performing rolling circle amplification on the modified and immobilized recognition elements thereby generating amplified recognition elements;(f) introducing a labeled nucleotide or nucleotide analog thereof from a set of labeled nucleotides or nucleotide analogs thereof to the amplified recognition elements to, wherein each of the labeled nucleotides or nucleotide analogs thereof comprises a detectable label, under conditions for incorporating the labeled nucleotide or nucleotide analog thereof into an extension reaction using amplified recognition elements as templates to produce amplified modified recognition elements;(g) introducing an unlabeled nucleotide or nucleotide analog thereof from a set of unlabeled nucleotides or nucleotide analogs thereof to the amplified modifiedrecognition elements under conditions for incorporating the unlabeled nucleotide or nucleotide analog thereof into the extension reaction using the amplified recognition elements as the template;(h) performing imaging of the solid surface to detect a signal from the incorporation of the labeled nucleotide or nucleotide analog thereof into the extension reaction; and(i) iteratively performing (f), (g), and (h), thereby determining the sequence of the code of the amplified modified recognition elements and the presence of the one or more targets from the set of targets.
2. The method of claim 1, wherein (f) and (g) occur substantially simultaneously.
3. The method of claim 1, wherein (f) and (g) occur sequentially.
4. The method of claim 1, wherein the labeled nucleotide or nucleotide analog thereof comprises a reversible terminator moiety.
5. The method of claim 4, wherein the reversible terminator moiety comprises a 3 ’ -O-alkyl hydroxylamino group, a 3’-phosphorothioate group, a 3’-O-malonyl group, a 3’-O-benzyl group, a 3’-O-Azidomethyl group, a 3’-O-Azide group, a 3’-O-Azido group, or a 3’-O-methyl group.
6. The method of claim 1, wherein the unlabeled nucleotide or nucleotide analog thereof comprises a reversible terminator moiety.
7. The method of claim 6, wherein the reversible terminator moiety comprises a 3 ’-O-alkyl hydroxylamino group, a 3’-phosphorothioate group, a 3’-O-malonyl group, a 3’-O-benzyl group, a 3’-O-Azidomethyl group, a 3’-O-Azide group, a 3’-O-Azido group, or a 3’-O-methyl group.
8. The method of any one of claims 4-7, wherein the reversible terminator moiety prevents incorporation of an additional labeled nucleotide or nucleotide analog thereof or an additional unlabeled nucleotide or nucleotide analog thereof at an N+l position, wherein N is the position of the labeled nucleotide or nucleotide analog thereof or the unlabeled nucleotide or nucleotide analog thereof in an extension reaction.
9. The method of any one of claims 4-8, further comprising subjecting the amplified modified recognition elements to a deblocking reaction sufficient to remove the reversibleterminator moiety from the labeled nucleotide or nucleotide analog thereof and remove the reversible terminator moiety from the unlabeled nucleotide or nucleotide analog thereof.
10. The method of any one of claims 1-9, further comprising removing the label from the labeled nucleotide or nucleotide analog thereof.
11. The method of any one of claims 4-10, further comprising:(j) subjecting the amplified modified recognition element to a deblocking reaction sufficient to remove the reversible terminator from the labeled nucleotide or nucleotide analog thereof and remove the reversible terminator from the unlabeled nucleotide or nucleotide analog thereof; and(k) removing the label from the labeled nucleotide or terminator moiety.
12. The method of any one of claims 4-10, further comprising:(j) subjecting the amplified modified recognition elements to a deblocking reaction sufficient to remove the reversible terminator from the labeled nucleotide or nucleotide analog thereof and remove the reversible terminator moiety from the unlabeled nucleotide or nucleotide analog thereof;(k) removing the label from the labeled nucleotide or terminator moiety; and(l) iteratively performing (f), (g), (h), (j), and (k) for at least two consecutive cycles.
13. The method of any one of claims 9-12, wherein the deblocking reaction comprises a buffer having one or more of:(i) a phosphine compound comprising Tris(2-carboxyethyl)phosphine, bis-sulfo triphenyl phosphine or Tri(hydroxyproyl)phosphine;(ii) tetrakis(triphenylphosphine)palladium(0) (Pd(P(C6H5)3)4), (CJLjsNH, or 2,3-Dichloro-5,6-dicyano-l,4-benzo-quinone;(iii) palladium on carbon;(iv) a thiol group comprising beta-mercaptoethanol or dithiothritol;(v) potassium carbonate in MeOH, triethylamine in pyridine, or with Zn in acetic acid; or(vi) tetrabutylammonium fluoride, pyridine-HF, ammonium fluoride, or triethylamine trihydrofluoride.
14. The method of any one of claims 9-12, wherein detecting the incorporation occurs prior to removal of the reversible terminator and removal of the detectable label.
15. The method of any one of claims 9-12, wherein detecting the incorporation occurs substantially simultaneously with removal of the reversible terminator and removal of the detectable label.
16. The method of any one of claims 9-12, wherein detecting the incorporation occurs substantially simultaneously with removal of the detectable label and prior to removal of the reversible terminator.
17. The method of any one of claims 9-12, wherein detecting the incorporation occurs substantially simultaneously with removal of the reversible terminator and prior to removal of the detectable label.
18. The method of any one of claims 1-17, wherein the code of the recognition elements comprises a soft decodable code.
19. The method of any one of claims 1-18, further comprising decoding the codes of the amplified modified recognition elements.
20. The method of claim 19, wherein decoding the codes comprises:(i) recording a signal produced in response to interrogation of each segment of the code; and(ii) upon completion of the interrogation, determining a probability of the presence of each of the codes by applying a soft-decision probabilistic decoding algorithm to the recorded signal, wherein the presence of the code is indicative of the presence of the target.21 . The method of claim 19, wherein decoding the codes comprises decoding the codes by a soft decision decoding algorithm.
22. The method of claim 20 or 21, wherein a percentage of the amplified modified recognition elements that comprise codes that are decoded by the soft decoding algorithm is greater than 10% of the amplified modified recognition elements.
23. The method of claim 20 or 21, wherein a percentage of the amplified modified recognition elements that comprise codes that are decoded by the soft decoding algorithm is greater than 50% of the amplified modified recognition elements.
24. The method of claim 1, wherein the label of the labeled nucleotide comprises an optical label.
25. The method of claim 1, wherein the label of the labeled nucleotide comprises a fluorescent label.
26. The method of claim 25, wherein the set of labeled nucleotides or nucleotide analogs thereof comprises a plurality of species, wherein the fluorescent label of each species of the plurality of species emits a different wavelength than other fluorescent labels of other species of the plurality of species.
27. The method of claim 26, wherein the plurality of species comprises adenine (A), cytosine (C), Guanine (G), or Thymine (T).
28. The method of claim 27, wherein the set of labeled nucleotides or nucleotide analogs thereof comprises A, C, G and T, and wherein the detectable label for each of the A, the C, the G and the T is different.
29. The method of claim 27, wherein the set of labeled nucleotides or nucleotide analogs thereof comprises two or more species of A, C, G and T, wherein the label for each of the two or more species of A, C, G, and T is different.
30. The method of claim 1, wherein the set of labeled nucleotides or nucleotide analogs thereof comprises one species comprising A, C, G or T.31 . The method of claim 30, wherein each of the labeled nucleotides or nucleotide analogs thereof of the set of labeled nucleotides or nucleotide analogs thereof comprises the same detectable label.
32. The method of claim 1, wherein the set of unlabeled nucleotides or nucleotide analogs thereof comprises a plurality of species, wherein the plurality of species comprises A, C, G or T.
33. The method of claim 1, wherein the set of targets comprises nucleic acid targets.
34. The method of claim 33, wherein the nucleic acid targets comprise one or more of deoxyribonucleic acid (DNA), bisulfite converted DNA, ribonucleic acid (RNA), methylated nucleic acid targets, amino acid targets, or polypeptide targets.
35. The method of claim 1, wherein the recognition elements comprises a padlock probe configuration or a molecular inversion probe configuration.
36. The method of claim 1, wherein the recognition elements further comprise:(i) one or more target recognition element sequences;(ii) one or more sequencing primer binding site sequences;(iii) one or more amplification primer binding site sequences;(iv) a unique molecular identifier sequence;(v) a sample index sequence;(vi) a restriction enzyme site sequence; or(vii) any combination of (i) to (vi).
37. The method of claim 36, wherein the one or more amplification primer binding site sequences comprise a universal primer binding site sequence that is common to all of the recognition elements of the set of recognition elements.
38. The method of any one of claims 1-37 wherein each code from the set of codes is the same length.
39. The method of any one of claims 1-37, wherein at least a subset of recognition elements from the set of recognition elements comprises codes of the same length.
40. The method of any one of claims 1-39, wherein each code from the set of codes comprises a nucleic acid sequence having a length of 25 or fewer nucleotides.
41. The method of claim 40, wherein each code from the set of codes comprises a nucleic acid sequence having a length of 20 or fewer nucleotides.
42. The method of claim 40, wherein each code from the set of codes comprises a nucleic acid sequence having a length of 15 or fewer nucleotides.
43. The method of claim 40, wherein each code from the set of codes comprises a nucleic acid sequence having a length of 10 or fewer nucleotides.
44. The method of claim 40, wherein each code from the set of codes comprises a nucleic acid sequence having a length of 5 or fewer nucleotides.
45. The method of any one of claims 1-44, wherein the rolling circle amplification is performed on the solid surface, wherein the solid surface does not comprise a covalent attachment molecule.
46. The method of claim 1, wherein the solid surface is a charged solid surface.
47. The method of claim 46, wherein the charged solid surface comprises a cation-coating layer bound to the solid surface.
48. The method of claim 47, wherein the cation-coating layer comprises poly -L-ly sine.
49. The method of any one of claims 1-48, wherein the rolling circle amplification generates a concatemer comprising multiple copies of the modified recognition element.
50. The method of claim 1, wherein the method is performed in vitro.
51. A method for identifying the presence of a target nucleic acid from a nucleic acid sample, the method comprising: a) providing the nucleic acid sample and a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a code, a 5’ recognition region, a 3 ’ recognition region, and one or more sequencing primer binding sites; b) hybridizing the 5 ’ recognition region and the 3 ’ recognition region to sequences that are complementary to the 5 ’ recognition region and the 3 ’ recognition region in the target nucleic acid from the nucleic acid sample to generate a hybridized recognition element; c) ligating and circularizing the hybridized recognition element thereby generating circularized recognition elements; d) immobilizing the circularized recognition element on a sequencing substrate; and e) identifying the code of the circularized recognition element by sequencing and using the determined code sequence to identify the presence of the target nucleic acid from the nucleic acid sample.
52. The method of claim 51, wherein the target nucleic acid is DNA, RNA or cDNA.
53. The method of claim 51, wherein the plurality of recognition elements comprises a subset of recognition element regions complementary to a wild type sequence in the target nucleic acid from the nucleic acid sample.
54. The method of claim 51, wherein the plurality of recognition elements comprises a subset of recognition element regions complementary to a variant sequence in the target nucleic acid from the nucleic acid sample.
55. The method of claim 51, wherein each recognition element from the plurality of recognition elements further comprises one or more of : (i) a universal primer binding site sequence, (ii) a unique molecule identifier sequence, (iii) a restriction endonuclease recognition sequence, (iv) a cleavage sequence, or (v) any combination of (i) through (iv).
56. The method of claim 51 , wherein the sequencing comprises next generation sequencing.
57. The method of claim 56, wherein the next generation sequencing comprises sequence by synthesis.
58. The method of claim 54, wherein the variant sequence is a single nucleotide polymorphism, an indel, or a copy number variant.
59. The method of claim 51, wherein the 5’ recognition region and the 3’ recognition region of the recognition element hybridize to adjacent sequences of the target nucleic acid.
60. The method of claim 51 , wherein the 5 ’ recognition region and the 3 ’ recognition region of the recognition element hybridize to non -adjacent sequences of the target nucleic acid.
61. A computer-implemented system comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device to create an application comprising a software module configured to perform a soft decision decoding of a plurality of the signals detected in a method of any one of claims 1-60.
62. A system comprising a computer processor, wherein the computer processor is programmed to execute a soft decision decoding of a plurality of the signals detected in a method of any one of claims 1-60.
63. A system for conducting an assay for a set of targets or target analytes, the system comprising:(a) a reaction vessel;(b) a reagent dispensing module; and (c) software to execute a method comprising a soft decision decoding of a plurality of the signal detected in a method of any one of claims 1-60, wherein the method is executed robotically.
64. A kit for conducting an assay for a set of targets, the kit comprising:(a) a set of recognition elements, wherein each recognition element of the set of recognition elements comprises a 5’ target-specific binding site and a 3’ target specific binding site and a code associated with a target from the set of targets;(b) a plurality of the labeled nucleotides or nucleotide analogs thereof;(c) a plurality of the unlabeled nucleotide or nucleotide analogs thereof; and(d) instructions for use of the kit in an assay, wherein the instructions comprise operations for performing the method of any one of claims 1-60 or using the system of any one of claims 61-63.
Citation Information
Patent Citations
Encoded Dual-Probe Endonuclease Assays
US20230257801A1
Encoded Endonuclease Assays
US20230295739A1
Encoded assays
US20240124920A1
Treating prostate disorders
US62636116P0
Encoded nucleic acid methylation assays
WO2023096671A1
Cited By
Methods and compositions for increasing detection of structural variants
WO2026161646A1