Methods and systems for target identification
By employing recognition elements with complementary sequences and decoding unique codes, the method enhances the sensitivity and efficiency of DNA library preparation for NGS, addressing inefficiencies in current methods and improving target detection in complex samples.
Patent Information
- Application Number
- PCT/US2025/037669
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-13
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Current DNA library preparation methods for next-generation sequencing (NGS) are inefficient, leading to increased sequencing costs and reduced coverage of biologically relevant regions due to off-target sequences and non-specific amplification, particularly in applications involving low abundance targets or complex sample matrices.
The use of recognition elements with complementary 5' and 3' end sequences that hybridize to target nucleic acids, followed by ligation, amplification, and decoding of a unique code to selectively enrich for specific nucleic acid targets, enhancing analytical sensitivity and efficiency.
This approach improves the signal-to-noise ratio and reduces sequencing costs by focusing on user-defined genomic sequences, enabling more accurate detection of rare targets in applications like oncology, inherited disease diagnostics, and pathogen detection.
Smart Images

Figure US2025037669_22012026_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR TARGET IDENTIFICATIONCROSS-REFERENCE
[0001] The present application claims the benefit of United States Provisional Application Serial No. 63 / 672,658, filed July 17, 2024, and United States Provisional Application Serial No. 63 / 771,542, filed March 13, 2025, each of which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Sequencing, for example next generation sequencing (NGS), has revolutionized genomics by enabling rapid, high-throughput analysis of nucleic acids. However, the sensitivity and specificity of NGS is predicated on the quality and relevance of the DNA library that feeds into any given sequencing workflow. Library preparation methods can involve random fragmentation and non-specific amplification which can result in libraries composed of a fair amount of off-target sequences. This inefficiency not only increases sequencing costs but can also reduce the effective depth of coverage for regions of biological or clinical interest, particularly in applications involving low abundance targets or complex sample matrices.
[0003] Accordingly, there is a need for improved DNA library preparation assays that are highly selective for user defined genomic sequences. A targeted DNA library preparation approach can allow for focused interrogation of specific loci, variants, and genes of interest while enabling higher signal-to-noise ratios, more efficient data generation and improved analytical sensitivity. Such targeted assays can be particularly valuable in areas such as oncology, inherited disease diagnostics, pathogen detection, minimal residual disease monitoring and agriculture, for example where accurate detection of rare or low frequency nucleic acid targets is needed.
[0004] The present disclosure addresses these challenges by providing for preparing highly targeted DNA libraries optimized for NGS. Enriching for specific nucleic acid targets during library construction has the potential to significantly enhance analytical performance while conserving sequencing resources and enabling scalable, cost-effective genetic analysis across a wide range of research and clinical applications.SUMMARY
[0005] Aspects disclosed herein, in some embodiments, provide methods for determining whether to screen a sample for a target of interest, comprising: a) hybridizing a sample with a recognition element, wherein the recognition element comprises; i) a 5’ end sequence and a 3’ end sequence, wherein the 5’ end sequence and the 3’ end sequence of the recognition element hybridize to complementary target sequences of interest in the sample to generate a hybridized recognition element; and ii) a code that is unique to the recognition element; b) ligating the 5’ end and the 3 ’ end of the hybridized recognition element, thereby generating a circularized recognition element; c) amplifying the circularized recognition element, thereby generating an amplified recognition element; d) detecting and decoding the code of the amplified recognition element; and e) based at least in part on the decoding of d), determine whether to perform screening for the presence of the target of interest. In some embodiments, the sample is an extracted nucleic acid sample. In some embodiments, the sample is a tissue sample, a cell sample or a blood sample. In some embodiments, the sample is derived from cell-free DNA, genomic DNA, cDNA, or RNA. In some embodiments, the 5’ end sequence and 3’ end sequence of the recognition element adjacently hybridize to the complementary target sequences of interest. In some embodiments, the code comprises at least 10 nucleotides. In some embodiments, the code comprises one or more nucleic acid segments. In some embodiments, the code comprises two or more nucleic acid segments. In some embodiments, the code comprises two to ten nucleic acid segments. In some embodiments, the amplifying comprises exponential amplification, rolling circle amplification, multiple strand displacement amplification, or extension amplification. In some embodiments, the method further comprises amplifying a target region of interest in the sample prior to a). In some embodiments, the amplifying the target region of interest in the sample comprises polymerase chain reaction. In some embodiments, the recognition element is a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5 ’ end and the 3 ’ end of the recognition element. In some embodiments, the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof. In some embodiments, the detecting the code comprises performing a hybridization event by hybridizing one or more fluorescently labelled detection polynucleotides to all or a portion of the code and imaging the hybridization event. In some embodiments, the hybridization event and the imaging of the hybridization event occurs more than one time such that a pattern of fluorescent images is generated which is indicative of the presence of the code. In some embodiments, the decoding the code comprises soft decision decoding. In someembodiments, the soft decision decoding generates a probability of the presence of the code and wherein the presence of the code indicates the presence of the target of interest in the sample. In some embodiments, the probability is high for the presence of the code and the presence of the target of interest in the sample such that additional screening of the sample for the target of interest is indicated. In some embodiments, the probability is low for the presence of the code and the presence of the target of interest in the sample such that additional screening of the sample for the target of interest is not indicated. In some embodiments, the screening for the presence of the target of interest comprises sequencing, microarrays, qPCR, or real-time PCR. In some embodiments, the sequencing comprises sequencing by synthesis. In some embodiments, the target of interest is indicative of the presence of a disease state or an abnormal condition. In some embodiments, the disease state or the abnormal condition is one or more of a cancer, a neurological disease, a cardiac disease, an immunological disease, a gastrointestinal disease, a drug metabolism disease, or an aneuploidy. In some embodiments, the target of interest is used for determining tumor mutational burden of a subject.
[0006] Aspects disclosed herein, in some embodiments, provide methods for determining a sequence of a target of interest, comprising: a) providing an extracted nucleic acid sample, wherein the nucleic acid sample comprises a target of interest; b) amplifying the target of interest from the extracted nucleic acid sample to produce an amplified target of interest; c) splitting all or a portion of the amplified target of interest into a first reaction and a second reaction; d) providing a recognition element to the first reaction, wherein the recognition element comprises i) a 5’ end and a 3 ’ end that are complementary to the target of interest and ii) a code unique to the recognition element, wherein the recognition element and the target of interest hybridize such that the recognition element is ligated to form a circularized recognition element; e) detecting and decoding the code of the circularized recognition element, thereby identifying a sequence of the target of interest; and f) providing a pair of sequencing library preparation primers to the second reaction and amplifying and adapterizing the amplified target of interest in the second reaction, thereby generating a library preparation, and sequencing the library preparation, thereby determining the sequence of a target of interest. In some embodiments, d), e), andf) are performed sequentially. In some embodiments, d), e), and f) are performed concurrently. In some embodiments, the extracted nucleic acid sample is derived from a tissue sample, a cell sample, or a blood sample. In some embodiments, the extracted nucleic acid sample comprises cell-free DNA, genomic DNA, cDNA, or RNA. In some embodiments, the 5’ end and 3 ’ end of the recognition element hybridize adjacently to the complementary target of interest. In some embodiments, the code comprises at least 5 nucleotides. In some embodiments, the code comprises one or more nucleic acid segments. Insome embodiments, the code comprises two or more nucleic acid segments. In some embodiments, the code comprises two to ten nucleic acid segments. In some embodiments, d) further comprises amplifying the circularized recognition element prior to the detection and decoding of e). In some embodiments, the amplifying the circularized recognition element is performed by primer extension, rolling circle amplification or strand displacement amplification. In some embodiments, amplifying the target of interest comprises PCR. In some embodiments, the recognition element is a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5’ end and the 3’ end of the recognition element. In some embodiments, the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof. In some embodiments, the detecting comprises performing a hybridization event by hybridizing one or more fluorescently labelled detection polynucleotides to all or a portion of the code and imaging the hybridization event. In some embodiments, the hybridization event and the imaging of the hybridization event occurs more than once, such that a pattern of fluorescent images is generated. In some embodiments, the decoding the code comprises soft decision decoding. In some embodiments, the soft decision decoding provides a probability of the presence of the code and wherein the presence of the code indicates a presence of the target of interest in the sample. In some embodiments, sequencing the library preparation of d) comprises next generation sequencing. In some embodiments, the next generation sequencing comprises sequencing by synthesis.
[0007] Aspects disclosed herein, in some embodiments, provide methods for determining a sequence of one or more samples, comprising: a) providing a reversibly partitioned first substrate for immobilizing one or more samples; b) immobilizing the one or more samples in one or more partitions of the reversibly partitioned first substrate, wherein the one or more partitions comprises one or more of the one or more samples; c) adding to each partition comprising one of the one or more partitions one or more recognition elements, wherein each recognition element of the one or more recognition elements comprises: i) a 5’ end and a 3 ’ end that are complementary to a target sequence in the one or more samples; and ii) a code that is unique to a recognition element of the one or more recognition elements, wherein the code serves as a proxy for the presence of the target sequence in the one or more samples; d) hybridizing each recognition element of the one or more recognition elements to its complementary target sequence, such that the hybridizing generates a recognition element comprising a padlock probe configuration; e) ligating the 5’ end and the 3 ’ end of the one or more recognition elements to generate one or more circularized recognition elements; f)amplifying the one or more circularized recognition elements; g) removing the reversible partitions from the first substrate; h) affixing a second substrate over the first substrate, thereby sealing the second substrate to the first substrate and generating a flow cell, wherein the flow cell comprises a first reagent inlet port and a second reagent outlet port; i) flowing a plurality for detection polynucleotides through the flow cell, wherein the detection polynucleotides hybridize to the code, or a portion thereof; and j) detecting the hybridization of the detection polynucleotides to the code, or the portion thereof, and decoding the detecting, thereby determining the sequence of the one or more samples. In some embodiments, the one or more samples are extracted nucleic acid samples. In some embodiments, the one or more samples are nucleic acid samples derived from one or more of a tissue sample, a cell sample, or a blood sample. In some embodiments, the extracted nucleic acid samples comprises cell-free DNA, genomic DNA, cDNA, or RNA. In some embodiments, the first substrate comprises optically clear glass. In some embodiments, the second substrate comprises an adhesive film, glass, or plastic. In some embodiments, the partition of the one or more partitions comprise a material selected from the group consisting of polystyrene, polycarbonate, plastic, rubber, silicone, poly ether ketone, and cyclic olefin copolymer. In some embodiments, the second substrate is reversibly applied to the first substrate using one or more interfaces selected from the group consisting of a single sided adhesive, acrylates, silicones, foams, a temperature sensitive adhesive, a fluorinated compound, dimethyl siloxanes, neoprene, and butyl rubber. In some embodiments, a surface of the first substrate comprises a nucleic acid immobilization compound selected from the group consisting of polylysine or an isomer thereof, PEG, cationic polymers, amine coating, amino silanes, silane, thiol / gold chemistry, reversible addition fragmentation chain transfer, and hydrogels. In some embodiments, the code comprises at least 5 nucleotides. In some embodiments, the code comprises one or more nucleic acid segments. In some embodiments, the code comprises two or more nucleic acid segments. In some embodiments, the code comprises two to ten nucleic acid segments. In some embodiments, the amplifying is performed by primer extension, rolling circle amplification or strand displacement amplification. In some embodiments, the recognition element comprises a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5’ end and the 3’ end of the recognition element. In some embodiments, the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof. In some embodiments, the detection polynucleotides are fluorescently labelled. In some embodiments, the detection further comprises imaging the hybridization of the detection polynucleotides to the code, or the portion thereof. In some embodiments, the flowingof i), the detecting of j), and the imaging is repeated more than once, such that a pattern of symbols is generated. In some embodiments, the decoding the detecting comprises soft decision decoding. In some embodiments, the soft decision decoding provides a probability of the presence of the code, wherein the code serves as a proxy for the presence of the sequence of the one or more samples.
[0008] Aspects disclosed herein, in some embodiments, provide methods for determining a sequence of a target of interest, comprising: a) generating a sequencing library from a first portion of a sample comprising a target of interest; b) performing whole exome sequencing on the sequencing library, thereby determining the sequence of the target of interest; c) providing one or more recognition elements to a second portion of the sample comprising the target of interest, wherein each recognition element of the one or more recognition elements comprises i) a 5’ end and a 3’ end which are complementary to the target of interest, and ii) a code that is unique to each recognition element, wherein the code serves as a proxy for the presence of the target of interest; d) hybridizing the one or more recognition elements to their complementary sequence in the target of interest to produce hybridized recognition elements; e) ligating the 5’ end and the 3’ end of the hybridized recognition elements to generate circularized recognition elements; f) amplifying the circularized recognition elements to generate amplified recognition elements; and g) detecting and decoding the code of the amplified recognition element; wherein the sequence of the target of interest is determined by combining whole exome sequencing data from b) and decoding data from g). In some embodiments, the target of interest is derived from a tissue sample, a cell sample, or a blood sample. In some embodiments, the target of interest is derived from cell-free DNA, genomic DNA, cDNA, or RNA. In some embodiments, the target of interest comprises a copy number variance, high GC content, one or more homopolymeric regions, an insertion, or a deletion. In some embodiments, the target of interest is a star allele and comprises a homolog or a pseudogene, and wherein determining the sequence of the target of interest differentiates the target of interest from the homolog or the pseudogene.
[0009] Aspects disclosed herein, in some embodiments, provide compositions comprising: a) a recognition element, wherein the recognition element comprises a 5’ end and a 3 ’ end, wherein the 5’ end and the 3’ end are hybridized to non -adjacent target nucleic acid sequences; and b) a third oligonucleotide that is hybridized to all or a portion of a nucleic acid sequence that is located between the non-adjacent target nucleic acid sequences that are hybridized such that a 3 ’ end of the third oligonucleotide is hybridized adjacent to the 5 ’ end of the recognition element and a 5’ end of the third oligonucleotide is hybridized adjacent to the 3’ end of the recognition element; and wherein the third oligonucleotide, when hybridized to the target nucleic acid sequences, forms a loop structure such that the loop structure is not hybridized to the nucleicacid sequence that is located between the non -adjacently hybridized target nucleic acid sequences. In some embodiments, a portion of the loop structure is a stem-loop structure. In some embodiments, a portion of the loop structure is not a self -hybridizing loop structure. In some embodiments, the third oligonucleotide comprises one or more of a primer binding site, a sample index, a unique molecular identifier and a cleavage site. In some embodiments, the recognition element comprises a code and one or more of a primer binding site, a unique molecular identifier, a sample index and a cleavage site. In some embodiments, the 5 ’ end of the recognition element and the 3 ’ end of the third oligonucleotide and the 3 ’ end of the recognition element and the 5’ end of the third oligonucleotide are ligated together to form a circularized recognition element comprising the recognition element and the third oligonucleotide.
[0010] Aspects disclosed herein, in some embodiments, provide methods for determining a presence of a target of interest, comprising: a) hybridizing a first amplification primer to a ligated recognition element, wherein the ligated recognition element comprises a complementary sequence of the target of interest, a hypercode, a sample index, and optionally a unique molecular identifier, wherein the first amplification primer comprises a first sequence which is capable of being captured on a flow cell; b) performing primer extension of the first amplification primer that is hybridized to the ligated recognition element, thereby generating a plurality of amplicons, wherein each amplicon of the plurality of amplicons comprises a sequence of the target of interest, or a complement thereof, the hypercode, the sample index, and optionally the unique molecular identifier and the first sequence which is capable of being captured on the flow cell; c) hybridizing a second amplification primer to the plurality of amplicons, wherein the second amplification primer comprises a second sequence which is capable of being captured on the flow cell; d) performing primer extension of the second amplification primer that is hybridized to the plurality of amplicons, thereby generating a plurality of double-stranded amplicons wherein each double-stranded amplicon of the plurality of double-stranded amplicons is flanked on a first end by the first sequence and on a second end by the second sequence; e) denaturing the plurality of double-stranded amplicons to produce denatured amplicons and capturing a plurality of single strands of the denatured amplicons on the flow cell; and f) sequencing at least the hypercode, the sample index, and optional the unique molecular identifier of the captured amplicons, or complements thereof, thereby determining the presence of the target of interest. In some embodiments, the target of interest is derived from a sample. In some embodiments, the target of interest is a variant target from a list consisting of a single nucleotide polymorphism, an insertion, a deletion, a splice variant and a copy number variant. In some embodiments, the sample is a tumor sample, a tissue sample, a cell sample, a blood sample, or a lysate. In some embodiments, the sample is obtained or derived from aeukaryote, a prokaryote, a mammal, a veterinary office, an agricultural sample, a companion animal, or a livestock animal. In some embodiments, the ligated recognition element is generated by practicing the method of Fig. 2. In some embodiments, the primer extension of b) or d) is performed for a plurality of cycles, thereby generating a plurality of amplicons. In some embodiments, a sequence of the hypercode is used as a proxy for identifying the presence of the target of interest. In some embodiments, the first sequence of the first amplification primer and the second sequence of the second amplification primer are further sequencing primer binding sites for hybridizing complementary sequencing primers for initiating sequence by synthesis. In some embodiments, the first sequence is captured on the flow cell and the second sequence is hybridized to the complementary sequencing primer for initiating sequence by synthesis. In some embodiments, the second sequence is captured on the flow cell and the first sequence is hybridized to the complementary sequencing primer for initiating sequence by synthesis. In some embodiments, the sequencing comprises sequence by synthesis. In some embodiments, the sequencing comprises performing bridge amplification on the single strands of the denatured amplicons, thereby forming colonies of single -stranded amplicons on the surface of the flow cell. In some embodiments, the sample index is alternatively located on the first amplification primer or the second amplification primer.
[0011] Aspects disclosed herein, in some embodiments, provide methods for determining a presence of a target of interest, comprising: a) hybridizing an amplification primer to a ligated recognition element, wherein the ligated recognition element comprises an amplification primer binding site, a hypercode, a sample index, and a sequence that is capable of being captured on a flow cell, wherein the amplification primer comprises a complement of the primer binding site and a sequence primer binding site which is complementary to a sequencing primer; b) performing primer extension on the amplification primer that is hybridized to the ligated recognition element, thereby generating an amplicon comprising the hypercode, the sample index, the sequence that is capable of being captured on the flow cell, and the sequence primer binding site which is complementary to the sequencing primer; c) capturing the amplicon on the flow cell; and d) sequencing at least the hypercode and the sample index of the captured amplicon, thereby determining the presence of the target of interest. In some embodiments, the target of interest is derived from a sample. In some embodiments, the target of interest is a variant target from a list consisting of a single nucleotide polymorphism, an insertion, a deletion, a splice variant and a copy number variant. In some embodiments, the sample is a tumor sample, a tissue sample, a cell sample, a blood sample or a lysate. In some embodiments, the sample is obtained or derived from a eukaryote, a prokaryote, a mammal, a veterinary office, an agricultural sample or a livestock sample. In some embodiments, the ligated recognition elementis generated by practicing the method of Fig. 2. In some embodiments, the primer extension is performed for one cycle, thereby generating one amplicon. In some embodiments, a sequence of the hypercode is used as a proxy for identifying the present of the target of interest. In some embodiments, the ligated recognition element alternatively comprises the sequence primer binding site which is complementary to the sequencing primer and the amplification primer comprises the sequence that is capable of being captured on the flow cell. In some embodiments, sequencing comprises providing two sequencing primers. In some embodiments, one of the two sequencing primers is complementary to the sequence primer binding site and the second of the two sequencing primers is complementary to the amplification primer binding site. In some embodiments, the sequencing comprises sequence by synthesis. In some embodiments, the sequencing comprises performing bridge amplification on the amplicon captured on the flow cell, thereby forming colonies of the amplicon on the flow cell. In some embodiments, the ligated recognition element further comprises a unique molecular identifier. In some embodiments, the ligated recognition element is a RNA ligated recognition element. In some embodiments, the primer extension is reverse transcription of the RNA ligated recognition element, thereby generating a cDNA amplicon.INCORPORATION BY REFERENCE
[0012] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The novel features of the inventive concepts are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying drawings of which:
[0014] Fig. 1 shows an illustration of an example of a recognition element as disclosed herein.
[0015] Fig. 2 shows an illustration of an example of a workflow wherein a linear recognition element is circularized and ligated if hybridized to a target nucleic acid followed by amplification of the ligated recognition element to generate a detectable concatemeric amplification product.
[0016] Fig. 3 A shows an example of detection of a code of a recognition element and an example of a detection polynucleotide.
[0017] Fig. 3B shows an example of detection of a code of a recognition element and an example of a detection polynucleotide hybridized to a portion of an example of a concatemeric amplification product.
[0018] Fig. 4 is an example of an instrument that comprises a computer system for use with the methods described herein.
[0019] Fig. 5 is an example of an application provision system for use with the methods described herein.
[0020] Fig. 6 is an example of an application provision system for use with the methods described herein.
[0021] Fig. 7 shows a schematic of an example of a soft decision decoding workflow for decoding a code of a recognition element.
[0022] Fig. 8 shows an example of a substrate comprising removable partitions.
[0023] Fig. 9A shows an example of a flow cell and a first substrate with removable partitions and circularized nucleic acids in the wells.
[0024] Fig. 9B shows an example of a flow cell and the first substrate of Fig. 9A with the partitions removed and a second substrate affixed to the first substrate to generate a flow cell with circularized nucleic acids in the flow cell.
[0025] Fig. 10A shows an example of a flow cell and a first substrate with removable partitions and circularized nucleic acids in wells.
[0026] Fig. 10B shows an example of a flow cell and the first substrate with the partitions removed and a second substrate with non-removable partitions.
[0027] Fig. 10C shows an example of a flow cell and the second substrate with nonremovable partitions affixed to the first substrate to generate a flow cell, in this example, a five- channel flow cell with circularized nucleic acids in the flow cell.
[0028] Fig. 11A shows a top-down view of the flow cell from Figs. 10A-10C and a first substrate with a removable partition tool for generating a plurality of discrete locations on which circularized nucleic acids can be immobilized on the first substrate,
[0029] Fig. 11B shows a top-down view of the flow cell from Figs. 10A-10C and the first substrate with the removable partition tool removed.
[0030] Fig. 11C shows a top-down view of the flow cell from Figs. 10A-10C and the first substrate with a second substrate affixed to the first substrate thereby generating a flow cell, in this example a four chambered flow cell, retaining the discrete locations on which circularized nucleic acids can be immobilized.
[0031] Fig. 12 shows an example of Strategy 1 and Strategy 2 for digitally partitioning a sample.
[0032] Fig. 13 shows an example of Strategy 3 for digitally partitioning a sample.
[0033] Fig. 14 shows an example of Strategy 4 for digitally partitioning a sample.
[0034] Figs. 15A-15C shows a first example strategy for decoding by sequencing directly from a recognition element.
[0035] Fig. 15A shows an example of a ligation recognition element.
[0036] Fig. 15B shows an example of a recognition element in an extension reaction.
[0037] Fig. 15C shows an example of an amplicon applied to a sequencing flow cell.
[0038] Figs. 16A-16C shows a second example strategy for decoding by sequencing directly from a recognition element.
[0039] Fig. 16A shows an example of a ligation recognition element.
[0040] Fig. 16B shows an example of an amplicon.
[0041] Fig. 16C shows an example of an amplicon applied to a sequencing flow cell.DETAILED DESCRIPTION
[0042] Identification of the presence or absence of a target of interest from a sample can sometimes be difficult, such that not one assay will serve the purpose. For example, an assay that identifies the actual base sequence of a target of interest can be challenged if the target of interest has high GC content, has homopolymer regions, or has other challenges which can render direct sequence of a target of interest problematic. However, complementing a sequence assay with an indirect method for identifying the presence of a target sequence of interest can supply the whole picture, with regards to confidently determining whether a target of interest is present in a sample. Such confidence is needed if the target of interest is associated with a disease state or a medical condition and if the purpose of identifying the target of interest is to screen a sample from a subject suspected of having a disease, cancer, or other adverse medical condition.
[0043] As such, the present disclosure provides for assays, and combinations of assays that can be used in a synergistic manner for identifying targets of interest and devices or options associated therewith.
[0044] The present disclosure provides methods and systems for determining the presence of a target nucleic acid. Recognition elements comprising codes which correlate to a target nucleic acid are used in combination with sequencing chemistries, wherein the code sequences either alone or in combination with sequencing methodologies enable the identification of the presence or absence of a target nucleic acid in a sample of interest.
[0045] The methods and compositions as disclosed herein for determining a nucleic acid sequence comprise an assay. In some embodiments, the assay is a solution -based assay. In some embodiments, the assay is a surface-bound assay. In some embodiments, the assay is a hybrid assay that includes a surface-bound component and a solution-based component. In some embodiments, the assay is performed in one or more tubes or in a plate -based format, such as a multi-well plate. A multi-well plate may include, for example, a 12, 24, 48, 96, 384, or 1536 well plate. In some embodiments, a multi-well plate may include, for example, an array of wells. In some embodiments, the assay may be performed on a microfluidics device. In some embodiments, the assay may be performed partially in one or more tubes and partially in a multi -well plate.
[0046] In some embodiments, a recognition element is used in the assay. A recognition element used in the assay may include sequences that are complementary to a target sequence of interest, wherein the complementary sequence can hybridize to a target sequence of interest, a code sequence that can be used to identify the target sequence of interest that has hybridized to its complement on the recognition element, and one or more functional sequences such as sequencing primer binding sites, one or more amplification primer binding sites, unique molecular identifier sequences (UMIs), sample indexes, or combinations thereof. In some embodiments, an amplification primer binding site may be adjacent to the code in a recognition element. The amplification primer binding site(s) may, in some cases, comprise universal primer sequence(s) that are common to all recognition elements in a set of recognition elements. Amplification primer binding site sequences may also be a code sequence or a portion thereof. A code sequence can be a combination of a number of subsequences, called nucleic acid segments, wherein their combination can identify a target sequence of interest that has hybridized to a recognition element. In some embodiments, an amplification primer can be a nucleic acid segment or a portion thereof. Unique identifier sequences (UMIs) and sample indexes, which may find utility in next generation sequencing (NGS) reactions for counting, error correction and sample identification purposes, may also be part of the code such as one or more nucleic acid segments.
[0047] Once a recognition element has recognized and hybridized to its target of interest, the recognition element may be circularized and ligated to generate a circular, ligated recognitionelement. The circular and ligated recognition element can be amplified in anticipation of a decoding event to identify the code associated with the original target of interest that hybridized to the recognition element. The amplification may be by any method of amplification, including for example, nucleic acid extension, polymerase chain reaction (PCR), isothermal amplification, rolling circle amplification (RCA), ultrarapid amplification, and the like. Surface based amplification may be performed using surface-anchored primers (e.g., Illumina bridge amplification technology), or recombinase polymerase amplification (RPA) (e.g., ExAmp technology).
[0048] In one embodiment, the amplification operation may comprise a rolling circle amplification (RCA) reaction to generate concatemeric amplification products.
[0049] In one embodiment, a recognition element may include a sequence which may prevent RCA of the recognition element while allowing for linear double -stranded PCR products. The non-extendable sequence may, for example, be located between a pair of amplification primer binding site sequences present on the recognition element. In one embodiment, a recognition element may include a restriction enzyme site that may be cleaved to yield a linear DNA molecule.
[0050] In some embodiments, a concatemeric amplification product may be sequenced to determine the nucleotide sequence of the code associated with the target molecule of interest. Any sequencing technology may be used to sequence the product. Non -limiting examples of sequencing technologies that may be used include sequencing by synthesis (SBS), avidity sequencing, sequencing by hybridization, sequencing by ligation, nanopore sequencing, and the like.
[0051] In some embodiments, a sequencing library may be generated from a set of recognition elements or complements or amplicons thereof. The sequencing library may be sequenced to determine the code of the recognition element associated with a target molecule of interest. The code sequence may be used as a digital count of the target molecule specific decoding event. In one embodiment, a sequencing library may be generated from a circularized recognition element. In another embodiment, a sequencing library may be generated from a concatemeric amplification product of a recognition element. In one embodiment, a concatemeric amplification product or a portion thereof that includes at least the code may be directly sequenced to determine each base of the nucleotide sequence of the code associated with the target molecule of interest.
[0052] Fig. 2 is an example of an encoded assay for use with the detection polynucleotides disclosed herein. A linear recognition element 210 may comprise a 5' end 220a which may be complementary to a portion of a target nucleic acid of interest 222, a 3' end 220b which may becomplementary to another portion of a target nucleic acid of interest 222 from a sample, a code 216, and additional functional sequences 212, 214, 218 such as amplification primer binding sites, capture sequences, cleavage sites, sequencing primer binding sites, unique molecule identifiers (UMIs), and the like. A target nucleic acid of interest 222 which may be complementary to the 5' 220a and 3' 220b ends of the linear recognition element 210 may hybridize to the linear recognition element, thereby bringing the ends in proximity for ligating to generate a circular and ligated recognition element 225. The circular and ligated recognition element 225 may be subjected to extension amplification using, for example, one of the functional sequences 212, 214, 218, or the code 216 or a portion thereof, as a primer binding site. The result may be a concatemeric amplification product 230 which can be decoded using the detection polynucleotides disclosed herein for identifying and determining the presence of the target nucleic acid of interest from a sample.
[0053] Additional examples of encoded assays can be found in WO2022 / 109496A2 and WO2023 / 096674A1 , both of which are incorporated herein by reference in their entireties.
[0054] The methods and compositions described herein may include providing recognition elements to an encoded assay for identifying the presence of a target molecule of interest from a sample. In some embodiments, a plurality of recognition elements is provided. In some embodiments, each recognition element in the plurality of recognition elements comprises one or more target recognition regions. The target recognition regions of the recognition elements may comprise one or more nucleic acid sequence(s) configured to hybridize to a target nucleic acid molecule. In some embodiments, the one or more nucleic acid sequences may hybridize to one or more target nucleic acid sequences of the target nucleic acid molecule. In some embodiments, the target recognition region may be configured to hybridize to one or more regions of the target nucleic acid molecule flanking a target of interest (e.g., SNP, insertion - deletion mutations (indel), and so on). In some embodiments, the target recognition region may be configured to hybridize to a variant nucleic acid of interest (e.g., the target recognition region base pairs with the SNP when the target of interest is a SNP).
[0055] As shown in Fig. 1, an example of a recognition element used in encoded assays described herein may comprise two target recognition regions, one at the 5’ prime end and another at the 3’ end. A recognition element may further comprise a code. As shown in Fig. 1, the example code may comprise four nucleic acid segments. However, the number of nucleic acid segments is not limiting. The number of segments may depend on the complexity of the assay (e.g., how many targets of interest are to be identified at one time). In some embodiments, a code may comprise more than or equal to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 segments. In some embodiments, a code maycomprise less than or equal to 30, 29, 28, 27, 26, 25, 24, 24, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 segments. In some embodiments, a target nucleic acid molecule of interest itself comprises the code, or an additional code or sequence that identifies another target molecule of interest such as a protein. In some embodiments, the code may be detected as a proxy for the target molecule, be it a nucleic acid or a protein, or both.
[0056] In some embodiments, the structure of the recognition element may vary. In some embodiments, the structure of the recognition element may configure into a specific structure when hybridized to a target nucleic acid. Non-limiting examples of a recognition element configuration may include a padlock probe, a molecular inversion probe, a hairpin oligonucleotide, a single-stranded oligonucleotide, a double-stranded oligonucleotide, a combination thereof, and the like. In some embodiments, the recognition element is linear prior to hybridization to its complementary target nucleic acid of interest. In some embodiments, the linear recognition element is circularized once hybridized to the respective target nucleic acid of interest and ligated thereafter. In some embodiments, the recognition element is circular prior to hybridization to the target nucleic acid of interest. In some embodiments, the target nucleic acid of interest may serve as a primer for an extension reaction, for example to initiate rolling circular amplification of the recognition element once hybridized to the recognition element.
[0057] In some embodiments, the recognition element is configured to be a padlock probe once the recognition element is hybridized to the target nucleic acid of interest. Padlock probes may be referred to as linear oligonucleotides whose ends may be complementary to adjacent target sequences, or to non-adjacent target sequences thereby leaving a gap between the ends of the hybridized recognition element. Upon hybridization to a target nucleic acid, the two ends (e.g., the 5’ end and the 3 ’ end) of the recognition element may be adjacently located, generating a padlock probe configuration for subsequent ligation. Alternatively, the two ends of the recognition element may be brought in proximity to, but not directly adjacent to, each other upon hybridization to a target nucleic acid. In this example, a gap may be leftbetween the 5' and 3' hybridized ends of the recognition element which can be filled in several ways, for example by extension of the 5' end until it is adjacent to the 3' end, or by hybridizing a third oligonucleotide that fills the gap. In any scenario, the recognition element ends may be ligated together if hybridization, or hybridization and gap fill, occurs thereby generating circular and ligated recognition elements that are indicative of the hybridization event.
[0058] In some embodiments, a recognition element may further comprise one or more functional sequences. Functional sequences may include, but are not limited to, primer binding sites, cleavage sites, unique molecular identifiers (UMIs), capture sequences, or combinations thereof. In some embodiments, a primer binding site and / or a cleavage site may be universal innature, such that a plurality of recognition elements may share the same sequence(s). A unique molecular identifier maybe included in a recognition element to identify a source of material, for example, for error correction.
[0059] The methods described herein may relate to the use of a code for identifying a target nucleic acid of interest from a sample that hybridized to a recognition element to initiate a ligation event. A code in a recognition element may be used to associate the recognition element 5' and 3' end regions with a target nucleic acid of interest, thereby determining the presence of a target nucleic acid of interest from a sample without having to directly assay the target molecule itself. As such, a code in a recognition element may uniquely identify the presence of a target molecule from a sample. Using codes, any number of recognition elements can be multiplexed in one encoded assay as each code is unique and correlates to the presence of one target molecule.
[0060] In some embodiments, the code is selected from a set of codes wherein the set of codes make up a “code space”. In some embodiments, the code comprises a plurality of nucleic acid segments, where each nucleic acid segment corresponds to one or more computational symbols, or colors, that are used in a decoding process. Fig. 1 shows an example of a recognition element where four nucleic acid segments make up the code of the recognition element, wherein each of the nucleic acid segments can be decoded using detection polynucleotides as disclosed herein and the combination of the decoded nucleic acid segments thereby builds the full code that is unique to the target nucleic acid of interest from a sample. The codes may be detected as proxies, thereby serving as an indirect analysis of the presence of a target molecule from a sample as the code correlates with the presence of the target molecule that hybridized to the recognition element allowing ligation, amplification and decoding. If there is no hybridization of a target of interest to its complementary sequences of a recognition element, there may be expected to be no ligation (e.g., as the 5’ and the 3’ ends of the recognition element are not expected to be adjacent), no amplification and subsequently nothing to decode. As such, if there is no amplification product to decode, that may be an indication that the target molecule of interest may be potentially absent from the sample, or at such a low incidence that hybridization resulted in too few amplification products to cross the threshold for detection by decoding.
[0061] In some embodiments, each code from a set of codes may be from a predetermined set of codes. In some embodiments, each code from the set of codes may be selected to ensure that the selected code differs from other codes in the set of codes. As such, in some embodiments, several selection criteria may be implemented to generate a set of codes, wherein each code of a set of codes comprises from one to more than one nucleic acid segment. In someembodiments, selection of the codes, or nucleic acid segments that make up a code, may incorporate a Hamming distance criterion.
[0062] In some embodiments, to generate a code selected from a set of codes for use in a recognition element, a Hamming distance (HD) selection criterion may be implemented between any two codes of the set of codes, and also between any two nucleic acid segments that may be used in a code. A Hamming distance between two codes in a set of codes may refer to the number of symbols, or nucleotides, that differ between the two codes in the set of codes. The Hamming distance may measure the number of changes that may need to be made to a first code sequence to change the string of symbols, in this case nucleotides, to the second code. As such, a Hamming distance criterion used to select a code may not be greater than the length of the code. For example, if the length of a code is measured by the number of cycles or flows of decoding runs or queries and that number being eight cycles, and if each cycle corresponds to one symbol or color, therefore eight symbols or colors, then the maximum Hamming distance is eight. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2-10. In some embodiments, the Hamming distance may be between 2-7. In some embodiments, the Hamming distance may be between 3-5. The Hamming distance may increase as the number of codes that can be used decreases, as one purpose of the code is to impart a way to uniquely identify one target molecu le from another target molecule.
[0063] The code may have a certain length in nucleotides. In some embodiments, the code has a length of greater than or equal to about three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous nucleotides. In some embodiments, the code has a length of fewer than or equal to about 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 contiguous nucleotides. In some embodiments, the length of the code is about 5 to 200, about 10 to 150, about 15 to 100, about 20 to 90, or about 30 to 80 contiguous nucleotides. The length of the code may be further divided into a number of discrete nucleic acid segments, which can be combined to yield a code, as shown, for example, in Fig. 1.
[0064] In some embodiments, each code from a set of codes is generated using a 4 -ary nucleotide alphabet of A, C, G, and T. In some embodiments, each code of a set of codes is generated using a 3 -ary nucleotide alphabet of a set of three of A, C, G, and T. In some embodiments, the codes can be generated from arbitrary symbols, or colors, 1 to 4, corresponding to fluorophores that are associated with, and help to detect, a unique string ofnucleotides. The numbers are used in the abstract but serve as a mechanism to numerate colors or mixes of colors that are utilized to query the codes, thereby building a code profile for decoding.
[0065] A code of a recognition element can comprise one or more nucleic acid segments. For example, as shown in Fig. 1, the example of the code comprises four nucleic acid segments.
[0066] In some embodiments, the recognition elements provided herein may comprise a code comprising one or more nucleic acid segments. The one or more nucleic acid segments, or the complements thereof, within the code may be used as a proxy for detection of the target molecule recognized by the recognition element.
[0067] The number of nucleic acid segments present in a code of a recognition element may be considered in the design of the recognition element. The number of segments in a code may help to determine the nucleotide length of the recognition element. For example, a recognition element that includes a code comprising five segments may comprise a greater nucleotide length than a recognition element that includes a code of two segments. A recognition element with a larger nucleotide length may run up against synthesis limits and may be at a greater risk of synthesis errors. Alternatively, a recognition element with a smaller nucleotide length may avoid synthesis limits and risks in synthesis errors. A recognition element with a larger nucleotide length may include less space for other portions of the recognition element, such as the target recognition regions, functional sequences, universal sequences, etc.
[0068] In some embodiments, the code comprises about 2 to 10 nucleic acid segments. In some embodiments, the code comprises about 2 to 8 nucleic acid segments. In some embodiments, the code comprises about 3 to 5 nucleic acid segments. In some embodiments, the code comprises at least 1 nucleic acid segments, at least 2 nucleic acid segments, at least 3 nucleic acid segments, at least 4 nucleic acid segments, at least 5 nucleic acid segments, at least 6 nucleic acid segments, at least 7 nucleic acid segments, at least 8 nucleic acid segments, at least 9 nucleic acid segments, at least 10 nucleic acid segments, at least 11 nucleic acid segments, at least 12 nucleic acid segments, at least 13 nucleic acid segments, at least 14 nucleic acid segments, or at least 15 nucleic acid segments.
[0069] In some embodiments, each nucleic acid segment may comprise a length in nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 30 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 25 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 15 to 20 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of 2 or more nucleotides, 3 or more nucleotides, 4 or more nucleotides, 5 or more nucleotides, 6 or more nucleotides, 7 or more nucleotides, 8 or more nucleotides, 9 ormore nucleotides, 10 or more nucleotides, 11 or more nucleotides, 12 or more nucleotides, 13 or more nucleotides, 14 or more nucleotides, 15 or more nucleotides, 16 or more nucleotides, 17 or more nucleotides, 18 or more nucleotides, 19 or more nucleotides , 20 or more nucleotides, 21 or more nucleotides, 22 or more nucleotides, 23 or more nucleotides, 24 or more nucleotides, or 25 or more nucleotides. In some embodiments, each of the nucleic acid segments in a code are of the same length. In some embodiments, each of the nucleic acid segments in a code are not the same length.
[0070] A Hamming distance selection criterion may be implemented between any two nucleic acid segments of a code. A Hamming distance between two nucleic acid segments in a code may refer to the number of symbols that differ between the segments. The Hamming distance may measure the number of changes that may need to be made to a first nucleic acid segment to change the string of symbols, or nucleotides, to a second nucleic acid segment. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance maybe from about 2-20, about 3-19, about 4-18, about 5-17, about 6-16, about 7-15, about 8-14, about 9-13, or about 10-12. In some embodiments, a minimum Hamming distance may be greater than or equal to about 2, greater than or equal to about 3 , greater than or equal to about 4, greater than or equal to about 5, greater than or equal to about 6, greater than or equal to about 7, greater than or equal to about 8, greater than or equal to about 9, greater than or equal to about 10, greater than or equal to about 11, greater than or equal to about 12, greater than or equal to about 13 , greater than or equal to about 14, greater than or equal to about 15, greater than or equal to about 16, greater than or equal to about 17, greater than or equal to about 18, greater than or equal to about 19, or greater than or equal to about 20.
[0071] In some embodiments, a nucleic acid segment may comprise a universal primer binding site for amplification. For example, a segment may comprise an amplification primer binding site for performing rolling circle amplification (RCA) for generating a plurality of concatemeric amplification products.
[0072] In some embodiments, a nucleotide or nucleic acid sequence of each segment may correspond to one or more computational symbols, such as a detection color, for performing a decoding process. For example, one or more nucleic acid segments of a code may be detected with a first pool of detection polynucleotide complexes to produce one or more detectable binding complexes, for example by using a fluorescent label. In some embodiments, the one or more detectable binding complexes, once imaged, may produce one or more optical signals such as fluorescence in a particular wavelength. When all or substantially all segments of the code aredetected by iteratively applying additional pools of detection polynucleotide complexes to the amplification products, a series of optical signals may be observed and collated.
[0073] The application of detection polynucleotides to amplification products for decoding can be called a “flow” or “cycle” or “query”, wherein a flow, cycle or query is the number of times a particular segment of an amplification product is queried, or the numb er of times a detection polynucleotide is flowed over an amplification product in order to detect a nucleic acid segment sequence or a portion thereof. If a nucleic acid segment is present, a detection polynucleotide that comprises a sequence complementary to that nucleic acid segment or portion thereof may hybridize to its complementary nucleic acid segment and the attached detectable label is detected, for example imaged. In some embodiments, one or more optical signals observed from querying an amplification product with detection polynucleotide complexes translates to one or more computational symbols such that each optical signal can be imaged, and if multiple images are captured an image profile generated, and decoded. In some embodiments, a plurality of nucleic acid segments on a recognition element may correspond to at least three computational symbols. In some embodiments, the optical signal may be a color or a non-color. In some embodiments, the optical signal may be a combination of colors (e.g., when the detection polynucleotide complex comprises a plurality of detectable labels). In some embodiments, the computational symbols or colors can be referred to as numbers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.). In some embodiments, each detection polynucleotide complex comprises a detectable label, for example when four different fluorescent moieties are used as detectable labels there are four possible symbols, 1 to 4. However, the number of symbols can be larger depending on the combination of detectable labels with each unique detection polynucleotide complex. For example, for a set of 16 unique detection polynucleotide complexes wherein each detection polynucleotide complex has one of four fluorescent moieties, there may be 16 computational symbols used for decoding if all 16 unique detection polynucleotide complexes are used to decode an amplification product. However, additional ways to increase the number of computational symbols for decoding include, but are not limited to, adding levels of identifiability associated with a particular detectable signal such as whether a detectable signal is brighter or dimmer compared to a determined normal level of signal, or whether there is a combination of detectable colors, a profile, that is used to identify a particular nucleotide. As such, the number of computational symbols that may be used may be limited by practicality for any given assay.
[0074] In some embodiments, the methods described herein may use a number of computational symbols. The number of computational symbols used in the methods and systems described herein may be considered in the design of the recognition elements. For example, insome embodiments, a detection scheme using a larger number of computational symbols may lead to a larger code space and a greater number of codes that may be generated, which may allow for a greater amount of information that may be detected thereby allowing for a higher degree of assay target molecule multiplexing. In some embodiments, a detection scheme using a smaller number of computational symbols may be limited in the amount of information that can be detected. In some embodiments, using a larger numb er of computational symbols may result in a faster detection process (less time to determine a target molecule compared to using a smaller number of computational symbols). In some embodiments, a detection scheme using a larger number of computational symbols may require greater instrument complexity, which may lead to potential drawbacks such as color crosstalk, wherein the computational symbols used in the detection scheme may become difficult to distinguish from other computational symbols. Fluorescent emissions that overlap is one example of computational symbols that may be difficult to distinguish due to potential overlap in emission wavelength detected. In some embodiments, a greater number of computational symbols may require that a more complex detection tool be used.
[0075] In some embodiments, each nucleic acid segment may correspond to a combination of computational symbols. In some embodiments, each nucleic acid segment may correspond to one or more computational symbols, two or more computational symbols, three or more computational symbols, four or more computational symbols, five or more computational symbols, six or more computational symbols, seven or more computational symbols, eight or more computational symbols, nine or more computational symbols, or 10 or more computational symbols. In some embodiments, each nucleic acid segment may correspond to 10 or less computational symbols, nine or less computational symbols, eight or less computational symbols, seven or less computational symbols, six or less computational symbols, five or less computational symbols, four or less computational symbols, three or less computational symbols, or two or less computational symbols.
[0076] The methods described herein include amplification of a circularized and ligated recognition element. In some embodiments, a target nucleic acid molecule is amplified. In some embodiments, the target nucleic acid molecule is a combination of a recognition element and a target nucleic acid molecule. In some embodiments, the amplification is selective amplification. For example, in some embodiments, amplification is able to occur if a target recognition region of a recognition element recognizes and binds to a complementary target nucleic acid of interest. In some embodiments, amplification occurs if a primer is used that is complementary to one or more of a portion of a recognition element, a portion of a nucleic acid segment of a code, or another sequence in the recognition element that is complementary to a primer used foramplification. In some embodiments, the amplification is non -selective. For example, in some embodiments randomers can be used to prime amplification from a recognition element.
[0077] In some embodiments, the methods described herein may include selectively amplifying a subset of nucleic acids. For example, in some embodiments, a subset of a plurality of recognition elements hybridized to a plurality of target nucleic acid molecules may be amplified. The subset may comprise a percentage of the total amount of recognition elements hybridized to target nucleic acid molecules as described herein. In some embodiments, the subsetmay include 5% ormore, 10% or more, 15% or more, 20% or more, 25% or more, 30% or more, 35% ormore, 40% ormore, 45% or more, 50% or more, 55% or more, 60% or more, 65% ormore, 70% or more, 75% ormore, 80% or more, 85% ormore, 90% or more, or 95% or more of the total amount of recognition elements hybridized to target nucleic acid molecules. In some embodiments, the subset may include 95% or less, 90% or less, 85% or less, 80% or less, 75% orless, 70% or less, 65% or less, 60% orless, 55% or less, 50% or less, 45% or less, 40% or less, 35% orless, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the total amount of recognition elements hybridized to target nucleic acid molecules.
[0078] In some embodiments, the amplification may include rolling circle amplification (RCA). In some embodiments, the RCA may generate a concatemer as an amplification product, wherein the concatemer comprises multiple copies of a circularized ligated recognition element, including associated codes, target recognition regions, and any other sequences that are included in the circularized and ligated recognition element. In some embodiments, RCA may be performed while the circularized and ligated recognition element is in solution. In some embodiments, RCA may be performed on a circularized recognition element while the circularized recognition element is immobilized, either reversibly or non -reversibly, on a substrate or surface. In some embodiments, the substrate or surface is a solid support and includes, but is not limited to, a bead, a flow cell, a microwell, a nanowell, a well, a slide. In some embodiments, the substrate is glass such as optical glass of imaging quality. In some embodiments, the substrate is plastic, polycarbonate, etc. In some embodiments, the substrate is positively charged or negatively charged. In some embodiments, the substrate is an anionic substrate. In some embodiments, the substrate is a cationic substrate. In some embodiments, the substrate comprises an immobilization composition, such as polyacrylamide, branched PEI, linear PEI, poly(P-aminoester) and poly(amidoamine), PEG, a gel, poly -L-ly sine, silane, agarose, muscle mimetic catecholamine polymer, and the like. In some embodiments, the substrate has no charge. In some embodiments, the substrate has no immobilization composition. In some embodiments, a substrate comprises a cationic polymer coated surface. An RCA reaction may be performed in the presence of a cationic polymer coated surface, resulting insimultaneous immobilization and amplification of a ligated recognition element. RCA primers may be supplied in solution or bound to the cationic polymer-coated surface prior to, or concurrent with, performing the RCA reaction.
[0079] In some embodiments, amplification may include on-surface polymerase chain reaction (PCR), isothermal amplification, RCA, ultrarapid amplification, or a combination thereof. In some embodiments, amplification may include polymerase chain reaction (PCR). In some embodiments, PCR is multiplexed PCR. In some embodiments, PCR is ultrafast multiplexed PCR. The amplification methods disclosed herein may include isothermal amplification. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification (TMA), Loop -mediated isothermal amplification (LAMP), Helicase-dependent amplification (HD A), Nucleic Acid Sequence Based Amplification (NASBA), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Fakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci. 2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety.
[0080] The methods described herein may include introducing detection polynucleotides to amplified recognition elements. A detection probe may be a single stranded oligonucleotide, or it may be partially single stranded and partially double stranded as a detection polynucleotide. A detection oligonucleotide or a detection polynucleotide comprises a detectable label.
[0081] In some embodiments, the methods described herein may include introducing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 50 or more, 100 or more, 200 ormore, 300 ormore, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1,000 or more detection polynucleotides (or single stranded oligonucleotides) to an amplification product. In some embodiments, the methods described herein may include introducing 1,000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 50 or less, 25 or less, 20 or less, 15 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotides to an amplification product.
[0082] In some embodiments, a detection polynucleotide may comprise a detectable label (e.g., fluorescent molecule). In some embodiments, a detection polynucleotide comprises a single stranded oligonucleotide with a portion that is complementary to a code, or a portion of acode, and a detectable label. In some embodiments, a detection polynucleotide comprises two oligonucleotides. Fig. 3A shows a non-limiting example of a structure of a detection polynucleotide 300 comprising two oligonucleotides. As shown in Fig. 3A, a first oligonucleotide 330 may comprise a detectable label 340. A second oligonucleotide 350 may comprise a portion that is complementary to the first oligonucleotide 310 and a second portion 320 that is complementary to a code or a portion of a code 360 (e.g., a nucleic acid segment of a code). The first oligonucleotide 330 may hybridize to the second oligonucleotide 350, thereby generating a detection polynucleotide. For decoding, a portion of the second oligonucleotide 320 may hybridize to its code complement 360 as seen in the example in Fig. 3B, and a signal may be detected from the detectable label following the hybridization event, thereby identifying the code which in turn can be correlated backto the presence of a target nucleic acid of interest from a sample.
[0083] The detection polynucleotide and its components may comprise various nucleotide lengths. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 25 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 20 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 15 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 10 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 8 nucleotides.
[0084] In some embodiments, the detection polynucleotide may comprise a length of between about 5-100 nucleotides, between about 10-80 nucleotides, between about 20-60 nucleotides, between about 30-50 nucleotides, or between 15-30 nucleotides. In some embodiments, the detection polynucleotide may comprise one or more detectable labels. In some embodiments, the one or more detectable labels may comprise a fluorescent moiety. The fluorescent moiety may emit in the red, far-red, near-red, yellow, green, or blue wavelengths. In some embodiments, the fluorescent moiety comprises one or more of 6 -FAM (6- carboxyfluorescein), JOE (6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein), TAMRA (6- carboxytetramethylrhodamine), 5-Cy5 (5 -carb oxy rhodamine), 5-Cy5.5 (5-carboxylic acid succinimidyl ester), 5-Cy7 (5 -carboxyrhodamine), (hexachlorofluorescein), Alexa Fluor 488 (AF488), Alexa Fluor 514 (AF514), Texas Red, Cyanine 3, Cyanine 5, Pacific Blue, Tetramethyl rhodamine, Oxazole Yellow, Atto647N, and Rhodamine 6G (R6G). In some embodiments, the detection polynucleotide may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or 10 or more fluorescent moieties. In some embodiments, the detection polynucleotide may comprise10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less fluorescent moieties.
[0085] In some embodiments, the fluorescent moiety may comprise an organic dye, a biological fluorophore, a quantum dot, or a combination thereof. In some embodiments, the organic dye may comprise an organic molecule. In some embodiments, the organic dye may comprise a coumarin, a cyanine, a benzofuran, a quinoline, a quinazolinone, an indole, a benzazole, a borapolyazaindacene, a xanthene, or a combination thereof. The organic dye may correspond to a color. For example, the organic dye may correspond to a green, a yellow, a blue, an indigo, a red, an orange a purple, a pink, a violet, or a combination thereof. In some embodiments, the organic dye may correspond to no color. In some embodiments, the organic dye may correspond to a black color. In some embodiments, the organic dye may correspond to a white color.
[0086] The detectable moiety can be identified by imaging. When the detectable label is a fluorophore, the fluorophore may emit a color in the visible light spectrum which can be captured by fluorescent imaging and associated filters. In some embodiments, the fluorophore may emit in a wavelength in the range between about 400 nanometers (nm) and 900 nm. In some embodiments, the fluorophore may emit in a wavelength between about 400 nm and 475 nm, about 475 nm and 490 nm, about 490 nm and 530 nm, about 530 nm and 575 nm, about 575 nm and 600 nm, about 600 nm and 700 nm, or about 700 nm and 800 nm. In some embodiments, the fluorophore may emit a wavelength of 400 nm or more, 425 nm or more, 450 nm or more, 475 nm or more, 500 nm or more, 525 nm or more, 550 nm or more, 575 nm or more, 600 nm or more, 625 nm or more, 650 nm or more, 675 nm or more, 700 nm or more, 725 nm or more, 750 nm or more, 775 nm or more, 800 nm or more, 825 nm or more, 850 nm or more, 875 nm or more, or 900 nm or more. In some embodiments, the fluorophore may emit a wavelength of 900 nm or less, 875 nm or less, 850 nm or less, 825 nm or less, 800 nm or less, 775 nm or less, 750 nm or less, 725 nm or less, 700 nm or less, 675 nm or less, 650 nm or less, 625 nm or less, 600 nm or less, 575 nm or less, 550 nm or less, 525 nm or less, 500 nm or less, 475 nm or less, 450 nm or less, 425 nm or less, or 400 nm or less.
[0087] The detectable labels (e.g., fluorescent moieties) may be optically distinct. The number of optically distinct detectable labels used in the methods described herein can impact the amount of information that may be detected. For example, a detection scheme using a larger number of optically distinct detectable labels may allow for a higher amount of multiplexing of codes, which may in turn allow for a greater amount of target molecule related information to be detected and captured. A detection scheme using a smaller number of optically distinct detectable labels may allow for a lesser amount of target molecule related information to bedetected and captured. In some embodiments, using a larger number of optically distinct detectable labels may lead to a detection process that identifies a target molecule in less time as compared to using a fewer number of optically distinct detectable labels when querying a concatemeric amplification product. In some embodiments, a detection scheme using a larger number of optically distinct detectable labels may lead to greater instrument complexity, which may lead to fluorescence detection crosstalk, whereby the fluorescence emission spectra of the optically distinct fluorescent moieties may not yield distinct fluorescence signals. In some embodiments, a detection scheme using a greater number of optically distinct fluorescent moieties may require use of a more complex detection tool.
[0088] In some embodiments, the detection polynucleotides maybe provided in one or more detection pools. In some embodiments, the methods herein may use one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 45 or more, or 50 or more detection pools. In some embodiments, the methods herein may use 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection pools.
[0089] In some embodiments, each detection pool provided may comprise a number of detection polynucleotides. In some embodiments, each detection pool may comprise two or more, three or more, four or more, five or more, ten or more, 15 or more, 25 or more, 50 or more, 100 ormore, 150 or more, 250 or more, 500 or more, 1,000 or more, 1,500 or more, 2,500 or more, or 5,000 or more detection polynucleotides. In some embodiments, each detection pool may comprise 5,000 or less, 2,500 or less, 1,500 or less, 1,000 or less, 500 or less, 250 or less, 150 or less, 100 or less, 50 or less, 25 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotides.
[0090] The number of detection pools and the number of detection polynucleotides in each detection pool may be considered in the design of the recognition elements. For example, one advantage to using a smaller number of detection pools and detection polynucleotides in the methods described herein may be to lower design costs. Conversely, one advantage to using a larger number of detection pools and detection polynucleotides in the methods described herein may be the need for a higher degree of multiplexing for target molecule detection and larger amounts of information that may be detected.
[0091] The methods described herein may include imaging a plurality of detection polynucleotides. The detection polynucleotides may have hybridized to their complementarycode or a portion of a code in order to obtain identifiable signals which can be correlated back to the presence of a target molecule of interest. In some embodiments, the signals are associated with one or more segments of a code for each concatemeric amplification product. In some embodiments, the imaging is performed by an imaging system comprising a fluorescence detection system.
[0092] In some embodiments, the imaging may be conducted using an imaging system. The imaging system may comprise at the minimum a camera, a detector, an illuminator, a condenser, or a combination thereof. In some embodiments, the imaging may include images of fluorescence emission, luminescence, or a combination thereof. In some embodiments, the imaging system may comprise components or sub -systems of a larger system that may also include optics modules including when needed fluorescence filters, fluidics modules, temperature control modules, translation stages, robotic fluid dispensing and / or microplate handling, processors or computers, instrument control software, data analysis and display software, etc. In some embodiments, the imaging system may be a fluorescence imaging system. In some embodiments, the imaging may include fluorescent images from the fluorescent moieties present on the labeled probes.
[0093] In some embodiments, the image may comprise fluorescence information from one or more wavelengths. In some embodiments, the fluorescence information may comprise emission data from a wavelength from about 220-830 nanometers (nm), about 230-820 nm, about 240-8 lO nm, about 250-800 nm, about 260-790 nm, about 270-780 nm, about 280-770 nm, about 290-760 nm, about 300-750 nm, about 310-740 nm, about 320-730 nm, about 330- 720 nm, about 340-710 nm, about 350-700 nm, about 360-690 nm, about 370-680 nm, about 380-670 nm, about 390-660 nm, about 400-650 nm, about 410-640 nm, about 420-630 nm, about 430-620 nm, about 440-610 nm, about 450-600 nm, about 460-590 nm, about 470-580 nm, about 480-570 nm, about 490-560 nm, about 500-550 nm, about 510-540 nm, about 520- 530 nm, or a combination thereof.
[0094] Image detection and capture may relate to iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. In some embodiments, the iterative repetition of the operations is performed for each nucleic acid segment of a code one or more times, thereby generating a code profile which can be subsequently decoded.
[0095] In some embodiments, the iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing thedetection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof may comprise two or more iterative repetitions. For example, the methods described herein may comprise about 2-50 iterative repetitions, about 2-10 iterative repetitions, or about 2-8 iterative repetitions. In some embodiments, the method described herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more iterative repetitions. Each iterative repetition may comprise: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. As such, a code profile can be built and decoded that is indicative of the original target of interest presence in a sample.
[0096] In some embodiments, the number of iterative repetitions of the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof may correspond to the number of nucleic acid segments present in a code of the recognition element. In some embodiments, each nucleic acid segment of the code of the recognition element may undergo a number of iterative repetitions of the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. For example, the methods described herein may comprise iteratively repeating the operations described herein two or more times per segment, three or more times per segment, or four or more times per segment and compiling the images into a code profile. In some embodiments, the methods described herein may comprise iteratively repeating the operations described herein two or more times per segment, three or more times per segment, four or more times per segment, five or more times per segment, six or more times per segment, seven or more times per segment, eight or more times per segment, nine or more times per segment, 10 or more times per segment, 11 or more times per segment, 12 or more times per segment, 13 or more times per segment, 14 or more times per segment, 15 or more times per segment, 16 or more times per segment, 17 or more times per segment, 18 or more times per segment, 19 or more times per segment, or 20 or more times per segment and compiling the images into a code profile.
[0097] Additional methods for imaging a detection polynucleotide can be found in WO2023 / 158993A2, which is incorporated herein by reference in its entirety.
[0098] Several models may be used to identify a code that is associated with a target molecule based on the fluorescence signals generated and images captured and compiled from detection polynucleotide hybridization to code sequences. In one embodiment, decoding makes use of a hard decision decoding model. In another embodiment, decoding makes use of a soft decision decoding model.
[0099] For soft decision decoding, it is not necessary to identify each base specifically. For example, signals generated during each detection event may be detected and recorded to produce a data set, for example a code image profile, that may be used as input into a model to calculate a probability that a specific code is present without requiring that each base of a code be determined. Although it may not be necessary in a soft decision decoding model to make a hard decision (e.g., a nucleotide base call) about the identity of each nucleotide, a model may nevertheless include assigning a probability or identity to each nucleotide in the sequence of a code, wherein each nucleotide in the sequence of a code can be sequenced. Data gathered may include intensity readings for signals produced by the hybridized detection polynucleotide fluorescent moiety in various spectral bands. A set of intensity readings may be detected by imaging, compiled and stored and used as input into a soft decision decoding model for determining a probability that a particular code is present, and hence a target nucleic acid is present in the sample.
[0100] A model may be developed or trained using data from known codes, such as signal intensity data across a predetermined spectrum. The model may be used to calculate a set of probabilities across a set of one or more codes, indicating, for example, for each code, a probability that it is present in a concatemeric amplification product and hence that the target of interest was present in the original sample.
[0101] The probability that a particular code is present may be indicative of the probability that a particular target molecule associated with the code is present in the sample of interest. Data indicating the probability that a particular target is present may be, for example, to calculate probabilities relevant to diagnosis or screening of various medical conditions, or selection of drugs for treatment of various medical conditions.
[0102] A soft decoding decision model may include using an algorithm to predict the presence of target molecules from a sample. In some embodiments, the algorithm is a soft- decision decoding algorithm. In some embodiments, the algorithm is applied to the codes or code profile of the concatemeric amplification products for predicting the presence of a target molecule from a sample.
[0103] The methods disclosed herein may comprise soft decision decoding to predict the presence of the code in a recognition element or concatemeric amplification product thereof,wherein the presence of the code correlates and serves as a proxy for the presence of a target nucleic acid in a sample. In some embodiments, the methods described herein may use soft decision decoding. In some embodiments, the methods described herein may use hard decision decoding. For hard decision decoding, signals from queried concatemers may be extracted from images. This may be the same for soft decision decoding, in that signals that are generated and imaged are extracted from the images. For hard decision decoding, however, hard symbol base to base calls may be generated from the intensities of the signals, whereas with soft decision decoding no hard symbol calls may be necessary as all of the signal range is retained. The code assignment for hard decision decoding may be determined by matching symbol-by-symbol readouts of codewords to codes, whereas with soft decision decoding, the signals may be cross correlated against the expected signals and a code assigned using a probabilistic methodology. When using soft decision decoding, it is not necessary for the model to identify each symbol specifically. For example, signals (e.g., fluorescent signals) generated during each cycle of a detection process maybe detected and recorded to produce a data set that may be used as input into a model to calculate the probability that a specific code is present.
[0104] The permutation space on a recognition element is the totality of factors that determines the number of unique nucleotide possibilities at each nucleic acid segment. Factors may comprise the number of segments present on a recognition element, the number of incubation periods or times a segment is queried with a detection pool comprising detection polynucleotides, and the number of computational symbols or colors.
[0105] Fig. 7 details an example of a soft decision decoding workflow for determining the probability of the presence of a target molecule from a sample based on detection and decoding of a code, or a code profile, associated with the target molecule that originally hybridized to a recognition element. Images of the sample may be acquired, aligned, and processed to extract the intensity of the features, or signals of interest across the imaged field of view in multiple spectral channels. The corrected intensities of said features may be fed through a series of algorithms that make up the soft decoder. At first, the intensity profiles of the codes may be learned based on features of high confidence or high intensity. This trained model may provide a template for each code from which the rest of the features of interest may be compared to in the second operation. Third, a confidence score may be computed from the difference between the intensity profile of each feature and the trained profiles. Several filters may be applied to remove outliers, duplicates, and low confidence decoded concatemers. The output may comprise a table of decoded concatemers with an associated filter status, confidence score, and most likely assignment to one of the codes of one or more concatemeric amplification products.
[0106] In some embodiments, a recognition element comprises a larger code, for example a code with four segments instead of two or three. In some embodiments, a recognition element comprising a larger code may result in a detection scheme with better error correction as compared to a detection scheme having a recognition element comprising a smaller code. Additionally, in some embodiments, a larger code may result in a lower signal -to-noise ratio as compared to a detection scheme having a recognition element comprising a smaller code.
[0107] In some embodiments, a recognition element comprises a smaller code, for example a code with two segments, or one segment, as compared to a recognition element with four segments. In some embodiments, a recognition element comprising a smaller code may result in a detection scheme with lower error correction abilities as compared to a detection scheme with a larger code. Further, in some embodiments, a smaller code may result in a higher signal -to- noise ratio as compared to a larger code.
[0108] The systems disclosed herein relate to detecting a target molecule around, on or in a sample. In some embodiments, the systems comprise a plurality of recognition elements and a plurality of labeled probes as described herein. In some embodiments, the systems comprise a solid substrate configured to immobilize one or more of a circularized and ligated recognition element, a concatemeric amplification product, a labeled probe, and a hybridized complex of a concatemeric amplification product and a labeled probe. In some embodiments, the systems comprise a welled plate or a flow cell. In some embodiments, the systems comprise a fluid flow controller, a temperature controller, an imaging system, a computer system, or any combination thereof.
[0109] In some embodiments, the systems disclosed herein may include a solid substrate or a solid surface. The solid substrates and surfaces disclosed herein may be referred to as a substrate, a support, a solid support, or a surface. The substrate may be modified for immobilizing circularized and ligated recognition elements or concatemeric amplification products, or both. Example solid substrates include, but are not limited to, glass, modified or functionalized glass, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, inorganic glasses, plastics, optical fiber bundles, optically clear glass, and other polymers. In some embodiments, the plastic solid substrates may include acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, or polyurethanes. In some embodiments, the silica -based solid substrates may include silicon or modified silicon.
[0110] In some embodiments, the substrate may be a welled plate. In some embodiments, the substrate may be a 96-well plate. In some embodiments, the substrate may be a 4-well plate, a 6-well plate, an 8-well plate, a 12-well plate, a 24-well plate, a 48-well plate, a 384-well plate,an 864-well plate, or a 1,536-well plate. In some embodiments, the substrate may have greater than or equal to 96 wells. In some embodiments, the substrate may have less than or equal to 96 wells.
[0111] In some embodiments, the substrate may be a flow cell. In some embodiments, the flow cell may have two or more lanes. In some embodiments, the flow cell may have two or less lanes.
[0112] In some embodiments, the substrate may be a microarray, a slide, a chip, a microwell, a tube, a column, a particle, a bead, or a paramagnetic bead.
[0113] In some embodiments, the substrate may comprise a coating. In some embodiments, the coating may comprise a layer that may be charged. In some embodiments, the coating layer may be positively charged. In some embodiments, the coating layer may be negatively charged. In some embodiments, the coating may be non-charged. In some embodiments, the substrate may comprise a surface comprising a cation -coating layer. In some embodiments, the substrate may comprise a surface comprising an anion-coating layer. In some embodiments, the substrate may comprise a surface comprising a neutral-charged layer. In some embodiments, the substrate may be coated with streptavidin. In some embodiments, the substrate may be coated with avidin. In some embodiments, the substrate may be coated with one or more antibodies.
[0114] The systems disclosed herein may comprise a fluidics system. The fluidics system may comprise a fluid flow controller. In some embodiments, the fluid flow controller may comprise one or more pumps, valves, mixing manifolds, reagent reservoirs, waste reservoirs, or any combination thereof. In some embodiments, the fluidic system and subcomponents of the fluidics system are fluidically connected to the reaction vessel of the present disclosure. In some embodiments, the fluidic system and subcomponents of the fluidics system iteratively flow in reagents (e.g., buffers, detector polynucleotides, anchor polynucleotides, detection oligonucleotide complexes, etc.) to the reaction vessel. In some embodiments, the reaction vessel comprises a solid substrate configured to immobilize the circularized and ligated recognition elements or concatemeric amplification products thereof.
[0115] The systems disclosed herein may comprise a temperature system. The temperature system may comprise a temperature controller. The temperature controller may be incorporated into the systems described herein to facilitate accuracy of the methods and systems described herein. In some embodiments, the temperature controller may comprise temperature control components. Non-limiting examples of temperature control components include resistive heating elements, infrared light sources, heating or cooling devices, heat sinks, thermocouples, thermistors, or a combination thereof. In some embodiments, the temperature controller may provide changes in temperature over specified time intervals. In some embodiments, thetemperature controller may provide an increase in temperature. In some embodiments, the temperature controller may provide a decrease in temperature. In some embodiments, the temperature controller may provide for cycling of temperatures between two or more set temperatures so that thermocycling or amplification may be performed. In some embodiments, the temperature controller may provide a constant temperature.
[0116] The systems disclosed herein may comprise an imaging system. In some embodiments, signals produced by the labeled probes disclosed herein may be imaged by the imaging systems disclosed herein. The imaging system may comprise one or more light sources, one or more optical components, one or more filters, one or one or more imaging sensors for imaging and detection, or a combination thereof. In some embodiments, the one or more light sources may comprise light from a bulb. In some embodiments, the one or more optical components may comprise lenses, mirrors, digital mirror devices, prisms, optical filters, colored glass filters, narrowband interference filters, broadband interference filters, dichroic reflectors, diffraction gratings, apertures, optical fibers, optical waveguides, or a combination thereof. In some embodiments, the one or more imaging sensors may comprise a charge -coupled device (CCD) sensor or camera, a complementary metal -oxide-semiconductor (CMOS) imaging sensor or camera, a negative-channel metal-oxide semiconductor (NMOS) imaging sensor or camera, or a combination thereof.
[0117] Various operations of the methods and systems disclosed herein may be performed by a computer system of the present disclosure. Referring to Fig. 4, an example of a block diagram is shown depicting an example of a machine that includes a computer system 400 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in Fig. 4 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0118] Computer system 400 may include one or more processors 401, a memory 403, and a storage 408 that communicate with each other, and with other components, via a bus 440. The bus 440 may also link a display 432, one or more input devices 433 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 434, one or more storage devices 435, and various tangible storage media 436. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 440. For instance, the various tangible storage media 436 can interface with the bus 440 via storage medium interface 426. Computer system 400 may have any suitable physical form, including but not limited toone or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0119] Computer system 400 includes one or more processor(s) 401 (e.g., central processing units (CPUs), general purpose graphics processing units (GPUs), or quantum processing units (QPUs) that carry out functions. Processor(s) 401 optionally comprises a cache memory unit 402 for temporary local storage of instructions, data, or computer addresses. Processor(s) 401 are configured to assist in execution of computer readable instructions. Computer system 400 may provide functionality for the components depicted in Fig. 4 as a result of the processor(s) 401 executing non -transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 403, storage 408, storage devices 435, and / or storage medium 436. The computer-readable media may store software that implements particular embodiments, and processor(s) 401 may execute the software. Memory 403 may read the software from one or more other computer-readable media (such as mass storage device(s) 435, 436) or from one or more other sources through a suitable interface, such as network interface 420. The software may cause processor(s) 401 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 403 and modifying the data structures as directed by the software.
[0120] The memory 403 may include various components (e.g., machine readable media) including, but not limited to, a random-access memory component (e.g., RAM 404) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase - change random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 405), and any combinations thereof. ROM 405 may act to communicate data and instructions unidirectionally to processor(s) 401, and RAM 404 may act to communicate data and instructions bidirectionally with processor(s) 401. ROM 405 and RAM 404 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 406 (BIOS), including basic routines that help to transfer information between elements within computer system 400, such as during start-up, may be stored in the memory 403.
[0121] Fixed storage 408 is connected bidirectionally to processor(s) 401, optionally through storage control unit 407. Fixed storage 408 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 408 may be used to store operating system 409, executable(s) 410, data 411, applications 412 (application programs), and the like. Storage 408 can also include an optical disk drive, a solid-state memorydevice (e.g., flash-based systems), or a combination of any of the above. Information in storage 408 may, in appropriate cases, be incorporated as virtual memory in memory 403.
[0122] In one example, storage device(s) 435 may be removably interfaced with computer system 400 (e.g., via an external port connector (not shown)) via a storage device interface 425. Particularly, storage device(s) 435 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 400. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 435. In another example, software may reside, completely or partially, within processor(s) 401.
[0123] Bus 440 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 440 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example, and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI -Express (PCLX) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0124] Computer system 400 may also include an input device 433. In one example, a user of computer system 400 may enter commands and / or other information into computer system 400 via input device(s) 433. Examples of an input device(s) 433 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 433 may be interfaced to bus 440 via any of a variety of input interfaces 423 (e.g., input interface 423) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0125] In particular embodiments, when computer system 400 is connected to network 430, computer system 400 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 430. Communications to and from computer system 400 may be sent through network interface 420. For example, network interface 420 mayreceive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 430, and computer system 400 may store the incoming communications in memory 403 for processing. Computer system 400 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 403 and communicated to network 430 from network interface 420. Processor(s) 401 may access these communication packets stored in memory 403 for processing.
[0126] Examples of the network interface 420 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 430 or network segment 430 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 430, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0127] Information and data can be displayed through a display 432. Examples of a display 432 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 432 can interface to the processor(s) 401, memory 403, and fixed storage 408, as well as other devices, such as input device(s) 433, via the bus 440. The display 432 is linked to the bus 440 via a video interface 422, and transport of data between the display 432 and the bus 440 can be controlled via the graphics control 421. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0128] In addition to a display 432, computer system 400 may include one or more other peripheral output devices 434 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 440 via an output interface 424. Examples of an output interface 424 include, but are notlimited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0129] In addition, or as an alternative, computer system 400 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0130] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0131] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general - purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general -purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0132] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An example of a storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside inan ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0133] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, notepad computers, set -top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those of skill in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations, known to those of skill in the art.
[0134] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, andNovell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of non-limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX- like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Those of skill in the art will also recognize that suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Those of skill in the art will also recognize that suitable video game console operating systems include, by way of non -limiting examples, Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.
[0135] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non -transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In further embodiments, a computer readable storage medium is optionallyremovable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi -permanently, or non-transitorily encoded on the media.
[0136] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, which perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages.
[0137] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.
[0138] In some embodiments, the computer programs described herein may be used to perform at least one function. The computer programs described herein may perform functions related to storing data, receiving data, analyzing data, exporting data, or a combination thereof. In some embodiments, the computer programs described herein may perform functions related to applying selection criteria, including in silico selection criteria, functional selection criteria, or a combination thereof. In some embodiments, the computer programs may receive sequence information, including sequence information for nucleic acid segments. The sequence information may be configured as an array, a table, a list, or combination thereof. The sequence information may be formatted in a variety of ways, including, but not limited to a .txt file, a FASTA file, an .xls file, or a combination thereof. The computer programs described herein may apply selection criterion or selection criteria to a set of nucleic acid segments. The computer programs may sort the nucleic acid segments, determine or compute characteristics of thenucleic acid segments, perform calculations, reorder the nucleic acid segments, or a combination thereof. In some embodiments, the computer programs described herein may store information related to the nucleic acid segments. In some embodiments, the computer program may use information stored related to the nucleic acid segments to apply selection criteria to the nucleic acid segments. In certain embodiments, the computer program may receive information and / or data related to nucleic acid segments, selection criteria, or a combination thereof. In some embodiments, the computer programs may perform functions related to analyzing data from functional assays, including, but not limited to functional assays described herein. In some embodiments, analyzing data from functional assays may comprise image analysis, image quantification, intensity quantification, feature identification, or a combination thereof. The computer programs described herein may also export information. In some embodiments, the exported information may comprise images, files, data tables, documents, folders, or a combination thereof.
[0139] In some embodiments, a computer program includes a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non -limiting examples, relational, non-relational, object oriented, associative, XML, and document-oriented database systems. In further embodiments, suitable relational database systems include, by way of non -limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a webapplication is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non -limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.
[0140] Referring to Fig. 5, in a particular embodiment, an application provision system comprises one or more databases 500 accessed by a relational database management system (RDBMS) 510. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBMDB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system further comprises one or more application servers 520 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 530 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 540. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.
[0141] Referring to Fig. 6, in a particular embodiment, an application provision system alternatively has a distributed, cloud-based architecture 600 and comprises elastically load balanced, auto-scaling web server resources 610 and application server resources 620 as well as synchronously replicated databases 630.
[0142] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.
[0143] In view of the disclosure provided herein, a mobile application is created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of non- limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0144] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non -limiting examples, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments areavailable without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.
[0145] Those of skill in the art will recognize that several commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.
[0146] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program (s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.
[0147] In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Those of skill in the art will be familiar with several web browser plug-ins including, Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.
[0148] In view of the disclosure provided herein, those of skill in the art will recognize that several plug-in frameworks are available that enable development of plug-ins in variousprogramming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.
[0149] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of nonlimiting examples, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, andKDEKonqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non- limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non -limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ browser.
[0150] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non- limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or moremachines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.
[0151] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of nucleic acid segment sequences or analysis thereof information. In various embodiments, suitable databases include, byway of non -limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity -relation ship model databases, associative databases, XML databases, document-oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0152] In some embodiments, there is a value proposition when combining two different assays for identifying a target of interest that may, or may not, be present in a sample, for example a sample that is suspected of representing a disease state, a cancerous state or an adverse medical condition. For example, a screen for a target of interest can be the first line of detection to demonstrate the presence of a target of interest that might be aligned with a disease state, cancer or adverse medical condition. After determining at a high level the presence of the target of interest, a second and more in-depth assay can be performed to support and more completed identify characteristics for the target of interest. As such, combining two assays and different levels of specificity can provide a synergistic effect that can enhance the value of determining the presence of a target of interest and align that determination with a disease state, cancerous state of adverse medical condition.
[0153] In some embodiments, a first screening assay to determine at a higher level the presence of a target of interest can be the methods described herein for utilizing targeted recognition elements to hybridize to the target of interest. By detecting and decoding the final product of the hybridization event, the determination of the presence of the target of interest is based on detecting a proxy for that target, in this methodology determining the presence of a code from the recognition element that hybridized to the target of interest. In some embodiments, the decoding is a probabilistic method of soft decision decoding where the actual base by base sequence of the target of interest is not identified, but instead the probability that a code is present based on detection by hybridization of the code, or portions thereof, is identified,and as a code is unique to each recognition element its detection is aligned with the target of interest.
[0154] In some embodiments, a second screening assay to determine at a more in-depth level the presence of a target of interest can be the use of next generation sequencing, or microarrays. For example, sequencing can be performed on the recognition element comprising the target of interest, or a derivative or complement thereof, for determining the base-by-base sequence of the target of interest for a validation that the target of interest is present in the sample. In some embodiments, a target of interest is identified by microarray analysis, where a probe on the microarray is complementary to the target of interest and, if present, the target of interest can hybridize to the probe on the microarray and its presence determined.
[0155] In some embodiments, the present disclosure provides for a first screening for the presence of a target of interest using recognition elements and associated workflows as described herein in combination with either a sequencing or a microarray second screening method. In some embodiments, sequencing is next generation sequencing including, but not limited to, sequence by synthesis, nanopore sequencing, sequencing by hybridization, sequencing by ligation, and the like. Next generation sequencing is well known in the art and multiple commercial instruments are available for performing sequencing, such as sequencing by synthesis practiced by Illumina, Singular Genomics, Pacific Biosciences, Beijing Genomics Institute and Oxford Nanopore to name but a few. As such, the present disclosure provides for a rapid screening assay based on coded recognition elements and indirect target identification in combination with a direct method of determining the sequence of a target of interest, such as next generation sequencing or microarray if additional testing is suggested based on the outcome of the indirect assay.
[0156] In some embodiments, the methods described herein for detecting a target of interest by employing coded recognition elements canbe performed on a substrate thatis also amenable for use in next generation sequencing, for example a flow cell that is specific to a sequencing platform.
[0157] In some embodiments, the methods described herein for detecting a target of interest by employing coded recognition elements also provide for generating a sequencing library, such that the same assay provides for both detection by hybridization of the code of a recognition element and the mechanism to perform sequencing if warranted without having to go back and generate a sequencing library separately. In some embodiments, the detection by hybridization of a code of a recognition element is performed on the same instrument as sequencing can be performed if chosen to do so. As such, in some embodiments, detection of the code and detection of a sequence of a target of interest are performed on the same instrument. In someembodiments, while the same instrument is used for detecting the presence of a code and sequencing a target of interest, each assay is performed on the same type of substrate, for example a flow cell that is accommodated by the instrument being utilized. As such, one platform, in this instance a platform that can detect and image fluorescence, can be used for both code detection via hybridization using fluorescently labeled detection polynucleotides that are complementary to a code or a portion thereof and for next generation sequencing, such as sequencing by synthesis, where a base by base incorporation of fluorescently labeled nucleotides can be detected and reported out, or an electrical charge or change thereof can be detected when a nucleic acid is passed through a nanopore.
[0158] In some embodiments, minimal sequencing of a target of interest can first be performed and an indirect detection method can be used subsequent or concurrent with minimal sequencing to identify the presence of a target of interest. In some embodiments, minimal sequencing includes next generation sequencing such as sequence by synthesis, wherein the sequencing is performed on amplicons derived from a target of interest and / or whole exome sequencing. In some embodiments, a target of interest can be difficult to sequence, for example if a target of interest has a high GC content, has homopolymer regions, has variant sequences such as copy number variance, insertions, deletions and the like. In this instance where targets of interest may provide challenges for next generation sequencing, a follow up, or concurrent, assay utilizing code detection and decoding as described herein can be utilized to supplement sequencing data, thereby filling in data for the hard to sequence content of the target of interest. Additionally, a target of interest may have homologs or pseudogenes where sequencing does not fully differentiate the target of interest from a homolog or a pseudogene, as such a specially designed recognition element as illustrated in the example shown in Fig. 13 can be utilized for differentiation followed by detection and decoding, wherein the decoding data can support sequencing data in determining the presence of the target of interest in the background of a homolog or pseudogene.
[0159] In some embodiments, a sample that may or may not comprise a target of interest includes a tissue sample, a cell sample, a blood sample, a biological sample of human origin, a biological sample not of human origin, and the like. In some embodiments, nucleic acids are extracted from the sample. In some embodiments, nucleic acids can be cell free DNA, genomic DNA, cDNA derived from RNA, or RNA such as mRNA. In some embodiments, if RNA is extracted it can be subsequently transcribed into complementary DNA or cDNA. In some embodiments, such as in the case of blood, cell free DNA is extracted and may be additionally size selected for selection of cell free DNA over, for example, longer contaminating genomic DNA if present.
[0160] In some embodiments, the extracted nucleic acids can be selectively amplified for increasing the amount of a target of interest, for example the target of interest which may be present in extracted nucleic acids can be amplified for example via PCR prior to performing any assay to identify the presence of the target of interest. In some embodiments, the extracted and optionally amplified target(s) of interest, or a portion thereof, are split into two reactions such that two different detection methodologies can be performed on the same sample either concurrently or subsequently. In some embodiments, one of the reactions comprises adding recognition elements to a first reaction, allowing for the hybridization of the target sequence of interest to its complementary sequence as found in the 5’ and 3 ’ ends of the recognition element to occur, ligating the ends of the recognition element(s) that are hybridized to the target(s) of interest, amplifying the ligated and circularized recognition element(s), and detecting the presence of the code in the recognition elements, a code which is unique to the recognition element and hence the target of interest, via detection by hybridization using fluorescently labeled detection polynucleotides, and decoding the detected fluorescent images to identify a probability that the code is present, thereby determining the presence of a target of interest. Indirectly identifying the presence of the target of interest by detecting the presence of the proxy code can be done in a, relatively speaking, short period of time. For example, the whole indirect detection workflow can take approximately 3 hrs from hybridization and ligation of the recognition elements to a final readout including the probability of the presence of a code, thereby providing a quick determination for a first level screening option.
[0161] In some embodiments, the results from a first reaction as described herein may indicate the need for a more in-depth analysis of the sequence of a target of interest, if present. In some embodiments, a more in-depth analysis is not predicated on any particular indication but is instead performed to provide a more thorough analysis of a target of interest. In some embodiments, a second reaction comprises adapterization of the amplified sample where PCR primers are utilized to amplify a portion of the sample while concurrently adding adaptors to the 5’ and 3’ ends of the amplification products, thereby generating a sequence ready library. In some embodiments, the adaptors comprise sequences that are not complementary to the target of interest but are instead complementary to sequences of oligonucleotides that are affixed to a sequencing flow cell, such as P5 and P7 oligonucleotides that are affixed to Illumina specific flow cells for compatibility with Illumina sequencing platforms or SI and S2 oligonucleotides that are affixed to Singular Genomics flow cells for compatibility with Singular Genomics sequencing platforms. In some embodiments, the generated sequencing library is sequenced, via sequence by synthesis technologies, and a base-by-base sequence of the target of interest is determined. Sequencing from library preparation through data output can take over 30 hrs, assuch a much longer assay time from sample to data output. In some embodiments, a shorter assay for gathering initial data for providing initial information and insights is followed by a longer assay for supporting or verifying the results of the shorter assay. This type of one-two assay can provide diagnosticians and researchers with initial insights and results without having to wait for more results which may take a much longer time to output.
[0162] In some embodiments, the methods described herein occur on a substrate. In some embodiments, the substrate is an optically clear glass slide wherein detectable signals that are generated during a detection operation of an assay, such as by using fluorescently labeled detection polynucleotides or by using fluorescently labeled nucleotides in a sequence by synthesis reaction. Fig. 8 provides an example of an assay device for practicing the methods as disclosed herein. The assay device 800 comprises a glass substrate 801, surface chemistry for nucleic acid immobilization 802 and a temporary, or removable, structure comprising a mechanism for separating samples 803 thereby providing discrete reaction wells or partitions for sample assays on the substrate 800. In some embodiments, the substrate 801 comprises optically clear glass. In some embodiments, the substrate comprises a surface chemistry that allows for immobilization of targets of interest. In some embodiments, the surface chemistry 802 is a cationic polymer, an amine, amino silane, silane, a hydrogel, thiol / gold chemistry, PDMS, RAFT, etc. In some embodiments, the structure 803 for generating partitions for discrete sample assays comprises polystyrene, polycarbonate, rubber, plastic, silicone, and the like. In some embodiments, the removable structure 803 is removably affixed to the substrate 801 by an adhesive, an acrylate, silicone, a foam, a temperature sensitive adhesive, a hydrophobic material, dimethyl siloxane, a fluorinated compound such as Teflon or fluorinated silane, a gasket such as a silicone gasket, neoprene, butyl rubber, and the like. In some embodiments, the temporary structure 803 can be removed from the substrate 801 by the application of heat or by manual manipulation, or other method that serves to remove the temporary structure while maintaining the integrity of the substrate and the targets of interest that are immobilized thereon. In some embodiments, the assay device can be used for indirect detection of the presence of the target nucleic acid after removal of the temporary structure. In some embodiments, the assay device can be used for sequencing a target of interest, for example after removal of the temporary structure.
[0163] In some embodiments, an assay device comprises a first substrate and a second substrate, as shown in the example in Figs. 9A-9B. Referring to Figs. 9A-9B, in some embodiments, a first substrate 901 comprises glass, such as optically clear glass. In some embodiments, a first substrate comprises a surface chemistry 902 for immobilizing targets of interest as listed herein. In some embodiments, a first substrate comprises a temporary orremovable structure 903 for generating partitions on the substrate for discretely accommodating one or more samples 905 for performing one or more assays. In some embodiments, after an assay is performed and prior to detection operations, the temporary structure 903 is removed, and a second substrate 904 is adjacently affixed to the first substrate, generating a flow cell. In some embodiments, the second substrate 904 is an adhesive film, a temporary or permanent glass substrate, a temporary or a permanent plastic substrate. An assay includes methods as described herein, for example hybridization of a target nucleic acid to complementary regions of a recognition element that comprises a unique code, ligation of the hybridized recognition elements to generate a circularized recognition element, amplification of the circularized recognition element to generate concatemeric amplification products, which can be subsequently queried for identification of the presence of a code that serves as a proxy for the originally hybridized target nucleic acid. In some embodiments, amplification may be, for example, rolling circle amplification, primer extension amplification, isothermal amplification, PCR, or bridge amplification (if the target nucleic acid is destined for sequencing). In some embodiments, amplification occurs inside the partition while the temporary substrate is in place. In some embodiments, amplification occurs in the flow cell wherein the temporary structure is removed, and the second substrate is affixed. In some embodiments, hybridization, ligation and amplification occurs off the substrate, for example in a well or tube, and the amplification products are added to the partitions prior to their removal and replacement with the second substrate. As seen in the example shown in Fig. 9B, in some embodiments, the discrete locations generated by placement of the temporary structure remain after removal of the temporary structure, thereby maintaining each assay partition in a relative sense. The example of the flow cell shown in Fig. 9B, which may include in inlet port and an outlet port, can be utilized for flowing detection reagents into the assay device, detection can be imaged, and the process can be repeated a number of times until the code associated with the samples in a discrete location are sufficiently queried such that decoding provides a probability for the presence of the target of interest.
[0164] In some embodiments, the present disclosure provides for generating a flow cell that comprises multiple flow cell lanes for practicing methods of the present disclosure as illustrated in the example in Figs. 10A-10C. As in Figs. 8A-8B and Figs. 9A-9B, in some embodiments, a first substrate 1001, which may comprise a surface chemistry 1002 for immobilizing targets of interest, comprises temporary partitions as generated by a removable structure 1003. In some embodiments, targets of interest 1004 are located within the discrete partitions and assays, as described herein, can be performed on the targets of interest. In some embodiments, the temporary structure 1003 is removed, while maintaining the discrete locations of theimmobilized targets of interest. In some embodiments, for generating a flow cell from which detection for the presence of a target of interest can be performed, a second substrate 1005 is affixed to the first substrate, the second substrate comprising structures that provide for the generation of multiple lanes on the first substrate, in this figure five lanes are generated by affixing the second substrate 1005 to the first substrate. In some embodiments, the second substrate and the partitions thereof may be composed of plastic and can be affixed to the first substrate by, for example, an adhesive, acrylate, silicone, foam, temperature sensitive adhesive, a hydrophobic material, fluorinated compounds, dimethyl siloxane, a silicone gasket, neoprene, butyl rubber, or other adhesive material. As described for the examples shown in Figs. 9A-9B, any configuration of assay operations can occur on or off the first substrate. Detection for the target of interest, either by detection by hybridization or sequencing, can occur in the lanes of the flow cell.
[0165] In one embodiment, the partitioning as shown in the example of Figs. 8-10 is performed via a temporary adhesive or silicone gasket or other material that enabled liquid tight separation between adjacent wells generated bythe temporary structure. In some embodiments, the temporary structure is removed prior to flow cell assembly. In some embodiments, removal of the temporary structure can be performed by heat, by chemical mechanisms such as by dissolution, etching etc., by physical mechanisms such as using low tack adhesive when affixing the temporary structure to the first substrate, by utilizing a soft deformable hydrophobic material such as silicone, neoprene, PDMS and the like.
[0166] In some embodiments, a flow cell assembly for performing assays as described herein is shown as examples in Figs. 11A-11C with a top-down view of the assay device. As with the examples shown in Figs. 8-10, the example in Fig. 11A utilizes a temporary structure 1101 temporarily affixed to the first substrate 1100 for generating partitions for discrete locations on which one or more samples can be immobilized. The first substrate 1100 can be initially treated with a surface chemistry as described herein for immobilizing a plurality of samples. In some embodiments, a sample can go through a series of assay operations in each discrete location, for example recognition hybridization, ligation and amplification. In other embodiments, hybridization of a target of interest to targeted recognition elements, followed by ligation and circularization is performed off the substrate, for example in a well or tube that is not on the substrate, and the ligated and circularized products are added to each discrete location. In some embodiments, amplification of the circularized recognition elements is also performed off the substrate and the amplification products, such as concatemeric amplification products ready for detection are added to the discrete locations as provided by the temporary support. In some embodiments, after the samples, in whatever form, are aliquotedto one or moreof the temporary partitions, the temporary structure 1101 can be removed while leaving intact the discrete locations on which samples are immobilized 1102, as demonstrated in the example of Fig. 11B. In some embodiments, a second substrate 1103 can be affixed to the first substrate 1100 after removal of the temporary structure 1101. In some embodiments, the second substrate 1103 includes longitudinal partitions for generating one or more assay lanes 1104; in the example shown in Fig. 11C there are four example assay lanes 1104 generated by the second substrate 1103. Each assay lane as demonstrated in the example of Fig. 11C comprises an inlet 1105 and outlet 1106 port for reagent delivery andremoval. In some embodiments, detection by hybridization is performed in the flow cells for determining the presence of a target nucleic acid in one or more samples. For example, reagents include fluorescently labeled detection polynucleotides can be flowed into an inlet port such that the detection polynucleotides have the opportunity to hybridize to their target sequences in a code, or a portion thereof, in a concatemeric amplification product, wherein each code is aligned with a target of interest. In some embodiments, after the hybridization events are imaged, the reagents can be washed through the outlet port and new reagents with new detection polynucleotides can be flowed into the inlet ports, images acquired, and the process repeated two or more times if necessary, depending on the complexity of the codes in need of detection. Once the desired number of images for developing a pattern of fluorescence detection are acquired, they can be combined and a soft decision decoding analysis pipeline can be applied to the data for determining which codes are present, and thus which targets of interest are present based on their association with a code.
[0167] In some embodiments, if a physical separation of samples is not possible on a substrate the samples can be digitally partitioned instead. In one embodiment, digital partitioning comprises splitting the code space over a number of different samples. Fig. 12 demonstrates an example of this strategy (Strategy 1). In Strategy 1, two recognition elements that comprise different codes (hypercodes) are utilized to target the same target sequence, wherein the two recognition element data when combined can serve as a sample index, thereby rendering each sample distinct from each other without the need for a physical partition on a substrate.
[0168] In some embodiments, Strategy 2 as demonstrated in the example shown in Fig. 12, can be utilized for digital partitioning. In Strategy 2, two recognition elements with two different codes (hypercode 1 and hypercode 2) are used for a sample, wherein each of the two recognition elements further comprise the same index sequence distinctive to each sample (Index A for sample 1, Index B for sample 2, etc.). Additionally, two recognition elements may share the same hypercode, however their target complementary sequences for each sample will differ andthe sample indexes differ for each sample, thereby allowing for digital partitioning of two or more samples.
[0169] In some embodiments, Strategy 3 as demonstrated in the example shown in Fig. 13 can be utilized for digital partitioning of two or more samples on a substrate. In Strategy 3, a recognition element is modified to include additional aspects for use in digital partitioning. For example, a recognition element 1301 comprises a unique code and further comprise a PCR binding site (PBS). In this example the recognition element is shown hybridized to a target of interest 1302 from a sample, however the hybridization is not adjacent, and a gap is formed between the 5’ end hybridization event with the target of interest and the 3 ’ end hybridization event with the target of interest. In this example, the gap if filled by hybridization of a third oligonucleotide 1303, wherein the hybridization of the third oligonucleotide with the target of interest results in a loop out of a portion of the third oligonucleotide. In some embodiments, the loop out sequence comprises one or more functional sequences, such as a primer binding site PBS (which may be the same or different from the primer binding site introduced into the recognition element), and a sample index (SI). In some embodiments, the loop generated by the third oligonucleotide is not self-complementary (left side). In some embodiments, the loop generated by the third oligonucleotide comprises a portion that is self -complementary, setting up a hairpin stem -loop configuration (right side). A double ligation event (1305) comprising the 5’ end of the recognition element and the 3 ’ end of the third oligonucleotide and the 5 ’ end of the third oligonucleotide and the 3 ’ end of the third oligonucleotide as hybridized to the target of interest can be performed to generate a circularized and modified recognition element that comprises additional functional elements including a sample index which can be used to digitally partition two or more samples when physical partitions are not present. Some advantages of Strategy 3 include leveraging already developed recognition elements in combination with different third oligonucleotides for each sample, as the shorter oligonucleotides that can be a third oligonucleotide are easier and quicker to synthesize than redesigning each recognition element as found in Strategies 1 and 2.
[0170] In some embodiments, combinatorial pooling of samples, or Strategy 4, can be leveraged for digital partitioning of samples. As demonstrated in the example shown in Fig. 14, if the incident rate, or positive sample rate, is less than the allelic fraction detection limit, then pooling samples for a first assay for detection of a target of interest may be leveraged (left side). In this example, ten samples are pooled into one for a first assay (e.g., samples 1-10, 11-20, 21- 30 and 31-40). If a sample is detected as being positive for a target of interest (e.g., sample 37) following the methods described herein, a re-assay of that pool of samples is performed. For example, if the incident rate for a positive detection of a target of interest is approximately 1 %,and the allele fraction detection limit is approximately 10%, then approximately 100 samples can be pooled to run across, for example, ten wells. As such, for Strategy 4, a possible total of 20 wells can be used for performing assays on 100 different samples, which serves to save reagents and costs for a given assay. As seen in Fig 14, if a pooled well tests positive for the presence of the target nucleic acid via detection and decoding (pooled well of samples 31 -40), then that well can be re-assayed and each sample separated out and run individually (right side) for determining which sample in the original pool tested positive for the presence of the target of interest, in this example sample 37 tested positive for the target of interest by detection and decoding following the methods described herein.
[0171] In some embodiments, samples can be pooled such that one sample is included in two or more pools of multiple different samples. For example, sample 1 can be included in two pools, wherein each pool, in addition to including sample 1, also includes four, five, six, or more different samples, or portions thereof, pooled together. A first pool might include sample 1 or a portion thereof, along with sample 5 or a portion thereof, sample 10 or a portion thereof, sample 20 for a portion thereof, etc. A second pool might include sample 50 or a portion thereof, sample 9 or a portion thereof, sample 4 or a portion thereof, and again sample 1 or a portion thereof. In some embodiments, the same sample, or a portion thereof, is included in more than two pools, for example three pools, four pools, etc. To continue with the example, sample 1 or a portion thereof can be included in pool 1, pool 3, pool 5, pool 7 and the like. The example holds true for other samples, in that each sample or a portionthereofcanbe found in two or more pools, three or more pools, four or more pools, etc. In some embodiments, when samples are assayed in more than one pool, targets of interest can be identified by detecting a code of a recognition element that targets a specific target of interest wherein the code serves as a proxy for the target of interest following detection and decoding as described herein. When a positive result is identified, such that a target of interest is identified in two or more pools of samples, one may look at the sample number that is common to the pools that tested positive for the target of interest to identify the sample that has tested positive for the target of interest. As such, individual sample re-assaying as described herein and as shown in the example illustrated in Fig. 14 to identify the one sample from the pool that resulted in a positive assay for the target of interest, can be avoided if desired.
[0172] In some embodiments, the present disclosure provides methods that can significantly reduce the time from sample to answer using decoding by sequencing. For example, in some instances, the time from sample to answer by practicing the methods described herein can be less than eight hours, which comprises a preparation time of four to six hours and sequencing time of two hours. Next generation sequencing, for example sequencing by synthesis, can takemuch longer, typically > 13 hours or more depending on the sequencing depth needed and the assay used. As such, practicingthe methods described herein can significantly reduce that time, thereby providing data in a more timely manner.
[0173] In some embodiments, the methods practiced herein can be used for identifying variants such as single nucleotide polymorphisms, insertions, deletions and copy number variants, target methylation, proteomic analysis, RNA fusion events, pseudogenes, tandem repeats and phased variants. The present methods can utilize samples comprising direct DNA capture and direct RNA capture. In some embodiments, the present methods can find utility in the fields of cancer recurrence by determining tumor profiles and identifying minimal residual disease targets, pharmacogenomics, genotyping, agriculture, companion animal genetic determinations, early cancer detection, hereditary disease, early cancer screening, expression - based diagnostics, targeted research, non-invasive neonatal testing, pathogen detection, companion diagnostic research, neurological research, single cell genomics and proteomics and spatial transcriptomics and proteomics. In some embodiments, a method for decoding by sequencing comprises the use of a ligated recognition element. Referring to the example illustrated in Fig. 15 A, 1500 represents a ligated recognition element from an assay as illustrated, for example in Fig. 1 and Fig. 2, and as described herein. In some embodiments, each recognition element comprises, minimally, a ligated 5’ and 3’ end 1501 and 1502, respectively, a sample index 1503 which can find utility for identifying one sample from another sample when multiple samples are assayed in the same reaction (e.g., sample multiplexing), a hypercode 1504, an optional unique molecular identifier (UMI) 1505, and a primer binding site 1506. In some embodiments, the UMI 1505 comprises a random sequence for each recognition element.
[0174] In some embodiments, a strategy for decoding by sequencing comprises a ligated recognition element 1500 and a complementary primer 1506. In some embodiments, the primer binding site complementary primer 1506 comprises a portion that is complementary to the primer binding site 1506 of the recognition element and a 5’ portion that is non-complementary to any recognition element sequence 1507, however, is capable of hybridizing to a capture oligonucleotide affixed to a flow cell for next generation sequencing. In some embodiments, the index 1503 can alternatively be located in the primer complementary to 1506, for example between the non-complementary sequence 1507 and the complementary primer binding site portion 1506. In some embodiments, the primer binding site is not an additional site, but instead the primer binding site can be the 5’ end of the ligated recognition element, the 3’ end of the recognition element, the hypercode, the index, or portions thereof. In this example, a primer site is illustrated.
[0175] In some embodiments, a primer complementary to 1506 is used for primer extension 1508 in the presence of a DNA polymerase, dNTPs and buffers for performing an extension reaction (Fig. 15B). In some embodiments, there is no displacement activity of the DNA polymerase. In some embodiments, the primer extension reaction is performed multiple times, for example by cycling primer extension reaction temperatures such that there are multiple cycles of denaturation, primer hybridization and primer extension, thereby generating multiple copies of a non-exponentially amplified recognition element 1510. In some embodiments, a DNA polymerase useful in primer extension comprises a thermostable DNA polymerase. In some embodiments, a DNA polymerase includes, but is not limited to, Taq polymerase, Pfu polymerase, Pfx polymerase, Phusion polymerase, Tth polymerase, KOD polymerase, BST polymerase, Q5 DNA polymerase, or variants or derivatives thereof. As such, following linear extension amplification 1508 of 1500, a plurality of amplification products 1510 is generated comprises all of the portions of the original recognition element, additionally the 5’ non- complementary sequence 1507 on the 5’ end of the linear amplicon.
[0176] In some embodiments, the amplicon 1510 can be applied to a sequencing flow cell Fig. 15C, which comprises the sequence complementary to the 1507 flow cell capture sequence, at which point the sequence 1507 can hybridize to its complementary sequence which is immobilized on the surface of the flow cell 1511. In some embodiments, the amplicon 1510 is denatured prior to capture on a flow cell. In some embodiments, sequence 1507 can be used as a sequencing primer binding site for complementary sequencing primer 1507 if the sequence 1509 is hybridized to its complementary capture probe 1509 on a flow cell.
[0177] In some embodiments, a second primer 1509 can hybridize to and extend the amplicon 1510. In this example, the second primer 1509 comprises a portion 1503 that is complementary to the index sequence 1503, and a non-complementary portion 1509 that comprises a sequence that is complementary to a second oligonucleotide that is affixed to the flow cell 1509. In some embodiments, sequence 1509 can be used as a sequencing primer binding site for sequencing primer 1509 if the sequence 1507 is hybridized to a capture probe 1507 on a flow cell. In some embodiments, an index can be alternatively located on the second primer 1509, for example between the 1503 index complementary sequence and the non- complementary portion 1509, and not on the recognition element. In some embodiments, the index 1503 can be alternatively located on the first primer 1506, for example between the non- complementary end sequence 1507 and the portion 1506 that is complementary to the primer binding site on the recognition element 1506, and not on the recognition element. In some embodiments, there are three index sequences 1503, for example one on the recognition element 1500 as described herein, one on the first primer 1506 as described herein and the third on thesecond primer 1509 as described herein. The placement and number of the index sequences included in this scenario can be predicated on, for example, the complexity of the assay including, but not limited to, the number of samples desired to be multiplexed and the interactions that might be present on any of the three nucleic acid constructs which need to be addressed.
[0178] In some embodiments, if the amplicon 1510 is again extended using primer 1509, and complementary sequences to 1507 and 1509 are both affixed to a flow cell, then both strands of 1510 can hybridize to the flow cell via the 1507 and 1509 end sequences. In some embodiments, once the amplicons are captured on a flow cell 1511, complementary sequencing primers to 1509 and optionally 1507 can be added to the flow cell, hybridized to the 3’ end complementary sequences of the amplicons and sequence by synthesis extension 1512 can be performed (Fig. 15C). In some embodiments, more than one sequencing primer can be added to sequence different portion of the captured DNA strand 1510.
[0179] In some embodiments, sequencing can be performed for multiple cycles, for example from 6-200 cycles, thereby providing depth of sequencing depending on the desired needs. Regardless of cycle number and depth of sequencing, the data can be demultiplexed for identifying the samples in the assay using the index sequence, UMIs can be filtered for error correction and duplicates, and the hypercode can be decoded and matched with a recognition element and the target of interest. In addition, a digital hypercode count can be established.
[0180] As such, both the amplicon and its complement can be captured on a flow cell and decoding by sequencing can be performed for both strands of the amplicon thereby potentially doubling the amount of data available for analysis.
[0181] In some embodiments, a strategy for decoding by sequencing may be used and as illustrated in the example shown in Figs. 16A-16C. In the example strategy as illustrated in Figs. 16A-16C, a primer extension reaction akin to that in the example of Fig. 15 can be performed, however in this instance a primer extension reaction is performed for one cycle such that only one amplicon for each ligated recognition element is generated. Generating one amplicon per ligated recognition element can provide a mechanism to quantitate the amount of a recognition element that was present in an assay. As illustrated in the example shown in Fig.16A, a ligated recognition element 1600 is utilized for generating an amplicon 1610 as shown in Fig. 16B. In some embodiments, a ligated recognition element 1600 comprises a ligated 3’ end 1601 and 5’ end 1602. Additional sequences as illustrated may comprise an optional UMI 1603, a hypercode 1604 that is unique to a recognition element, a primer binding site 1605 and a sequence 1606 that is similar or the same, in part or in whole, to an oligonucleotide affixed on a sequencing flow cell.
[0182] In some embodiments, a primer for primer extension 1605’ that is complementary to a portion of a recognition element is added to the recognition element 1600. In some embodiments, the primer 1605’ comprises a portion complementary to the primer binding site 1605, a portion that comprises an index sequence 1607 for sample multiplexing and a sequencing primer binding site 1608, wherein sequences 1607 and 1608 are non-complementary to other recognition element sequences.
[0183] In some embodiments, the primer 1605’ can prime an extension reaction, or a onesided PCR reaction, on the ligated recognition element 1600, wherein the extension reaction is performed using a DNA polymerase, dNTPs and buffers for performing an extension reaction. In some embodiments, there is no displacement activity of the DNA polymerase and only one cycle of amplification is performed. In some embodiments, a DNA polymerase useful in primer extension linear amplification comprises a thermostable DNA polymerase. In some embodiments, a DNA polymerase includes, but is not limited to, Taq polymerase, Pfu polymerase, Pfx polymerase, Phusion polymerase, Tth polymerase, KOD polymerase, BST polymerase, Q5 DNA polymerase, or variants or derivatives thereof. As such, in some embodiments, following linear extension amplification, one amplicon is generated 1610 that comprises all of the portions of the original recognition element, additionally flanked on both ends by one of the sequencing related sequences 1608 and 1606’.
[0184] In some embodiments, the recognition element 1600 is a RNA recognition element. In some embodiments, if the recognition element in a RNA recognition element, then reverse transcription in lieu of a DNA extension reaction can be performed, such that the extension product 1610 can comprise complementary DNA, or cDNA. In such a case, reverse transcriptase in lieu of a DNA polymerase can be used to perform the extension reaction and the primer 1605’ can be an RNA transcription-based primer for performing a one-sided RT-PCR reaction.
[0185] In some embodiments, the amplicon 1610 can be applied to a sequencing flow cell 1609 (Fig. 16C) and the affixed oligonucleotide 1606 can capture the amplicon complementary sequence 1606’. In some embodiments, the separate samplescan be pooled, hybridized to a flow cell 1609, bridge amplified, and sequenced. In some embodiments, recombinase polymerase amplification is utilized to amplify the amplicon affixed to the flow cell. In this strategy, it is contemplated that only the amplicons are able to hybridize to the flow cell for bridge amplification as no sequences complementary to any ligated recognition elements can be present on the flow cell. As such, the example scenario of Fig. 16 can provide a sequencing library preparation workflow where no post extension clean-up to remove the template ligated recognition elements is needed and UMIs are optional. In some embodiments, following bridge amplification and addition of sequencing by synthesis related reagents, a sequencing primer1608’ can be added to sequence through the index 1607, and primer 1605 can be used as a sequencing primer to determine the sequence of the hypercode 1604’ and additional recognition element sequences, including the UMI 1603’ if present. In some embodiments, the sequencing primer 1608’ is added and the sequence of the index 1607 is read for a plurality of cycles, followed by removal of the sequencing primer 1608’ and addition of the primer 1605 for reading the recognition element sequences for a plurality of cycles which may be more than that of the sequencing primer 1608’. For example, the index sequence may be read for eight cycles and the recognition element sequences can be read to 18 cycles. In some embodiments, primer 1608’ is the only sequencing primer added into the sequencing reaction, and the number of cycles is predicated on generating sequencing reads for all the relevant amplicon elements, including portions 1607, 1605, 1604, and 1603, or the complements thereof .
[0186] In some embodiments, additional sequencing primer binding sites can be included between 1608 and 1607, thereby providing an alternative mechanism for sequencing through the index and optionally other recognition element sequences. In some embodiments, the sequences 1608 and 1606 may be switched, such that 1608 is included in the recognition element and 1606 is at the 5’ end of the primer 1605’. In such a scenario, 1608’ can be captured on a flow cell by a complementary capture oligonucleotide affixed to the flow cell and a sequencing primer 1606’ can be used to sequence the recognition element. In other embodiments, a recognition element can comprise additional sequences to those shown in the examples in Figs. 15A-15C and Figs. 16A-16C, including, but not limited to, functional sequences such as additional primer binding sites, cleavage sites, additional sequencing related sequences such as sequences complementary to additional capture oligonucleotides affixed to a flow cell or additional sequencing primer binding sites.
[0187] In some embodiments, sequencing of the library preparation amplicons illustrated in the examples of Figs. 15A-15C and Figs. 16A-16C can be performed for multiple cycles, for example from 1 -200 cycles, from 10-150 cycles, from 20-100 cycles, from 30-75 cycles, or from 40-60 cycles, thereby providing depth of sequencing depending on the desired needs. Regardless of cycle number and depth of sequencing, the data can be demultiplexed for identifying the samples using the index sequence, UMIs can be optionally filtered for error correction and duplicates if they were included in the recognition element, and the hypercode can be decoded and matched with a recognition element and the target of interest. In addition, a digital hypercode count can be established and recognition elements can be quantified.
Claims
CLAIMSWHAT IS CLAIMED IS1 . A method for determining whether to screen a sample for a target of interest, comprising: a) hybridizing a sample with a recognition element, wherein the recognition element comprises; i) a 5’ end sequence and a 3’ end sequence, wherein the 5’ end sequence and the 3’ end sequence of the recognition element hybridize to complementary target sequences of interest in the sample to generate a hybridized recognition element; and ii) a code that is unique to the recognition element; b) ligating the 5 ’ end and the 3 ’ end of the hybridized recognition element, thereby generating a circularized recognition element; c) amplifying the circularized recognition element, thereby generating an amplified recognition element; d) detecting and decoding the code of the amplified recognition element; and e) based at least in part on the decoding of d), determine whether to perform screening for the presence of the target of interest.
2. The method of claim 1, wherein the sample is an extracted nucleic acid sample.
3. The method of claim 2, wherein the sample is a tissue sample, a cell sample or a blood sample.
4. The method of claim 1, wherein the sample is derived from cell -free DNA, genomic DNA, cDNA, or RNA.
5. The method of claim 1, wherein the 5’ end sequence and 3’ end sequence of the recognition element adjacently hybridize to the complementary target sequences of interest.
6. The method of claim 1, wherein the code comprises at least 10 nucleotides.
7. The method of claim 1, wherein the code comprises one or more nucleic acid segments.
8. The method of claim 1, wherein the code comprises two or more nucleic acid segments.
9. The method of claim 1, wherein the code comprises two to ten nucleic acid segments.
10. The method of claim 1, wherein the amplifying comprises exponential amplification, rolling circle amplification, multiple strand displacement amplification, or extension amplification.11 . The method of claim 1, further comprising amplifying a target region of interest in the sample prior to a).
12. The method of claim 11, wherein the amplifying the target region of interest in the sample comprises polymerase chain reaction.
13. The method of claim 1, wherein the recognition element is a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5’ end and the 3’ end of the recognition element.
14. The method of claim 1, wherein the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof.
15. The method of claim 1 , wherein the detecting the code comprises performing a hybridization event by hybridizing one or more fluorescently labelled detection polynucleotides to all or a portion of the code and imaging the hybridization event.
16. The method of claim 15, wherein the hybridization event and the imaging of the hybridization event occurs more than one time such that a pattern of fluorescent images is generated which is indicative of the presence of the code.
17. The method of claim 1, wherein the decoding the code comprises soft decision decoding.
18. The method of claim 17, wherein the soft decision decoding generates a probability of the presence of the code and wherein the presence of the code indicates the presence of the target of interest in the sample.
19. The method of claim 18, wherein the probability is high for the presence of the code and the presence of the target of interest in the sample such that additional screening of the sample for the target of interest is indicated.
20. The method of claim 18, wherein the probability is low for the presence of the code and the presence of the target of interest in the sample such that additional screening of the sample for the target of interest is not indicated.21 . The method of claim 1, wherein the screening for the presence of the target of interest comprises sequencing, microarrays, qPCR, or real-time PCR.
22. The method of claim 21, wherein the sequencing comprises sequencing by synthesis.
23. The method of claim 1, wherein the target of interest is indicative of the presence of a disease state or an abnormal condition.
24. The method of claim 23, wherein the disease state or the abnormal condition is one or more of a cancer, a neurological disease, a cardiac disease, an immunological disease, a gastrointestinal disease, a drug metabolism disease, or an aneuploidy.
25. The method of claim 1, wherein the target of interest is used for determining tumor mutational burden of a subject.
26. A method for determining a sequence of a target of interest, comprising:a) providing an extracted nucleic acid sample, wherein the nucleic acid sample comprises a target of interest; b) amplifying the target of interest from the extracted nucleic acid sample to produce an amplified target of interest; c) splitting all or a portion of the amplified target of interest into a first reaction and a second reaction; d) providing a recognition element to the first reaction, wherein the recognition element comprises i) a 5 ’ end and a 3 ’ end that are complementary to the target of interest and ii) a code unique to the recognition element, wherein the recognition element and the target of interest hybridize such that the recognition element is ligated to form a circularized recognition element; e) detecting and decoding the code of the circularized recognition element, thereby identifying a sequence of the target of interest; and f) providing a pair of sequencing library preparation primers to the second reaction and amplifying and adapterizing the amplified target of interest in the second reaction, thereby generating a library preparation, and sequencing the library preparation, thereby determining the sequence of a target of interest.
27. The method of claim 26, wherein d), e), and f) are performed sequentially.
28. The method of claim 26, wherein d), e), and f) are performed concurrently.
29. The method of claim 26, wherein the extracted nucleic acid sample is derived from a tissue sample, a cell sample, or a blood sample.
30. The method of claim 26, wherein the extracted nucleic acid sample comprises cell -free DNA, genomic DNA, cDNA, or RNA.31 . The method of claim 26, wherein the 5 ’ end and 3 ’ end of the recognition element hybridize adjacently to the complementary target of interest.
32. The method of claim 26, wherein the code comprises at least 5 nucleotides.
33. The method of claim 26, wherein the code comprises one or more nucleic acid segments.
34. The method of claim 26, wherein the code comprises two or more nucleic acid segments.
35. The method of claim 26, wherein the code comprises two to ten nucleic acid segments.
36. The method of claim 26, wherein d) further comprises amplifying the circularized recognition element prior to the detection and decoding of e).
37. The method of claim 36, wherein the amplifying the circularized recognition element is performed by primer extension, rolling circle amplification or strand displacement amplification.
38. The method of claim 26, wherein amplifying the target of interest comprises PCR.
39. The method of claim 26, wherein the recognition element is a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5’ end and the 3’ end of the recognition element.
40. The method of claim 26, wherein the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof.41 . The method of claim 26, wherein the detecting comprises performing a hybridization event by hybridizing one or more fluorescently labelled detection polynucleotides to all or a portion of the code and imaging the hybridization event.
42. The method of claim 41, wherein the hybridization event and the imaging of the hybridization event occurs more than once, such that a pattern of fluorescent images is generated.
43. The method of claim 26, wherein the decoding the code comprises soft decision decoding.
44. The method of claim 43, wherein the soft decision decoding provides a probability of the presence of the code and wherein the presence of the code indicates a presence of the target of interest in the sample.
45. The method of claim 26, wherein sequencing the library preparation of d) comprises next generation sequencing.
46. The method of claim 45, wherein the next generation sequencing comprises sequencing by synthesis.
47. A method for determining a sequence of one or more samples, comprising: a) providing a reversibly partitioned first substrate for immobilizing one or more samples; b) immobilizing the one or more samples in one or more partitions of the reversibly partitioned first substrate, wherein the one or more partitions comprises one or more of the one or more samples; c) adding to each partition comprising one of the one or more partitions one or more recognition elements, wherein each recognition element of the one or more recognition elements comprises: i) a 5 ’ end and a 3 ’ end that are complementary to a target sequence in the one or more samples; and ii) a code that is unique to a recognition element of the one or more recognition elements, wherein the code serves as a proxy for the presence of the target sequence in the one or more samples;d) hybridizing each recognition element of the one or more recognition elements to its complementary target sequence, such that the hybridizing generates a recognition element comprising a padlock probe configuration; e) ligating the 5 ’ end and the 3 ’ end of the one or more recognition elements to generate one or more circularized recognition elements; f) amplifying the one or more circularized recognition elements; g) removing the reversible partitions from the first substrate; h) affixing a second substrate over the first substrate, thereby sealing the second substrate to the first substrate and generating a flow cell, wherein the flow cell comprises a first reagent inlet port and a second reagent outlet port; i) flowing a plurality for detection polynucleotides through the flow cell, wherein the detection polynucleotides hybridize to the code, or a portion thereof ; and j) detectingthe hybridization of the detection polynucleotides to the code, or the portion thereof, and decoding the detecting, thereby determining the sequence of the one or more samples.
48. The method of claim 47, wherein the one or more samples are extracted nucleic acid samples.
49. The method of claim 47, wherein the one or more samples are nucleic acid samples derived from one or more of a tissue sample, a cell sample, or a blood sample.
50. The method of claim 48, wherein the extracted nucleic acid samples comprises cell-free DNA, genomic DNA, cDNA, or RNA.
51. The method of claim 47, wherein the first substrate comprises optically clear glass.
52. The method of claim 47, wherein the second substrate comprises an adhesive film, glass, or plastic.
53. The method of claim 47, wherein the partition of the one or more partitions comprise a material selected from the group consisting of polystyrene, polycarbonate, plastic, rubber, silicone, polyether ketone, and cyclic olefin copolymer.
54. The method of claim 53, wherein the second substrate is reversibly applied to the first substrate using one or more interfaces selected from the group consisting of a single sided adhesive, acrylates, silicones, foams, a temperature sensitive adhesive, a fluorinated compound, dimethyl siloxanes, neoprene, and butyl rubber.
55. The method of claim 47, wherein a surface of the first substrate comprises a nucleic acid immobilization compound selected from the group consisting of poly lysine or an isomer thereof, PEG, cationic polymers, amine coating, amino silanes, silane, thiol / gold chemistry, reversible addition fragmentation chain transfer, and hydrogels.
56. The method of claim 47, wherein the code comprises at least 5 nucleotides.
57. The method of claim 47, wherein the code comprises one or more nucleic acid segments.
58. The method of claim 47, wherein the code comprises two or more nucleic acid segments.
59. The method of claim 47, wherein the code comprises two to ten nucleic acid segments.
60. The method of claim 47, wherein the amplifying is performed by primer extension, rolling circle amplification or strand displacement amplification.
61. The method of claim 47, wherein the recognition element comprises a padlock probe, a molecular inversion probe, or a probe that is circularizable upon ligation of the 5 ’ end and the 3 ’ end of the recognition element.
62. The method of claim 47, wherein the recognition element further comprises one or more of a sample index sequence, one or more amplification primer binding sites, one or more sequencing primer binding sites, a unique molecular identifier sequence, a cleavage site, or any combination thereof.
63. The method of claim 47, wherein the detection polynucleotides are fluorescently labelled.
64. The method of claim 47, wherein the detection further comprises imaging the hybridization of the detection polynucleotides to the code, or the portion thereof.
65. The method of claim 64, wherein the flowing of i), the detecting of j), and the imaging is repeated more than once, such that a pattern of symbols is generated.
66. The method of claim 47, wherein the decoding the detecting comprises soft decision decoding.
67. The method of claim 47, wherein the soft decision decoding provides a probability of the presence of the code, wherein the code serves as a proxy for the presence of the sequence of the one or more samples.
68. A method for determining a sequence of a target of interest, comprising: a) generating a sequencing library from a first portion of a sample comprising a target of interest; b) performing whole exome sequencing on the sequencing library, thereby determining the sequence of the target of interest; c) providing one or more recognition elements to a second portion of the sample comprising the target of interest, wherein each recognition element of the one or more recognition elements comprises i) a 5 ’ end and a 3 ’ end which are complementary to the target of interest, and ii) a code that is unique to each recognition element, wherein the code serves as a proxy for the presence of the target of interest; d) hybridizing the one or more recognition elements to their complementary sequence in the target of interest to produce hybridized recognition elements;e) ligating the 5’ end and the 3’ end of the hybridized recognition elements to generate circularized recognition elements; f) amplifying the circularized recognition elements to generate amplified recognition elements; and g) detecting and decoding the code of the amplified recognition element; wherein the sequence of the target of interest is determined by combining whole exome sequencing data from b) and decoding data from g).
69. The method of claim 68, wherein the target of interest is derived from a tissue sample, a cell sample, or a blood sample.
70. The method of claim 68, wherein the target of interest is derived from cell -free DNA, genomic DNA, cDNA, or RNA.
71. The method of claim 68, wherein the target of interest comprises a copy number variance, high GC content, one or more homopolymeric regions, an insertion, or a deletion.
72. The method of claim 68, wherein the target of interest is a star allele and comprises a homolog or a pseudogene, and wherein determining the sequence of the target of interest differentiates the target of interest from the homolog or the pseudogene.
73. A composition comprising: a) a recognition element, wherein the recognition element comprises a 5’ end and a 3 ’ end, wherein the 5’ end and the 3’ end are hybridized to non-adjacent target nucleic acid sequences; and b) a third oligonucleotide that is hybridized to all or a portion of a nucleic acid sequence that is located between the non-adjacent target nucleic acid sequences that are hybridized such that a 3’ end of the third oligonucleotide is hybridized adjacent to the 5’ end of the recognition element and a 5’ end of the third oligonucleotide is hybridized adjacent to the 3’ end of the recognition element; and wherein the third oligonucleotide, when hybridized to the target nucleic acid sequences, forms a loop structure such that the loop structure is not hybridized to the nucleic acid sequence that is located between the non -adjacently hybridized target nucleic acid sequences.
74. The composition of claim 73, wherein a portion of the loop structure is a stem -loop structure.
75. The composition of claim 73 , wherein a portion of the loop structure is not a self -hybridizing loop structure.
76. The composition of claim 73, wherein the third oligonucleotide comprises one or more of a primer binding site, a sample index, a unique molecular identifier and a cleavage site.
77. The composition of claim 73, wherein the recognition element comprises a code and one or more of a primer binding site, a unique molecular identifier, a sample index and a cleavage site.
78. The composition of claim 73, wherein the 5’ end of the recognition element and the 3’ end of the third oligonucleotide and the 3 ’ end of the recognition element and the 5 ’ end of the third oligonucleotide are ligated together to form a circularized recognition element comprising the recognition element and the third oligonucleotide.
79. A method for determining a presence of a target of interest, comprising: a) hybridizing a first amplification primer to a ligated recognition element, wherein the ligated recognition element comprises a complementary sequence of the target of interest, a hypercode, a sample index, and optionally a unique molecular identifier, wherein the first amplification primer comprises a first sequence which is capable of being captured on a flow cell; b) performing primer extension of the first amplification primer that is hybridized to the ligated recognition element, thereby generating a plurality of amplicons, wherein each amplicon of the plurality of amplicons comprises a sequence of the target of interest, or a complement thereof, the hypercode, the sample index, and optionally the unique molecular identifier and the first sequence which is capable of being captured on the flow cell; c) hybridizing a second amplification primer to the plurality of amplicons, wherein the second amplification primer comprises a second sequence which is capable of being captured on the flow cell; d) performing primer extension of the second amplification primer that is hybridized to the plurality of amplicons, thereby generating a plurality of double-stranded amplicons wherein each double-stranded amplicon of the plurality of double-stranded ampliconsis flanked on a first end by the first sequence and on a second end by the second sequence; e) denaturing the plurality of double-stranded amplicons to produce denatured amplicons and capturing a plurality of single strands of the denatured amplicons on the flow cell; and f) sequencing at least the hypercode, the sample index, and optional the unique molecular identifier of the captured amplicons, or complements thereof, thereby determining the presence of the target of interest.
80. The method of claim 79, wherein the target of interest is derived from a sample.
81. The method of claim 79, wherein the target of interest is a variant target from a list consisting of a single nucleotide polymorphism, an insertion, a deletion, a splice variant and a copy number variant.
82. The method of claim 80, wherein the sample is a tumor sample, a tissue sample, a cell sample, a blood sample, or a lysate.
83. The method of claim 80, wherein the sample is obtained or derived from a eukaryote, a prokaryote, a mammal, a veterinary office, an agricultural sample, a companion animal, or a livestock animal.
84. The method of any one of claims 79-83, wherein the ligated recognition element is generated by practicing the method of Fig. 2.
85. The method of any one of claims 79-84, wherein the primer extension of b) or d) is performed for a plurality of cycles, thereby generating a plurality of amplicons.
86. The method of any one of claims 79-85, wherein a sequence of the hypercode is used as a proxy for identifying the presence of the target of interest.
87. The method of any one of claims 79-86, wherein the first sequence of the first amplification primer and the second sequence of the second amplification primer are further sequencing primer binding sites for hybridizing complementary sequencing primers for initiating sequence by synthesis.
88. The method of claim 87, wherein the first sequence is captured on the flow cell and the second sequence is hybridized to the complementary sequencing primer for initiating sequence by synthesis.
89. The method of claim 87, wherein the second sequence is captured on the flow cell and the first sequence is hybridized to the complementary sequencing primer for initiating sequence by synthesis.
90. The method of any one of claims 79-89, wherein the sequencing comprises sequence by synthesis.
91. The method of any one of claims 79-90, wherein the sequencing comprises performing bridge amplification on the single strands of the denatured amplicons, thereby forming colonies of single-stranded amplicons on the surface of the flow cell.
92. The method of claim 79, wherein the sample index is alternatively located on the first amplification primer or the second amplification primer.
93. A method for determining a presence of a target of interest, comprising: a) hybridizing an amplification primer to a ligated recognition element, wherein the ligated recognition element comprises an amplification primer binding site, a hypercode, a sample index, and a sequence that is capable of being captured on a flow cell, wherein the amplification primer comprises a complement of the primer binding site and a sequence primer binding site which is complementary to a sequencing primer;b) performing primer extension on the amplification primer that is hybridized to the ligated recognition element, thereby generating an amplicon comprising the hypercode, the sample index, the sequence that is capable of being captured on the flow cell, and the sequence primer binding site which is complementary to the sequencing primer; c) capturing the amplicon on the flow cell; and d) sequencing at least the hypercode and the sample index of the captured amplicon, thereby determining the presence of the target of interest.
94. The method of claim 93, wherein the target of interest is derived from a sample.
95. The method of claim 93, where the target of interest is a variant target from a list consisting of a single nucleotide polymorphism, an insertion, a deletion, a splice variant and a copy number variant.
96. The method of claim 94, wherein the sample is a tumor sample, a tissue sample, a cell sample, a blood sample or a lysate.
97. The method of claim 94, wherein the sample is obtained or derived from a eukaryote, a prokaryote, a mammal, a veterinary office, an agricultural sample or a livestock sample.
98. The method of any one of claims 93-97, wherein the ligated recognition elementis generated by practicing the method of Fig. 2.
99. The method of any one of claims 93-98, wherein the primer extension is performed for one cycle, thereby generating one amplicon.
100. The method of any one of claims 93-99, wherein a sequence of the hypercode is used as a proxy for identifying the present of the target of interest.
101. The method of any one of claims 93-100, wherein the ligated recognition element alternatively comprises the sequence primer binding site which is complementary to the sequencing primer and the amplification primer comprises the sequence that is capable of being captured on the flow cell.
102. The method of any one of claims 93-101, wherein sequencing comprises providing two sequencing primers.
103. The method of claim 102, wherein one of the two sequencing primers is complementary to the sequence primer binding site and the second of the two sequencing primers is complementary to the amplification primer binding site.
104. The method of any one of claims 93-103, wherein the sequencing comprises sequence by synthesis.
105. The method of any one of claims 93-104, wherein the sequencing comprises performing bridge amplification on the amplicon captured on the flow cell, thereby forming colonies of the amplicon on the flow cell.
106. The method of any one of claims 93-105, wherein the ligated recognition element further comprises a unique molecular identifier.
107. The method of any one of claims 93-106, wherein the ligated recognition element is a RNA ligated recognition element.
108. The method of claim 107, wherein the primer extension is reverse transcription of the RNA ligated recognition element, thereby generating a cDNA amplicon.
Citation Information
Patent Citations
Encoded assays
WO2022109496A2
Encoded assays
WO2023096674A1
Multiscale lens systems and methods for imaging well plates and including event-based detection
WO2023158993A2
Methods and compositions for multimodal in SITU analysis
EP4012046A1
Multiplexed covid-19 padlock assay
US20230295692A1