Methods and compositions for preparing samples for sequencing

Probes with sensor and linearization regions enhance library preparation by enabling efficient identification and sequencing of genetic markers, improving sequencing outcomes through circularization and re-linearization processes.

WO2026096395A1PCT designated stage Publication Date: 2026-05-07ILLUMINA INC
View PDF 23 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-10-27
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current library preparation methods for sequencing are inefficient and lack the ability to effectively identify and amplify specific genetic markers, leading to suboptimal sequencing outcomes.

Method used

The use of probes with a sensor region for circularization and a linearization region for re-linearization, combined with amplification and sequencing regions, allows for the identification and sequencing of genetic markers, and can be attached to solid surfaces for enhanced library preparation.

Benefits of technology

This approach enables efficient enrichment and library preparation, facilitating the accurate sequencing of genetic markers and diseases/disorders by enhancing the sensitivity and specificity of nucleic acid analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052713_07052026_PF_FP_ABST
    Figure US2025052713_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are compositions and methods for preparing target nucleic acid for sequencing. In general, the method uses reversibly circularizable probes that target specific nucleic acid, including genetic markers. In some embodiments the probes include a sensor region and a linearization region. In other embodiments, the probes further include a code region. In yet other embodiments, the probes further include a code region and / or at least one primer region.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND COMPOSITIONS FOR PREPARING SAMPLES FOR SEQUENCINGBACKGROUND

[0001] Library preparation is the conversion of nucleic acid to an appropriate format for the chosen sequencing platform. This generally involves extracting nucleic acid, fragmenting the extracted nucleic acid, and adding adaptors compatible with the sequencing platform. Disclosed herein are compositions and methods for an improved library preparation method.SUMMARY

[0002] Probes are disclosed herein that include a sensor region and a linearization region. The sensor region can be used to circularize the probe. The linearization region can be used to re-linearize circular probes. In some embodiments, the probes include regions for amplifying and / or sequencing the probes. In other embodiments, the probes include code regions to identify various aspects of the probe such as the target sequence or the sample. In some embodiments, the amplifying and / or sequencing regions are located near the linearization region such that when the probe is linearized, the amplifying and / or sequencing regions are located at opposing ends of the probe.

[0003] The probes herein can be used to identify target nucleic acid from a sample. The probes disclosed herein can also be used to identify various genetic markers. For example, a panel of probes can used to identify various diseases or disorders. In this embodiment, the sensor region is designed to hybridize to a genetic marker of interest. Once hybridized, the probes are circularized. The circularized probes are linearized and the genetic marker of interest is sequenced.

[0004] The probes herein can be attached to a solid surface such as beads, particles, or the surface of a flow cell.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Fig. 1Aand 1B illustrate non-limiting examples of the probes described herein.

[0006] Fig. 2A, 2B, 2C, 2D, and 2E illustrate non-limiting examples of the probes described herein in both circular and linear form.

[0007] Fig. 3A, 3B, 3C, and 3D illustrate non-limiting examples of the methods described herein.

[0008] Fig. 4 illustrates a non-limiting example of the probes described herein.DETAILED DESCRIPTIONDefinitions

[0009] Terms used herein will be understood to take on their ordinary meaning in the relevant art unless specified otherwise. Several terms used herein and their meanings are set forth below.

[0010] As used herein, the singular forms “a,” “an,” and “the” refer to both the singular as well as plural, unless the context clearly indicates otherwise. The term “comprising” as used herein is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps.

[0011] Reference throughout the specification to “one example,” “another example,” “an example,” and so forth, means that a particular element (e.g., feature, structure, composition, configuration, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.

[0012] The terms “substantially" and “about” used throughout this disclosure, including the claims, are used to describe and account for small fluctuations, such as due to variations in processing. For example, these terms can refer to less than or equal to ±5% from a stated value, such as less than or equal to ±2% from a stated value, suchas less than or equal to ±1% from a stated value, such as less than or equal to ±0.5% from a stated value, such as less than or equal to ±0.2% from a stated value, such as less than or equal to ±0.1% from a stated value, such as less than or equal to ±0.05% from a stated value.

[0013] Amplification Some embodiments further comprise amplifying and / or replicating one or more nucleic acid templates, Including fragments thereof. The amplifying and / or replicating comprises use of one or more of a bridge amplification reaction, an isothermal bridge amplification reaction, a rolling circle amplification (RCA) reaction, a modified rolling circle multiple displacement amplification, a helicase-dependent amplification reaction, a recombinase-dependent amplification reaction, a singlestranded DNA binding (SSB) protein mediated Isothermal amplification, a PCR reaction, a strand-displacement reaction, a ligase chain reaction, a transcription-mediated reaction, a loop-mediated amplification reaction, other suitable reactions, and combinations thereof.

[0014] Clusters: Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster” or “colony”) refers to a localized group or collection of DNA or RNAon a nucleotide-sample support, such as a flow cell, particle, polymer scaffold, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, the cluster can comprise one or more concatemers, such as, for example, a polony or a nanoball. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-patterned flow cell. In typical embodiments, a cluster is the product of an amplification reaction. A cluster of oligonucleotides can be imaged utilizing one or more light signals, changes in pH, changes in conductance, and other signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent labeled nucleotides incorporated into oligonucleotides, fluorescent labeled nucleotides bound but not incorporated into oligonucleotides, and otherfluorescent labeled complexes associated with incorporated or bound nucleotides from one or more clusters on a flow cell. Examples of other sequencing procedures are set forth herein. In some embodiments, a cluster can be monoclonal or polyclonal.

[0015] Amplification Domain A portion of an adapter having a universal nucleotide sequence, such as a P5 or P7 sequence or a complement thereof, that can serve as a starting point for template amplification and cluster generation.

[0016] Corresponds with: When one primer “corresponds with” an amplification domain, the primer and amplification domain may have the same sequence, so that a copy of the amplification domain generates a sequence complementary to the primer.

[0017] Depositing: Any suitable application technique, which may be manual or automated, and, in some instances, results in modification of the surface properties. Generally, depositing may be performed using vapor deposition techniques, coating techniques, grafting techniques, or the like. Some specific examples include chemical vapor deposition (CVD), spray coating (e.g., ultrasonic spray coating), spin coating, dunk or dip coating, doctor blade coating, puddle dispensing, flow through coating, aerosol printing, screen printing, microcontact printing, inkjet printing, or the like.

[0018] Depression: A discrete concave feature in a substrate or a layer of a substrate (e.g., a patterned resin) having a surface opening that is at least partially surrounded by interstitial region(s) of the substrate or the layer. Depressions can have any of a variety of shapes at their opening in a surface including, as examples, round, elliptical, square, polygonal, star shaped (with any number of vertices), etc. The cross-section of a depression taken orthogonally with the surface can be curved, square, polygonal, hyperbolic, conical, angular, etc. The depression may also have more complex architectures, such as ridges, step features, etc.

[0019] DNA Sample: A polymeric form of nucleotides that includes deoxyribonucleotides, deoxyribonucleotide analogs, or complementary deoxyribonucleotides derived from an RNA (ribonucleic acid) sample. The DNA sample is double stranded. The DNA sample may include naturally occurring DNA, which includes a nitrogen containing heterocyclic base (a nucleobase such as adenine, thymine, cytosine and / or guanine), a sugar (specifically deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2’ position in ribose), and a backbone containing phosphodiester bonds. An analogstructure can have an alternate backbone linkage including any of a variety known in the art.

[0020] The sample may be genomic DNA (gDNA) that can be isolated from one or more cells, bodily fluids (e.g., whole blood, blood spots, saliva) or tissues. gDNA can be prepared by lysing a cell that contains the DNA. The cell may be lysed under conditions that substantially preserve the integrity of the cell's gDNA. In one particular example, thermal lysis may be used to lyse a cell. In another particular example, exposure of a cell to alkaline pH can be used to lyse a cell while causing relatively little damage to gDNA. Any of a variety of basic compounds can be used for lysis including, for example, potassium hydroxide, sodium hydroxide, and the like. Additionally, relatively undamaged gDNA can be obtained from a cell lysed by an enzyme that degrades the cell wall. Cells lacking a cell wall either naturally or due to enzymatic removal can also be lysed by exposure to osmotic stress. Other conditions that can be used to lyse a cell include exposure to detergents, mechanical disruption, sonication heat, pressure differential such as in a French press device, or Dounce homogenization. Agents that stabilize gDNA can be included in a cell lysate or isolated gDNA sample including, for example, nuclease inhibitors, chelating agents, salts, buffers and the like.

[0021]

[0078] Each’. When used in reference to a collection of items, each identifies an individual item in the collection, but does not necessarily refer to every item in the collection. Exceptions can occur if explicit disclosure or context clearly dictates otherwise.

[0022] Flow Cell A vessel having an open or enclosed flow channel where a reaction can be carried out, an inlet for delivering reagent(s) to the flow channel, and an outlet for removing reagent(s) from the flow channel. The vessel with an open flow channel may be referred to herein as an open wafer flow cell. In some examples, the flow cell enables the detection of the reaction that occurs therein. For example, the flow cell can include one or more transparent surfaces allowing for the optical detection of arrays, optically labeled molecules, or the like.

[0023] Flow channel: An area that is defined between two bonded or otherwise attached components or that is defined within a lane so that it is open to the surrounding environment. The flow channel can selectively receive a liquid sample. In someexamples, the flow channel may be defined between two patterned sequencing surfaces or a patterned sequencing surface and a lid, and thus may be in fluid communication with one or more components of the sequencing surface(s).

[0024] Immobilization'. The term “immobilized”, “affixed” and “attached” are used interchangeably herein and both terms are intended to encompass direct or indirect, covalent or non-covalent attachment unless indicated otherwise, either explicitly or by context.

[0025] Non-limiting exemplary covalent attachment includes, for example, those that result from the use of click chemistry techniques. Exemplary non-covalent attachment includes, but are not limited to, non-specific interactions (e.g. hydrogen bonding, ionic bonding, van der Waals interactions etc.) or specific Interactions (e.g. affinity interactions, receptor-ligand interactions, antibody-epitope interactions, avidin-biotin interactions, streptavidin-biotin interactions, lectin-carbohydrate interactions, etc.). Exemplary attachments are set forth in U. S. Pat. Nos. 6,737,236 B1; 7,259,258 B2; 7,375,234 B2and 7,427,678 B2; and U. S. Pat. Pub. 2011 / 0059865 A1, each of which is incorporated herein by reference in its entirety.

[0026] In certain embodiments, the molecules (e.g. nucleic acids, enzymes) remain immobilized or attached to the solid support under the conditions in which it is intended to use the solid support, for example in applications requiring nucleic acid amplification and / or sequencing. In other embodiments, the molecules are reversibly immobilized and can be removed from the solid support through the use of cleavable sites, linkers, and the like.

[0027] Nanoballs: Some embodiments further comprise rolling circle amplification / replication used to form nucleic acid nanoballs. The term “nucleic acid nanoball” may be a concatemer comprising multiple copies of a target nucleic acid molecule. These nucleic acid copies may be arranged one after another in a continuous linear strand of nucleotides. These nucleic acid copies may result in a nanoball folding configuration. The multiple copies of a target nucleic acid molecule in a nucleic acid nanoball may each contain an adaptor sequence of known sequence to facilitate amplification or sequencing. The adaptor sequence of each target nucleic acid molecule may be the same or different. The nucleic acid nanoball can be loaded on the surfaceof solid support. The nanoball can be attached to the surface of solid support by any suitable method. Non-limiting examples of such methods include nucleic acid hybridization, biotin streptavidin binding, thiol binding, photoactive binding, covalent binding, antibody-antigen, physical constraints via hydrogels or other porous polymers, etc., or combinations thereof. In some cases, the nanoball can be digested with an enzyme (nuclease, etc.) to produce a smaller nanoball or a fragment from the nanoball.

[0028] Nucleic Acid Sample: A sample, typically derived from any organism, including but not limited to animals, plants, fungi, and microbes. For example, such samples may be derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like.Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. Alternatively, the sample may be microbial such as bacteria, viral, or fungal. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test" samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0029] Patterned / Random-. In some embodiments, the solid support comprises a patterned surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in an ordered pattern. A “patterned surface” refers to an arrangement of different regions in or on an exposed layer of a solid support. The features can be separated by interstitial regions that contribute to the pattern. In some embodiments, the interstitial regions can be a different height, creating wells or raised platform patterns. In other embodiments, the interstitial regions can have a different surface charges. In yet other embodiments, the interstitial regions can have a different attachment moieties. In some embodiments, the pattern can be any suitable pattern, such as a grid patterns, radial patterns, and combinations thereof. In some embodiments, a patterned surface can contain pre-determined locations of features but the features are not arrayed in a repetitive pattern. Examples of grid patterns include rectangular patterns, hexagonal patterns, triangular, and other suitable grid patterns. The regions for immobilization of molecules may be depressed regions, elevated regions, or planar regions relative to the interstitial regions. The regions may be fabricated as is generally known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, microetching techniques, and combinations thereof. As will be appreciated by those in the art, the technique used will depend on the composition and shape of the regions. For example, the regions for immobilization of molecules of a patterned surface may be wells, pits, channels, posts, pillars, ridges, stripes, swirls, lines, and other suitable topographies. For example, the wells may have any opening in any shape, such as circular, oval, polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.). Exemplary patterned surfaces that can be used in the methods and compositions set forth herein are described in U. S. Pat. No. 8,778,849 B2, which is incorporated herein by reference in its entirety.

[0030] In some embodiments, the solid support comprises a surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in a random distribution over the solid support. Exemplary random distribution over a solid support is described in U. S. Pat. No. 8,241,573 B2, which is incorporated herein by reference in its entirety.

[0031] Polonies Some embodiments further comprise rolling circle amplification / replication used to form polonies. The term “polony” or “polonies” used herein refers to a nucleic acid library molecule clonally amplified in-solution or on-support to generate an amplicon that can serve as a template molecule for sequencing. In some aspects, a linear library molecule can be circularized to generate a circularized library molecule, and the circularized library molecule can be clonally amplified insolution or on-support to generate a concatemer. In some aspects, the concatemer can serve as a nucleic acid template molecule which can be sequenced. The concatemer is sometimes referred to as a polony. In some aspects, a polony includes nucleotide strands.

[0032] Probe: A single stranded nucleic acid molecule designed to hybridize to a target nucleic acid sequence. In some examples, only a portion of the probe hybridizes to the target nucleic acid. In these examples, one or both ends may remain single stranded while the middle portion of the probe hybridizes to the target nucleic acid. In other examples, both ends may hybridize to the target nucleic acid while the middle portion of the probe remains single stranded.

[0033] Primer. A single stranded nucleic acid molecule that can hybridize to a target sequence, to enable copying of the target sequence. As one example, a flow cell surface bound primer can serve as a starting point for fragment amplification and cluster generation. As another example, a primer (e g., a sequencing primer) may be introduced that can hybridize to fragments or fragment amplicons in order to prime synthesis of a new strand that is complementary to the fragments or fragment amplicons. Any primer can include any combination of nucleotides or analogs thereof. In some examples, the primer is a single-stranded oligonucleotide or polynucleotide. The primer length can be any number of bases long. In an example, each of the flow cell surface bound primer and the sequencing primer is a short strand, ranging from 10 to 60 bases, or from 20 to 40 bases.

[0034] Sequencing Procedures: The term “read” or “sequence read" (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a shortsequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques.

[0035] Embodiments described herein can be used with any suitable sequencing chemistry, such as sequencing by synthesis (SBS), sequencing by binding, sequencing by ligation, or nanopore sequencing.

[0036] SBS can be with or without the use of reversible terminators. For example, SBS can be initiated by contacting the target nucleic acids with one or more nucleotides (e.g., labelled, synthetic, modified, ora combination thereof), DNA polymerase, etc. Those features where a primer is extended using the target nucleic acid as template will incorporate a labeled nucleotide that can be detected. The incorporation time used in a sequencing run can be significantly reduced using the altered polymerases described herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery steps. Thecycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems, and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U. S. Pat. Nos.7,057,026 B2, 7,329,492 B2, 7,211,414 B2, 7,315,019 B2, 7,405,281 B2, and 8,343,746 B2. Sequence reads can be generated using instruments such as MiniSeq™, MiSeq™, NextSeq™, HiSeq™, and NovaSeq™ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0037] One example of SBS is termed sequencing by binding. One implementation of sequencing by binding includes cycles of initiating sequencing of a template with a reversible blocker on the 3’ end to prevent additional bases from incorporating, interrogating the template by flooding the flow cell with fluorescently tagged bases that do not include a blocker and measuring an emitted signal of bound bases, activating the 3’ end via removal of the reversible blocker, and incorporating the complementary base from unlabeled, blocked nucleotides. Reads using sequencing by binding can be generated from using instruments such as Onso™ sequencing instruments from Pacific Biosciences of California, Inc. (Menlo Park, CA). Another implementation of sequencing by binding could be sequencing by avidity. In sequencing by avidity, fluorescent dye labeled cores termed avidites are used. One potential cycle of sequencing by avidity includes providing a reagent of polymerase and reversibly terminated nucleotides to templates immobilized on a solid surface, de-blocking the incorporated nucleotides, flowing a set of four types of avidites, washing away unbound avidites, detecting the incorporated bases / nucleotides, and removing the bound avidites. The steps in the cycle of sequencing by avidity may be performed in other orders. Sequencing by avidity is described in Arslan, S., Garcia, F. J., Guo, M. et al. Sequencing by avidity enables high accuracy with low reagent consumption. Nat Biotechnol 42, 132-138 (2024). https: / / doi.org / 10.1038 / s41587-023-01750-7, which is incorporated by reference in its entirety. Reads using sequencing by avidity can be generated using instruments such as Aviti™ sequencing instruments from Element Biosciences (San Diego).

[0038] One example of SBS using an open flow cell and without using reversible terminators is disclosed in Almogy, G.(2022) “Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform" https: / / doi.org / 10.1101 / 2022.05.29.493900, which is incorporation by reference in its entirety. Sequence reads using an open flow cell can be generated using instruments such as UG 100TM Sequencer from Ultima Genomics, Inc. (Fremont, CA).

[0039] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are described in U. S. Pat. Nos. 8,262,900 B2, 7,948,015 B2, 8,349,167 B2, and U. S. Pat. Pub. 2010 / 0137143 A1, which are incorporated by reference in its entirety.

[0040] Sequence reads can be generated using instruments such as DNBSEQTM sequencing instruments from MGI Tech Co., Ltd. (Shenzhen, China) and as SURFSeq™, FASTASeq™, and GenoLab™ sequencing instruments from GeneMind Biosciences Co., Ltd. (Shenzhen, China).

[0041] Some embodiments can use methods involving the real-time monitoring of DNA polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate-labeled nucleotides, or with zeromode waveguides. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), which are incorporated by reference in its entirety. Techniques sequencing using zeromode waveguides is described in U. S. Pat. No. 6,917,726 B2, which is incorporated by reference in its entirety.

[0042] Solid Support The terms “solid support,” “solid surface,” and other grammatical equivalents herein refer to any substrate that is appropriate for or can be modified to be appropriate for the attachment of enzymes, nucleic acids, and complexes thereof. As will be appreciated by those in the art, the number of possible substrates is very large.Possible substrates include, but are not limited to, glass and modified or functionalized glass, polymers (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, plastics, optical fiber bundles, quartz, metal oxides, inorganic oxides, other suitable transparent materials, other suitable non-transparent materials, other suitable translucent materials, and combinations thereof. The composition and geometry of the solid support can vary with its use.

[0043] In some embodiments, the solid support or solid surface is a planar structure, such as a flowcell, slide, chip, microchip, array, microarray, wafer, panel, charge pad, and / or web. The planar structure can be a single surface structure having a single surface of sample / reaction sites. The planar structure can be a dual surface structure. One example of a dual surface structure includes a top substrate having a top surface of sample / reactions sites, a bottom substrate having a bottom surface of sample / reactions sites, and a spacer layer separating the top substrate and the bottom substrate. The solid support or solid surface can be open to direct application of a fluid. One example of an open solid support or open solid surface is an open flow cell having a single surface structure without an inlet port. In some embodiments, the solid support is not necessarily planar, such as, for example, the surface of a well, tube, or other vessel. Nonlimiting examples include the surface of a microcentrifuge tube, a well of a multiwell plate, and the like.

[0044] In some embodiments, the solid support comprises one or more surfaces of a flowcell or flow cell. The term “flowcell” or “flow cell” as used herein refers to a solid surface across which one or more fluid reagents can be flowed. Examples of flowcells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U. S. 7,057,026 B2; WO 91 / 06678; WO 07 / 123744; U. S. 7,329,492 B2; U. S. 7,211,414 B2; U. S. 7,315,019 B2; U. S. 7,405,281 B2, and U. S. Pat. Pub. 2008 / 0108082 A1, each of which is incorporated herein by reference in its entirety. In some embodiments, the flowcells can be one or more flow lanes. For flowcells having a plurality of flow lanes, each of the flow lanes can be independently accessed or two or more flow lanes can be accessed as a group.

[0045] In some embodiments, the solid support or solid surface is a non-planar structure, such beads, microspheres, and / or inner and / or outer surface of a tube or vessel. The terms “beads”, ‘'microspheres,” or “particles” or grammatical equivalents herein is refer to small discrete particles. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, thoria sol, carbon graphite, titanium dioxide, latex, polysaccharide (e.g. Dextran™, Sepharose™, cellulose, nylon, cross-linked micelles, Teflon™, as well as any other materials outlined herein for solid supports may all be used. “Microsphere Detection Guide” from Bangs Laboratories, Fishers Ind. is a helpful guide. In certain embodiments, the microspheres are magnetic microspheres or beads. The beads need not be spherical; irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead sizes range from nanometers, i.e. 100 nm, to millimeters, i.e. 1 mm, with beads from about 0.2 micron to about 200 microns being preferred, and from about 0.5 to about 5 micron being particularly preferred, although in some embodiments smaller or larger beads may be used.Compositions

[0046] Disclosed herein are probes for preparing nucleic acid for sequencing. In general, the probes are circularizable, that is when properly hybridized to a target nucleic acid the ends of the probes can be joined together. In other embodiments, the probes are reversibly circularizable in that after circularization, the probes can be relinearized. The reversibly circularizable probes are referred to as “rc-probes”. In some embodiments, the re-linearized probes have different 5’ and 3’ ends when compared to the pre-circularized probe. Re-linearizing rc-probes enables the rc-probes to be seeded onto a flow cell for amplification and sequencing.

[0047] Fig. 1A and 1B are non-limiting examples of the rc-probes. In general, the rc-probes, 10 and 10a include one or more of the following regions: a sensor region 15 and 15’, a re-linearization region 20, one or more primer regions 30 and 30’, and a coderegion 40. In general, the sensor region is used to circularize the probe and the code region uniquely identifies each probe.

[0048] Figs. 2A-2E are non-limiting examples of the circularization and relinearization of the rc-probes described herein.

[0049] Fig. 2A illustrates rc-probe 10a hybridized to a target nucleic acid 55. Sensor regions 15 and 15’ hybridize to the target nucleic acid 55 and the rest of the rc-probe, e.g., code region 40 and primer regions 30 and 30’, does not hybridize to the target nucleic acid 55. In this form, enzymes or click chemistry can be used to join ends 15 and 15’ together at junction 50. Non-limiting enzymes include polymerases, reverse transcriptases, ligases, and the like. In some embodiments, 15 and 15’ can hybridize next to each other along target 55. In these embodiments, ligase or click chemistry can be used to join the ends at junction 50. In other embodiments, there can be a gap between ends (not illustrated) when 15 and 15’ are hybridized to target 55. In these embodiments, the gap can be 1 or more bases, 2 or more bases, 3 or more bases, 4 or more bases, 5 or more bases, 6 or more bases, 7 or more bases, 8 or more bases, 9 or more bases, or 10 or more bases. In these embodiments, polymerases or polymerases combined with ligases can be used to join the ends at junction 50.

[0050] Fig. 2B illustrates a circularized form 10b of the rc-probes.

[0051] Fig. 2C illustrates a re-linearized form 10c of the rc-probes. The rc-probe is linearized using the linearization region 20. Linearization can be accomplished using enzymatic or chemical means.

[0052] Fig. 2D illustrates a double stranded embodiment of the circularized rc-probes 10d. In this embodiment, second strand 60 can be generated after circularization of the rc-probe. In this embodiment, an intercalating dye can be used to visualize the generation of the circularized rc-probes. In this embodiment, the inner strand (solid line) can be selectively linearized.

[0053] Fig. 2E illustrates a partially double stranded rc-probe 10e. In this embodiment, all or a portion of the non-sensor regions of the rc-probe can be double stranded. In this embodiment, oligo 65 (dotted line) can be used to make the rc-probe partially double stranded. Although illustrated as one oligo, oligo 65 can also be several shorter oligos that completely (end to end) or partially (gaps between the ends) span the desiredregion of rc-probe 10e. In this embodiment, the double stranded region can be extended to generate the double stranded circularized rc-probe shown in Fig. 2D. This embodiment may prevent the probe from randomly sticking or hybridizing to the wrong target.

[0054] In some embodiments, the sensor region is designed to hybridize both upstream and downstream of a genetic marker 60. Non-limiting examples of genetic markers include single polymorphism nucleotides (SNPs), restriction fragment length polymorphisms (RFLPs), variable number of tandem repeats (VNTRs), microsatellites, and copy number variants (CNVs). Upon hybridizing both sensor regions of the rc-probe to the target nucleic acid, a polymerase is used to extend and ligate 50 the ends of the probe to form the circularized probe 10b. The re-linearization region 20 is used to open the circularized rc-probe into a second linear form 10c.

[0055] The re-linearization region of the probe can be located either 5’ or 3’ of the code region as shown in Fig. 1A and 1B. The rc-probes can be linearized using various methods. Non-limiting examples of linearization methods are shown in Table 1, below. When using double stranded restriction enzymes, a separate oligonucleotide can be hybridized to the recognition sequence portion of the probe to create a double stranded recognition sequence.Table 1

[0056] The primer region can include one or more amplification primers and / or one or more sequencing primers.

[0057] In some embodiments the primer region includes two amplification regions located 5’ and 3’ of the re-linearization region 20. The amplification regions can be used to amplify the rc-probe in solution or on a flow cell. In some embodiments, the amplification regions are compatible with P5 and P7, shown in Table 2, below.

[0058] In some embodiments, the primer region includes at least one sequencing region. In one embodiment, the at least one sequencing region can be the only primer region (no illustrated). In these embodiments, the sequencing region can be located anywhere in the probe and is generally located such to enable sequencing of the code region. In another embodiment, the at least one sequencing region can be in addition to the two amplification regions located 5’ and 3’ of the re-linearization region 20.Generally, the sequencing regions can be used to sequence the entire or a portion of the rc-probe using known methods, including (but not limited to) sequence by synthesis (SBS). In some embodiments, the sequencing regions are compatible with A14 and B15, shown in Table 2, below.

[0059] In some embodiments, the primer region includes two amplification regions and two sequencing. In these embodiments, the amplification regions are outer to the sequencing regions.Table 2

[0060] The code region of the rc-probe can be used to identify the probe. The code region can be one single code region or can include one or more subcode regions. Non-limiting examples of subcode regions include barcodes, indexes, and / or UMIs. The one or more subcode regions can be in any order. The subcode regions can be used to identify various aspects of sequencing. For example, barcodes can be used to identify the sensor region, indexes can be used to identify the sample, and / or UMI's can be used to identify individual molecules. In general, the code region can be 5 to 100 bases in length or 10 to 50 bases in length or 10-20 bases in length.Flow Cells

[0061] The rc-probes described herein can be attached to the surface of a flow cell. In some embodiments, the rc-probes are attached after circularization. Commonly known binding protein conjugates can be used to attach the rc-probes to the surface of the flow cell. Non-limiting examples include biotin and streptavidin. Alternatively, the rc-probes can be attached by hybridizing to a capture probe that is complementary to a portion of the rc-probe. In yet another embodiment, the rc-probe can be attached via a mechanism introduced during the circularization step. For example, incorporation of a biotin labeled nucleotide during the circularization step.Methods

[0062] The method disclosed herein includes a workflow for preparing a target nucleic acid for sequencing. The method combines enrichment and library prep into one method.

[0063] Nucleic acid is extracted from a sample and denatured. Non-limiting examples of single stranded nucleic acid include genomic DNA, bisulfite converted DNA, RNA, aptamers, or oligos attached to monoclonal antibodies.

[0064] The extracted single stranded nucleic acid is hybridized to a panel of rc-probes. The hybridized rc-probes are circularized. Non-hybridized rc-probes are removed. Anon-limiting method of removing non-hybridized rc-probes includes digestion with enzymes.

[0065] In an alternate embodiment, the circularized rc-probes are copied, forming double stranded circularized rc-probes as illustrated in

[0066] The circularized rc-probes are re-linearized by using the appropriate method for the re-linearization region.

[0067] The re-linearized rc-probes can be seeded onto a flow cell for clustering and sequencing.

[0068] Embodiments of the disclosure are further described by the following numbered paragraphs:

[0069] 1. A probe comprising: a sensor region; and a linearization region; wherein the sensor region is used to circularize the probe; and wherein the linearization region is used to re-linearize the circularized probe.

[0070] 2. The probe of numbered paragraph 1, wherein the re-linearized probe has different 5’ and 3 ends compared to the pre-circularized probe.

[0071] 3. The probe of numbered paragraph 1, wherein the linearization region comprises a restriction recognition site, an 8 oxoguanine, UdT, Allyl-T, or diol mods.

[0072] 4. The probe of any one of the previous numbered paragraphs, further comprising at least one code region.

[0073] 5. The probe of numbered paragraph 4, wherein the code region comprises one or more of a barcode, and index, or a unique molecular identifier.

[0074] 6. The probe of any one of the previous numbered paragraphs, further comprising at least one primer region.

[0075] 7. The probe of numbered paragraph 6, wherein the at least one primer region comprises at least one amplification region.

[0076] 8. The probe of numbered paragraph 7, wherein the amplification region is compatible with P5 and / or P7.

[0077] 9. The probe of any one of numbered paragraphs 6 to 8, wherein the amplification is done in solution.

[0078] 10. The probe of any one of numbered paragraphs 6 to 9, wherein the amplification is done on a flow cell.

[0079] 11. The probe any one of the previous numbered paragraphs, wherein the at least one primer region comprises at least one sequencing region.

[0080] 12. The probe of numbered paragraph 11, wherein the sequencing method is sequence by synthesis.

[0081] 13. The probe of numbered paragraph 11, wherein the sequencing method is nanopore based.

[0082] 14. The probe of numbered paragraph 11, wherein the sequencing primer region is compatible with A14 and / or B15.

[0083] 14A. The probe of any one of the previous numbered paragraphs, wherein the probe is attached to a solid surface is a bead, particle, or flow cell.

[0084] 14B. The probe of numbered paragraph 1 A, wherein the solid surface is a bead, particle, or

[0085] 15. A flow cell comprising the probes of numbered paragraphs 1 to 14.

[0086] 16. The flow cell of numbered paragraph 15, wherein the probes are attached in a linear or re-linearized form.

[0087] 17. The flow cell of numbered paragraph 15, wherein the probes are attached in a circular form.

[0088] 18. A method for preparing a sample for sequencing comprising: hybridizing the probe of anyone of numbered paragraphs 1 to 13 to nucleic acid from a sample; circularizing the probe by ligating the ends of the hybridized probe; separating the circular probes from non-circularized probes; and re-linearizing the circularized probes.

[0089] 19. The method of numbered paragraph 18, wherein the sample nucleic acid comprises genomic DNA, bisulfite converted DNA, RNA, aptamers, or oligos attached to monoclonal antibodies.

[0090] 20. The method of numbered paragraphs 18 or 19, wherein ligating comprises extending and ligating.

[0091] 21. The method of numbered paragraph 20, wherein the extending and ligating uses biotinylated nucleotides.

[0092] 22. The method of any one of numbered paragraphs 18 to 21, wherein the separating comprises capturing circularized probes on a flow cell surface.

[0093] 23. The method of any one of numbered paragraphs 18 to 21, wherein the separating comprises digesting non-circularized probes.

[0094] 24. The method of any one of numbered paragraphs 18 to 23, further comprises amplifying the re-linearized probes.

[0095] 25. The method of any one of numbered paragraphs 18 to 24, further comprising sequencing the re-linearized probes.Additional Notes

[0096] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. It should also be appreciated that terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.

[0097] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.

[0098] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.

Claims

CLAIMS1. A probe comprising:a sensor region; anda linearization region;wherein the sensor region is used to circularize the probe; andwherein the linearization region is used to re-linearize the circularized probe.

2. The probe of claim 1, wherein the re-linearized probe has different 5’ and 3’ ends compared to the pre-circularized probe.

3. The probe of claim 1, wherein the linearization region comprises a restriction recognition site, an 8 oxoguanine, UdT, Allyl-T, or diol mods.

4. The probe of any one of the previous claims, further comprising at least one code region.

5. The probe of claim 4, wherein the code region comprises one or more of a barcode, and index, or a unique molecular identifier.

6. The probe of any one of the previous claims, further comprising at least one primer region.

7. The probe of claim 6, wherein the at least one primer region comprises at least one amplification region.

8. The probe of claim 7, wherein the amplification region is compatible with P5 and / or P7.

9. The probe of any one of claims 6 to 8, wherein the amplification is done in solution.

10. The probe of any one of claims 6 to 9, wherein the amplification is done on a flow cell.

11. The probe any one of the previous claims, wherein the at least one primer region comprises at least one sequencing region.

12. The probe of claim 11, wherein the sequencing method is sequence by synthesis.

13. The probe of claim 11, wherein the sequencing method is nanopore based.

14. The probe of claim 11, wherein the sequencing primer region is compatible with A14 and / or B15.

15. A flow cell comprising the probes of claims 1 to 14.

16. The flow cell of claim 15, wherein the probes are attached in a linear or relinearized form.

17. The flow cell of claim 15, wherein the probes are attached in a circular form.

18. A method for preparing a sample for sequencing comprising:hybridizing the probe of anyone of claims 1 to 13 to nucleic acid from a sample; circularizing the probe by ligating the ends of the hybridized probe; separating the circular probes from non-circularized probes; andre-linearizing the circularized probes.

19. The method of claim 18, wherein the sample nucleic acid comprises genomic DNA, bisulfite converted DNA, RNA, aptamers, oroligos attached to monoclonal antibodies.

20. The method of claim 18 or 19, wherein ligating comprises extending and ligating.

21. The method of claim 20, wherein the extending and ligating uses biotinylated nucleotides.

22. The method of any one of claims 18 to 21, wherein the separating comprises capturing circularized probes on a flow cell surface.

23. The method of any one of claims 18 to 21, wherein the separating comprises digesting non-circularized probes.

24. The method of any one of claims 18 to 23, further comprises amplifying the relinearized probes.

25. The method of any one of claims 18 to 24, further comprising sequencing the relinearized probes.

Citation Information

Patent Citations

  • Polymerase enzymes and reagents for enhanced nucleic acid sequencing

    US20080108082A1

  • Methods and apparatus for measuring analytes

    US20100137143A1

  • Modified Molecular Arrays

    US20110059865A1

  • Bioconjugation of macromolecules

    US6737236B1

  • Zero-mode clad waveguides for performing spectroscopy with confined effective observation volumes

    US6917726B2