High-throughput DNA methylation beadchip assay
Patent Information
- Application Number
- EP2024886910
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-10-31
- Publication Date
- 2026-09-09
AI Technical Summary
Current DNA methylation analysis methods, such as microarrays and next-generation sequencing, face challenges in accurately measuring subtle differences in DNA methylation patterns, which are often associated with complex diseases, requiring larger sample sizes and more comprehensive genomic coverage.
The development of a high-throughput DNA methylation BeadChip assay featuring an array of polynucleotide probes with diverse nucleotide sequences, strategically designed to detect methylation levels at multiple genomic loci, including CpG islands and adjacent regions, using bisulfite conversion and advanced probe designs.
This approach enables highly accurate and cost-effective quantification of DNA methylation, facilitating population-scale epigenome-wide association studies (EWAS) with reduced sample sizes, while maintaining exploratory power and detecting known trait associations.
Smart Images

Figure IMGF000014_0001 
Figure IMGF000015_0001 
Figure IMGF000016_0001
Description
HIGH-THROUGHPUT DNA METHYLATION BEADCHIP ASSAYCROSS-REFERNCE TO RELATED APPLICATIONS
[0001] This application claims priority of US Provisional application number 63 / 594,868, filed October 31, 2023, and US Provisional application number 63 / 596,091, filed November 3, 2023, the entire contents of each being incorporated herein by reference as though set forth in full.REFERENCE TO SEQUENCE LISTING
[0002] The present application incorporates by reference a Sequence Listing filed in U.S. Prov. App. No. 63 / 594,868 filed October 31, 2023. The Sequence Listing is filed in electronic format. The Sequence Listing is provided as a file entitled SEQLIST_ILLINC8O8PR, created October 31, 2023, which is approximately 349,426,430 bytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.FIELD OF THE INVENTION
[0003] Some embodiments of the methods and compositions provided herein relate to measuring methylation levels in DNA. Some embodiments include arrays of polynucleotide probes for detecting methylated CpGs at certain genomic loci; and methods, kits and systems comprising such arrays.BACKGROUND OF THE INVENTION
[0004] DNA methylation is one of the most studied epigenetic modifications in human cells. Changes in DNA methylation patterns play a critical role in development, differentiation and diseases such as multiple sclerosis, diabetes, schizophrenia, aging, and multiple forms of cancer. Over the past decade, interest in DNA methylation has grown rapidly and expanded across new areas of research. Consequently, DNA methylation analysis methods have undergone dramatic changes. Many microarray and next-generation sequencing based technologies have emerged, and analyses that were previously restricted to specific loci in a limited number of genes can now be performed on a genome-wide scale. Accordingly, thereis a need for improved methods and compositions for determining the methylation status of DNA.SUMMARY OF THE INVENTION
[0005] Some embodiments of the methods and compositions provided herein include an array comprising a plurality of polynucleotide probes attached to a surface of a substrate, wherein the polynucleotide probes each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386.
[0006] In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In some embodiments, the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another.
[0007] In some embodiments, the substrate comprises a plurality of beads. In some embodiments, the substrate comprises a flow cell. In some embodiments, the plurality of polynucleotides are biotinylated.
[0008] Some embodiments of the methods and compositions provided herein include a method for measuring a methylation level for a plurality of loci of target nucleic acids, comprising: (a) obtaining a sample comprising target nucleic acids; (b) converting conversion-sensitive cytosine residues of the target nucleic acids to another base residue to obtain converted nucleic acids; (c) hybridizing the converted nucleic acids to any one of the arrays provided herein ; and (d) detecting the presence or absence of methylated loci in the converted nucleic acids in the array.
[0009] In some embodiments, the converting comprises bisulfite conversion.
[0010] Some embodiments also include amplifying the converted nucleic acids. In some embodiments, the amplifying is performed before the hybridizing. In some embodiments, the amplifying comprises adding indexes to the converted nucleic acids.
[0011] In some embodiments, the detecting comprises extending the polynucleotide probes. In some embodiments, the detecting comprises sequencing the converted nucleic acids. Some embodiments also include aligning sequences of the convertednucleic acids with a reference sequence. Some embodiments also include mapping a methylated cytosine residue on a sequence of a converted nucleic acid on a reference sequence.
[0012] In some embodiments, the sample comprises cell free DNA, genomic DNA. In some embodiments, the sample comprises a formalin fixed paraffin embedded (FFPE) sample.
[0013] Some embodiments of the methods and compositions provided herein include a system comprising: any one of the arrays provided herein; and a detector.
[0014] Some embodiments of the methods and compositions provided herein include a kit comprising a plurality of polynucleotides, wherein the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386. In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In some embodiments, the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another. In some embodiments, the plurality of polynucleotides are biotinylated.
[0015] Some embodiments of the methods and compositions provided herein include a method of making an array, comprising: (a) providing a substrate; and (b) attaching a plurality of polynucleotides to a surface of the substrate, wherein the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386. In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In some embodiments, the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another. In some embodiments, the substrate comprises a plurality of beads. In some embodiments, the substrate comprises a flow cell. In some embodiments, the plurality of polynucleotides are biotinylated.BRTEF DESCRIPTION OF THE DRAWINGS
[0016] FIG. 1A depicts an Infinium I methylation assay (Illumina, Inc., San Diego CA) scheme in which two bead types correspond to each CpG locus: one bead type — to methylated (C), another bead type — to unmethylated (T) state of the CpG site. Probe design assumes same methylation status for adjacent CpG sites. Both bead types for the same CpG locus can incorporate the same type of labeled nucleotide, determined by the base preceding the interrogated “C” in the CpG locus, and therefore may be detected in the same color channel in certain embodiments.
[0017] FIG. IB depicts an Infinium II methylation assay (Illumina, Inc., San Diego CA) scheme in which one bead type corresponds to each CpG locus. Probe can contain up to 3 underlying CpG sites, with degenerate R base corresponding to C in the CpG position. Methylation state can be detected by single-base extension. Each locus may be detected in two colors in certain embodiments. In some embodiments of an Infinium II methylation assay design, labeled “A” may be incorporated at unmethylated query site (“T”), and “G” may be incorporated at methylated query site (“C”).
[0018] FIG. 2A depicts gene regions that may be targeted by probes for methylation sites and includes functional regions: TSS200 is the region from transcription start site (TSS) to 200 nt upstream of the TSS; TSS 1500 covers -200 to -1500 nt upstream of the TSS; 5' UTR, 1st exon, gene body and 3' UTR.
[0019] FIG. 2B depicts CpG islands and adjacent regions that may be targeted by probes for methylation sites. CpG islands longer than 500 bp were divided into separate bins. The 2 kb regions immediately upstream and downstream of the CpG island boundaries, or “CpG island shores”, and the 2 kb regions upstream and downstream of the CpG island shores, referred to here as “CpG island shelves,” were also targeted separately.
[0020] FIG. 3 depicts an embodiment for a workflow (WF) for measuring and / or detecting methylation in target nucleic acids.
[0021] FIG. 4 depicts an embodiment for a workflow (WF) for measuring and / or detecting methylation in target nucleic acids.
[0022] FIG. 5 depicts an embodiment for a design process and probe summary including screening and probe selection for a methylation screening array (MSA) in which probes were selected for designability and association with diverse target biology groups.
[0023] FIG. 6A depicts a Venn diagram for genomic coordinate overlap of probes between the MSA and EPICv2.
[0024] FIG. 6B depicts a bar graph fro number of probes in the MSA and EPICv2 for SNP and CpH probes.
[0025] FIG. 7 depicts target biology groups and a comparison with EPICv2 with graphs of probe count and odds rations. Enrichment of MSA probes in EWAS traits. MSA was enriched in methylation EWAS hits. Enrichment was tested for MSA, EPICv2 and a random selection of equal size to MSA against collated EWAS hits from multiple databases. Across diverse trait groups, MSA was highly enriched reflecting the targeted selection of known trait associated CpGs.
[0026] FIG. 8 depicts cell specific marker types and probe counts represented on MSA with graphs for loglO CpG counts for ‘pan tissue’, ‘brain cells’, and ‘immune cells’. Darker portions of the bars depict hypomethylation; lighter portions of the bars depict hypermethylation. Single cell and sorted WGBS data were analyzed to identify cell specific CpG markers. Thousands of signatures were identified delineating hundreds of cell types and cell type comparisons. When possible, a balanced selection of hypo and hypermethylated markers were selected. The novel cell-specific probe designs may enhance cell deconvolution and enable investigations into selectively vulnerable populations across diverse diseases.
[0027] FIG. 9A and FIG. 9B depict cell signature and full stack ChromHMM comparison between MSA and EPICv2. FIG. 9A is a graph depicts levels of depletion and enrichment for certain probe types in various platforms including the MSA and EPICv2. FIG. 9B is a graph for number of CpGs / cell type contrast. MSA had greater cell type marker coverage per cell type comparison branch compared to EPICv2 despite the lower overall probe count. Both platforms were highly depleted from heterochromatic, quiescent and alignment gap states. MSA was more enriched than EPICv2 in cell specific enhancer elements and some bivalent I downstream promoters.
[0028] FIG. 10A and FIG 10B depict bar graphs for cis-regulatory element coverage. ENCODE cis regulatory elements were included in the target biology design groups. MSA had higher enhancer coverage with a decrease in promoter and CpG island representation compared to EPICv2.
[0029] FIG. 11 depicts a bar graphs for additional target biology enrichments. MSA was highly enriched in several additional methylation features. These included: CpGs with methylation correlated with gene expression, CoRSIVs, monoallelic methylation I imprinting loci, and CpGs shown to be hydroxy methylated at the pan tissue and tissue specific level.
[0030] FIG. 12 depicts a graph for MSA probe detection rate. MSA was generating high quality data, with greater than 90% probe detection rate on non- FFPE processed samples.
[0031] FIG. 13 depicts an MSA correlation rate with methylation titrations. Methylation readings at most probes showed a high correlation with DNA methylation titrations of 0, 50, and 100%.DETAILED DESCRIPTION
[0032] Some embodiments of the methods and compositions provided herein relate to measuring methylation levels in DNA. Some embodiments include arrays of polynucleotide probes for detecting methylated CpGs at certain genomic loci; and methods, kits and systems comprising such arrays.
[0033] DNA methylation is a stable covalent chemical modification to DNA that shapes cell identity and impacts diverse processes from genome stability to gene expression. Infinium Methylation BeadChips (Illumina, Inc. San Diego, CA) contain probes for preselected CpG sites and can be used for highly accurate quantification of CpG methylation at selected loci. Such microarrays are a cost effective and computationally tractable alternative and have been successfully utilized for population scale epigenome- wide association study (EWAS) studies.
[0034] Despite the successes of previous BeadChip versions, some traits associate with only subtle differences in DNA methylation, necessitating ever larger sample sizes. Therefore, a highly consolidated BeadChip strategically designed for EWAS studies would be beneficial for the population epigenetics community. As provided herein, a compact Infinium Methylation BeadChip was designed containing -250K probes, approximately 30% of the size of the newly released EPICv2 array (Illumina, Inc. San Diego, CA) which includes more than about 935k probes.
[0035] The methylation status of nucleic acids is important information that is useful in many biological assays and studies. Very often, it is of particular interest to identifypattems of methylation at specific regions in the genome. Also, it is often of particular interest to identify the methylation status of specific CpG dinuclcotidcs. The methylation level and pattern of a locus in a nucleic acid sample can be determined using any of a variety of methods described herein and capable of distinguishing presence or absence of a methyl group on a nucleotide base of the nucleic acid. In the case of DNA, methylation, when present, typically occurs as 5-methylcytosine (5-mCyt) in CpG dinucleotides. Methylation of CpG dinucleotide sequences or other methylated motifs in DNA can be measured using any of a variety of techniques used in the ail for the analysis of specific CpG dinucleotide methylation status.
[0036] A commonly-used method of determining the methylation level and / or pattern of DNA requires methylation status-dependent conversion of cytosine in order to distinguish between methylated and non-methylated CpG dinucleotide sequences. For example, methylation of CpG dinucleotide sequences can be measured by employing cytosine conversion-based technologies, which rely on methylation status-dependent chemical modification of CpG sequences within isolated genomic DNA, or fragments thereof, followed by DNA sequence analysis. Chemical reagents that are able to distinguish between methylated and non-methylated CpG dinucleotide sequences include hydrazine, which cleaves the nucleic acid, and bisulfite treatment. Bisulfite treatment followed by alkaline hydrolysis specifically converts non-methylated cytosine to uracil, leaving 5-methylcytosine unmodified as described by Olek A., Nucleic Acids Res. 24:5064-6, 1996 or Frommer et al., Proc. Natl. Acad. Sci. USA 89:1827-1831 (1992), each of which is incorporated herein by reference in its entirety. The bisulfite-treated DNA can subsequently be analyzed by conventional molecular techniques, such as PCR amplification, sequencing, and detection comprising oligonucleotide hybridization.
[0037] Some embodiments of the invention include to use of bisulfite conversion conditions and bisulfite-resistant cytosine analogs. One consequence of bisulfite-mediated deamination of cytosine is that the bisulfite treated cytosine is converted to uracil, which reduces the complexity of the genome. Specifically, a typical 4-base genome (A,T,C,G) is essentially reduced to a 3-base genome (A,T,G) because uracil is read as thymine during downstream analysis techniques such as PCR and sequencing reactions. Thus, the only cytosines present are those that were methylated prior to bisulfite conversion.
[0038] Certain aspects disclosed in the following references are useful with some embodiments of the methods and compositions provided herein: Liu, L., et al., (2018) Annals of Oncology 29:1445-1453; Bibikova, M. et al., (2011) Genomics 98:288-295; U.S. Pat. No. 7,899,626; U.S. Pat. No. 8,541,207; and WO 2023 / 028478, which are each incorporated by reference in its entirety.Definitions
[0039] As used herein, reference to determining the methylation status and like terms refers to at least one or more of the following: 1) determining the level or amount of cytosine methylation in a sample, 2) determining the position of methylated cytosine residues within a sequence, 3) determining the pattern of methylated cytosine in a sequence, and / or 4) determining the whole sequence including the specific position and identity of methylated residues in the context of the sequence.
[0040] As used herein, "converted," when used in reference to a nucleic acid or portion thereof, refers to nucleic acid or a portion thereof which has been treated under conditions sufficient to convert cytosine to another base. As used herein, "bisulfite-converted", "bisulfite-treated" and like terms, when used in reference to a nucleic acid or portion thereof, refer to nucleic acid or a portion thereof which has been treated with sodium bisulfite under conditions sufficient to convert cytosine to uracil. Thus, for example, in some embodiments, template nucleic acid will have at least one cytosine residue that is not methylated and which is converted to uracil by bisulfite treatment. However, the template nucleic acid need not comprise a non-methylated cytosine, either because all cytosines are methylated or because no cytosine residues are present in the template nucleic acid.
[0041] As used herein, "non-converted," when used in reference to a nucleic acid or portion thereof, refers to a nucleic acid or portion thereof where one or more of the cytosines, if present, are not converted to another base, such as uracil, after conversion treatment, such as treatment with sodium bisulfite. Thus, for example, a non-converted complementary copy is a nucleic acid that comprises one or more bisulfite -resistant cytosine analogs that prevent the conversion of cytosine to uracil.
[0042] Substrates of the present disclosure that contain nucleic acid arrays can be used for any of a variety of purposes. A particularly desirable use for the nucleic acids is toserve as capture probes that hybridize to target nucleic acids having complementary sequences. The target nucleic acids once hybridized to the capture probes can be detected, for example, via a label recruited to the capture probe. Methods for detection of target nucleic acids via hybridization to capture probes are known in the art and include, for example, those described in U.S. Pat. Nos.7,582,420; 6,890,741; 6,913,884 or 6,355,431 or U.S. Pat. Pub. Nos. 2005 / 0053980 Al; 2009 / 0186349 Al or 2005 / 0181440 Al, each of which is incorporated herein by reference.
[0043] Nucleic acid sequencing can be used to determine a nucleotide sequence of a polynucleotide by various processes known in the art. In a preferred method, sequencing - by-synthesis (SBS) is utilized to determine a nucleotide sequence of a polynucleotide attached to a surface of a substrate (e.g., via any one of the polymer coatings described herein). In such a process, one or more nucleotides are provided to a template polynucleotide that is associated with a polynucleotide polymerase. The polynucleotide polymerase incorporates the one or more nucleotides into a newly synthesized nucleic acid strand that is complementary to the polynucleotide template. The synthesis is initiated from an oligonucleotide primer that is complementary to a portion of the template polynucleotide or to a portion of a universal or non-variable nucleic acid that is covalently bound at one end of the template polynucleotide. As nucleotides are incorporated against the template polynucleotide, a detectable signal is generated that allows for the determination of which nucleotide has been incorporated during each step of the sequencing process. In this way, the sequence of a nucleic acid complementary to at least a portion of the template polynucleotide can be generated, thereby permitting determination of the nucleotide sequence of at least a portion of the template polynucleotide.
[0044] Flow cells provide a convenient format for housing an array that is produced by the methods of the present disclosure and that is subjected to a sequencing-by-synthesis (SBS) or other detection technique that involves repeated delivery of reagents in cycles. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc., can be flowed into / through a flow cell that houses a nucleic acid array made by methods set forth herein. Those sites of an array where primer extension causes a labeled nucleotide to be incorporated can be detected. Optionally, the nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moiety can beadded to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference in its entirety.Certain methylation assays
[0045] Bead microarrays, such as Infinium Methylation BeadChips (Illumina, Inc. San Diego, CA), enable highly accurate and quantitative assessment of DNA methylation at the single-CpG-site level. Infinium Methylation BeadChips apply the Illumina Infinium assay to detect epigenetic modifications, delivering broad coverage and high-throughput capabilities for large-scale epigenome- wide association studies (EWAS). By combining Infinium I and Infinium II probe chemistries, Infinium Methylation BeadChips can provide expanded coverage of an expert-defined selection of CpG targets.
[0046] The Infinium methylation assay uses beads displaying target-specific probes designed to interrogate individual CpG sites within a DNA sample. Infinium I and Infinium II chemistries differ in the number of probes needed to query a single CpG locus. The Infinium I assay uses two probes per CpG, whereas the Infinium II assay requires one probe per CpG locus due to its ability to measure both unmethylated and methylated DNA states. As a result, Infinium I and II assays offer complementary strengths that enhance the breadth of coverage of the array.
[0047] The Infinium I assay employs two probes per CpG locus; one probe for unmethylated and one probe for methylated DNA states (FIG. 1A). The 3' terminus of each probe is designed to match either the protected cytosine, which indicates methylation, or the thymine base resulting from bisulfite conversion and whole-genome amplification of unmethylated cytosine. Probe designs for Infinium I assays are based on the assumption thatmethylation is regionally correlated within a 50 bp span. Thus, underlying CpG sites are treated as in phase with the methylated (C) or unmcthylatcd (T) query sites. This comcthylation theory is supported in a study in which bisulfite sequencing of chromosomes 6, 20, and 22 showed that over 90% of CpG sites within 50 bases had the same methylation status (Eckhardt F, et al. Nat Genet. 2006;38:1378-1385). A second study showed that, in general, methylation status at adjacent sites tends to be correlated, suggesting that correlation depends upon the cell types or nearby polymorphic sites (Shoemaker R, et al., Genome Res. 2010;20:883-889).
[0048] The Infinium II assay design employs only one probe per locus (FIG. IB). The 3' terminus of the probe complements the base directly upstream of the query site. A single base extension results in the addition of a labeled G or A base, complementary to either the methylated C or the unmethylated T bases. For the opposite strand, a labeled C or T base would be added.
[0049] In some embodiments, a single 50-mer probe can be used to determine methylation state, making an all-or-none approach inapplicable. In some embodiments, underlying CpG sites may be represented by degenerate R-bases or degenerate Y -bases for the opposite strand. Infinium II probes can have up to three underlying CpG sites within the 50- mer probe sequence without compromising data quality. This feature enables the methylation status at a query site to be assessed independently of assumptions on the status of neighboring CpG sites. In some embodiments, the use of only a single bead type enables increased capacity for the number of CpG sites that can be queried.Certain compositions
[0050] Some embodiments of the methods and compositions provided herein include arrays for measuring and / or detecting methylation in target nucleic acids. Some such arrays include a plurality of polynucleotide probes attached to a surface of a substrate. In some embodiments, the substrate comprises a plurality of distinct sites at which the polynucleotides are attached. In some embodiments, the polynucleotide probes each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386. In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In someembodiments, the polynucleotide probes comprise at least 500, 1 ,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another.
[0051] In some embodiments, the substrate comprises a plurality of beads. In some embodiments, the substrate comprises a flow cell. In some embodiments, the plurality of polynucleotides are biotinylated. In some such embodiments, the plurality of polynucleotides can be attached to the substrate via a biotin - streptavidin binding pair.
[0052] Some embodiments of the methods and compositions provided herein include kits for measuring and / or detecting methylation in target nucleic acids. Some such kits include a plurality of polynucleotides, wherein the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386. In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In some embodiments, the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another. In some embodiments, the plurality of polynucleotides are biotinylated.
[0053] Some embodiments of the methods and compositions provided herein include systems for measuring and / or detecting methylation in target nucleic acids. Some such systems include any one of the arrays provided herein, and a detector. Some such detectors can measure a signal from a nucleic acid, such as a converted nucleic acid, hybridized to a polynucleotide of an array.Certain methods for measuring and / or detecting methylation in target nucleic acids
[0054] Some embodiments of the methods and compositions provided herein include methods for measuring and / or detecting methylation in target nucleic acids. Some such methods include measuring a methylation level for a plurality of loci of target nucleic acids. Some embodiments include (a) obtaining a sample comprising target nucleic acids; (b) converting conversion-sensitive cytosine residues of the target nucleic acids to another base residue to obtain converted nucleic acids; (c) hybridizing the converted nucleic acids to any one of the arrays provided herein; and (d) detecting the presence or absence of methylated lociin the converted nucleic acids in the array. In some embodiments, the converting comprises bisulfite conversion. Some embodiments also include amplifying the converted nucleic acids. In some embodiments, the amplifying is performed before the hybridizing. In some embodiments, the amplifying comprises adding indexes to the converted nucleic acids. In some embodiments, the detecting comprises extending the polynucleotide probes. In some embodiments, the detecting comprises sequencing the converted nucleic acids. Some embodiments also include aligning sequences of the converted nucleic acids with a reference sequence. Some embodiments also include mapping a methylated cytosine residue on a sequence of a converted nucleic acid on a reference sequence. In some embodiments, the sample comprises cell free DNA, genomic DNA. In some embodiments, the sample comprises a formalin fixed paraffin embedded (FFPE) sample.
[0055] An embodiment for a workflow for measuring and / or detecting methylation in target nucleic acids is depicted in FIG. 3. FIG. 3 depicts a semi- automated workflow including bisulfite conversion of target nucleic acids in a DNA source, such as cell line, blood, or formalin fixed paraffin embedded (FFPE); target preparation; and a beadchip assay. DNA is subjected to bisulfite conversion, then amplified by whole genome amplification (WGA). Amplified DNA is fragmented, precipitated, resuspended, then hybridized to beadchips ‘EX48 MSA’ or ‘EX24’. The hybridized nucleic acids are stained and the beadchip is imaged and the data analyzed.
[0056] Another embodiment for a workflow for measuring and / or detecting methylation in target nucleic acids is depicted in FIG. 4. The workflow depicts in FIG. 4 is substantially the same as the workflow depicted in FIG. 3, and also includes certain reagents used in certain steps which are listed in TABLE 1.TABLE 1Certain methods for making an array
[0057] Some embodiments of the methods and compositions provided herein include methods for measuring and / or detecting methylation in target nucleic acids. Some such methods include (a) providing a substrate; and (b) attaching a plurality of polynucleotides to a surface of the substrate. In some embodiments, the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01—382,386.
[0058] In some embodiments, the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386. In some embodiments, the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972. In some embodiments, the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another. In some embodiments, the substrate comprises a plurality of beads. In some embodiments, the substrate comprises a flow cell. In some embodiments, the plurality of polynucleotides are biotinylated.EXAMPLESComparative Example 1 — High density DNA methylation array with single CpG site resolution
[0059] An advantage of Infinium technology includes assay complexity is limited by the number of beads which are assembled on the slide section. The Infinium methylation array uses beads with long target- specific probes designed to interrogate individual CpG sites DNA methylation is measured using quantitative genotyping of bisulfite-converted genomic DNA. A genome-wide DNA methylation BeadChip array for high-throughput methylation profiling of the human genome was designed. The array included targets listed in TABLE 2. As shown in FIG. 2A and FIG. 2B, gene and CpG island regions were subdivided according to UCSC classifications and each subcategory was targeted individually (Fujita, P.A., et al., (2011) Nucleic Acids Res. 39: D876-D882; and Rhead, B. et al., (2010) Nucleic Acids Res. 38:D613-D619).TABLE 2
[0060] A correlation of methylation beta-values measured by the Infinium Methylation assay was correlated with results from whole genome bisulfite sequencing (WGBS) data generated on a HiSeq2000 (Illumina, Inc, San Diego, CA) using next-generation sequencing technology. Two comparisons were run, one with a normal lung tissue and the other with a lung tumor sample. WGBS data were filtered to include corresponding array loci covered with a minimum of 10 and maximum of 121 aligned reads, resulting in a total of 189,821 and 167,996 loci for comparison in the normal and tumor samples, respectively. The observed beta value correlations were 0.95 and 0.96 for the normal and tumor samples, respectively. These results indicate that the beta values generated by the array and whole genome bisulfite sequencing were consistent in reporting DNA methylation state across queried CpG loci.MethodsArray design
[0061] Probe performance assessment experiments were run to determine optimal probe design parameters for both Infinium I and II designs. Assay probes were selected with the goal of providing the most complete coverage possible across the identified content categories. Among these content categories, gene regions and CpG islands were given top priority. Each target region was allocated a maximum loci count which was inversely related to its level of priority (e.g. gene promoter regions and CpG islands were allotted a higher number of loci than other target regions). Large regions such as large CpG islands were subdivided into separate sub-regions to ensure even coverage. After each round of target selection and probe design, empirical analytical testing removed poorly -performing probes. Subsequent rounds of selection were then run until the pool size was exhausted. This approach ensured strong probe performance as well as an optimal balance between coverage density in the highest priority regions and breadth of coverage across remaining targeted regions.Bisulfite conversion of genomic DNA
[0062] DNA samples for Infinium Methylation assay were bisulfite converted using EZ DNA methylation kit (Cat. #D5001) from Zymo Research (CA, USA). 500 ng ofgDNA was denatured by addition of Zymo M-Dilution buffer (contains NaOH) and incubated for 15 min at 37 °C. CT-convcrsion reagent (bisulfitc-containing) was added to the denatured DNA and incubated for 16 h at 50 °C in a thermocycler and denatured every 60 min by heating to 95 °C for 30 s. DNA samples for the whole-genome bisulfite sequencing were bisulfite converted using EpiTect Bisulfite conversion kit (Cat. #59104) from QIAGEN (Valencia, CA) following manufacturer's recommendations with modifications.Infinium methylation assay
[0063] In brief, 4 pl of bisulfite-converted DNA (-150 ng) was used in the wholegenome amplification (WGA) reaction. After amplification, the DNA was fragmented enzymatically, precipitated and re-suspended in hybridization buffer. All subsequent steps were performed following the standard Infinium protocol (User Guide part #15019519 A). Fragmented DNA was dispensed onto the HumanMethylation450 Bead-Chips, and hybridization performed in hybridization oven for 20 h. After hybridization, the array was processed through a primer extension and an immunohistochemistry staining protocol to allow detection of a single-base extension reaction. Finally, BeadChips were coated and then imaged on an Illumina iScan. Methylation level of each CpG locus was calculated in GenomeStudio ® Methylation module as methylation beta-value (P=intensity of the Methylated allele (M) / (intensity of the Unmethylated allele (U)+ intensity of the Methylated allele (M)+100).Whole-genome bisulfite sequencing
[0064] For the whole-genome DNA methylation analysis at single nucleotide resolution, 2-5 pg of lung normal and lung tumor genomic DNA was fragmented using Covaris shearing. The fragmented DNA was end polished, and a single ‘A’ nucleotide was added to the 3' ends of the blunt fragments. The fragments were ligated with Illumina methylated forked adaptors, and 200-400 bp fragments were selected by gel electrophoresis and purified using a QIAquick Gel Extraction Kit (QIAGEN). Purified DNA fragments were treated with bisulfite using the EpiTect Bisulfite Kit (QIAGEN) for approximately 14 h to ensure maximal conversion rate. The bisulfite-treated DNA was enriched by 4 cycles of PCR with Pfu Turbo Cx DNA polymerase (Stratagene Products, Agilent, La Jolla, CA). The libraries weresequenced on Illumina HiSeq2000 sequencing instrument according to standard Illumina cluster generation and sequencing protocols with a 2x75 bp read length.Data analysis
[0065] Infinium methylation data was processed with Methylation Module of GenomeStudio software using HumanMethylation450 manifest vl.l. Whole genome bisulfite sequencing data was analyzed using pipeline developed at the Salk Institute. Briefly, raw sequencing data was processed using Illumina pipeline and FastQ output data was generated and aligned to the human genome (hgl9) using Bowtie alignment algorithm. Methylation status for each aligned site was calculated at a minimum of lOx coverage per site. All CpG sites with a p-value less than 0.01 on the 450k array were mapped to WGBS data on the same strand and coordinate. For each of these mapped sites a beta value was calculated by taking the number of methylated tags (C) and dividing by the sum of methylated (C) and unmethylated (T) tag counts. WGBS sites with less than lOx and greater than 120x coverage were removed. Scatter plots and r-squared statistics were then calculated by comparing array vs. sequencing beta values for the remaining matching sites.Comparative Example 2 — Methylation sequencing assay targeting 9223 CpG sites
[0066] A next-generation sequencing (NGS) targeted methylation sequencing assay to simultaneously measure the methylation status of 9223 CpG (50-C-phosphate-G-30) sites known to be hypermethylated in cancer was developed.
[0067] Pan-cancer methylation panel targets were particularly selected for hypermethylated CpG sites in tumor versus normal tissue based on The Cancer Genome Atlas (TCGA) data. A total of 10 888 CpG sites in 34major cancer types and subtypes were selected. Probe sequences for targeted CpG sites were selected from the Infinium HM450 array (Illumina, San Diego, CA). Probes were individually synthesized and 50-biotinylated at Illumina. Analysis of 20 normal plasma samples resulted in removal of 1235 (11%) sites with mean methylation levels >2% and 430 (3.9%) sites with coverage below the fifth percentile in at least 10 samples, resulting in a total of 9223 CpG sites on the panel. Quality control was imposed after sequencing by excluding any samples with zero coverage at more than 15% of panel sites.
[0068] Methylation microarray data from about 10,000 cancer samples from TCGA were collected for analysis. A set of cancer typc-spccific hypcrmcthylation sites was selected for each targeted cancer type, although it was observed that some cancer types, such as stomach and colorectal, had higher overall methylation signatures compared with others. Candidate sites were filtered by requiring low ( <20%) methylation levels in normal blood cells. Annotation of the targeted CpG sites showed that 77% are located in CpG islands and that they were evenly distributed across gene regions, except for the 3' UTR, which was underrepresented. It was also observed that the targeted CpG sites were associated with known regulatory regions such as promoters (27.5%), enhancers (18.2%), and cell-type-specific regions (16.4%) (UCSC hgl9).
[0069] Extracted cfDNA was treated with bisulfite reagents followed by whole genome library preparation to amplify the bisulfite-converted DNA and incorporate sequencing adaptors.
[0070] An algorithm was developed to integrate the pan cancer methylation sequencing data into a single methylation score to identify cancer samples. Briefly, z-scores were calculated for individual CpG sites; these sites were normalized by transforming to P- values; and a weighted sum of these P-values was calculated for the final sample-specific methylation score.
[0071] Methylation scores from plasma cfDNA samples from patients with advanced breast cancer, colorectal cancer, NSCLC and melanoma accurately classified the presence of cancer in 83.8% of cases with 100% specificity. In addition, methylation scores from plasma cfDNA accurately predicted cancer type in 78.9% of cases (breast cancer, 72.7%; colorectal cancer, 88.5%; NSCLC, 81.8%; melanoma, 55.6%). Methylation scores were highly predictive of tumor origin in colorectal cancer, but much less in melanoma.Example 3 — High-throughput DNA methylation BeadChip assay for population- scale EWAS studies
[0072] DNA methylation is a stable covalent chemical modification to DNA that shapes cell identity and impacts diverse processes from genome stability to gene expression. Infinium Methylation BeadChips (Illumina, Inc. San Diego, CA) contain probes for preselected CpG sites and can be used for highly accurate quantification of CpG methylationat selected loci. Such microarrays are a cost effective and computationally tractable alternative and have been successfully utilized for population scale cpigcnomc-widc association study (EWAS) studies.
[0073] Despite the successes of previous BeadChip versions, some traits associate with only subtle differences in DNA methylation, necessitating ever larger sample sizes. Therefore, a highly consolidated BeadChip strategically designed for EWAS studies would be beneficial for the population epigenetics community. To that end, a compact Infinium Methylation BeadChip was designed containing -250K probes, approximately 30% of the size of the newly released EPICv2 array (Illumina, Inc. San Diego, CA) which includes more than about 935k probes.
[0074] EWAS hits of high statistical significance were incorporated along with new CpGs to facilitate discovery and scientific rigor in EWAS. Approximately 50% of probes were derived from mining EWAS databases and literature. The methylation of these CpGs are linked to 8511 biological features from broad categories including cardiovascular, metabolic, neurodegenerative / psychiatric, autoimmune, genetic, environmental exposure and infection- related traits and diseases that were studied using existing Infinium array platforms (Illumina, Inc. San Diego, CA). The remaining 50% were novel probe designs obtained from integrative analysis of >300 public bulk and single-cell whole genome bisulfite sequencing (WGBS) datasets. These probes targeted DNA methylation associated with cell type, gene expression, chromatin accessibility and mono-allelic expression. Additional cytosines were selected from recent regulatory region annotations throughout the human genome to further enhance discovery power. While the vast majority of EWAS studies have focused exclusively on 5mC, the identification and incorporation of tissue specific and variable regions of 5- hydroxymethylcytosine provided deeper insights into the epigenetic landscape. Altogether, the EWAS array contained both previously identified markers and candidate CpG sites and is a valuable, more scalable tool for population level EWAS studies. TABLE 3 lists probe sequences of the EWAS array.TABLE 3Example 4 — Detection rates for BeadChip array
[0075] Genomic DNA samples obtained from a cell line, blood, and a FFPE sample were prepared and analyzed in a method substantially the same as the workflow depicted in FIG 3 and FIG. 4. BeadChip types included an ‘EX48 MSA 1.0’ BeadChip and a ‘EX48 MSA 0.3’ BeadChip. The data was analyzed with two calling methods. Probe detection rate passed specifications with genomic DNA (cell line) input of 250 ng on both EX48 MS A 1.0 and EX24 MSA0.3 beadchips. Probe detection rate passed specifications with genomic DNA (cell line and blood samples) input of 50 ng on both EX48 MSA1.0 and EX24 MSA0.3 BeadChips. Probe Detection Rate passed specifications with FFPE DNA (ACq < 1.1) input of 250 ng on EX48 MSA1.0 BeadChip. Results are summarized in TABLE 4.TABLE 4Example 5 — Methylation screening array (MSA): a high-throughput DNA methylation BeadChip assay for epigenome-wide association studies
[0076] DNA methylation is a stable covalent chemical modification to DNA that shapes cell identity and impacts diverse processes from genome stability to gene expression. Infinium Methylation BeadChips (Illumina, Inc. San Diego, CA) contain probes for preselected CpG sites and can be used for highly accurate quantification of CpG methylation at selected loci. Such microarrays are a cost effective and computationally tractable alternative and have been successfully utilized in epigenome-wide association studies (EWAS). Despite the successes of previous BeadChip versions, some traits associate with only subtle differences in DNA methylation, necessitating ever larger sample sizes. Therefore, a highly consolidated BeadChip strategically designed for EWAS studies would be beneficial for the population epigenetics community.
[0077] To address this need, a targeted Infinium Methylation Screening Array was designed which contained -272K unique CpG, CpH and SNP sites, approximately 30% of the size of the newly released Infinium Methylation EPICv2.0 BeadChip. See e.g., Noguera- Castells, A. et al., Epigenetics. 2023 Dec;18(l):2185742 which is incorporated by reference in its entirety.
[0078] Known EWAS hits of high statistical significance were incorporated along with new CpGs to facilitate discovery and scientific rigor in EWAS. Approximately 50% of probes were derived from mining EWAS databases and literature. The methylation of these CpGs were linked to thousands of biological features from broad categories including cardiovascular, metabolic, neurodegenerative / psychiatric, autoimmune, genetic, environmental exposure and infection-related traits and diseases that were studied using historical and existing Infinium array platforms. The remaining 50% were novel probe designs obtained from integrative analysis of >300 public bulk and single-cell WGBS datasets. These probes target DNA methylation that have been associated with cell type, gene expression,chromatin accessibility and mono-allelic expression. Additional cytosines were selected from regulatory region annotations throughout the human genome to further enhance discovery power. While the vast majority of EWAS studies have focused exclusively on 5mC, the identification and incorporation of tissue specific and variable regions of 5- hydroxymethylcytosine for the array developed herein provides deeper insights into the epigenetic landscape. Altogether, the newly developed EWAS array contained both previously identified markers and candidate CpG sites and is a valuable, more scalable tool for large scale EWAS studies..Design process and probe summary
[0079] An embodiment for screening and probe selection process for the methylation screening array (MSA) is depicted in FIG. 5. Probes were selected for designability and association with diverse target biology groups.
[0080] With regard to MSA size and probe type distribution, the MSA had 281 ,806 probes, 108,964 of which were not found on previous array platforms. TABLE 5 lists the distribution of certain probe types, including CpH dinucleotides (where H=A, C or T). FIG. 6A depicts genomic coordinate overlap of probes between the MSA and EPICv2. As depicted in FIG. 6B, the MSA had a higher proportion of SNP and CpH probes, as compared to the EPICv2 array.TABLE 5Target biology groups and EPICv2 comparison
[0081] A summary of enrichment of MSA probes in EWAS traits is depicted in FIG. 7. MSA was enriched in methylation EWAS hits. The enrichment of MSA, EPICv2 and a random selection of equal size to MSA was tested against collated EWAS hits from multiple databases. Across diverse trait groups, MSA was highly enriched reflecting the targeted selection of known trait associated CpGs.
[0082] Cell specific marker types and probe counts represented on MSA are depicted in FIG. 8. Single cell and sorted WGBS data were analyzed to identify cell specific CpG markers. Thousands of signatures were identified delineating hundreds of cell types and cell type comparisons. When possible, a balanced selection of hypo and hypermethylated markers were selected. The novel cell-specific probe designs enhance cell deconvolution and provide investigations into selectively vulnerable populations across diverse diseases.
[0083] Cell signature and full stack ChromHMM comparison between MSA and EPICv2 are depicted in FIG. 9A and FIG. 9B. MSA had greater cell type marker coverage per cell type comparison branch (FIG. 9B) compared to EPICv2 despite the lower overall probe count. Both platforms were highly depleted from heterochromatic (HET), quiescent (Quies) and alignment gap(GapArtf) states (FIG. 9A). MSA was more enriched than EPICv2 in cell specific enhancer elements and some bivalent / downstream promoters.
[0084] Cis-regulatory element coverage is depicted in FIG. 10A and FIG. 10B. ENCODE cis regulatory elements were included in the target biology design groups (Moore, I.E. et al., Nature. 2020 Jul;583(7818):699-710). MSA had higher enhancer coverage with a decrease in promoter and CpG island representation compared to EPICv2.
[0085] Additional target biology enrichments are depicted in FIG. 11. The MSA was highly enriched in several additional methylation features. These included: CpGs with methylation correlated with gene expression, correlated regions of systemic interindividual variation (CoRSIVs), monoallelic methylation I imprinting loci, and CpGs shown to be hydroxy methylated at the pan tissue and tissue specific level.Technical performance and validation
[0086] As depicted in FIG. 12, the MSA generated high quality data, with greater than 90% probe detection rate on non- FFPE processed samples. MSA correlation rate with methylation titrations is depicted in FIG. 13. Methylation readings at most probes showed a high correlation with DNA methylation titrations of 0, 50, and 100%.Conclusion
[0087] A new methylation BeadChip was designed that incorporated CpG probes from previous arrays and -110K new probes. The array was highly enriched in previouslyidentified EWAS hits and target biology groups including cell specific 5mC and 5hmC signatures, gene expression correlated CpGs, CoRSIVs, monoallclic methylation sites and more. The reduced size coupled with the targeted probe selections provides cost effective large scale EWAS studies while maintaining exploratory power and the ability to detect known trait associations.
[0088] The term “comprising” as used herein is synonymous with “including,” “containing,” or “characterized by,” and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps.
[0089] The above description discloses several methods and materials of the present invention. This invention is susceptible to modifications in the methods and materials, as well as alterations in the fabrication methods and equipment. Such modifications will become apparent to those skilled in the art from a consideration of this disclosure or practice of the invention disclosed herein. Consequently, it is not intended that this invention be limited to the specific embodiments disclosed herein, but that it cover all modifications and alternatives coming within the true scope and spirit of the invention.
[0090] All references cited herein, including but not limited to published and unpublished applications, patents, and literature references, are incorporated herein by reference in their entirety and are hereby made a part of this specification. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
Claims
WHAT IS CLAIMED IS:
1. An array comprising a plurality of polynucleotide probes attached to a surface of a substrate, wherein the polynucleotide probes each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386.
2. The array of claim 1, wherein the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386.
3. The array of claim 1, wherein the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972.
4. The array of any one of claims 1-3, wherein the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another.
5. The array of any one of claims 1-4, wherein the substrate comprises a plurality of beads.
6. The array of any one of claims 1-5, wherein the substrate comprises a flow cell.
7. The array of any one of claims 1-6, wherein the plurality of polynucleotides are biotinylated.
8. A method for measuring a methylation level for a plurality of loci of target nucleic acids, comprising:(a) obtaining a sample comprising target nucleic acids;(b) converting conversion- sensitive cytosine residues of the target nucleic acids to another base residue to obtain converted nucleic acids;(c) hybridizing the converted nucleic acids to the array of any one of claims 1- 7; and(d) detecting the presence or absence of methylated loci in the converted nucleic acids in the array.
9. The method of claim 8, wherein the converting comprises bisulfite conversion.
10. The method of claim 8 or 9, further comprising amplifying the converted nucleic acids.
11. The method of claim 10, wherein the amplifying is performed before the hybridizing.
12. The method of claim 10 or 11 , wherein the amplifying comprises adding indexes to the converted nucleic acids.
13. The method of any one of claims 8-12, wherein the detecting comprises extending the polynucleotide probes.
14. The method of any one of claims 8-13, wherein the detecting comprises sequencing the converted nucleic acids.
15. The method of claim 14, further comprising aligning sequences of the converted nucleic acids with a reference sequence.
16. The method of claim 14 or 15, further comprising mapping a methylated cytosine residue on a sequence of a converted nucleic acid on a reference sequence.
17. The method of any one of claims 8-16, wherein the sample comprises cell free DNA, genomic DNA.
18. The method of any one of claims 8-17, wherein the sample comprises a formalin fixed paraffin embedded (FFPE) sample.
19. A system comprising: the array of any one of claims 1-7; and a detector.
20. A kit comprising a plurality of polynucleotides, wherein the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01—382,386.
21. The kit of claim 20, wherein the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386.
22. The kit of claim 20, wherein the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972.
23. The kit of any one of claims 20-22, wherein the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another.
24. The kit of any one of claims 20-23, wherein the plurality of polynucleotides are biotinylated.
25. A method of making an array, comprising:(a) providing a substrate; and(b) attaching a plurality of polynucleotides to a surface of the substrate, wherein the polynucleotides each comprise a different nucleotide sequence from one another selected from any one of SEQ ID NOs: 01 — 382,386.
26. The method of claim 25, wherein the polynucleotide probes comprise pairs of sequences selected from any one of SEQ ID NOs: 213,973 — 382,386.
27. The method of claim 25, wherein the polynucleotide probes comprise sequences selected from any one of SEQ ID NOs: 01 — 213,972.
28. The method of any one of claims 25-27, wherein the polynucleotide probes comprise at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 250,000, 300,000 or 350,000 different nucleotide sequences from one another.
29. The method of any one of claims 25-28, wherein the substrate comprises a plurality of beads.
30. The method of any one of claims 25-29, wherein the substrate comprises a flow cell.
31. The method of any one of claims 25-30, wherein the plurality of polynucleotides are biotinylated.