Methylated region wide association studies
The MRWAS method systematically identifies DMRs by regressing methylation and phenotype data to classify regions associated with phenotypes, addressing the inconsistency in DMR definition and improving the understanding of biological processes and disease states.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2026-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Existing methods lack a consistent and rigorous approach to identify differentially methylated regions (DMRs) in genomic sequences, which are indicative of phenotypic effects, due to varying definitions of what constitutes a DMR across individuals and studies.
A method and system for performing methylated region-wide association analysis (MRWAS) by regressing methylation level data against phenotype data to generate z-values, determining combined values for consecutive CpGs, and classifying regions that pass a threshold as differentially methylated based on these values.
Provides a systematic and reliable method to identify DMRs associated with phenotypes, enhancing the understanding of biological processes and potential disease states like cancer by accurately categorizing methylated regions.
Smart Images

Figure US2026012148_30072026_PF_FP_ABST
Abstract
Description
METHYLATED REGION WIDE ASSOCIATION STUDIESREFERENCE TO ELECTRONIC SEQUENCE LISTING
[0001] The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on January 21, 2026, is named “ILUM0161 xml” and is 12,470 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to the systems and methods for identifying methylated regions of a genome having an association with phenotypic effects and, in particular, to rigorous and consistent techniques for doing so.BACKGROUND
[0003] Aspects of the present disclosure relate generally to devices, systems, and methods providing biological or chemical analysis. Various protocols in biological or chemical research involve performing a large number of controlled reactions on local support surfaces or within predefined reaction chambers. The designated reactions may then be observed or detected, and subsequent analysis may help identify or reveal properties of chemicals involved in the reaction. For example, in some multiplex assays, an unknown analyte having an identifiable label (e.g., fluorescent label) may be exposed to thousands of known probes under controlled conditions. Each known probe may be deposited into a corresponding well of a flow cell channel. Observing any chemical reactions that occur between the known probes and the unknown analyte within the wells may help identify or reveal properties of the analyte. Other examples of such protocols include known DNA sequencing processes, such as sequencing-by-synthesis (SBS) or cyclic-array sequencing.
[0004] While a variety of devices, systems, and methods have been made and used to perform biological or chemical analysis, it is believed that no one prior to the inventor(s) has made or used the devices and techniques described herein.
[0005] With the preceding in mind, the various processes in the human body are complex in nature, having both genetic and epigenetic components. By way of example, cancer development is one such complex process, as are processes related to embryonic and fetal development. With respect to epigenetic mechanisms, one such mechanism is methylation, which often occurs via the addition of a methyl group at the 5' carbon of cytosine residues. This methyl addition to cytosine residues often occurs at CpG dinucleotides. Regions of a genome that contain CpG dinucleotides at a higher frequency than expected are often characterized as CpG islands. Such CpG islands may be present in gene promoter regions. Such methylation is generally understood or believed to serve a regulatory or control function with respect to gene expression, with methylation or increased methylation of such CpG islands often being associated with repression of expression of the gene in question (i.e., turning the gene “off’). Correspondingly, abnormalities in the methylation associated with epigenetic control processes may be relevant to numerous disorders and disease states, including but not limited to cancer.
[0006] In practice, methylation affects biological processes in a locally-additive way (i.e., there is “consistent local directionality” to the methylated signal. This is in contrast to mutation of a nucleotide (e.g., a germline mutation from an A to a T), which does not a priori affect an individual in the same way as the same type of mutation in a neighboring nucleotide. Hence, regions of a DNA sequence where methylation is more highly present may be indicative of biological process or pathway effects or, more generally, phenotypic effects. In the context of such methylated regions, differentially methylated regions (DMRs) where the methylation differs between cells may be of particular interest and may be indicative of a phenotypic effect associated with the methylated region. However, individuals, organizations, studies, and so forth may differ in how they characterize what constitutes a differentially methylated region. Correspondingly, work based upon the presence or absence of such differentially methylated regions may be impacted by the absence of a singular understanding of what constitutes such a region.SUMMARY
[0007] A summary of certain embodiments disclosed herein is set forth below. It should be understood that these aspects are presented merely to provide the reader with a brief summary of these certain embodiments and that these aspects are not intended to limit the scope of this disclosure. Indeed, this disclosure may encompass a variety of aspects that may not be set forth below.
[0008] In accordance with certain embodiments, a method is provided for categorizing methylated regions as differentially methylated regions. In accordance with this method, a methylation sequencing data set and a phenotype data set are accessed or acquired. Methylation level data is regressed against phenotype data of the phenotype data set, per CpG site, to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set. A methylated region-wide association analysis is performed on the set of z-values by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and, based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
[0009] In accordance with further embodiments, a processor-based system is provided. In accordance with this embodiment, the processor-based system comprises: at least one processor configured to execute stored routines and one or more tangible storage media storing processorexecutable routines. The processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising: accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and, based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
[0010] In accordance with additional embodiments, one or more non-transitory computer-readable media are provided. The one or more non-transitory computer-readable media store processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform steps comprising: accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and, based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype
[0011] Various refinements of the features noted above may exist in relation to various aspects of the present disclosure. Further features may also be incorporated in these various aspects as well. These refinements and additional features may exist individually or in any combination. For instance, various features discussed below in relation to one or more of the illustrated embodiments may be incorporated into any of the above-described aspects of the present disclosure alone or in any combination. The brief summary presented above is intended only to familiarize the reader with certain aspects and contexts of embodiments of the present disclosure without limitation to the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Various aspects of this disclosure may be better understood upon reading the following detailed description and upon reference to the drawings in which:
[0013] FIG. 1 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0014] FIG. 2 depicts a cross-sectional view of an example of a flow cell that may be used in the system of FIG. 1, in accordance with aspects of the present disclosure.
[0015] FIG. 3 depicts a schematic view of another example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.
[0016] FIG. 4 depicts a schematic view of an example of a networked system in which the system of FIG. 1 or 3 may be incorporated, in accordance with aspects of the present disclosure.
[0017] FIG. 5 depicts a schematic view of an example of a base calling arrangement that may be carried out using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0018] FIG. 6 depicts a schematic view of an example of a base caller training technique that may be implemented using the system of FIG. 1, 3, or 4, in accordance with aspects of the present disclosure.
[0019] FIG. 7 depicts examples of corresponding nucleic acid sequences (SEQ ID NOS: 1, 2, 3, 4, 5, 6,1, 7, and 1, respectively) encompassing a methylated region having cytosine sites, some of which are methylated, in accordance with aspects of the present disclosure.
[0020] FIG. 8 depicts a view of methylation level data of samples illustrated as a two-dimensional matrix in conjunction with a phenotype matrix, in accordance with aspects of the present disclosure.
[0021] FIGS. 9A-9C depicts whole-genome methylation sequencing (WGMS) data in a matrix and illustrates sequential regression analysis of the methylation state of CpG sites with respect to phenotype data, in accordance with aspects of the present disclosure.
[0022] FIGS. 10A-10C depicts processing of regression outputs of the process of FIG. 9 against a threshold to characterize sequential CpGs sites as differentially methylated or not, in accordance with aspects of the present disclosure.
[0023] FIGS. 11A-11C depicts processing of regression outputs of the process of FIG. 9 against a threshold to characterize sequential pairs of CpGs sited as differentially methylated or not, in accordance with aspects of the present disclosure.
[0024] FIG. 12A-12B depicts processing of regression outputs of the process of FIG. 9 against a threshold to characterize sequential triplets of CpGs sited as differentially methylated or not, inaccordance with aspects of the present disclosure.
[0025] FIG. 13 depicts an illustrative Manhattan plot of methyl ome-wi de association study (MWAS) data.
[0026] FIG. 14 depicts an illustrative example of methylated region-wide association study (MRWAS) data, in accordance with aspects of the present disclosure.
[0027] FIG. 15 illustrates the three-dimensional nature of the MRWAS data of FIG. 14, in accordance with aspects of the present disclosure.DETAILED DESCRIPTION
[0028] The following detailed description of certain examples will be better understood when read in conjunction with the appended drawings. To the extent that the figures illustrate diagrams of the functional blocks of various examples, the functional blocks are not necessarily indicative of the division between hardware components. Thus, for example, one or more of the functional blocks (e.g., processors or memories) may be implemented in a single piece of hardware (e.g., a general -purpose signal processor or random-access memory, hard disk, or the like). Similarly, the programs may be stand-alone programs, may be incorporated as subroutines in an operating system, may be functions in an installed software package, and the like. It should be understood that the various examples are not limited to the arrangements and instrumentality shown in the drawings.
[0029] I. Overview of System for Biological or Chemical Analysis
[0030] Examples described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction (e.g., methylation). Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterativecycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No.7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety; and / or U.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety.
[0031] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.
[0032] II. Examples of Fluidics Devices and Fluid Flow Paths - Example of System with Higher Volume Throughput
[0033] FIG. 1 illustrates a schematic diagram of an example of a system (100) that may be used to perform an analysis on one or more samples of interest. In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shown, system (100) is configured to receive a flow cell cartridge assembly (102) including a flow cell assembly (103) and a sample cartridge (104). System (100) includes a flow cell receptacle (122) that receives flow cell cartridge assembly (102), a vacuum chuck (124) that supports flow cell assembly (103), and a flow cell interface (126) that is used to establish a fluidic coupling between system (100) and flow cell assembly (103). Flow cell interface (126) may include one or more manifolds. System (100) further includes a sipper manifold assembly (106), a sample loading manifold assembly (108), and a pump manifold assembly (110). System (100) also includes a drive assembly (112), a controller (114), an imaging system (116), and a waste reservoir (118). Controller (114) is electrically and / or communicatively coupled to drive assembly (112) and to imaging system (116); and is configured to cause drive assembly (112) and / or the imaging system (116) to perform various functions asdisclosed herein.
[0034] In the present example, flow cell assembly (103) includes a flow cell (128) having a channel (130) and defining a plurality of first openings (132), which are fluidically coupled to the channel (130) and arranged on a first side (134) of the channel (130). Flow cell (128) further includes a plurality of second openings (136) fluidically coupled to the channel (130) and arranged on a second side (138) of the channel (130). Fluid may thus flow through flow cell (128) via channel. While the flow cell (128) is shown including one channel (130), flow cell (128) may include two or more channels (130). Flow cell assembly (103) also includes a flow cell manifold assembly (140) coupled to flow cell (128) and having a first manifold fluidic line (142) and a second manifold fluidic line (144). Flow cell manifold assembly (140) may be in the form of a laminate including a plurality of layers as discussed in more detail below.
[0035] In the implementation shown, first manifold fluidic line (142) has a first fluidic line opening (146) and is fluidically coupled to each of the first openings (132) of flow cell (128); and second manifold fluidic line (144) has a second fluidic line opening (148) and is fluidically coupled to each of the second openings (136). As shown, flow cell assembly (103) includes gaskets (150) coupled to flow cell manifold assembly (140) and fluidically coupled to fluidic line openings (146, 148). In some implementations where flow cell (128) includes a plurality of channels (130), flow cell manifold assembly (140) may include additional fluidic lines (152) that couple first fluidic line openings (146) to a single manifold port (154). In such implementations, a single gasket (150) may be coupled to flow cell manifold assembly (140) that surrounds the manifold port (154) and is in fluidic communication with a plurality of channels (130). In operation, flow cell interface (126) engages with corresponding gaskets (150) to establish a fluidic coupling between system (100) and flow cell (128). The engagement between flow cell interface (126) and gaskets (150) reduces or eliminates fluid leakage between flow cell interface (126) and flow cell (128).
[0036] In the implementation shown, first manifold fluidic line (142) has a portion (156) that is substantially parallel to a longitudinal axis (158) of channel (130); and second manifold fluidic line (144) has a portion (160) that is substantially parallel to longitudinal axis (158) of channel (130). Additionally, first manifold fluidic line (142) is shown being at least partially adjacent a first end (162) of flow cell (128) and spaced from a second end (164) of flow cell (128); and second manifold fluidic line (144) is shown being at least partially adjacent second end (164) of flow cell(128) and spaced from first end (162). Other arrangements of manifold fluidic lines (142, 144) may prove suitable, however.
[0037] In the implementation shown, system (100) includes a sample cartridge receptacle (166) that receives sample cartridge (104) that carries one or more samples of interest (e.g., an analyte). System (100) also includes a sample cartridge interface (168) that establishes a fluidic connection with sample cartridge (104). Sample loading manifold assembly (108) includes one or more sample valves (170). Pump manifold assembly (110) includes one or more pumps (172), one or more pump valves (174), and a cache (176). Valves (170, 174) and pumps (172) may take any suitable form. Cache (176) may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system (100). While cache (176) is shown being included in pump manifold assembly (110), cache (176) may alternatively be located elsewhere (e.g., in sipper manifold assembly (106) or in another manifold downstream of a bypass fluidic line (178), etc.).
[0038] Sample loading manifold assembly (108) and pump manifold assembly (110) flow one or more samples of interest from sample cartridge (104) through a fluidic line (180) toward flow cell cartridge assembly (102). In some implementations, sample loading manifold assembly (108) may individually load or address each channel (130) of flow cell (128) with a respective sample of interest. The process of loading channel (130) with a sample of interest may occur automatically using system (100). As shown in FIG. 1, sample cartridge (104) and sample loading manifold assembly (108) are positioned downstream of flow cell cartridge assembly (102). In the implementation shown, sample loading manifold assembly (108) is coupled between flow cell cartridge assembly (102) and pump manifold assembly (110). To draw a sample of interest from sample cartridge (104) and toward pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward pump manifold assembly (110). Sample cartridge (104) may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves (170). To individually flow the sample of interest toward channel (130) of flow cell (128) and away from pump manifold assembly (110), sample valves (170), pump valves (174), and / or pumps (172) may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly (102) and into respective channels (130) of flow cell (128).
[0039] Drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow one or more reagents that interact with the sample within flow cell (128). In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system (116) excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system (100). Examples of features and functionalities that may be incorporated into imaging system (116) will be described in greater detail below.
[0040] After the image data is obtained, drive assembly (112) interfaces with sipper manifold assembly (106) and pump manifold assembly (110) to flow another reaction component (e.g., a reagent) through flow cell (128) that is thereafter received by waste reservoir (118) via a primary waste fluidic line (182) and / or otherwise exhausted by system (100). Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.
[0041] The primary waste fluidic line (182) is coupled between pump manifold assembly (110) and waste reservoir (118). In some implementations, pumps (172) and / or pump valves (174) of pump manifold assembly (110) selectively flow the reaction components from flow cell cartridge assembly (102), through fluidic line (180) and sample loading manifold assembly (108) to primary waste fluidic line (182). Flow cell cartridge assembly (102) is coupled to a central valve (184) via flow cell interface (126). Central valve (184) is coupled with flow cell interface (126) via a fluidic line (185). An auxiliary waste fluidic line (186) is coupled to central valve (184) and to waste reservoir (118). In some implementations, auxiliary waste fluidic line (186) receives excess fluid of a sample of interest from flow cell cartridge assembly (102), via central valve (184), and flows the excess fluid of the sample of interest to waste reservoir (118) when back loading the sample of interest into flow cell (128), as described herein.
[0042] Sipper manifold assembly (106) includes a shared line valve (188) and a bypass valve(190). Shared line valve (188) may be referred to as a reagent selector valve. Central valve (184) and the valves (188, 190) of sipper manifold assembly (106) may be selectively actuated to control the flow of fluid through fluidic lines (192, 194, 196). Sipper manifold assembly (106) may be coupled to a corresponding number of reagent reservoirs (198) via reagent sippers (200). Reagent reservoirs (198) may contain fluid (e.g., reagent and / or another reaction component). In some implementations, sipper manifold assembly (106) includes a plurality of ports. Each port of sipper manifold assembly (106) may receive one of the reagent sippers (200). Reagent sippers (200) may be referred to as fluidic lines. Some forms of reagent sippers (200) may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly (106). Reagent reservoirs (198) may be provided in a cartridge, and the tubes of reagent sippers (200) may be configured to be inserted into corresponding reagent reservoirs (198) in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir (198) into the sipper manifold assembly (106).
[0043] Shared line valve (188) of sipper manifold assembly (106) is coupled to central valve (184) via shared reagent fluidic line (196). Different reagents may flow through shared reagent fluidic line (196) at different times. In some versions, when performing a flushing operation before changing between one reagent and another, pump manifold assembly (110) may draw wash buffer through shared reagent fluidic line (196), central valve (184), and flow cell cartridge assembly (102).
[0044] Bypass valve (190) of sipper manifold assembly (106) is coupled to central valve (184) via dedicated reagent fluidic lines (194, 196). Each of the dedicated reagent fluidic lines (194, 196) may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines (194, 196) may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.
[0045] Bypass valve (190) is also coupled to cache (176) of pump manifold assembly (110) via bypass fluidic line (178). One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line (178). The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly (102). Thus, the operations using bypass fluidic line (178) may occur during, for example, incubation of one ormore samples of interest within flow cell cartridge assembly (102). That is, shared line valve (188) may be utilized independently of bypass valve (190) such that bypass valve (190) may utilize bypass fluidic line (178) and / or cache (176) to perform one or more operations while shared line valve (188) and / or central valve (184) simultaneously, substantially simultaneously, or offset synchronously perform other operations.
[0046] Drive assembly (112) includes a pump drive assembly (202) and a valve drive assembly (204). Pump drive assembly (202) may be adapted to interface with one or more pumps (172) to pump fluid through flow cell (128) and / or to load one or more samples of interest into flow cell (128). Valve drive assembly (204) may be adapted to interface with one or more of the valves (170, 174, 184, 188, 190) to control the position of the corresponding valves (170, 174, 184, 188, 190).
[0047] Controller (114) of the present example includes a user interface (206), a communication interface (208), one or more processors (210), and a memory (212) storing instructions executable by the one or more processors (210) to perform various functions including the disclosed implementations. User interface (206), communication interface (133), and memory (212) are electrically and / or communicatively coupled to the one or more processors (210). User interface (206) may be adapted to receive input from a user and to provide information to the user associated with the operation of system (100) and / or an analysis taking place. User interface (206) may include a touch screen, a display, a keyboard, a speaker(s), a mouse, a track ball, and / or a voice recognition system.
[0048] Communication interface (208) is adapted to enable communication between system (100) and a remote system(s) (e.g., computers) via a network(s) (e.g., the Internet, an intranet, a local-area network (LAN), a wide-area network (WAN), a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line (DSL) network, a cellular network, a Bluetooth connection, a near field communication (NFC) connection, etc.). Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system (100). Some of the communications provided to system (100) may be associated with a fluidics analysis operation, patient records, and / or a protocol(s) to be executed by system (100).
[0049] The one or more processors (210) and / or system (100) may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors (210) and / or system (100) includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.
[0050] Memory (212) may include one or more of a semiconductor memory, a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable readonly memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in which information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).
[0051] III. Examples of Flow Cell Structures
[0052] As noted above, a system (100) may execute reactions in a flow cell (128) and / or perform analysis on one or more samples of interest in a flow cell (128). The following describes examples of forms that such flow cells (128) may take, it being understood that flow cells (128) may take various other forms and have various other features in addition to or in lieu of the features described below.
[0053] Example of Single-Surface Patterned Flow Cell
[0054] FIG. 2 shows an example of a flow cell (400) that includes a patterned substrate (402), which includes depressions (404) separated by interstitial regions (406), and surface chemistry (410, 412) positioned in the depressions (404). Depressions (404) may be in the form of microwells or nanowells. Depressions (404) may be configured to contain nucleic acid strands or other oligonucleotides and thereby provide a reaction site for SBS and / or for other kinds ofprocesses. In some versions, each depression (404) has a cylindraceous configuration, with a generally circular cross-sectional profile. In some other versions, each depression (404) has a polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.) cross-sectional profile. Alternatively, depressions (404) may have any other suitable configuration. It should also be understood that depressions (404) may be arranged in any suitable pattern, including but not limited to a grid pattern.
[0055] Surface chemistry (410, 412) of the present example includes functionalized coating layer (410) and primers (412). While not shown, it is to be understood that the depressions (404) may also have surface preparation or treatment chemistry (e g., silane or a silane derivative) positioned between the substrate (402) and the functionalized coating layer (410). This same surface preparation or treatment chemistry may also be positioned on the interstitial regions (406). In the present example, a hydrogel (440) is applied before lid (420) is bonded to substrate (402). Hydrogel (440) covers surface chemistry (410, 412) in depressions (404), and at least a portion of the patterned substrate (402) (e.g., those interstitial regions (406) that are not also bonding regions (422)). By way of example only, hydrogel (440) may comprise PAZAM, crosslinked polyacrylamide, agarose gel, etc.
[0056] Flow cell (400) of this example further includes a lid (420) bonded to bonding region(s) (422) of patterned substrate (402). In the example shown in FIG. 2, lid (420) includes a top portion (424) that is connected to several sidewalls (426), and these components (424, 426) define a portion of each of the six flow channels (430A, 430B, 430C, 430D, 430E, 430F). The respective sidewalls (426) isolate one flow channel (430A, 430B, 430C, 430D, 430E, 430F) from each adjacent flow channel (430A, 430B, 430C, 430D, 430E, 430F). Each flow channel (430A, 430B, 430C, 430D, 430E, 430F) is in selective fluid communication with a respective set of depressions (404).
[0057] Lid (420) may be bonded to bonding region (422) of substrate (402) using any suitable technique, such as laser bonding, diffusion bonding, anodic bonding, eutectic bonding, plasma activation bonding, glass frit bonding, or other methods known in the art. In some versions, a spacer layer (428) may be used to bond lid (420) to bonding region (422). Spacer layer (428) may comprise any material that will seal at least some of interstitial regions (404) (e.g., bonding region (422)) of substrate (402) and lid (420) together. While not shown, lid (420) or the patternedsubstrate (402) may include inlet and outlet ports that are to fluidically engage other ports (not shown), such as those of sample cartridge interface (168), for directing fluid(s) into the respective flow channels (430A, 430B, 430C, 430D, 430E, 430F) (e.g., from a reagent cartridge or other fluid storage system) and out of the flow channel (e.g., to waste reservoir (118) or another waste removal system). Flow channels (430A, 430B, 430C, 430D, 430E, 430F) may serve to, for example, selectively introduce reaction components or reactants to hydrogel (440) and the underlying surface chemistry (410, 412) in order initiate designated reactions in / at depressions (404).
[0058] While flow cell (400) includes a pattern of depressions (404) to provide an array of reaction sites, other variations may provide reaction sites on or at various other kinds of structural features, including but not limited to continuously planer surfaces and / or protruding surfaces, etc. By way of further example only, flow cell (400) may be constructed and operable in accordance with at least some of the teachings of U.S. Pat. No. 10,919,033, entitled “Flow Cells with Hydrogel Coating,” issued February 16, 2021, the disclosure of which is incorporated by reference herein, in its entirety.
[0059] IV. Examples of Imaging System Features
[0060] As noted above, system (100) includes an imaging system (116) that excites one or more identifiable labels (e.g., a fluorescent label) in samples in reaction sites provided by depressions (404, 462, 464) of a flow cell (128, 400); and thereafter obtains image data for the identifiable labels. This image data is used to identify nucleotides as part of a nucleic acid sequencing process. Alternatively, the image data may be used for various other purposes. The following description provides details on how some versions of imaging system (116) may be configured and operable.
[0061] FIG. 3 illustrates a schematic diagram of another example of a system (500) that may be used to perform an analysis on one or more samples of interest. Except as otherwise described below, system (500) of this example may be configured and operable like systems (100) described above. System (500) is configured to perform a large number of parallel reactions within a flow cell ( 10). Flow cell (510) may be configured and operable like flow cells (400) described above or may have any other suitable configuration. Flow cell (510) may thus include one or more flowchannels that receive a solution from system (500) and direct the solution toward reaction sites of flow cell (510).
[0062] System (500) includes a system controller (520) that may communicate with the various components, assemblies, and sub-systems of the system (500). Controller (520) may be configured and operable like controllers (114) described above. An imaging assembly (522) of system (500) includes a light emitting assembly (550) that emits light that reaches reaction sites on flow cell (510). Light emitting assembly (550) may include an incoherent light emitter (e.g., emit light beams output by one or more excitation diodes), or a coherent light emitter such as emitter of light output by one or more lasers or laser diodes. In some implementations, light emitting assembly (550) may include a plurality of different light sources (not shown), each light source emitting light of a different wavelength range. Some versions of light emitting assembly (550) may also include one or more collimating lenses (not shown), a light structuring optical assembly (not shown), a projection lens (not shown) that is operable to adjust a structured beam shape and path, epifluorescence microscopy components, and / or other components. Although system (500) is illustrated as having a single light emitting assembly (550), multiple light emitting assemblies (550) may be included in some other implementations.
[0063] In the present example, the light from light emitting assembly (550) is directed by dichroic mirror assembly (546) through an objective lens assembly (542) onto a sample of a flow cell (510), which is positioned on a motion stage (570). In the case of fluorescent microscopy of a sample, a fluorescent element associated with the sample of interest fluoresces in response to the excitation light, and the resultant light is collected by objective lens assembly (542) and is directed to an image sensor of camera system (540) to detect the emitted fluorescence. In some implementations, a tube lens assembly may be positioned between the objective lens assembly (542) and the dichroic mirror assembly (546) or between the dichroic mirror (546) and the image sensor of the camera system (540). A moveable lens element may be translatable along a longitudinal axis of the tube lens assembly to account for focusing on an upper interior surface or lower interior surface of the flow cell (510) and / or spherical aberration introduced by movement of the objective lens assembly (542).
[0064] In the present example, a fdter switching assembly (544) is interposed between dichroic mirror assembly (546) and camera system (540). Filter switching assembly (544) includes one ormore emission filters that may be used to pass through particular ranges of emission wavelengths and block (or reflect) other ranges of emission wavelengths. For example, emission filters may be used to direct different wavelength ranges of emitted light to different image sensors of the camera system (540) of imaging assembly (522). For instance, the emission filters may be implemented as dichroic mirrors that direct emission light of different wavelengths from flow cell (510) to different image sensors of camera system (540). In some variations, a projection lens is interposed between filter switching assembly (544) and camera system (540). Filter switching assembly (544) may be omitted in some versions.
[0065] System (500) further includes a fluid delivery assembly (590) that may direct the flow of reagents (e g., fluorescently labeled nucleotides, buffers, enzymes, cleavage reagents, etc.) to (and through) flow cell (510) and waste valve (580). Fluid delivery assembly (590) may be configured and operable like the various fluid delivery components described herein. System (500) of the present example also includes a temperature station actuator (530) and heater / cooler (532) that may optionally regulate the temperature of conditions of the fluids within the flow cell (510). In some implementations, the heater / cooler (532) may be fixed to sample stage (570), upon which the flow cell (510) is placed, and / or may be integrated into sample stage (570).
[0066] Flow cell (510) may be removably mounted on sample stage (570), which may provide movement and alignment of flow cell (510) relative to objective lens assembly (542). Sample stage (570) may have one or more actuators to allow sample stage (570) to move in any of three dimensions. For example, actuators may be provided to allow sample stage (570) to move in the x, y, and z directions relative to objective lens assembly (542), tilt relative to objective lens assembly (542), and / or otherwise move relative to objective lens assembly (542). Movement of sample stage (570) may allow one or more sample locations on flow cell (510) to be positioned in optical alignment with objective lens assembly (542). Movement of sample stage (570) relative to objective lens assembly (542) may be achieved by moving sample stage (570) itself, by moving objective lens assembly (542), by moving some other component of imaging assembly (522), by moving some other component of system (500), or any combination of the foregoing. For instance, in some implementations, the sample stage (570) may be actuatable in the x and y directions relative to the objective lens assembly (542) while a focus component (562) or z-stage may move the objective lens assembly (542) along the z direction relative to the sample stage (570).
[0067] In some implementations, a focus component (562) may be included to control positioning of one or more elements of objective lens assembly (542) relative to the flow cell (510) in the focus direction (e.g., along the z-axis or z-dimension). Focus component (562) may include one or more actuators physically coupled to the objective lens assembly (542), the optical stage, the sample stage (570), or a combination thereof, to move flow cell (510) on sample stage (570) relative to the objective lens assembly (542) to provide proper focusing for the imaging operation. In the present example, the focus component (562) utilizes a focus tracking module (560) that is configured to detect a displacement of the objective lens assembly (542) relative to a portion of the flow cell (510) and output data indicative of an in-focus position to the focus component (562) or a component thereof or operable to control the focus component (562), such as controller (520), to move the objective lens assembly (542) to position the corresponding portion of the flow cell (510) in focus of the objective lens assembly (542).
[0068] In some implementations, an actuator of focus component (562) or for sample stage (570) may be physically coupled to objective lens assembly (542), the optical stage, sample stage (570), or a combination thereof, such as, for example, by mechanical, magnetic, fluidic, or other attachment or contact directly or indirectly to or with the stage or a component thereof. The actuator of focus component (562) may be configured to move objective lens assembly (542) in the z-direction while maintaining sample stage (570) in the same plane (e.g., maintaining a level or horizontal attitude, perpendicular to the optical axis). In some implementations, sample stage (570) includes an x direction actuator and a y direction actuator to form an x-y stage. Sample stage (570) may also be configured to include one or more tip or tilt actuators to tip or tilt sample stage (570) and / or a portion thereof, to account for any slope in its surfaces.
[0069] Camera system (540) may include one or more image sensors to monitor and track the imaging (e.g., sequencing) of flow cell (510). Camera system (540) may be implemented, for example, as a CCD or CMOS image sensor camera, but other image sensor technologies (e.g., active pixel sensor) may be used. By way of further example only, camera system (540) may include a dual-sensor time delay integration (TDI) camera, a single-sensor camera, a camera with one or more two-dimensional image sensors, and / or other kinds of camera technologies. While camera system (540) and associated optical components are shown as being positioned above flow cell (510) in FIG. 3, one or more image sensors or other camera components may be incorporatedinto system (500) in numerous other ways as will be apparent to those skilled in the art in view of the teachings herein. For instance, one or more image sensors may be positioned under flow cell (510), such as within the sample stage (570) or below the sample stage (570); or may even be integrated into flow cell (510).
[0070] V. Examples of Data Processing Features
[0071] A. Example of Networked Data Processing Arrangement
[0072] As noted above, a system (100, 500) may include a controller (114, 520) that is configured to process data, execute algorithms, etc., as needed to perform a sequencing operation or other kind of operation. In some scenarios, system (100, 500) may be coupled with other devices via a network to perform further data processing, data storage, execution of algorithms, etc. FIG.4 shows an example of such an arrangement. In particular, FIG. 4 shows a networked system (800) that includes a sequencing device (810), a server device (820), a client device (830), and a local device (840), with all devices (810, 820, 830, 840) being coupled together via a network (850). Network (850) may take any suitable form as will be apparent to those skilled in the art in view of the teachings herein.
[0073] As shown in FIG. 4, sequencing device (810) comprises a computing device and a sequencing device system (812) for sequencing a genomic sample or other nucleic-acid polymer. In some versions, by executing sequencing device system (812) using a processor, sequencing device (810) analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on sequencing device (810). More particularly, sequencing device (810) receives nucleotide- sample slides (e.g., flow cells (128, 400, 510)) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments. It should be understood that sequencing device (810) may represent a version of systems (100, 500) described above.
[0074] In some versions, the sequencing device (810) utilizes SBS to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. By executing sequencing device system (812), sequencing device (810) may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and sendthe BCL file to the local device (840) and / or the server device(s) (820). Sequencing device (810) may communicate the BCL file and / or other data to local device (840) and / or client device (830) via network (850) or directly (i.e., bypassing network (850)).
[0075] In some scenarios, local device (840) is located at or near a same physical location of sequencing device (810). For instance, local device (840) and sequencing device (810) may be integrated into a single computing device. Local device (840) may run sequencing system (814) to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. By executing software in the form of sequencing system (814), local device (840) may align nucleotide reads with a structural variation graph genome (824) and determine genetic variants based on the aligned nucleotide reads. Local device (840) may also send data to client device (830), including a variant call file (VCF) or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0076] Server device(s) (820) may be located remotely from the local device (840) and sequencing device (810). Server device(s) (820) may comprise a distributed collection of servers, where server device(s) (820) include a number of server devices distributed across network (850) and located in the same or different physical locations. Similar to local device (840), server device(s) (820) may include a version of sequencing system (814). Accordingly, server device(s) (820) may generate, receive, analyze, store, and transmit digital data, such as by receiving basecall data or determining variant calls based on analyzing such base-call data. As indicated above, sequencing device (810) may send (and server device(s) (820) may receive) base-call data from sequencing device (810). Server device(s) (820) may also send data to client device (830), including VCFs or other sequencing related information.
[0077] As indicated above, as part of server device(s) (820) or local device (840), sequencing system (814) may generate or implement a structural variation graph genome with alternate contiguous sequences representing structural variant haplotypes. For instance, system (814) may identify candidate structural variants of a threshold frequency (or that otherwise satisfy another occurrence threshold) within a genomic sample database. From among the candidate structural variants, sequencing system (814) selects structural variant haplotypes based on one or both of satisfying another occurrence threshold and finding flanking variants adjacent to particularstructural variant haplotypes. Sequencing system (814) may likewise select reference haplotypes of genomic regions corresponding to the selected structural variant haplotypes from a reference genome. Based on the selected haplotypes, sequencing system (814) generates a structural variation graph genome comprising both alternate contiguous sequences representing the structural variant haplotypes and reference sequences representing the reference haplotypes. Based on comparing nucleotide reads of a genomic sample with alternate contiguous sequences representing structural variant haplotypes, sequencing system (814) can determine nucleobase calls for the genomic sample.
[0078] By executing a sequencing application (832), client device (830) may generate, store, receive, and send digital data. In particular, client device (830) may receive sequencing data from local device (840) or receive call files (e.g., BCL) and sequencing metrics from sequencing device (810). Furthermore, client device (830) may communicate with local device (840) or server device(s) (820) to receive a VCF comprising nucleobase calls and / or other metrics, such as a basecall-quality metrics or pass-filter metrics. Client device (830) may accordingly present or display information pertaining to variant calls or other nucleobase calls within a graphical user interface of sequencing application (832) to a user associated with client device (830). For example, client device (830) may present structural variant calls and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of sequencing application (832).
[0079] As shown in FIG. 4, sequencing application (832) is included in client device (830). Sequencing application (832) may include a web application or a native application stored and executed on client device (830) (e g., a mobile application, desktop application). Sequencing application (832) may include instructions that (when executed) cause client device (830) to receive data from sequencing system (814) and present, for display at client device (830), basecall data or data from a VCF. Furthermore, sequencing application (832) may instruct client device (830) to display summaries for multiple sequencing runs.
[0080] As further illustrated in FIG. 4, a version of sequencing system (814) may be located and implemented (e.g., entirely or in part) on client device (830) or sequencing device (810). In some versions, sequencing system (814) is implemented by one or more other components of networked system (800), such as local device (840). In particular, sequencing system (814) may be implemented in a variety of different ways across sequencing device (810), local device (840),server device(s) (820), and client device (830). For example, sequencing system (814) may be downloaded from server device(s) (820) to sequencing system (814) and / or local device (840) where all or part of the functionality of sequencing system (814) is performed at each respective device within networked system (800).
[0081] B. Examples of Base Calling Schemes
[0082] FIG. 5 illustrates a system (900) that employs two or more base callers for base calling operations on the raw images ( e., sensor data) output by image sensors in a sequencing machine sequencing machine (910). Sequencing machine (910) of this example includes a flow cell (912), which includes a plurality of tiles (914). Each tile (914) includes a plurality of clusters (916). Sequencing machine (910) may be understood to represent a version of systems (100, 500) or sequencing device (810) described above; while flow cell (912) may be understood to represent a version of flow cells (128, 400, 510) described above. Sequencing machine (910) thus outputs sensor data (920) comprising raw images from the tiles (914) of flow cell (912).
[0083] In the present example, system (900) comprises a first base caller (922) and a second base caller (926), though some variations may include more than two base callers (922, 926). Each base caller (922, 926) of this example outputs corresponding base call classification information. For example, first base caller (922) outputs first base call classification information (924); and second base caller (926) outputs second base call classification information (928). A base calling combining module (930) generates final base calls (932), based on one or both first base call classification information (924) and / or second base call classification information (928). In some versions, first base caller (922) is a neural-network based base-caller; while second base caller (926) is a non-neural network based base-caller. For example, first base caller (922) may include a non-linear system employing one or more neural network models for base calling. The first base caller (922) may also be referred to as a DeepRTA (Deep Real Time Analysis) base caller or Deep Neural Network base caller.
[0084] By way of further example only, second base caller (926) may include, at least in part, a linear system used for base calling. For example, some versions of second base caller (926) do not employ a neural network for base calling (or use a smaller neural network model for base calling, compared to a larger neural network model used by first base caller (922)). Second basecaller (926) may also be referred to as an RTA (Real Time Analysis) base caller. An RTA base caller may use linear intensity extractors to extract features from sequencing images for base calling. In some such versions, RTA performs a template generation step to produce a template image that identifies locations of clusters (916) on a tile (914) using sequencing images from some number of initial sequencing cycles called template cycles. The template image is used as a reference for subsequent registration and intensity extraction steps. The template image is generated by detecting and merging bright spots in each sequencing image of the template cycles, which in turn involves sharpening a sequencing image (e.g., using the Laplacian convolution), determining an “on” threshold by a spatially segregated Otsu approach, and subsequent five-pixel local maximum detection with subpixel location interpolation.
[0085] In another example, locations of clusters (916) on a tile (914) are identified using fiducial markers. A solid support upon which a biological specimen is imaged may include such fiducial markers, to facilitate determination of the orientation of the specimen or the image thereof in relation to probes that are attached to the solid support. Examples of fiducials include, but are not limited to, beads (with or without fluorescent moieties or moieties such as nucleic acids to which labeled probes can be bound), fluorescent molecules attached at known or determinable features, or structures that combine morphological shapes with fluorescent moieties.
[0086] RTA then registers a current sequencing image against the template image. This is achieved by using image correlation to align the current sequencing image to the template image on a sub-region, or by using non-linear transformations (e.g., a full six-parameter linear affine transformation). RTA generates a color matrix to correct cross-talk between color channels of the sequencing images. RTA implements empirical phasing correction to compensate noise in the sequencing images caused by phase errors. After different corrections are applied to the sequencing images, RTA extracts signal intensities for each spot location in the sequencing images. For example, for a given spot location, signal intensity may be extracted by determining a weighted average of the intensity of the pixels in a spot location. For example, a weighted average of the center pixel and neighboring pixels may be performed using bilinear or bicubic interpolation. In some implementations, each spot location in the image may comprise a few pixels (e.g., 1-5 pixels). RTA then spatially normalizes the extracted signal intensities to account for variation in illumination across the sampled imaged. For example, intensity values may be normalized suchthat a 5th and 95th percentiles have values of 0 and 1, respectively. The normalized signal intensities for the image (e.g., normalized intensities for each channel) may be used to calculate mean chastity for the plurality of spots in the image.
[0087] In some implementations, RTA uses an equalizer to maximize the signal-to-noise ratio of the extracted signal intensities. The equalizer may be trained (e.g., using least square estimation, adaptive equalization algorithm) to maximize the signal-to-noise ratio of cluster intensity data in sequencing images. In some implementations, the equalizer is a lookup table (LUT) bank with a plurality of LUTs with subpixel resolution, also referred to as “equalizer filters” or “convolution kernels.” By way of example only, the number of LUTs in the equalizer may depend on the number of subpixels into which pixels of the sequencing images can be divided. For example, if the pixels are divisible into n-by-n subpixels (e.g., 5 x 5 subpixels), then the equalizer generates H2LUTs (e.g., 25 LUTs).
[0088] In some implementations of training the equalizer, data from the sequencing images is binned by well subpixel location. It should be understood that a “well” may include depressions (404) of a flow cell (400) or any other kind of reaction site (e.g., in a flow cell or otherwise). In an example of sequencing images being binned by well subpixel location, for a 5 x 5 LUT, 1 / 25th of the wells have a center that is in bin (1,1) (e.g., the upper left comer of a sensor pixel), l / 25th of the wells are in bin (1,2), and so on. The equalizer coefficients for each bin may be determined using least squares estimation on the subset of data from the wells corresponding to the respective bins. This way, the resulting estimated equalizer coefficients are different for each bin. Each LUT / equalizer filter / convolution kernel has a plurality of coefficients that are learned from the training. The number of coefficients in a LUT may correspond to the number of pixels that are used for base calling a cluster. For example, if a local grid of pixels (image or pixel patch) that is used to base call a cluster is of size p xp (e.g., 9x 9 pixel patch), then each LUT has p1coefficients (e g., 81 coefficients). The training may produce equalizer coefficients that are configured to mix / combine intensity values of pixels that depict intensity emissions from a target cluster being base called and intensity emissions from one or more adjacent clusters in a manner that maximizes the signal-to-noise ratio. The signal maximized in the signal-to-noise ratio is the intensity emissions from the target cluster, and the noise minimized in the signal-to-noise ratio is the intensity emissions from the adjacent clusters, i.e., spatial crosstalk, plus some random noise (e.g.,to account for background intensity emissions). The equalizer coefficients are used as weights and the mixing / combining includes executing element-wise multiplication between the equalizer coefficients and the intensity values of the pixels to calculate a weighted sum of the intensity values of the pixels, i.e., a convolution operation.
[0089] RTA then performs base calling by fitting a mathematical model to the optimized intensity data. Suitable mathematical models that can be used include, for example, a k-means clustering algorithm, a k-means-like clustering algorithm, expectation maximization clustering algorithm, a histogram-based method, and the like. Four Gaussian distributions may be fit to the set of two-channel intensity data such that one distribution is applied for each of the four nucleotides represented in the data set. In some implementations, an expectation maximization (EM) algorithm may be applied. As a result of the EM algorithm, for each X, Y value (referring to each of the two channel intensities respectively) a value may be generated which represents the likelihood that a certain X, Y intensity value belongs to one of four Gaussian distributions to which the data is fitted. Where four bases give four separate distributions, each X, Y intensity value will also have four associated likelihood values, one for each of the four bases. The maximum of the four likelihood values indicates the base call. For example, if a cluster is “off’ in both channels, the base call is G. If the cluster is “off’ in one channel and “on” in another channel the base call is either C or T (depending on which channel is on), and if the cluster is “on” in both channels the base call is A.
[0090] In some implementations of RTA, the base calling errors get averaged out across many training examples. In some other implementations, the ground truth may be sourced using aligned genomic data, which may provide better quality because aligned genomic data may use reference genome and truth information that incorporate the knowledge gained from multiple sequencing platforms and sequencing runs to average out the noise. The ground truth may include basespecific intensity values (or feature values) that reliably represent intensity profiles of bases A, C, G, and T, respectively. A base caller like RTA base caller (926) base calls clusters by processing the sequencing images and producing, for each base call, color-wise intensity values / outputs. The color-wise intensity values may be considered base-wise intensity values because, depending on the type of chemistry (e g., 2-color chemistry or 4-color chemistry), the colors map to each of the bases A, C, G, and T. The base with the closest matching intensity profile is called.
[0091] A trainer may train base caller (926) and generate the trained coefficients of the sharpening masks using various training techniques. FIG. 6 shows one implementation of an adaptive technique that may be used to train base caller (926), e.g., using an offline or online mode. Here, the logic is y = x.h + d, where x is the input pixel intensities, h is the sharpening mask coefficients, d is the DC offset. In some implementations, x and h are row and column vectors respectively, with length 81. This vector model is equivalent to a dot product of 9 x 9 matrices representing input pixels and coefficients. The cost is the expected value of error squared. The gradient update moves each coefficient in a direction that reduces the expected value of error squared. Applying this update generates a new estimate of the coefficients that moves them in a direction that (on average) reduces the mean squared error (MSE). In some implementations, Mu is a small constant used to change the adaptation rate / convergence speed. A DC term update can be calculated in a similar way. A gain term update also can be calculated in a similar way.
[0092] In some implementations, since linear interpolation is applied on the coefficient sets, the updates are applied slightly differently in the following manner:
[0093] h(q, n+1) = h(q, n) + lambda q. mu.x(n). e(n)
[0094] In the equation above, h(q, n) is weight q at cycle n, lambda q is the linear interpolation weight for a particular set of coefficients and can include four updates per output due to linear interpolation in two dimensions. The recursive least-squares technique extends the least squares technique to a recursive algorithm.
[0095] C. Examples of Structural Variation Graph Genome Generation
[0096] In some scenarios, a secondary analysis may be performed iteratively while sequence reads are generated by a sequencing system such as systems (100, 500, 814) described herein. Secondary analyses may encompass both alignment of sequence reads to a reference sequence (e.g., the human reference genome sequence) and utilization of this alignment to detect differences between a sample and the reference. Secondary analyses may enable detection of genetic differences, variant detection and genotyping, identification of single nucleotide polymorphisms (SNPs), small insertions and deletion (indels) and structural changes in the DNA, such as copy number variants (CNVs) and chromosomal rearrangements.
[0097] By performing secondary analyses while sequence reads are generated, system (100, 500, 814) may determine preliminary variant calls iteratively in real-time (or with zero or low latency). Final results of variant determinations may be available soon after (or immediately after) the end of a sequencing run. Alternatively, a sequencing run may be terminated early if variant calls are available with sufficient confidence during the run. In some scenarios, only information related to variant determinations (e.g., variant calls) is transferred off the sequencing system (100, 500, 814). This may decrease, or minimize, the data bandwidth required in comparison to performing the variant determinations in a system that is external. In addition, only variant information may be sent to a computing system (e.g., a cloud computing system) for further processing. In this example, sequencing runs may be terminated prior to completion of an entire sequencing process. For example, if the identity of a pathogen of interest is determined after a number of sequencing cycles of a sequencing run, the sequencing run may be terminated. Thus, the time to a particular answer (e g., pathogen identification) may be decreased. In some implementations, outputs and intermediate results of system (100, 500, 814) may include histograms of duplicates, exact matches, single and double SNPs, and single and double indels.
[0098] VI. Examples of Methylated Region-Wide Association Study
[0099] The sequencing systems and methodologies described herein may be utilized in various contexts for the generation of methylation (e.g., epigenetic) data with respect to a given genome or sample. As used herein, methylomics may be understood to relate to the characterization of DNA regions where the degree of methylation is associated with a phenotype. As discussed herein, such regions may be referred to as differentially methylated regions (DMRs). Conventionally, to identify such regions, DNA from multiple people exhibiting different degrees of the phenotype (whether binary or continuous) is extracted and whole-genome methylation sequencing is performed). In practice, however, the number of samples needed may prove problematic for satisfactory results. Further, the lack of a clear standard in defining what was or was not a DMR was equally, if not more, problematic in terms of comparative results and meaningful study.
[0100] As discussed herein, sequencing-based methylation data may be generated using various techniques, one of which is bisulfite sequencing (BiS). In bisulfite sequencing approaches methylated cytosines may be detected in DNA (e.g., genomic DNA) by treating a DNA sample with sodium bisulfite prior to sequencing. In response to the bisulfite treatment, unmethylatedcytosines are deaminated to uracils, which upon sequencing convert to thymidines. Conversely, methylated cytosines are not deaminated and are read as cytosines. The location of the methylated cytosines is then determined by comparison of bisulfite treated sequences to a reference genome or to untreated sequences. In this manner, methylation data for a sample may be obtained using sequencing techniques.
[0101] By way of a further example, enzymatic methyl sequencing (EMS or EM-Seq) is another approach by which methylation data may be obtained using sequencing techniques. In EMS approaches, non-destructive enzymatic reactions are employed which address certain technical biases that may be introduced by BiS techniques. In EMS approaches, a two-step conversion process is employed. In the first reaction TET2 and T4-BGT convert methylated cytosine (and hydroxymethylcytosine) into products that are not deaminated by APOBEC3A (i.e., methylated cytosines are modified to be protected from deamination). In the second reaction, APOBEC3A is used to deaminate unmodified cytosines to uracils. Methylations may then be determined using sequencing techniques and systems as described herein.
[0102] With this in mind, conventional methylation analysis approaches work in a site-specific manner, identifying methylation status site-by-site, even within a larger region. Such a base-level approach may provide data utilized as part of a methylome-wide association study (MWAS) that may be used to identify methylation sites associated with certain traits or conditions (e.g., phenotypes). As discussed herein, however, methylation affects tend to be locally additive. Correspondingly, such site-by-site analyses may be insufficient for assessing the full effects of methylation within a region as these effects relate to a phenotype of interest. With this in mind, and as discussed herein, the presently described techniques relate to a methylated region-wide association study (MRWAS) that may be useful in assessing region-wide methylation effects and, in particular, characterizing differentially methylated regions (DMRs) that may be indicative of phenotype effects that may be associated with differentially methylated regions at different levels (e.g., individual, organ, tissue, single-cell, and so forth).
[0103] As discussed herein, the presently described MRWAS techniques and systems may be used in characterizing DMRs in a defined manner. For example, in contrast to MWAS techniques, the presently described techniques facilitate analysis in a regional manner, as opposed to a sitespecific manner, thereby helping to assess the additive effects of localized methylation. Inparticular, MRWAS techniques as described herein significantly increase sensitivity of the methylation analysis relative to MWAS techniques. Further, the presently described techniques allow for the definition of differentially methylated regions in a rigorous and consistent manner, allowing for the standardization of otherwise subjective terminology. In addition, the MRWAS techniques as presently disclosed do not require any additional prior information, such as gene coordinates or gene sets with pathway descriptions.
[0104] It may be appreciated that, at first impression, a computationally rigorous technique for identifying DMRs would be computationally impossible due to the number of operations involved. For example, the typical number of CpGs is 2.6 x 107, which gives 3.6 x 1013methylated regions and approximately IO20computational operations, which would require approximately 10,000 years to perform. Conversely, in accordance with the techniques disclosed and described herein, all DMRs within a sample can be called within minutes. In particular, as discussed herein, the presently disclosed techniques leverage several concepts including, but not limited to, use of cumulative z-values, implementation of a rolling algorithm that drops CpGs determined not to be part of a more significant DMR than other CpGs, and so forth as discussed below. In this manner, the presently disclosed techniques effectively define DMRs in a statistically rigorous and consistent manner as a consecutive set of CpGs, the significance of association of which pass a statistical threshold for significance (e.g., a Bonferroni corrected threshold for significance).
[0105] With the preceding in mind, sequencing technology systems and methods as described above may be utilized as part of assessing differentially methylated regions within a subject. Such differentially methylated regions and their correspondence to a phenotype of interest may in turn be used for generating a diagnosis and, correspondingly, recommending, selecting, or implementing a treatment plan, such as a pharmacological, radiological, surgical, treatment plan or a suitable combination of such treatment techniques. Further, methylation assessment as discussed herein may be ongoing over time so as to allow or facilitate monitoring of treatment efficacy, adjustment to an ongoing treatment plan, and / or confirmation of a remission state for a subject.
[0106] By way of example, methylation may be a useful tool in identifying and quantifying markers for tumors. In particular, low tumor mutation burden (TMB) and / or low shedding tumors are difficult to detect solely using DNA-based techniques due to the scarcity of signal. As a result,methylation has become a popular tool, such as within the liquid biopsy field, due to the relative increase in signal relative to DNA-based techniques since the methylation signal is not dictated by TMB. Indeed, abnormal methylation patterns may be associated with non-mutated sites. Further, methylation may serve as a tissue-specific marker, which may be of particular importance in cancer diagnosis and in treatment selection and parameterization.
[0107] Conventional approaches for assessing methylation look at a single site of methylation (e.g., a single CpG site) in contrast to the presently described approaches which quantify methylation across a region. With this in mind, the techniques described herein may, in certain implementations, be used to characterize the association of methylation within a region with a phenotype (as opposed to associating methylation at single sites with the phenotype), thereby allowing the identification and characterization of differentially methylated regions.
[0108] It should be understood that the epigenetic information obtainable by the present techniques may be used in a multitude of contexts where regulatory pathways may be implicated in an abnormality or disorder or may otherwise be useful in monitoring (either once or periodically over a period of time) a physiological process. By way of example, the presently described techniques may be useful in the context of reproductive health (e.g., assessing methylation changes or abnormalities during fetal development), in the context of cardiovascular disease (e.g., assessing methylation changes or abnormalities associated with cardiac or vascular function at a moment in time or over a period of time), in the context of neurodegeneration, including Alzheimer’s (e.g., assessing methylation changes or abnormalities associated with neural function, cognition, and / or memory at a moment in time or over a period of time), and / or in the context of metabolic disorders, such as diabetes (e.g., assessing methylation changes or abnormalities associated with metabolic or anabolic function at a moment in time or over a period of time). Further, the presently described techniques may be used in the context of monitoring transplant acceptance (or rejection) where methylation markers may be indicative of the subject’s adaptation to and acceptance of the transplanted tissue or organ. More generally, the presently described techniques may be used in contexts where abnormal methylation patterns are deemed to be of interest or of significance. Additionally, the presently described techniques may be of particular usefulness in situations or uses where a genetic signal (e g., a signal based directly or indirectly on DNA sequence) is low (e.g., at or near a limit of detectability) and / or is unclear, such as due to noise, artifacts, or likelyerror (e.g., sequencing error). In such situations, methylation analysis may help refine the analysis and allow one to observe alterations earlier in time.
[0109] In the oncology context, the presently described techniques may be particularly useful in enabling certain clinical application, such as cancer screening. In particular, such applications may be difficult due to the need to detect abnormalities from a limited sample. In such contexts a clinician may be looking for a small population of abnormal molecules. As discussed herein, methylation may be regulated in regions, not single bases, which is one benefit of a region-based approach as described.
[0110] Turning to FIG. 7, a schematic is shown depicting example corresponding sequences encompassing a methylated region 1000 having three cytosine sites, some of which are methylated. The corresponding sequences shown are listed as being from different DNA molecules, which in turn may be grouped based on sample source (e.g., individuals, tissues, replicates, and so forth). Cytosines within the depicted example sequences are shown by the conventional “C”, with methylated cytosines shown as “5mC”. In this example, the number of methylated molecules may be represented by m while the total number of molecules may be represented by n, with the Yflmethylation level of an individual sample and CpG given by: / ? = — .
[0111] In this example, each cytosine site is depicted as undergoing a respective regression step 1004 (e.g., beta-binomial regression). The output 1008 of each regression step 1004 is a respective A and z / , wherein A corresponds to the observed differential methylation for an individual CpG site corresponding to a sample site or index, z, and z / corresponds to the respective z-value for the null hypothesis that dt = 0. The respective regression outputs 1008 may be combined at step 1012 to generate a respective z-value for the association of the methylated region 1000 with phenotype, as shown by equation 1 (illustrated by reference number 1016 in FIG. 7):
[0112] While the preceding illustrates at a high-level a conceptualized implementation of thepresent techniques, the following figures and discussion provide more detailed walk-through of an example implementation. With this in mind, whole-genome methylation sequencing (WGMS) may be performed to generate methylation sequencing data. In certain implementations, strandspecificity may be ignored or discounted and the most frequent methylation modification sites (e.g., 5mC in CpG dinucleotides) considered. The WGMS data for an individual can be represented as a matrix with four metrics parameters: (1) chromosome, (2) CpG position, (3) number of methylated reads, M, and (4) total number of reads, N.
[0113] With this in mind, and turning to FIG. 8, methylation level data 1050 for all samples is represented as a two-dimensional (2D) matrix iy in which sample indices i run along the vertical axis and CpG indices j run along the horizontal axis. For simplicityis shown as representing methylation level, however in this example it is a combination of two matrices: (1) the number of methylated reads, My, and (2) the total number of reads, Ny. As may be appreciated, the sample or samples may correspond to one individual (or other sample type, such as tissue, replicates, single-cells, and so forth). In this example, each row corresponds to WGMS data of an individual (or other sample categorization) and each column contains methylation information for a single CpG. Also shown in the figures is a matrix corresponding to phenotype data 1054 or categories, which may be binary (e.g., diseased / healthy) or continuous (e.g., age-at-onset) for a group of individuals. As may be appreciated, a continuous phenotype is illustrated in this example. The WGMS matrices 1050 and phenotype matrix 1054 together represent a technical context for which consecutive sets of CpGs are to be identified where the methylation levels are associated with a phenotype. In particular, in certain aspects an objective is to identify or locate a consecutive set of CpGs where methylation levels are associated with a phenotype, referred to herein as a differentially methylated region (DMR) 996
[0114] Turning to FIGS. 9A-9C, WGMS matrices 1050 for a group of individuals (or other sample types) are depicted where each row corresponds to WGMS data of one individual (or other sample type, such as tissue, replicates, single-cells, and so forth) and each column contains methylation information for a single CpG. Also shown is a phenotype 1054 as discussed above. Based on this data methylome-wide association study (MWAS) techniques may be employed in which a regression 1004 is performed, yielding respective 3 values 1058 and z values 1062 for each column (i.e., CpG site) of WGMS data. As discussed herein, the z values 1062 conveysignificance of association of a phenotype with methylation levels for individual CpGs while the t) values 1058 correspond to the respective effect size.
[0115] In accordance with aspects of the presently described techniques, the vector of z values 1062 generated using MWAS techniques is used as an input to subsequent processes for performing methylated region-wide association studies (MRWAS) as discussed herein. Correspondingly, because the MRWAS technique uses MWAS summary statistics as an input, MRWAS may be understood to be a universal method for calling or otherwise characterizing differentially methylated regions (DMRs) for any phenotype, whether binary or continuous, and any sample hierarchy (e.g., different tissues of the same individual, individuals that are closely related, individuals belonging to different ethnic groups, and so forth).
[0116] In particular, as described in this example, the z values 1062 for individual CpGs that are output from the MWAS analysis are received as an input. The methylated region-wide association analysis of these z values 1062 consider some or all possible methylated regions (MRs) and calculate a z value for each MR based on a sum of z-values of the corresponding individual CpGs (see equation 1). Methylated regions passing a threshold for significance (e.g., a Bonferroni corrected threshold for significance) are selected as DMRs (i.e., MRs associated with a phenotype), which is in turn output, such as in the form of a list of DMRs. This output may be used for diagnostic purposes, for devising or selecting treatment options, including personalized treatment options and so forth.
[0117] Turning now to FIGS. 10A-10C, this process is visually illustrated. In this example individual CpGs in are processed in accordance with present technique. In particular, z-values 1062 of each CpG are initially evaluated and their significance calculated using equation 1 (step 1090). If the z value generated using equation 1 for a possible MR passes the threshold for significance (step 1094) (e.g., a Bonferroni corrected threshold for significance), the MR in question is determined to be associated with a phenotype (i.e., the MR is determined to be a DMR) (i.e., a DMR is detected (block 1098)). Otherwise, the MR is not categorized as a DMR and may be skipped (block 1102). Though FIGS, 10A-10C only illustrate the first three CpGs being evaluated, in practice the analysis extends to each individual CpG represented in the z-value vector 1062.
[0118] While FIGS. 10A-10C illustrate an iteration in which the individual CpGs are evaluated, the process continues to extend to multiple CpGs in the vector. For example, turning to FIGS. 11A-11C, each adjacent pair of CpGs is next evaluated using the process as described with respect to FIGS. 10A-10C. As in the preceding discussion, if the z value generated using equation 1 for a possible MR passes the threshold for significance (step 1094) (e.g., a Bonferroni corrected threshold for significance), the MR in question is determined to be associated with a phenotype (i.e., the MR is determined to be a DMR) (i.e., a DMR is detected (block 1098)). Otherwise, the MR is not categorized as a DMR and may be skipped (block 1102).
[0119] Similarly, and turning to FIGS. 12A-12B, an iteration is depicted in which triplets of CpGs in the matrix as analyzed in this manner. For example, turning to FIGS. 12A-12B, each adjacent triplet of CpGs is evaluated using the process as described with respect to FIGS. 10A-10C and FIGS. 11A-11C. As in the preceding discussion, if the z value generated using equation 1 for a possible MR passes the threshold for significance (step 1094) (e.g., a Bonferroni corrected threshold for significance), the MR in question is determined to be associated with a phenotype (i.e., the MR is determined to be a DMR) (i.e., a DMR is detected (block 1098)). Otherwise, the MR is not categorized as a DMR and may be skipped (block 1102). Though the examples illustrated only depict calculations and comparisons up to the triplet level, in practice the described operation may proceed through the length I of the z-value vector 1062 (i.e., proceeding through all possible groups of three, four, five, six, and so forth groupings up to / . If two statistically significant MRs overlap with each other, the MR with lowest statistical significance is skipped.
[0120] While the preceding demonstrates suitable techniques for characterizing DMRs in a rigorous and objective manner, it may be appreciated that such approaches improve granularity of the analysis of methylated regions, thereby improving the usefulness of the results obtained. As noted herein, conventional base-level approaches may be characterized as methylome-wide association studies (MWAS) that may be used to identify methylation sites associated with certain traits or conditions (e.g., phenotypes). Such site-by-site analyses may be insufficient for assessing the full effects of methylation within a region as these effects relate to a phenotype of interest. Conversely, as discussed herein, techniques described herein relate to a methylated region-wide association study (MRWAS) that may be useful in assessing region-wide methylation effects and, in particular, characterizing differentially methylated regions (DMRs) that may be indicative ofphenotype effects that may be associated with differentially methylated regions at different levels (eg., individual, organ, tissue, single-cell, and so forth). As discussed herein, the presently described MRWAS techniques and systems may be used in characterizing DMRs in a defined manner. For example, in contrast to MWAS techniques, the presently described techniques facilitate analysis in a regional manner, as opposed to a site-specific manner, thereby helping to assess the additive effects of localized methylation. In particular, MRWAS techniques as described herein significantly increase sensitivity of the methylation analysis relative to MWAS techniques.
[0121] This may be visualized using Manhattan plots, as shown in FIGS. 13, 14, and 15. As may be appreciated, such Manhattan plots may be used to summarize and visualize association signals seen across a genome (or methylome). Typically, the y-axis corresponds to or is derived using the -loglO(p-value) while the x-axis corresponds to chromosome positions. The y-values of the points in the Manhattan plot are thus inversely related to their p-values. With this in mind, FIG. 13 depicts an MWAS plot derived for a semi-artificial dataset having 1,000 DMRS. As may be observed based on the dotted lines 1150 conveying statistical significance, none of the MWAS data plotted in this manner reaches the level of statistical significance.
[0122] Conversely, and turning to FIG. 14, a Manhattan plot of the MRWAS results for the same data is shown. In FIG. 14, the upper plot depicted the results for chromosomes 1 through 22 while the lower plot shows the results for chromosome 20 in greater detail. As with FIG. 13, the same level of statistical significance (i.e. Bonferroni corrected threshold) is conveyed by dotted lines 1150. As may be appreciated, the MRWAS results show greater granularity and account for regional additive effects, allowing significant DMRs to be readily characterized and localized down to the chromosome region at which they are present.
[0123] In addition, it may be appreciated that the MRWAS is actually three-dimensional in nature, thus the plots of FIG. 14 may be understood to be merely projections of the three-dimensional data into two-dimensional space. This is visually illustrated in FIG. 15, in which the genome and chromosome plots are illustrated in conjunction with a three-dimensional representation of a DMR regions identified on chromosome 20. As may be noted, the three-dimensional nature of the plot is due to the identified region not being a single CpG site, but a regional location having a start-site and a different end-site on the chromosome, thus havingcorresponding data over the range covered by the start site and end-site.
[0124] VII. Examples of Combinations
[0125] The following examples relate to various non-exhaustive ways in which the teachings herein may be combined or applied. The following examples are not intended to restrict the coverage of any claims that may be presented at any time in this application or in subsequent fdings of this application. No disclaimer is intended. The following examples are being provided for nothing more than merely illustrative purposes. It is contemplated that the various teachings herein may be arranged and applied in numerous other ways. It is also contemplated that some variations may omit certain features referred to in the below examples. Therefore, none of the aspects or features referred to below should be deemed critical unless otherwise explicitly indicated as such at a later date by the inventors or by a successor in interest to the inventors. If any claims are presented in this application or in subsequent filings related to this application that include additional features beyond those referred to below, those additional features shall not be presumed to have been added for any reason relating to patentability.
[0126] Example 1
[0127] A method for categorizing methylated regions as differentially methylated regions, the method comprising: accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values. The methylated region-wide association analysis on the set of z-values is performed by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
[0128] Example 2
[0129] The method of Example 1, further comprising generating the methylation sequencingdata set using whole-genome methylation sequencing.
[0130] Example 3
[0131] The method of Examples 1 or 2, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
[0132] Example 4
[0133] The method of Examples 1, 2, or 3, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
[0134] Example 5
[0135] The method of Examples 1, 2, 3, or 4, wherein the samples comprise one or more of individuals, tissues, replicates, or single cells.
[0136] Example 6
[0137] The method of Examples 1, 2, 3, 4, or 5, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
[0138] Example 7
[0139] The method of Examples 1, 2, 3, 4, 5, or 6, wherein the respective sum of z-values is determined for each possible methylated region from a length of 1 to a length of I, wherein / corresponds to the total length of the set of z-values.
[0140] Example 8
[0141] The method of Examples 1, 2, 3, 4, 5, 6, or 7, wherein the threshold comprises a Bonferroni corrected threshold for significance.
[0142] Example 9
[0143] The method of Examples 1, 2, 3, 4, 5, 6, 7, or 8, wherein respective methylated regions that do not pass the threshold are skipped or discarded.
[0144] Example 10
[0145] A processor-based system, the processor-based system comprising: at least one processor configured to execute stored routines and one or more tangible storage media storing processor-executable routines. The processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising: accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values. The methylated region-wide association analysis on the set of z-values is performed by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
[0146] Example 11
[0147] The processor-based system of Example 10, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
[0148] Example 12
[0149] The processor-based system of Examples 10 or 11, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
[0150] Example 13
[0151] The processor-based system of Examples 10, 11, or 12, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
[0152] Example 14
[0153] The processor-based system of Examples 10, 11, 12, or 13, wherein the respective sumof z-values is determined for each possible methylated region from a length of 1 to a length of / , wherein / corresponds to the total length of the set of z-values.
[0154] Example 15
[0155] The processor-based system of Examples 10, 11, 12, 13, or 14, wherein respective methylated regions that do not pass the threshold are skipped or discarded.
[0156] Example 16
[0157] One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform steps comprising: accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values. The methylated region-wide association analysis on the set of z-values is performed by performing acts comprising: determining a respective combined value of z-values for methylated regions comprising consecutive CpGs; comparing each respective combined value of z-values to a threshold; and based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
[0158] Example 17
[0159] The one or more computer-readable media of Example 16, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
[0160] Example 18
[0161] The one or more computer-readable media of Examples 16 or 17, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
[0162] Example 19
[0163] The one or more computer-readable media of Examples 16, 17, or 18, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
[0164] Example 20
[0165] The one or more computer-readable media of Examples 16, 17, 18, or 19, wherein the respective sum of z-values is determined for each possible methylated region from a length of 1 to a length of / , wherein I corresponds to the total length of the set of z-values.
[0166] VIII. Miscellaneous
[0167] While the foregoing examples are provided in the context of a system (100) that may be used in nucleotide sequencing processes, and in particular the sequencing of methylation data, the teachings herein may also be readily applied in other contexts, including in systems that perform other processes (i.e., other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes.
[0168] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in the description herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,” “comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
[0169] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understoodas indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.
[0170] Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.
[0171] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departing from its scope. While the dimensions, types of materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U.S.C. §112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.
[0172] The following claims recite aspects of certain examples of the disclosed subject matterand are considered to be part of the above disclosure. These aspects may be combined with one another.
Claims
What is claimed is:
1. A method for categorizing methylated regions as differentially methylated regions, comprising:accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; andperforming a methylated region-wide association analysis on the set of z-values by performing acts comprising:determining a respective combined value of z-values for methylated regions comprising consecutive CpGs;comparing each respective combined value of z-values to a threshold; and based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
2. The method of claim 1, further comprising:generating the methylation sequencing data set using whole-genome methylation sequencing.
3. The method of claim 1, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
4. The method of claim 1, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
5. The method of claim 1, wherein the samples comprise one or more of individuals, tissues, replicates, or single cells.
6. The method of claim 1, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
7. The method of claim 1, wherein the respective sum of z-values is determined for each possible methylated region from a length of 1 to a length of I, wherein / corresponds to the total length of the set of z-values.
8. The method of claim 1, wherein the threshold comprises a Bonferroni corrected threshold for significance.
9. The method of claim 1, wherein respective methylated regions that do not pass the threshold are skipped or discarded.
10. A processor-based system, comprising:at least one processor configured to execute stored routines;one or more tangible storage media storing processor-executable routines, wherein the processor-executable routines, when executed by the at least one processor, cause acts to be performed comprising:accessing or acquiring a methylation sequencing data set and a phenotype data set;regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z- value for each CpG site of the methylation sequencing data set; and performing a methylated region-wide association analysis on the set of z-values by performing acts comprising:determining a respective combined value of z-values for methylated regions comprising consecutive CpGs;comparing each respective combined value of z-values to a threshold; andbased upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
11. The processor-based system of claim 10, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
12. The processor-based system of claim 10, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
13. The processor-based system of claim 10, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
14. The processor-based system of claim 10, wherein the respective sum of z-values is determined for each possible methylated region from a length of 1 to a length of / , wherein / corresponds to the total length of the set of z-values.
15. The processor-based system of claim 10, wherein respective methylated regions that do not pass the threshold are skipped or discarded.
16. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform steps comprising:accessing or acquiring a methylation sequencing data set and a phenotype data set; regressing, per CpG site, methylation level data against phenotype data of the phenotype data set to generate a set of z-values comprising a respective z-value for each CpG site of the methylation sequencing data set; andperforming a methylated region-wide association analysis on the set of z-values by performing acts comprising:determining a respective combined value of z-values for methylated regions comprising consecutive CpGs;comparing each respective combined value of z-values to a threshold; and based upon the comparison, classifying respective methylated regions that pass the threshold as differentially methylated regions that are associated with a phenotype.
17. The one or more computer-readable media of claim 16, wherein the methylation sequencing data set is represented as one or more matrices conveying chromosome, CpG position, number of methylated reads, and total number of reads.
18. The one or more computer-readable media of claim 16, wherein the methylation sequencing data set comprises methylations levels at different CpG sites for a plurality of samples.
19. The one or more computer-readable media of claim 16, wherein the z-values correspond to significance of association of a respective phenotype with methylation levels for respective CpG sites.
20. The one or more computer-readable media of claim 16, wherein the respective sum of z-values is determined for each possible methylated region from a length of 1 to a length of I, wherein I corresponds to the total length of the set of z-values.